Why our agent OS runs on Gleam and Zig, not Python or Rust
Agent infrastructure—for example, an agent operating system like WunderOS—does two kinds of work and their requirements run in opposite directions:
- Coordination: sessions, tool calls, retries, timeouts, admission, and the steady arrival of failures from networks, models and tenant code. I use the BEAM/OTP and Gleam for this.
- Computation: inference, vector operations, retrieval kernels and storage I/O, where the requirements are throughput, memory layout and exact reproducibility. I use Zig for this.
This post explains why I chose this stack and where its split falls, then dives deep on the interface between the two runtimes, because that boundary is where systems like this usually fail. I generate that boundary with a homegrown compiler of sorts and describe what the compiler must guarantee and why a hand-written boundary isn’t good enough.
Contents: The Gleam/Zig split · Why not Python and Rust · Ports as records of functions · Pin the model, never adapt to the workload · Compiling the BEAM–Zig boundary · Hot code reload across two runtimes · Determinism is a ladder · What this stack costs
The Gleam/Zig split
If it’s expected to fail, it runs on the BEAM; if it’s got to be fast and exact, it runs in Zig.
An agent session is a process. It owns its state, is supervised, and when it crashes it is restarted from its last checkpoint by a supervisor that knows nothing about agents. OTP has handled this pattern in telephone switches for thirty years. WhatsApp and Discord run on the BEAM for similar reasons. A POTS call, a chat connection and an agent session have one similar shape: long-lived, stateful, and subject to failure from things they do not control. That resemblance is why I chose the BEAM.
Session state never leaves the BEAM. Native work crosses through generated NIFs, each declared with its scheduling class. Tenant code runs in venues that the platform can hibernate and restore, but never treats as durable.
In Gleam, the owner of one session’s lifecycle is a process whose message type the compiler checks:
type Message {
Hibernate(Subject(Result(OperationRef, Failure)))
AcquireVisit(
String,
process.Pid,
Int,
fn(trace.Sample) -> Nil,
Subject(Acquisition),
)
ReleaseLease(Int, Int, String, Subject(Result(Nil, Failure)))
Deliver(Int, Int, String, BitArray, Subject(Result(Nil, Failure)))
// …
Finished(Int, Int, Step, Result(Outcome, String), Int)
Permit(Int, Int, Bool, Subject(Result(Nil, String)))
WorkerLost
LeaseLost
}
Its loop is one case over that type. A loop that forgets a message won’t
compile, which matters more here than in most Erlang code because session
protocols change often. Failure of the worker that drives the session arrives as
an ordinary message, WorkerLost, through a monitor:
fn loop(s: State, inbox: Subject(Message)) -> Nil {
let selector = process.new_selector() |> process.selecting(inbox, fn(m) { m })
let selector = case s.worker {
Some(worker) ->
process.selecting_process_down(selector, worker.monitor, fn(_) {
WorkerLost
})
None -> selector
}
// …
case process.select_forever(selector) {
Inspect(reply) -> {
process.send(reply, projection(s))
loop(s, inbox)
}
// …
Hibernate(reply) -> {
let #(next, answer) = hibernate_request(s, inbox)
process.send(reply, answer)
loop(next, inbox)
}
// …
LeaseLost -> loop(State(..s, lease: None), inbox)
// …
}
}
Tenant code gets handler, not process, semantics. It runs in a per-tenant microVM or, for native WunderOS agents written in TypeScript or JavaScript, in an interpreter that runs as one of my OTP applications. It never lands on the VM as BEAM modules or native libraries. Memory inside the tenant’s venue is not the system of record. The session owner above can hibernate an idle venue: freeze it, snapshot it, and restore it later, so a warm venue resumes where it stopped. A venue can also be lost, so anything the tenant must remember goes through the platform’s memory interface. That is a programming-model constraint: tenant code is written as request handlers. The constraint buys strong isolation and BEAM-native hot code reload at once, which otherwise exclude each other.
Zig gets the work where a garbage collector, a scheduler preemption or an unpredictable allocation would cost either throughput or reproducibility: SIMD kernels over 64-byte-aligned structures, io_uring storage paths, an agent-native database, and the model inference engine.
Why not Python and Rust
Python is the default language of AI work and Rust is the default for infrastructure that must be fast and safe, so the question of why I use neither deserves an answer.
Python’s best place is research, model development and, often, deployment. However, as an agent operating system control plane, it lacks what the BEAM provides: preemptive scheduling across millions of lightweight processes, a separate heap per process so that one session’s garbage collection never pauses another, and supervision as a property of the runtime rather than of a library.
Rust is a stronger alternative to Gleam than Python and I considered it seriously. Its async runtimes are fast, but their scheduling is cooperative: a task that does not yield stalls its executor thread. That’s the failure the BEAM rules out for Erlang code, and the one I had to engineer around at the NIF boundary. In Rust, isolation and supervision are assembled from libraries. On the BEAM they are the runtime’s own model. For this kind of work its concurrency model is unsurpassed so far: no other runtime I know of offers the same guarantees with thirty years of production use behind them.
For the data plane, again, I considered Rust and Zig, and the deciding difference is the ecosystem, perhaps counter-intuitively. Rust’s crate ecosystem is large and high quality; an ordinary service pulls in hundreds of transitive dependencies. Who wrote a dependency matters less every year. What matters is that I cannot review it in full and must trust it again at every update. Attacks on package registries are now routine, and a system’s attack surface grows with its software bill of materials.
Zig’s ecosystem is relatively small, which obliges me either to write what I need or to redesign the system until I don’t need it. Five years ago that would have been a poor trade. Agentic software engineering has made writing code much cheaper, while reviewing someone else’s code costs about what it always has, so the balance has shifted. A dependency once saved weeks of work; it now saves far less and costs nearly as much in trust. An SBOM short enough to read in full is a security property in its own right.
The trade is not free. Code I write myself is code I must test and maintain, and a small dependency graph exchanges other people’s bugs for my own. I accept that exchange because my own bugs fall under my own review, determinism and replay tooling, and other people’s don’t.
All that said, a pure Rust operating system for agents is a perfectly sane choice, but it wasn’t mine.
Ports as records of functions
Gleam has no type classes, no interfaces and no traits, and Rust or Haskell engineers may take that as a gap. But for hexagonal architecture, it’s not. A port is a record whose fields are functions, and an adapter is any function that builds such a record. The session owner above is written against this port:
/// Resource operations belong to the concrete microVM-controller backend.
pub type Backend {
Backend(
observe_trace: fn(trace.Sample) -> Nil,
record_lifecycle: fn(List(BitArray)) -> Result(journal.RecordDurability, String),
attach_owner: fn(process.Pid) -> Result(Nil, String),
check: fn(Admission) -> Result(Nil, String),
freeze: fn(Int) -> Result(cp.Checkpoint, String),
prepare: fn(cp.Checkpoint) -> Result(Nil, String),
release_vm: fn() -> Result(Nil, String),
// …
fence: fn() -> Nil,
stop: fn() -> Result(Nil, String),
// …
)
}
The production adapter turns each field into a call to the actor that drives the microVM:
pub fn operations(handle: Handle, config: Config) -> owner.Backend {
owner.Backend(
observe_trace: config.observe_trace,
// …
freeze: fn(id) {
use answer <- result.try(call(handle, Freeze(id)))
case answer {
Frozen(token) -> Ok(token)
_ -> Error("missing checkpoint")
}
},
prepare: fn(token) { call(handle, Prepare(token)) |> result.replace(Nil) },
release_vm: fn() { call(handle, Release) |> result.replace(Nil) },
// …
stop: fn() { call(handle, Stop) |> result.replace(Nil) },
// …
)
|> visit_lifecycle_tape.record_owner(config.lifecycle_mode, _)
}
The test adapter is a record of closures that report each effect to the test process, with no VM anywhere:
let effect = fn(name) {
process.send(events, name)
Ok(Nil)
}
owner.Backend(
observe_trace: fn(_) { Nil },
// …
prepare: fn(_) { effect("snapshot") },
release_vm: fn() { effect("release") },
load: fn(_) { effect("load") },
// …
stop: fn() { effect("stop") },
// …
)
Record update syntax makes fault injection one line. This test replaces one field with a failure and checks that the owner marks the session for recovery:
let bad =
start(owner.Backend(..backend(events), stop: fn() { Error("VM remains") }))
let assert owner.Ready(lease) = owner.acquire(bad, "fire")
owner.terminate_lease(lease) |> should.be_error
let assert Ok(state) = owner.status(bad)
state.recovery_required |> should.be_true
Ports also compose. The record_owner call at the end of the production adapter
is a decorator: in recording mode it wraps one field so that every lifecycle
record also goes to the replay tape, and in every other mode it returns the port
unchanged.
pub fn record_owner(
mode: visit_replay.Mode,
backend: visit_owner.Backend,
) -> visit_owner.Backend {
case mode {
visit_replay.Recording(recorder) | visit_replay.Branching(_, recorder) ->
visit_owner.Backend(..backend, record_lifecycle: fn(records) {
use _ <- result.try(backend.record_lifecycle(records))
visit_replay.record_lifecycle_batch(recorder, records)
|> result.map(fn(_) { visit_journal.TraceFlushed })
|> result.replace_error("lifecycle tape recording failed")
})
_ -> backend
}
}
I use the pattern only on the Gleam side. The Zig side selects implementations at compile time, and repeating the hexagonal pattern there would add indirection to the hot path for no benefit. The vector algebra, for example, binds to AVX-512 or NEON kernels when the library is built:
const impl = if (config.explicit_simd or (builtin.cpu.arch == .x86_64))
switch (builtin.cpu.arch) {
// AMD EPYC / Intel Xeon (Utilizes 512-bit ZMM registers)
// 16,384 bits = 32 instructions
.x86_64 => @import("algebra_avx512.zig"),
// Apple Silicon (M-Series) / Standard ARM (Utilizes 128-bit NEON registers)
// 16,384 bits = 128 instructions
.aarch64 => @import("algebra_neon.zig"),
else => @compileError("Explicit SIMD requested but architecture is not explicitly mapped."),
}
else
// Fallback to the LLVM auto-vectorized scalar implementation
@import("algebra_scalar.zig");
Pin the model, never adapt to the workload
The small model inference engine, Zug, does not adapt to the traffic it serves. A model enters service as a serving variant: a registry entry that pins the engine artifact and the weight bundle by digest. The control plane pins a variant digest on every request, and the execution plane resolves exactly that variant:
//! The execution plane resolves a serving variant's native bundle by the
//! exact `variant_digest` the control plane pinned on the request. Unknown
//! fields are ignored; other `schema_version` values are refused whole.
// …
pub const Variant = struct {
name: []const u8,
variant_digest: []const u8,
engine_artifact: []const u8,
bundle_digest: []const u8,
bundle_path: []const u8,
execution: execution.Values,
};
A model promotion changes data, not code. The engine is stable and loads the promoted bundle at run time. Promotion atomically moves the active variant reference: new requests use the new variant, and requests already in flight finish on the old one. The promotion record holds both digests. Rollback is the same operation in the other direction. Only a new engine is a code change. A code change goes through the hot-reload procedure described below.
The GPU engine is pinned the same way. Before it loads the CUDA library, the loader checks the library’s hash and the runtime it was built against:
pub const Contract = struct {
library_blake3: []const u8,
cuda_runtime: u32,
cublas: u32,
sm: u32,
abi: u32,
kernel_policy: u32,
};
The reason for pinning is reproducibility. An engine that retunes itself under load, by changing batch shapes, kernel choice or reduction order, produces outputs that depend on what else was running at the time. Whatever value this has, it’s not a luxury I allow myself in WunderOS, which is heavily predicated on the value of providing deterministic, replayable execution of agent fleets as one of the primary distinctions between an OS for fleets of agents and a harness for one agent.
Reduction order is the hazard that is easiest to miss. Floating-point addition
is not associative, so a dot product summed 16 lanes at a time on AVX-512 and 4
lanes at a time on NEON can differ in its last bits. My CPU engine avoids this
by never reducing across lanes. Each lane owns an output row, and each row is
summed in the order the scalar loop uses. @setFloatMode(.strict) stops the
compiler from reassociating the sums or fusing the multiply into the add:
fn linear(w: []const f32, x: []const f32, y: []f32) void {
@setFloatMode(.strict);
// Each lane owns a row. Never reduce across columns: D0 requires the
// scalar multiply/add order, including separate rounding of the product.
const lanes = 8;
var row: usize = 0;
while (row + lanes <= y.len) : (row += lanes) {
var sum: @Vector(lanes, f32) = @splat(0);
for (x, 0..) |v, column| {
var weights: @Vector(lanes, f32) = undefined;
inline for (0..lanes) |lane| weights[lane] = w[(row + lane) * x.len + column];
sum += weights * @as(@Vector(lanes, f32), @splat(v));
}
y[row..][0..lanes].* = sum;
}
for (y[row..], row..) |*out, tail_row| {
var sum: f32 = 0;
for (x, w[tail_row * x.len ..][0..x.len]) |v, weight| sum += weight * v;
out.* = sum;
}
}
The vector width then changes the speed and not the bits. A test compares the result with the scalar loop, bit for bit, over row and column counts chosen to hit the tails.
Compiling the BEAM–Zig boundary
A native implemented function (NIF) runs inside the BEAM’s address space, on one of its scheduler threads, with none of its protections. Three failures follow from that. Everyone who writes NIFs by hand meets all three.
- A crash is total. A segfault in a NIF takes down the VM and every supervisor with it. The fault-tolerance argument for the BEAM does not extend across this line.
- A slow call is contagious. The schedulers are preemptive only for Erlang code. A NIF that runs for 20 milliseconds holds its scheduler for 20 milliseconds, and every process queued there waits. The documented budget for a regular NIF is about one millisecond.
- Lifetimes cross the line silently. A resource handed to Erlang outlives the call that made it, and may outlive the version of the library that allocated it. A stale pointer here is a use-after-free with scheduler privileges.
None of these is hard to handle once. Each is hard to handle in every one of 648 exported functions, written by different people, revised under deadline.
So I stopped writing the boundary and started generating it. My NIF compiler
reads one declaration file, nifs.zon, and emits
- Zig wrappers and argument marshalling,
- resource-type table,
- Erlang stub module,
- raw Gleam bindings,
- a typed registry of the surface, and
- its documentation.
Engineers write the kernel and the declaration. The few places that need
something the generator cannot emit, such as a resource type with a stop
callback for enif_select, are hand-written instead.
Declaring the scheduling class
Every exported function states how it behaves with respect to the scheduler:
.{
.name = "zug_batch_at_boundary",
.args = .{
.{ .name = "model", .type = "resource:zug_batch_model" },
},
.return_type = "bool",
.scheduler = "normal",
.tier = "none",
.status = "blessed",
},
.{
.name = "zug_batch_step",
.args = .{
.{ .name = "batch", .type = "resource:zug_batch" },
.{ .name = "signals", .type = "bytes" },
},
.return_type = "binary",
.scheduler = "dirty_cpu",
.chunk_us = 200000,
.tier = "none",
.status = "blessed",
},
.{
.name = "owner_trace_wal_put",
.args = .{
.{ .name = "wal", .type = "resource:owner_trace_wal" },
.{ .name = "bytes", .type = "bytes" },
},
.return_type = "binary",
.scheduler = "dirty_io",
.tier = "none",
.status = "blessed",
},
.{
.name = "slab_ingest_bulk",
.args = .{
.{ .name = "slab_ptr", .type = "int64" },
.{ .name = "tsv_data", .type = "binary" },
},
.return_type = "any",
.scheduler = "dirty_cpu",
.prod = false,
.tier = "multi",
.status = "blessed",
},
The classes map to what the runtime offers. Of the 648 functions, 357 run on
normal schedulers and must finish inside the one-millisecond convention. 112
block on storage and run on the dirty I/O schedulers. 179 are long CPU work on
the dirty CPU schedulers, and these declare chunk_us, the longest they may run
between reports to the VM. A CI gate measures each budgeted function and fails
the pull request when its p99 exceeds 1.5 times its declared budget. prod =
false marks development tooling, which may never be hot-reloaded.
The compiler refuses a declaration it cannot honor:
validate_budget_placement(nif) catch |err| switch (err) {
error.ChunkUsOnNormalScheduler => {
std.debug.print(
"nif_gen: ADR-054 §5 — NIF '{s}' is scheduler=\"normal\" but declares chunk_us; " ++
"the Fast tier has no per-yield budget. Remove chunk_us or move it to dirty_cpu.\n",
.{nif.name},
);
std.process.exit(1);
},
};
// …
const status = ratchet_check(nif, budget_baseline.contains(nif.name)) catch |err| switch (err) {
error.NewUnbudgetedDirtyCpu => {
std.debug.print(
"nif_gen: ADR-054 §5 — new dirty_cpu NIF '{s}' has no chunk_us budget and is not grandfathered. " ++
"Declare chunk_us (with an enif_consume_timeslice yield) or prod=false. " ++
"Grandfathering is closed (PEN-3621).\n",
.{nif.name},
);
std.process.exit(1);
},
// …
The budget rule arrived after a few hundred thousand lines of code, so it’s a ratchet. 101 older dirty CPU functions are listed in a baseline file without a budget. The file may only shrink: a new function without a budget fails the build, and so does a budgeted function whose name is still in the file.
Chunked kernels
Dirty schedulers are a finite pool, and a kernel parked on one for a long time
is a kernel nobody can interrupt. My answer is to keep long work divisible. The
kernel’s state lives in a resource that Erlang holds. Each call does one bounded
unit of work, reports it to the VM with enif_consume_timeslice, and returns. A
Gleam process owns the loop and decides whether to call again.
Zug’s token generator is the largest example. One call advances every live sequence in a batch by one step:
pub fn step(env: ?*e.ErlNifEnv, resource: *Resource, flags: []const u8) !e.ErlNifBinary {
if (flags.len % 16 != 0 or flags.len > 64 * 16) return error.InvalidSignals;
var signals: [64]owned.Signal = undefined;
for (signals[0 .. flags.len / 16], 0..) |*signal, i| {
const mask = std.mem.readInt(u64, flags[i * 16 + 8 ..][0..8], .little);
if (mask > 3) return error.InvalidSignals;
signal.* = .{ .handle = std.mem.readInt(u64, flags[i * 16 ..][0..8], .little), .cancelled = mask & 1 != 0, .expired = mask & 2 != 0 };
}
try lock(resource);
defer unlock(resource);
if (resource.gpu) |state| {
defer resource.model.gpu_flight.store(state.owner.flight_count != 0, .release);
return state.step(signals[0 .. flags.len / 16]);
}
const owner = resource.owner orelse return error.Closed;
var binary: e.ErlNifBinary = undefined;
if (e.enif_alloc_binary(frame.capacity(owner.count - owner.waiting), &binary) == 0) return error.OutOfMemory;
errdefer e.enif_release_binary(&binary);
const view = try owner.advance(resource.model.weights, signals[0 .. flags.len / 16]);
frame.encode(view, binary.data[0..binary.size]);
if (env) |actual| _ = e.enif_consume_timeslice(actual, 100);
return binary;
}
Cancellation and deadline expiry arrive as signals with the next step, so the native side never has to call back into the BEAM to learn that a caller has gone away.
Where the work is not something that I can divide unilaterally, such as a query
stream whose producer is a worker thread, the continuation is hand-written with
enif_schedule_nif. Each re-entry waits at most one millisecond for the worker
and then reschedules itself:
fn gleamalog_stream_continue(env: ?*e.ErlNifEnv, argc: c_int, argv: [*c]const e.ERL_NIF_TERM) callconv(.c) e.ERL_NIF_TERM {
// …
if (stream.pending_terminal and state.done.load(.acquire) != 0)
return emit_stream_terminal(env, state) catch make_error_atom(env, "out_of_memory");
// Bound each dirty-IO continuation. No normal scheduler waits for a worker.
std.Thread.Futex.timedWait(&state.done, 0, std.time.ns_per_ms) catch {};
_ = state.continuations.fetchAdd(1, .monotonic);
return e.enif_schedule_nif(env, "gleamalog_stream_continue", e.ERL_NIF_DIRTY_JOB_IO_BOUND, gleamalog_stream_continue, 2, argv);
}
Both shapes share one property: no state lives on the C stack across a return. That costs a little. It’s the property that makes the next section possible. Between any two calls, the library can be replaced, and the work continues in new code with its old state.
Resource identity across versions
When the library is upgraded, existing resources must keep their type. Every
resource type is opened with ERL_NIF_RT_CREATE | ERL_NIF_RT_TAKEOVER, so the
same function serves first load and upgrade. On upgrade, the takeover binds the
new library’s destructor to resources the old library created, and nothing is
reallocated:
fn open_all_resource_types(env: ?*e.ErlNifEnv) c_int {
core.WMCognitiveContextType = e.enif_open_resource_type(env, null, "WMCognitiveContext", wm_destructor, e.ERL_NIF_RT_CREATE | e.ERL_NIF_RT_TAKEOVER, null) orelse return 1;
// PEN-2505 multi-resource: every type declared in nifs.zon's
// `.resource_types` block (generator already emits CREATE | TAKEOVER).
if (bp.open_resource_types(env) != 0) return 1;
// …
if (!core.vsock.open_resource_type(env)) return 1;
return 0;
}
Resource types are only half the problem. The other half is the library’s own
live state: the io_uring thread, workspaces, write-ahead logs. A new library
image starts with its globals reset, so that state rides across the upgrade in a
bundle passed through the VM’s priv_data. The bundle’s field types mirror the
globals with @TypeOf, so it cannot drift from what it carries:
pub const NifState = struct {
bg: @TypeOf(core.global_bg) = null,
sinkhorn_ws: @TypeOf(core.global_sinkhorn_ws) = null,
gw_ws: @TypeOf(core.global_gw_ws) = null,
// …
The upgrade callback adopts the bundle and then transfers ownership of it:
fn nif_upgrade(
env: ?*e.ErlNifEnv,
priv_data: [*c]?*anyopaque,
old_priv_data: [*c]?*anyopaque,
load_info: e.ERL_NIF_TERM,
) callconv(.c) c_int {
_ = load_info;
const raw = old_priv_data[0] orelse return 1;
const old: *nif_state.NifState = @ptrCast(@alignCast(raw));
if (!old.accepts_upgrade()) return 1;
// …
// §3 inv2: rebuild this image's deterministic dispatch tables in new code.
encoder.init();
oracle.init();
// §3 inv1: re-bind resource types to new-code destructors via RT_TAKEOVER.
if (open_all_resource_types(env) != 0) return 1;
// …
old.restore_to_globals(); // new globals + registries point at old's objects
// …
priv_data[0] = @ptrCast(old);
old_priv_data[0] = null;
The last line is the one hand-written upgrade code most often gets wrong. The VM
runs the old library’s unload after the swap. If the old slot still pointed at
the bundle, the old library would free objects the new library now holds.
Nulling it makes the old unload free nothing.
The Gleam side
Gleam calls Erlang through @external, so the compiler also emits the Erlang
module that loads the library in -on_load. Each function appears twice in it:
zug_batch_step(Arg0, Arg1) -> 'replay@activity_wrap':c1(13583521933491346705, ?MODULE, zug_batch_step_raw, [Arg0, Arg1]).
zug_batch_step_raw(_Arg0, _Arg1) -> erlang:nif_error(nif_library_not_loaded).
load_nif replaces only the _raw stub. The public function keeps its Erlang
body, which sends every call through the replay recorder with a stable activity
ID. When no recorder is attached, that path returns at once. The replay section
below depends on this: every crossing of the boundary can become a trace record,
and no engineer has to remember to make it so.
The Gleam bindings are generated from the same declarations:
@external(erlang, "wunderblock_nif", "zug_batch_step")
pub fn zug_batch_step(arg_0: Dynamic, arg_1: BitArray) -> Dynamic
Arity and argument types come from the declaration, so a change to a function’s
arguments breaks the Gleam build at every raw call site. Results cross as
Dynamic and are decoded in one hand-written module per family, which turns a
malformed result into a typed error at that point:
use bytes <- result.try(native_result(
nif_bindings.zug_batch_step(batch.native, flags),
dynamic.bit_array,
))
use decoded <- result.try(
frame.decode(bytes) |> result.replace_error("invalid native batch frame"),
)
Hot code reload across two runtimes
The BEAM can replace running code without stopping, and fleet-scale customer-agent sessions that last days make that a capability worth having. If WunderOS succeeds, it may never be the case that in an enterprise tenant there is any quiet time for a rolling upgrade. Much like the telephone system then and cellular network now.
The difficulty is that only half my system is Erlang code. Reload has to hold across three layers, each with its own mechanism.
The Gleam control plane uses the OTP’s own mechanism: two versions of a module
may be loaded at once, and a process migrates its state through code_change/3.
Gleam’s type system cannot name the old state type after an upgrade, so my rule
is that upgradable state is a tagged union with an explicit version, and each
migration is a separate module that decodes the old version from Dynamic. A
migration may not emit trace records. Gleam’s actor library does not expose
code_change/3. I will generate it for each actor that declares an upgrade
path; but that generator doesn’t exist yet.
The data plane uses the library’s upgrade callback from the previous section,
and that part is built. So is the counter the drain waits on. Every generated
wrapper for a reload-eligible function brackets the call:
export fn nif_zug_batch_at_boundary(env: ?*erl.ErlNifEnv, argc: c_int, argv: [*c]const erl.ERL_NIF_TERM) erl.ERL_NIF_TERM {
core.inflight.enter();
defer core.inflight.exit();
// …
pub fn enter() void {
_ = in_flight.fetchAdd(1, .monotonic);
}
pub fn exit() void {
const prev = in_flight.fetchSub(1, .monotonic);
std.debug.assert(prev > 0); // never decrement below zero (leak/double-dec)
}
What remains is the release handler that runs the swap. Its procedure is:
- Stop dispatching new work to the affected kernels.
- Wait for the in-flight count and the io_uring submission queue to reach zero. The wait is bounded because every kernel has a budget. The target is a p99 under 100 milliseconds; past one second the upgrade is refused and the refusal is written to the trace.
- Upgrade the control plane first, then load the new library, which runs
upgrade. - Write a discontinuity marker to the trace and resume dispatch. Chunked work re-enters in new code with its old state.
Until the release handler lands, nothing in production calls enif_upgrade; the
callback is wired and tested, and waits for its caller.
The third layer is replay, and it was built very early on in the current WunderOS dev lifecycle. Here I refuse to make reload transparent. Deterministic replay is defined against fixed code, so a replay across an upgrade is a different simulation. Every trace record carries the version of the code that produced it:
pub type ActivityRecord {
ActivityRecord(
activity_id: ActivityId,
execution_position: Int,
// …
result_hash: Blake3Hash,
timestamp_micros: Int,
code_version: GitSha,
schema_version: Int,
lineage_pointer: Blake3Hash,
)
}
When replay reaches a discontinuity marker, it compares the marker’s two versions with the binary that is running, and reports a cross-version replay at the top of its output instead of failing somewhere in the middle.
Determinism is a ladder
“Deterministic” is used loosely about model inference, and the loose usage hides real engineering choices. I scope it as five rungs, each stronger than the last.
| Rung | Property |
|---|---|
| D0 | The same variant, hardware, batch, input, and decoding state produce the same result. |
| D1 | Batch composition and batch size do not change one request result. |
| D2 | Prefill and decode phase boundaries do not change one token result. |
| D3 | CPU and GPU placement do not change the result. |
| D4 | ISA, device generation, and vendor do not change the result. |
Each rung forbids something. D1 forbids reduction orders that change with batch size. D2 forbids a numeric split between the single-token and multi-token code paths, which several popular runtimes have. It also permits a great deal: speculative decoding, prefix caching and n-gram drafting all compute the same token by a different route, so under D2 they change latency and nothing else.
I do not promise every rung everywhere. Ordinary serving requires D0. Re-inference for audit requires D2, and a variant certified at D2 can also serve ordinary traffic. CPU and GPU serve as separate variants, so I do not promise D3 or D4 today.
Sampling is the place where most implementations break the ladder without noticing. A shared random stream makes output depend on arrival order. Zug gives each sequence its own stream, keyed by a hash of the request’s seed, the serving variant’s digest and the sequence’s index:
/// Key derivation, fixed by the harness: BLAKE3-32 over
/// `seed LE64 || serving_variant_digest[32] || sequence LE32`, first 8 bytes
/// little-endian. Two sequences of one request never share a stream, and a
/// re-served variant never replays another variant's stream.
pub fn deriveKey(seed: u64, variant_digest: [32]u8, sequence: u32) u64 {
var msg: [8 + 32 + 4]u8 = undefined;
std.mem.writeInt(u64, msg[0..8], seed, .little);
@memcpy(msg[8..40], &variant_digest);
std.mem.writeInt(u32, msg[40..44], sequence, .little);
var h: [32]u8 = undefined;
std.crypto.hash.Blake3.hash(&msg, &h, .{});
return std.mem.readInt(u64, h[0..8], .little);
}
The seed is part of the request’s contract, alongside temperature. The generator’s state is part of a sequence’s checkpoint, so a restored sequence continues with the same draws:
pub const Checkpoint = struct { produced: usize, rng: [4]u64 };
The purpose of all this is replay. A customer agent’s decision can be audited only if it can be reproduced, and it can be reproduced only if every layer under it is pinned: the code version stamped into each trace record, the model variant addressed by its digest, and the inference beneath both deterministic to whichever rung the deployment promises.
What this stack costs
This stack has costs, and a reader deciding whether to adopt it should weigh them.
- Gleam is young. The OTP library covers actors and supervision well and
leaves gaps elsewhere, of which
code_change/3is one. I fill them with generated code or with Erlang. Both are maintenance costs that I own. - Zig is pre-1.0. The language and standard library still change between releases, and every upgrade of the toolchain is a small porting project, though in practice this cost has been negligible so far.
- The hiring pool is small for any one of the three languages, and smaller for engineers comfortable in all of them. But many startups have found this is as much a help as a hindrance.
- The boundary needs discipline the languages do not enforce. Neither compiler knows about scheduler budgets or resource lifetimes across library versions. My NIF compiler exists because that discipline didn’t survive being left to docs-and-convention.
My conclusion is that the costs are worth paying when the two kinds of work are both large. A system that is mostly coordination should stay on the BEAM and call out to a service for its numeric work. A system that is mostly computation does not need the BEAM.
An operating system for agents is both, and for that case the split described here, with a generated boundary between its halves, is the best arrangement I have found.