Filed from the urus v0.3 endpoint work. start_child spawns and moves on,
so a later sibling can whereis an earlier named child before that child's
actor has run. Notes why blocking spawn is not the fix ('has begun
executing' != 'has bound its name', plus a per-accept round-trip tax and
every spawn becoming a context-switch point), that OTP has the same async
spawn and synchronises one level up in gen_server:start_link, the
readiness-ack shape if scheduled, and the structural workaround urus uses
today (registrar spawns its own consumers).
Root cause behind the "pin the endpoint" gotcha and the trapping-wrapper
pattern: a gen_server had two lifetime authorities — its refs (last one
dropped → inbox closes → exit) and, when supervised, its supervisor. OTP has
one: a process lives until it stops, is shut down, or is killed; a pid is an
address. Root exit now shutting down every forest root removes the reason the
ref-governed idiom existed (a forgotten server no longer hangs the run), so
adopt the one rule:
- The server/machine loop holds one inbox sender for its life; the inbox
never closes. GenServerRef / GenStatemRef are addresses. Explicit close is
`shutdown()`; a forgotten one is swept at root exit.
- `NamedGenServerBuilder::run()` / `gen_statem::run_named(name, m)` run the
loop inline as the current actor: a server is a direct ChildSpec child,
gets the supervisor's shutdown as handle_shutdown / a shutdown row, re-binds
its name on restart, and is addressed by name. The wrapper in
examples/graceful_shutdown.rs is gone.
- gen_statem gains GenStatemName + whereis_machine/send/call/shutdown by name
(parity with gen_server); the macro gets `Sm::new`.
- Root-exit sweep records `Event::RootSweep { target, trapping }` under
smarm-trace ("root_sweep shutdown|stopped"): unsupervised leftovers are
visible rather than silently owned-by-refs.
- Named start() name-clash path stops the spawned actor instead of relying
on ref drop.
Tests: tests/gen_server_lifetime.rs, tests/gen_statem_lifetime.rs,
tests/root_sweep_trace.rs (feature-gated); three existing tests that used
drop-closes-inbox now use shutdown(). Docs/README/ROADMAP/Deep Dive updated.
Mirrors the gen_server surface in gen_statem's event model:
- Cx::trap_exit() (in the initial enter): a shutdown request then arrives
as the Shutdown event, routed by state through `shutdown` rows (default
for a state with no row: stop); linked-peer deaths as `exit <pat>` rows
(default: drop). Non-trapping machines are stopped outright, as before.
- Cx::stop(): normal self-exit after the current event; `stop` tail keyword
is sugar for { cx.stop(); prev }.
- Machine::terminate (optional `terminate { … }` macro block), run from a
Drop guard on every exit path; the guard also drains armed timers.
- Machine::shutdown_ev / exit_ev (defaults None) so hand-written machines
keep compiling; GenStatemRef::shutdown() is graceful and waits.
- Loop selects exits > timers > inbox.
Tests: tests/gen_statem_shutdown.rs.
The RFC 014 root-exit sweep hard-stopped every live slot once nothing was
runnable. That deferral privileged queued work over parked-with-a-pending-
wake work (a sleeper was killed, a queued cast was drained) and any attempt
to widen the notion of pending wake (timers, fd readiness) re-wedges the
run on the periodic-timer daemon the sweep exists to end.
Root exit now means "the program is done": finalize_actor delivers
request_shutdown to every forest root — each live actor whose parent is
the run (ROOT_PID) or is dead — synchronously, before the live-count
decrement. Supervisors cascade per child Shutdown policy; trapping actors
may Continue/drain with working timers and end the run when they stop
themselves; non-trapping actors are stopped outright. No forcing sweep.
Removes root_exited/root_swept, Pop::RootDrain and the idle-verdict
condition; adds tests/root_exit.rs.
Lift OTP's `exit(Pid, shutdown)` + child-spec `shutdown` wholesale.
scheduler / runtime
- `request_shutdown(pid)`: the polite stop. A target trapping exits gets an
`ExitSignal { reason: DownReason::Shutdown }` on its trap inbox and keeps
running; a non-trapping target is stopped as by `request_stop`, which is
now documented as the hard stop (`exit(Pid, kill)`). Dead pid: no-op.
- `RuntimeHandle::request_shutdown` for the off-runtime (signal thread) path;
`from == ROOT_PID` there.
- `DownReason::Shutdown` — appears only in ExitSignal, never in Down (a
complying target exits *normally*).
supervisor
- `ChildSpec::shutdown(Shutdown::{BrutalKill, Timeout(d), Infinity})`,
default Timeout(5s). Every supervisor-initiated stop (ordered shutdown and
OneForAll/RestForOne sibling cycling) is: request_shutdown → await the
child's Signal up to the grace → request_stop → await. Sequential, reverse
start order.
- The supervisor traps exits; a Shutdown ExitSignal runs the ordered
shutdown and `run()` returns normally, so `request_shutdown(root_sup)`
tears a whole tree down top-down with each child's grace period.
- FIX: a hard `request_stop` on a supervisor previously orphaned its
children (the ordered shutdown lived after the loop, and the unwind
skipped it). `Live` (the by_pid map) now carries a drop guard that
fire-and-forget hard-stops live children when unwinding.
gen_server
- `GenServerCtx::trap_exit()` opt-in in `init`; the trap inbox becomes arm 0
of the loop's select. Shutdown ExitSignal → `handle_shutdown() ->
ShutdownAction::{Exit, Continue}` (default Exit: loop breaks, `terminate`
runs on the normal path and may block). Other ExitSignals →
`handle_exit(sig)`.
- `GenServerCtx::stop_handle() -> StopHandle`, `stop()` ends the server
after the current message with a *normal* exit — the missing
`{stop, normal, State}`; `request_stop(self_pid())` was the only self-exit
and it is abnormal (Transient restarts it).
- `GenServerRef::shutdown()` / `gen_server::shutdown(name)` now go through
`request_shutdown`.
Tests: tests/shutdown.rs, tests/supervisor_shutdown.rs,
tests/gen_server_shutdown.rs. Full suite green; fmt + clippy --lib clean.
Wakes issued from a non-scheduler OS thread were silent no-ops. Every
off-runtime wake primitive (unpark, unpark_at, request_stop) reaches the
runtime through the RUNTIME thread-local, which is unset on any foreign
thread — so a send from a plain std::thread enqueued its message but never
woke the parked receiver, and there was no way to drive a stop into a
runtime from an application thread (e.g. an OS-signal handler). The former
strands a parked recv forever; the latter is why a downstream server must
poll a shutdown flag instead of parking on it.
Generalize RFC 018's rule — a producer reaches the runtime through a Weak it
holds — from the IO backend to channel senders and to a new handle:
- A receiver captures a Weak<RuntimeInner> (provably live at that moment)
alongside its (pid, epoch) when it parks. send() and the last-sender drop
wake through scheduler::unpark_at_via, which takes the thread-local path
when on a scheduler thread (preemption-gated, slot-eligible) and the
captured Weak otherwise — the same cross-context wake the IO threads do.
- Runtime::handle() returns a Send + Sync RuntimeHandle carrying that Weak;
RuntimeHandle::request_stop drives a cooperative stop from any thread and
is a no-op once the runtime is dropped.
The in-runtime wake paths (recv/select timers) are unchanged; only the sites
reachable from a foreign thread route through the Weak. RuntimeHandle exposes
request_stop only: send-wake needs no user-facing handle, and off-runtime
unpark is covered because request_stop drives unpark on the upgraded inner.
tests/cross_thread_wake.rs: a foreign-thread send wakes a parked receiver; a
foreign-thread request_stop wakes and stops a parked actor; a RuntimeHandle
held across and beyond run never blocks all-done and degrades to a no-op.
Expands the intro into an overview/limitations structure, documents
preemption, tracing/causal profiling, URUS, and adds a 'Coming up' and
'A note on open source' section.
allocate_slot() panics on a full slab; for a load-shedding caller (an
accept loop spawning one actor per connection) that panic lands in the
spawning actor, which then crash-loops under Restart::Transient into the
still-full slab until its restart budget is spent — and the service stops
accepting entirely. Observed live (urus slowloris scaling, 2026-08-10).
A full slab is a routine overload condition for such callers, not an
invariant violation.
- RuntimeInner::try_allocate_slot() -> Option<u32>: the non-panicking
core; a single pop under the free-list lock, so the claim is atomic
(claim-or-report — no check-then-spawn TOCTOU, no headroom margin).
allocate_slot() is now a thin panicking wrapper over it.
- scheduler::try_spawn / try_spawn_under_with -> Result<JoinHandle,
SpawnError>: parity with spawn/spawn_under_with except a full slab
returns Err(SpawnError::AtCapacity) instead of panicking. Minimal
surface per the agreed strategy; the remaining _with/_addr mirrors are
trivial wrappers if ever needed.
- Slot-first ordering on the try path (reverse of spawn's stack-first):
under overload Err is the hot path, and a rejection costs one mutex
pop — no mmap/pool-pop + init + recycle per shed unit of work. A
drop-guard returns the claimed slot if stack allocation panics in the
claim-to-install window (would otherwise leak and trip run()'s
teardown slot-leak debug_assert).
- SpawnError: non_exhaustive, Display + std::error::Error.
- spawn and every existing call site untouched: the panic remains the
correct loud invariant check at internal/bounded spawn sites.
tests/try_spawn.rs: parity when slots free; exact slab accounting at
capacity (Err, no panic, repeatable); custom-shape try refuses before
stack allocation; self-heal after slots free; plain spawn still panics
(surfaced via JoinError payload); 4-thread race for the last slots
claims exactly the free count; SpawnError impl checks.
Design doc: smarm-suggestion-try-spawn.md. Downstream consumer change
(canned 503 on AtCapacity in urus's accept loop) is urus scope, not
smarm.
(cherry picked from commit 36de4b36aeaa72b2a5f9f3797b9854652656dcf6)
The terminal record existed for watches that raced their target's death, but
e43c673 scoped its stamp to named tenancies — and the pid-identity watch
surface (§4 Slice 3) targets arbitrary actors, including anonymous ones whose
pids cross the boundary in contract replies. The first wild pid-face hit
(width-20 soak, pid_watch_test.exs:47, 1/600 full-suite: a monitor installed
while the child was alive delivered :noproc instead of {:smarm_exit, :panic})
is exactly the residual a0ba9be's commit body deferred.
ever_named becomes `watchable`, with a second set-site: mark_watchable(pid),
which the bridge calls wherever a smarm pid is encoded across the boundary —
BEAM can only watch pids it holds, and can only hold pids that crossed.
Anonymous never-exported churn (holder threads, egress tasks) stays
ineligible, preserving e43c673's LIFO-eviction protection unchanged.
mark_watchable takes the cold lock before the liveness screen: finalize
publishes Done and reads the bit under the same lock, so the mark either
lands before the death stamps or observes the tenancy dead and no-ops —
no lost-stamp window, and marking a corpse cannot invent history (pinned
in the test alongside the mark-while-alive stamp).
Discovered wiring the bridge consult: with an unconditional stamp, the record
for the very death being raced was the shortest-lived data in the runtime.
Every green thread is a slot tenant, the free list is LIFO — so the slot a
named server's death frees is the first one recycled, and the next throwaway
exit (monitor holders, chain-runner work, anything) overwrote the record
before a raced watch could consult it. Deterministic bridge repro: the
corpse resolved fine, terminal_reason read None every time.
register_with now flags the tenancy (ever_named, reset at reclaim) before
the binding lands — set outside the registry lock, so no successfully
registered actor can die unflagged and a failed register's overshoot is
harmless — and finalize stamps only flagged tenancies. Watchable identities
are exactly the named ones (the bridge's pid-identity path deliberately
keeps Erlang's raw :noproc), so nothing consultable is lost.
Contract test updated: the three death modes now self-register; a new
anonymous control pins that unregistered deaths neither stamp nor evict.
A watch installed after its target's death has, until now, only NoProc to
report — but the bridge's proxies install their native watch asynchronously
after acquire returns, so a link established before a crash (from the BEAM's
view) could still lose the panic's translated reason to that blanket NoProc
(width-20 soak signature 4: link_test.exs:26, 1/600 full-suite, 3/2000
link-only, all whereis-miss; deterministic repro in the bridge suite).
Two primitives, no change to monitor()'s own Erlang-faithful stale-pid
semantics — the upgrade is the caller's deliberate act:
- finalize_actor stamps the slot with (generation, DownReason) under the same
cold-lock block that publishes the outcome. The record survives reclaim,
registry pruning, and the next tenant's install; only the slot's next death
overwrites it. terminal_reason(pid) reads it generation-matched.
- resolve_name(name) is whereis with the corpse kept: the dead-holder arm
returns the stored pid it prunes (NameResolution::Corpse) instead of
discarding the only evidence of who died — whereis itself prunes on the way
out, so a whereis-then-lookup consumer would find the evidence already
destroyed. Live/Unbound match whereis's Some/None; the name heals exactly
as before.
Contract pinned in tests/terminal_outcome_after_death.rs: one record per way
of dying (Exit/Panic/Stopped), no record while live, corpse capture + heal on
resolve_name, record independence from registry pruning, survival across slot
re-tenancy, overwrite at the next tenancy's death.
Per-actor stack shapes on every spawn surface (SpawnOpts stack_reserve/
guard_size, Config defaults, pool rule: only default-shaped recycle);
sampled stack high-water + MADV_FREE shrink at actor-park (THRESHOLD
256 KiB, COOLDOWN 64 parks, redzone 1 page); pool-recycle MADV_DONTNEED
above the retained 64 KiB entry end; SIGSEGV overflow diagnostics
(two-tier: in-guard definitive / 1 MiB overshoot 'stepped over', prior
handler chained for foreign faults) with per-scheduler sigaltstack; and
the per-actor introspection surface (ActorInfo.stack: reserve, guard,
sampled depth, parks_since_shrink, shrinks).
Amendments ratified during implementation, for the RFC changelog:
- DEFAULT_STACK_GUARD 64 KiB -> 1 MiB, following the kernel's post-Stack-
Clash stack_guard_gap convention; PROT_NONE width is VA-only and free.
- §7's motivating segfault was a cargo-vendored gz build, not SQLite as
the RFC text says (cc-built C lacks -fstack-clash-protection; distro
libraries have it — the risky class is vendored builds).
- §4 hibernate() deferred to the jar (bolt-on: force-flag on the §3
shrink path, ~10 lines when wanted).
Gates (jobrunner box, 2026-08-08): reclaim gate PASS at c3 and again at
tip (3.0 MiB LazyFree -> kernel reclaim -> Rss to one live page ->
re-spike bit-identical, live data intact; MADV_PAGEOUT stands in for
memcg — cgroup2 is RO in the job container — driving the same reclaim
path). E1 interleaved A/B vs v0.5.0: every ka cell (the E1 subject)
within +0.3..+2.9% at tip; close-mode control cells within noise except
t8-c4 close, which is bistable (~40-44k vs ~46-49k modes for BOTH
variants, base self-disagrees by 11% across rounds); 6 rounds across two
runs are inconclusive there and a 10-round focused run is noted in the
handoff as deferred follow-up, accepted for this release.
No breaking API changes since v0.5.0: SpawnOpts fields and ActorInfo
gained members (exhaustive-construction downstream will need the new
ActorInfo.stack field; urus does not construct it).
- introspect::StackInfo { reserve, guard, depth_high_water,
parks_since_shrink, shrinks } as ActorInfo.stack; re-exported at crate
root beside ActorInfo.
- All reads lock-free: geometry from the c6 diag slot atomics, depth =
top - hwm (the §2 sampled high-water; doc spells out sampled-not-exact
and that 0 means never-descheduled-at-depth), counters straight off the
§3 atomics. Coherence for the incarnation rides read_slot's existing
generation check, same as overruns/messages_received.
- Slot::stack_introspect(): one pub(crate) tuple accessor beside the other
counter accessors.
- Exact RSS deliberately absent per RFC (mincore = debug tooling only,
never a runtime path); stack_shape(pid) untouched (cold-lock exact
variant from c2).
- tests/introspect.rs: defaults surface (64 KiB reserve / 1 MiB guard /
sampled ~32 KiB depth / gate park counted / zero shrinks) + live shrink
counters (spike visible pre-shrink; shrinks>=1, cooldown counter reset,
hwm reset after crossing COOLDOWN) read mid-run -- post-join the slot
reclaim correctly hides the incarnation, which the first draft of the
test learned the hard way.
FLAGGED (Claude-solo calls):
- Nested StackInfo struct over five flat ActorInfo fields (grain break;
the five fields are one concern and ActorInfo is already 12 fields).
- Field names reserve/guard/shrinks (RFC says stack_reserve/stack_guard/
shrink count; the stack_ prefix is redundant inside StackInfo).
- src/signal.rs: process-global SA_SIGINFO|SA_ONSTACK handler installed once
at runtime::init (before any scheduler thread -> unracing PRIOR save);
per-scheduler-thread 64 KiB sigaltstack registered at schedule_loop entry
(a guard hit leaves no stack to handle on). Async-signal-safe throughout:
classification is plain loads (const-init TLS Cell + slot atomics), print
is fixed-buffer itoa + one write(2), death is SIG_DFL + refault at the
same instruction (core-dumpable, correct wait status).
- Two-tier classification (agreed): in-guard = definitive; OVERSHOOT window
below the guard = 'unprobed (FFI?) frame stepped over it' probable
attribution -- the RFC's motivating incident (cargo-vendored gz, not
SQLite as the RFC text says) faults there under a small guard. Pure
classify() fn, 5 adversarial units incl. saturation at low addresses.
- DEFAULT_STACK_GUARD 64 KiB -> 1 MiB (agreed): kernel stack_guard_gap
anchor post-Stack-Clash; PROT_NONE is VA-only (no RSS, no page tables,
no overcommit charge) so width is free at any actor count.
- Unclassified faults reinstate the PRIOR sigaction and refault (agreed):
std's own OS-thread overflow diagnostics survive our presence.
- Slot: diag_{stack_top,stack_reserve,stack_guard,pid} atomics written in
install_actor pre-publish; readable without the cold lock (Stack lives
under it); only consulted while CURRENT_SLOT points at the slot, so
never stale where read. preempt::current_slot_ptr ungated from
smarm-causal (now also the classifier's anchor).
- build.rs + cc (agreed Q3): canary/canary.c, 96 KiB local touched low-end
first, -fno-stack-clash-protection pinned so hardened toolchains don't
probe the canary into uselessness.
- tests/stack_diag.rs: subprocess x4 -- Rust recursion tier-1; FFI canary
tier-1 at defaults (1 MiB guard catches the jump); tier-2 at guard=4 KiB
('stepped over', reproduces the incident); clean at reserve=256 KiB
(the §1 knob is the fix, same frame).
FLAGGED (Claude-solo calls):
- OVERSHOOT_SLOP = 1 MiB (matches guard default/kernel gap; beyond it
attribution would be dishonest).
- Altstack 64 KiB, mmap'd once per OS thread, never freed (bounded by
thread count; reused across run()s via TLS flag).
- Foreign-fault reinstate permanently deregisters our handler; accepted --
the process is dying either way.
- Diag geometry as 4 slot atomics (install-time cost only) over a per-switch
TLS snapshot (hot-path stores).
- stack::retain_range: pure checked span fn (retain page-up = zap less;
None when retain covers the reserve, so the 64 KiB default config never
pays a syscall) + 6 adversarial units mirroring shrink_range's.
- Stack::recycle_zap: advisory MADV_DONTNEED of [usable_base, top-RETAIN);
stack is unowned at the call site, synchronous eager zap races nothing.
- recycle_stack: zap OFF-LOCK before pool admission (acquire_stack's
no-syscall-under-the-pool-lock invariant); rare cap-overflow pays a
wasted zap ahead of munmap, accepted over a second lock round-trip.
- pub const RECYCLE_RETAIN = 64 KiB beside the shrink knobs, ratified-as-
constant rationale in doc.
- tests/stack_recycle.rs: mincore-based exact-zero-resident assert over
the zap span. smaps was tried first and over-counts: a neighboring rw
anon VMA can merge flush against the stack top (observed once under the
full-suite run); the PROT_NONE guard pins the usable base exactly.
FLAGGED (Claude-solo calls):
- RFC §6 'above the bottom RETAIN' is direction-ambiguous in address
terms; implemented as retain the ENTRY end (highest addresses, the
pages the next actor faults first), zap the cold deep span below.
- Const named RECYCLE_RETAIN (RFC says RETAIN) to sit beside SHRINK_*.
hwm: AtomicUsize lands beside sp on the slot: the single context-save
site min-updates it (one branch + at most one Relaxed store into the
line the sp store just dirtied), install resets it to the fresh top.
Advisory by construction — correctness never depends on it. The mod-doc
ordering chain gains a line: hwm piggybacks the existing
Relaxed-store-before-Release pattern and adds no edges.
Shrink hook in the YieldIntent::Park arm only, before the park_return
Release transition — the owned window (obligation 1's assert-comment at
the site): after the sp store, before Parked is published, scheduler on
its own stack, actor saved and unstealable. It runs on both arms of the
park_return race (a consumed unpark flag means one wasted-but-harmless
madvise). The preempt/yield path deliberately never checks: §4's
bounded, self-healing leak under saturation, when syscalls are least
affordable.
SHRINK_THRESHOLD = 256 KiB and SHRINK_COOLDOWN = 64 parks are pub
constants with the ratified doc rationale, not Config fields. The freed
span is shrink_range(hwm, sp, page): whole pages of [hwm, sp − 1-page
redzone), rounded inward, checked arithmetic — adversarial inputs
collapse to None (obligation 2). MADV_FREE marks lazily; the kernel's
reclaim-under-pressure IS the hysteresis, cancel-on-write is the safety
net. parks_since_shrink + shrink_count ride the slot for the cooldown
and the future introspect surface.
Tests: 7 adversarial shrink_range units (inverted/empty spans, redzone
underflow, unaligned ends, sp-crossing sweep); integration — 8 MiB
reserve, ~3 MiB spike sampled via yield-at-depth, parks gated on
introspected Parked state past the cooldown, then ≥ 2 MiB LazyFree
asserted inside the stack's smaps range with live data intact; and the
inverse guard — a shallow never-spiking actor ends at exactly 0
LazyFree (also proves the parser isn't vacuously zero via the first
test).
SpawnOpts { stack_reserve, guard_size } with Option<usize> fields, None
resolving to the Config defaults at spawn time — a deliberate deviation
from the RFC's plain-usize struct so struct-update syntax works without
a runtime handle in scope. Threaded across the five surfaces:
spawn_with, spawn_under_with, spawn_addr_with,
GenServerBuilder::stack_opts (mirrored on NamedGenServerBuilder), and
gen_statem::spawn_with (gen_statem has no builder, so the opts ride a
_with variant — Claude-solo surface call, flagged for review). Existing
spawns forward defaults; no call-site churn.
introspect::stack_shape(pid) pulled forward (agreed) as the first slice
of the RFC 019 introspection surface, giving tests an observable.
Tests (tests/spawn_opts.rs): override/partial-override/rounding on each
surface; obligation 4 from the outside — a dead custom stack is never
handed to the next default spawn (LIFO pool would expose it), and the
reverse (default stacks ARE recycled); 8 MiB reserve behaviorally
permits ~1 MiB recursion. Also: silence unused-Result in the c1
runtime test (join now unwrapped).
Stack takes an explicit (reserve, guard) shape, both page-rounded and
stored; usable_base derives from the stored guard. Guard default raised
4 KiB -> 64 KiB (DEFAULT_STACK_GUARD): probestack makes one page enough
for Rust frames, but an unprobed C frame can leap a page in one sub rsp
— the motivating SQLite segfault. Reserve default stays 64 KiB
(DEFAULT_STACK_RESERVE); ACTOR_STACK_SIZE retired.
Config::{stack_reserve, stack_guard} thread the runtime defaults into
RuntimeInner pre-rounded. All acquisition/recycling now goes through
acquire_stack/recycle_stack carrying the pool rule: only default-shaped
stacks are pooled (pooled ⇒ default-shaped by induction); custom shapes
mmap fresh and munmap at death. Pool lock still dropped before any mmap.
No public spawn API change (SpawnOpts is the next commit).
Tests: shape rounding + accessors, wide-guard faults at both ends
(subprocess), Config::stack_reserve permits >64 KiB recursion that
previously could only segfault.
Breaking API rename since v0.4.0: gen_server's ServerRef/ServerBuilder/
ServerCtx -> GenServerRef/GenServerBuilder/GenServerCtx, Watcher<G> is
now generic over its GenServer, and GenServer gained a required
associated Timer type for timer-fire payloads (arm_after/handle_timer).
Downstream consumers (urus) have been ported.
The swap (RFC 018). Schedulers no longer sleep on a shared level-triggered
wake pipe — the herd source that made the default 8-thread config 7x
slower than 2 threads (E1). They park on per-thread futex parkers via the
coordination layer; IO backends become producers behind a two-call
contract (make runnable, then the enqueue tail wakes exactly one parked
scheduler).
Deleted: the drain lock and the one-winner phase-1 drain; the shared
completions VecDeque; the wake pipe fds, poll_wake, drain_wake_pipe,
wake_scheduler, the FdReady/Blocking Completion enum; the 100us idle nap;
the per-pop io.lock liveness read; io.rs's as_millis timeout truncation.
Added:
- enqueue wake tail (fixes the silent enqueue): wake_one_if_idle, a fence
+ one Relaxed mask load when everyone is busy — the pure-compute hot
path pays almost nothing.
- driver-enqueues: the pool thread stashes its result in the slot,
decrements io_outstanding, unparks; the epoll thread removes+DELs the
waiter under the waiters lock and unparks. Both reach the runtime via a
Weak (no Arc cycle). The waiters map moves behind its own Arc<Mutex> so
the epoll thread never takes the runtime io lock (teardown holds it
while joining that thread).
- io_outstanding / io_fd_waiters atomics: the termination verdict reads
two atomics instead of taking io.lock on every pop.
- timekeeper idle path: at most one parked scheduler holds the timer
deadline (an expiry wakes one, not a herd); everyone else parks
indefinitely and is woken by the enqueue tail.
- busy-path timer due-check (ratified design point (a)): under saturation
nobody parks and no timekeeper exists, yet due timers must still fire —
one Relaxed load of the earliest-deadline snapshot per loop, clock read
only when a timer is armed. Maintained under the timers mutex.
- chain rule: a scheduler that pops with more work queued and a sibling
parked wakes one, so surplus runs in parallel rather than behind it.
tests/park_wake.rs pins the two new observable properties: timers fire
under full scheduler saturation, and sub-ms sleeps are prompt (the
as_millis truncation regression). Full suite + all loom models green;
clippy --lib clean.
Two integration-driven amendments ahead of the runtime swap:
wake_one_if_idle() realizes RFC 018's "empty-mask fast path is one
relaxed load" soundly: a bare relaxed load is a lost-wake in the Dekker
shape for the lock-free ring queues, so the producer publishes work,
fences (SeqCst), then reads the mask Relaxed — paired with a matching
fence between the consumer's bit-publish and its re-check in park().
The pure-compute hot path (mask 0) never takes the shared mask line
exclusive; the RMW read stays on the rare chain-rule path only. Loom
models 1/2 now drive the fenced pattern end to end.
next_deadline is the earliest KNOWN timer deadline, independent of
whether anyone is parked — which tk_armed cannot give: under saturation
nobody parks, nobody arms, yet due timers must still fire (ratified
design point (a): the busy-path due-check). Maintained under the timers
mutex (note_deadline on insert — which also carries the timekeeper
re-arm wake — refresh_deadline after pop/clear); read lock-free.
deadline_due() costs one Relaxed load and a branch when no timer exists;
the clock is read only when one does.
Schedulers get an IO-agnostic sleep/wake primitive of their own: one
futex Parker per scheduler thread (permit semantics, std::thread::park
shaped — closes the check-then-park race), an AtomicU64 idle mask with a
set-bit → re-check → wait park protocol, wake_one (highest-bit LIFO,
CAS-clear before unpark: exactly one wakeup per call by construction),
wake_all for the terminal path, and the timekeeper role — at most one
parked scheduler holds the timer deadline, with an atomic armed-deadline
snapshot for the busy-path due-check and an insert-side re-arm wake.
Deadlines travel as nanosecond timespecs end to end; the wake pipe's
as_millis truncation is unrepresentable here. The Dekker publish/re-check
shape is resolved by the same-location-RMW handshake (AcqRel), not SeqCst
loads; loom verifies exactly this in four models (no-lost-wake, chain
propagation, timekeeper handoff, termination), run with
LOOM_MAX_PREEMPTIONS=3 — unbounded exploration is impractical for the
looped models. Loom/non-Linux builds park on a Mutex+Condvar via
sync_shim.
Standalone until the runtime swap (next commit): nothing outside tests
constructs a Coordinator yet, hence the temporary dead_code allow in
lib.rs.
Desktop migration: the home-manager rust here ships without the clippy
component. Prefer an installed cargo-clippy; otherwise run clippy from an
ephemeral nix-shell with a separate target dir (mixed-compiler artifacts
are an E0514 hard error). MSRV keeps the shell's older toolchain a
legitimate gate.
Lead with the user's problem (learn when another actor dies without
it knowing you're watching), explain one-directional/one-shot
semantics and contrast briefly with link without assuming link.rs has
been read. Add a compiling doctest. Drop em-dashes. Correctness facts
about registration/death races and demonitor-after-fire safety kept,
reworded in plain terms and separated from the public item docs.
Was the worst offender for external-context references (RFC 016
Chunk 1/4, DECISION D1/D2, RFC 003/011), all removed. Lead with the
practical use cases (debugging, health checks, test assertions,
dashboards) for snapshot()/actor_info()/tree(), and explain the
per-actor-reads-not-a-world-freeze consistency model in plain terms
instead of citing a decision log. Add a compiling doctest.
Every public item was previously undocumented. Lead with why Mutex<T>
exists (a channel/gen_server is overkill for plain shared state) and
how it differs from std::sync::Mutex (parks the actor not the OS
thread, every lock is timeout-bounded by default). Add a compiling
doctest. Document new/lock/lock_timeout/try_lock/set_default_timeout/
MutexGuard/LockTimeout/DEFAULT_TIMEOUT. Drop em-dashes; keep wake
protocol mechanics as contributor-facing comments on private internals.
Replace the 'what changed' diff-against-a-prior-design framing with a
plain explanation of what the registry is for (naming an actor so
others can find and message it by name) and a compiling doctest
(register/whereis/send/unregister). Cut all RFC/decision-number/bug-id
references and em-dashes; move type-erasure and locking-discipline
detail into an Implementation notes section for contributors.
Lead with what a channel is and how to use it (compiling doctest for
channel()/send/recv/close), before any internal rationale. Document
every previously-undocumented public item (channel(), Sender, Receiver,
SendError, RecvError). Move the RawMutex-vs-std::sync::Mutex rationale
and lock-class discipline into an Implementation notes section. Drop
em-dashes throughout.
Lead with what an actor is and how to start one with run()/spawn(),
following gen_server.rs's example-first style. Add a compiling module
doctest. Drop RFC references and em-dashes; keep internal mechanics
(preemption gating, thread-local borrow rules) as plain contributor
comments rather than public-facing doc prose.
run_experiments snapshots the site registry once at entry. stage-x is
registered lazily (first causal_site! execution in the worker), so on a
1-core box the snapshot deterministically wins whenever a sibling test
has already paid the tsc_hz calibration — stage-x missed the sweep and
the summary assert tripped (silently, pre-propagation; the earlier
cascade attribution was incomplete). The worker now signals after its
first site entry and the test waits on it before starting the sweep.
Jarred separately: the lazy-registration trap is a lib UX hazard worth a
doc note or warm-up guidance.
A fresh/reused slot started with causal_delay=0 and causal_parked=true:
the first resume then 'forgave' the entire monotone global backlog, once
per spawn, booked as park forgiveness. Under close-mode conn churn (~95k
spawns/s) that is millions of phantom forgiven ms per 700ms window, even
in 0% cells — the books could not balance under actor churn while
ka-mode stayed plausible (few spawns). Impacts were unaffected: the
first resume always precedes the first check, so nobody ever spun.
reset_counters now installs Coz's new-thread rule: causal_delay starts
at the current global ledger and causal_parked starts false — a newborn
neither owes nor is forgiven the process's history, and delay injected
while it sits spawn-queued (runnable, not blocked) is owed and paid at
its first check, the semantics audit_zero_pct_window_absorbs_leftover_
debt pins. That test (born failing, unmasked by the run() propagation
fix) and the new actor_churn_between_experiments_forgives_nothing
regression test both go green.
The trampoline caught the root's panic, recorded it as Outcome::Panic on
the slot, and run() dropped the initial handle without reading it — every
assert inside run(), the standard test-suite pattern, was silently
vacuous (found live: a failing-first test passed). run() now reads the
root outcome before the handle drop and resume_unwinds the payload after
full teardown, so a caller's catch_unwind leaves the Runtime reusable;
Exit and Stopped return normally. The payload message is printed before
re-raising (the throw-site hook output was suppressed in-actor).
Correct the two tests this unmasked, both born failing and never run:
the select loser-arm test kept a closed arm in the set (the documented
closed-arm rule: a closed arm reports ready forever — observe the
disconnect and drop it); the send_after-to-dead test expected Ok(None)
from a closed+empty channel (documented: Err(RecvError), which proves
nothing-delivered even more strongly).
The probe now prints eff+offcpu next to eff — attributed plus the
counted runnable off-CPU gaps over ground-truth in-site time — the
per-window check that the located mechanism accounts for the whole
residual (~1.00 = books closed, no remaining silent loss). Audit line
gains the offcpu column, same delta-terms convention as the lib
renderer. Header doc rewritten from hypothesis to resolution.
GPU sweep decomposed the ~23ms/700ms @50 injection deficit: every
ledger bucket is ~zero (drop park 0, discards 0, drop yield ~0.3ms),
books balance at absorbed+forgiven = 4x injected in all 48 cells, and
the 0%-cell contamination signature is absent. The probe pins the
residual: eff 0.933-0.943, a constant 22-27µs missing per site entry
= ~4.9 slice-expiry yields/entry x ~5.6µs runqueue wait. The
"deficit" is runnable off-CPU time inside the site — wall time the
probe's ground truth counts but on-CPU attribution correctly skips
(Coz model: speeding the site's code does not shrink queue-wait).
Measure-only reclassification, no behaviour change: a yield in the
target site stashes (tsc, experiment epoch) on the slot; the next
on_resume counts the gap into OFFCPU_IN_SITE_{CYCLES,N} (would-be
delta terms, MAX_SAMPLE_CYCLES-capped) iff the epoch still matches
and the word is live — a gap straddling end()/a same-word begin()
(live in the probe's 50,50 schedule) is dropped, never a leaked
cooldown. Parks excluded: blocked time is forgiveness territory.
New offcpu column in render_ledger_audit; LedgerCounters and
ExperimentResult grow the two fields. Fidelity footer now states
the on-CPU basis (deliberate wording change to the pinned summary;
the substring pin test still holds). +2 tests (counted gap; epoch
straddle) + render assert.
- causal_pipeline: SMARM_CAUSAL_AUDIT=1 appends render_ledger_audit()
after the summary; pinned summary format untouched.
- causal_attrib_probe: per-pct audit line (absorbed/forgiven/drops/
discards) under the existing eff line, so in-site-vs-attributed and
the loss buckets land in one place for the sweep.
1-core smoke (work + wide): books balance — absorbed = 2x injected and
forgiven = 2x injected, i.e. owed = injected x (N-1) with N=5 actors,
zero outstanding. drop park = 0 even in wide mode (the bottleneck's
queue is never empty, so its in-guard recv never parks); drop yield is
~1500 events but ~0.1ms per window, confirming slice-expiry yields
sample at their own checkpoint. Deficit decomposition needs the
parallel box.
Measure-only counters for the deficit hunt (~23ms short per 700ms window
at 50% on the bottleneck site; superlinear vs 25%). Nothing here changes
injection or absorption; the sweep decides the fix.
Buckets, windowed per cell into new ExperimentResult fields (audit
snapshot taken at end() — spin/attribution freeze there, forgiveness
does not):
- spin_absorbed / park_forgiven: where owed delay actually went. Spin
during a 0% cell is the baseline-contamination signature — leftover
debt from a prior window being paid in a later one (checks gate on
the experiment word, so cooldowns pay nothing and debt carries over).
- drop_park / drop_yield (+counts): the deschedule path flushes no
sample tail — an in-target-site park or yield silently loses
[last sample -> now]; on_resume re-arms before the actor runs again.
New on_deschedule hook in all three intent arms (real park; explicit/
slice-expiry yield; requeued park counts as yield — it never blocked).
Slice-expiry yields sample at the descheduling checkpoint, so a fat
yield bucket points at explicit yield_now or requeued parks.
- discard_overmax (+count, in would-be delta terms so columns compare
against injected_cycles) / discard_unarmed: the attribute() clamps,
previously silent.
LedgerCounters + ledger_counters() expose cumulative totals (tests,
run-level prints); render_ledger_audit() is the per-cell companion to
render_summary, which stays byte-identical (pinned). ExperimentResult
now derives Default so literals survive future audit-field growth.
Tests: +7 (spin counted, forgiveness counted, in-site park drop, in-site
yield drop, overmax discard, 0%-window leftover absorption — synthesized
deterministically via inject_delay_cycles_for_test with no experiment
active — and the audit render). 22/22 causal.
One unconditional footer line whenever there are results:
"note: impacts are lower bounds — undershoot grows with speedup pct;
rankings unaffected" — surfacing the RFC 007 Validation fidelity statement
where users actually look, instead of only in the RFC. Wording pinned by
the summary test.
send_after_wall / send_after_named_wall (+ Timers::insert_send_wall) arm a
message-delivery timer that opts out of the RFC 007 virtual-time shift and
fires at its raw deadline regardless of injected delay — the Send-reason
sibling of sleep_wall, closing the jar item whose substrate efbc254 landed.
For deadlines that reflect the outside world (protocol timeouts, wall-clock
schedules) rather than workload pacing. cancel_timer is anchor-agnostic and
unchanged; without the feature the API exists and is identical to send_after.
The gen_server timer layer (send_after_to, RFC 015 §5) deliberately stays
virtual-only — an opt-out there means new options on the gen_server/statem
timeout API, out of scope for now.
Tests: wall send fires at raw deadline while a virtual sibling shifts;
cancel on a wall send with debt outstanding; featureless delivery/cancel
smokes through the public named API. 15/15 causal, 34/34 binaries both
feature configs, lib clippy clean both.
Occupancy probe on 24 cores: δ = 0.3µs/item (0.1% of the serialized
path); wide guard leaves the @50% cell unchanged. The +84-vs-+100
shortfall is controller-side (injected 327ms of the ideal 350ms over
the 700ms window, plus ~3% real-rate dip during experiments), not
unguarded stage time. urus's ~70µs/request remainder remains the
guard-placement case; the occupancy probe discriminates the two.
SMARM_CAUSAL_MODE selects the reserve stage's guard placement:
work (default, unchanged) | wide (guard over recv+work+send, the whole
serialized per-item path) | occupancy (no experiments; per-segment
timing of reserve's loop at baseline, reporting the unguarded
remainder δ and the impact ceiling it implies).
Discriminates the two candidate explanations for the demo's +84-vs-+100
@50% shortfall: physical recv/send time outside the guard (occupancy
sees δ≈30µs, wide recovers ~2x) vs. injection-side credit loss
(occupancy sees δ≈0, wide caps at ~+84 too — guard cadence identical).
New timer anchor: insert_sleep_wall / scheduler::sleep_wall (exported) opt a
Sleep entry out of the RFC 007 virtual-time shift, so it fires at its raw
deadline regardless of injected delay. Featureless config is unchanged (the
API exists but is identical to sleep).
The causal controller's window/cooldown sleeps and the tsc_hz calibration
sleep use it on the actor path (the OS-thread path was already wall). This
fixes the controller's own sleeps dilating under its own injection —
experiment windows stretched ~2x at 50% speedup (337ms -> 646ms injected/
window). Deltas were rate-normalized so results were unbiased; this fixes
sweep cost, not bias. The general wall-anchored-timer-semantics jar item
(user-facing opt-out) remains open; this lands the substrate.
Test: wall_timer_ignores_injected_delay — wall entry fires at raw deadline
while a virtual sibling in the same heap shifts. 13/13 causal, 34/34
binaries both feature configs.
Injected delays dilate virtual time for the workload, but timer deadlines
stayed wall-anchored: a sleep or receive-timeout fired early in virtual
terms, so timeout/retry behaviour sped up relative to the dilated world
(v1 known gap #1).
Every heap Entry now carries delay_stamp — the global delay ledger at
(re-)queue time, cfg-gated on smarm-causal. pop_due converts any debt
accrued since the stamp to wall time (tsc_hz) and shifts the effective
deadline; a not-yet-due entry is re-queued at the shifted deadline with a
fresh stamp, so it keeps chasing delay injected while it waits. seq is
preserved across re-queues, keeping send_after cancellation identity
intact (cancelled entries are discarded before any shift). Zero debt is
byte-identical to the old path; peek_deadline may under-report, costing
one spurious scheduler wake per injected chunk (documented).
This also makes the park-gated resume credit *correct* rather than
forgiving for sleepers: a sleeping actor now physically pays its debt by
sleeping longer, so the on_resume fast-forward reflects real payment
(sleeping_actor_pays_injected_delay pins this end-to-end through the
runtime).
New test hooks: inject_delay_cycles_for_test (deterministic ledger
driver, eagerly TSC-calibrating so conversion never stalls a scheduler
loop) and cycles_to_duration. The ledger is process-global, so the
delta-sensitive causal tests now serialize on a shared test mutex — they
were racy under the parallel test harness before this, in principle.
Samples were taken only when maybe_preempt's cold block happened to fire
in-site, so the interval between the last check and SiteGuard drop was
discarded on every site entry. Measured live on the 24-core box:
22-29us lost per entry, a constant attribution efficiency of ~0.93-0.94,
which under-reported every impact (+83.5% where theory says +100%; the
observed shortfall fits 1/(1-pct*eff)-1 at both 25% and 50%).
SiteGuard enter/drop now call site_transition(): leaving the target site
flushes the pending interval into the ledger (sample-only, never spins,
so safe under no-preempt regions); entering the target site re-arms the
sample clock so pre-site time is never attributed (the symmetric
over-attribution). Winner attribution is factored into attribute(),
shared by the cold check and the flush, with the same interval clamps.
Adds examples/causal_attrib_probe.rs (ground-truth in-site time vs
ledger attribution, the probe that confirmed the leak) and the
site_boundaries_flush_tail regression test (a site entry that never
hits a cold check must still be attributed). Also gates causal_probe
on smarm-causal in Cargo.toml - it never was, so featureless builds
of the examples were broken.
causal_site! scoped site guards per actor slot, progress! throughput
points, and a Coz-style virtual-speedup engine hooked into
maybe_preempt's amortized cold block: target-site samples grow a global
delay ledger; bystanders spin-absorb their debt at the next causal
check, with timeslice extension so injected delay is not charged
against the slice.
Resume credit (Coz's blocked-thread rule) is gated on a causal_parked
slot bit set only by a real park: crediting on every resume made any
yield-cadence actor delay-immune and every experiment inert (found
live on a 24-core run — dead-flat deltas across all sites).
Report normalization uses a measured TSC frequency (~50ms calibration
on first use) instead of the crate-wide 3 GHz assumption, which
uniformly inflated impact numbers on a 3.7 GHz box. impact_pct() is
the machine-readable form of the summary for programmatic checks.
examples/causal_pipeline.rs burns fixed *work* (calibrated LCG loop),
not fixed wall time — a timed busy-wait absorbs injected delay into
its own budget and reads as a no-op. Self-checking: exits nonzero if
causal separation fails; skips the verdict below 4 cores. Validated
on a 24-core box: reserve (true bottleneck) +29.3%@25/+83.5%@50;
serialize and background-compaction ~0%.
Known v1 gaps (jar): timer-heap deadlines unshifted, no-check!/no-alloc
actors undelayable, multi-scheduler coherence best-effort Relaxed,
off-CPU blame punted, Instant::now() uncorrected.
Zero-cost with the feature off; clippy -D warnings clean both ways;
full suite green with and without smarm-causal.
Root cause of soak20 signature 2 (refcount_test.exs 'watcher crash',
110235x fast {:error, :server_down} probes over the full await window):
by_name mapped name -> slot *index*, so a name whose holder died (no stop
path unregisters; prune is lazy) and whose slot was then re-tenanted read
as live-held: register failed NameTaken{holder: <unrelated tenant>} (which
the bridge macro's generated start() swallows -> start_server/1 reports :ok
for a server that never came up), while name resolution reached the
tenant's mailbox, missed on the message TypeId and failed fast WITHOUT
pruning — the wedge self-sustained for the tenant's lifetime. Name-
addressed send additionally judged liveness on the slot's *current*
mailbox pid, so a same-typed tenant would have received the message
(misdelivery) and a differently typed one a misleading NoChannel.
Fix: by_name: HashMap<&'static str, Pid> — every reader judges the
*stored* holder with the generation-checked live(), so a recycled slot's
tenant no longer impersonates a dead holder, and every touch (register /
whereis / resolve / send) prunes and heals a stale name. prune(index)
becomes prune_holder(pid): names bound to the holder go; the mailbox goes
only while still the holder's own (a tenant's replacement mailbox is left
untouched). Introspection matches names to mailboxes by full pid, so a
stale name never annotates a slot's new tenant.
Deterministic regression test added first and shown to fail pre-fix
(tests/stale_name_slot_reuse.rs: tiny slab forces re-tenanting; old slot
(1,0) died, tenant (1,1) took the index; register -> NameTaken pre-fix).
Post-fix it asserts the healed contract: whereis -> None (pruned), call ->
ServerDown, re-register -> Ok. Suite 33 ok-binaries, clippy gate clean.
In the wild the window opened at every splice_test teardown:
Splice.terminate -> exit_server('subtree') left the name bound; width 20
raised the re-tenant probability. Downstream (smarm_beam): install_child's
unregister-before-register workaround becomes dead code (removed there);
the #[smarm_server] macro's swallowed register error becomes truthful
idempotency (a NameTaken now really is a live holder).
Extra scheduler threads (slots 1..N-1) are now spawned via thread::Builder
with the name smarm-sched-{slot}, so they are identifiable in
/proc/<pid>/task/*/comm, stack dumps and debuggers. Thread 0 keeps its
caller-given name (an embedder names the thread that calls run — smarm_beam
names it smarm-runtime). A refused spawn still panics, matching the previous
thread::spawn semantics.
Motivation: the §17 scheduler-width knob in smarm_beam asserts the *live*
width by counting these named threads, and a 20-scheduler soak needs the
threads tellable apart in wedge captures.
Lost-wakeup: schedule_loop's phase-1 drain uses drain_lock.try_lock(), and
try_lock losers skip the completion drain entirely. Both schedulers park on
one shared wake pipe and, until now, drained ALL its bytes right after their
idle poll_wake returned — outside the drain lock. A loser could therefore
eat the byte announcing a completion the winner had not seen (the winner was
already past drain_completions when the epoll thread pushed it), and both
threads would park with the completion stranded. Because the bridge eventfd
is registered EPOLLONESHOT, the kernel had already disarmed it at
epoll_wait, so no later write could re-fire it: the runtime slept until an
unrelated timer deadline forced another phase-1 pass.
Fix: drain_wake_pipe() moves inside the drain guard, immediately before
drain_completions(); the two post-poll drains in the Pop::Idle arms are
removed. Producers push their completion before writing the byte, so a byte
consumed under the guard always has its completion visible to the drain that
follows. An unconsumed byte keeps the (level-triggered) idle poll returning
instantly, so a try_lock loser spins briefly until the winner releases —
it can no longer sleep through stranded work.
Found via smarm_beam's ingress-cap drain barrier flaking under CPU load
(5/25 loaded suite runs wedged; mid-wedge stacks showed both schedulers in
poll_wake with an FdReady stranded and the eventfd disarmed). Post-fix:
60/60 loaded runs green, tight 5.8-6.8s timing band, no stall tail.
Root-cause notes: smarm_beam outputs/flake-rootcause-egress-overload.md.
A queued Envelope::Call was stranded until the last Sender dropped, so a
caller parked in gen_server::call was never released with ServerDown when a
*named* server was request_stop'd — the registry's inbox Sender clone (lazy
prune) kept the channel Arc, and the queued reply_tx, alive indefinitely.
Receiver::Drop now drains the queue (items dropped after releasing the lock,
since a reply_tx drop reaches a different channel's lock + the scheduler),
restoring the documented ServerDown guarantee on every teardown path.
Adds tests/stop_with_queued_call.rs: deterministic pure-smarm reproducer.