Commit Graph
193 Commits
Author SHA1 Message Date
claude-asm-audit bc0a5e8656 scheduler: spawn_monitor / spawn_monitor_with — monitor registered on the child's slot before publish, so the Down always carries the real reason (spawn-then-monitor could race to NoProc); tests; channel test uses it 2026-08-21 12:23:46 +00:00
claude-asm-audit ea2b222cff runtime: wake_slot default ON (RFC 005 accepted, finding 18); supervisor stops survivors sequentially so reverse-order teardown is guaranteed; deflake gen_server/channel tests; baseline.json regenerated slot-on; ROADMAP notes 2026-08-21 12:23:33 +00:00
claude-asm-audit 7001f04b65 diag(runtime): wake-path counters in SchedulerStats + RuntimeStats::wake_diag; RQDIAG/DIAG lines in rq_runtime and general (target 5, finding 17) 2026-08-21 12:22:46 +00:00
claude-asm-audit 77938cd31d bench: split ping_pong_oneshot — spawn_pair_control + ping_pong_steady; refresh baseline (job e982b5b5)
general.rs sections 5/6: spawn_pair_control is ping_pong_oneshot with the
messages removed (2 spawn + 2 join per round); ping_pong_steady is one
persistent pair × 10k roundtrips over unbounded MPSC (smarm::channel vs
tokio::sync::mpsc::unbounded_channel). Box (5900X, 3729fff, rq-mpmc):
oneshot 804/control 633 → 79% spawn+join; steady 142 ns vs tokio 128
(0.90×) at 1T, 1.8 µs/roundtrip at 20T. history.md finding 16.

baseline.json regenerated from the same run (20T labels, 14 benches).
2026-08-21 12:22:46 +00:00
claude-asm-audit 5ddd122711 perf(preempt): rdtsc unserialised by default; causal attribution opts into lfence
reset_timeslice paid an lfence pipeline drain on every resume via the
shared rdtsc() helper. Nothing in preempt.rs needs it: the timeslice
arm/expiry compare against a ~1e5-cycle slice, and an early stamp only
makes the slice look more used. The consumer that does need it — causal
site attribution, where a speculative early read misattributes a site's
tail — now calls rdtsc_serialising() explicitly (cold_check sample,
SiteGuard enter/exit).

Measured vs 31dc26a (baseline commit), 24-core box, rq-mpmc default:
- switch_cost, taskset -c 2, 9 interleaved old/new pairs: mean_cyc
  149 -> 117 per roundtrip (-32, -21%); mean_ns 43.2 -> 34.9.
- sweep.py regress + run, 20T, two runs agree:
  yield_in_hot_loop 1T 40171 -> 30277/30497 us (-24%)
  yield_many        1T 12513 ->  9965/10013 us (-20%)
  ping_pong_oneshot unchanged (RFC 005 wake-slot handoffs skip
  reset_timeslice, so that path never paid the lfence).
  All other rows within the box's ~+-15% noise floor.
- 1-core sandbox: 218 -> 206 cyc (understates by ~3x).
- bench binary lfence count 24 -> 4 (survivors = the bench's own
  rdtscp;lfence bracket). Tests green.
Baseline (benches/baseline.json) not re-saved in this commit.
2026-08-21 12:22:46 +00:00
claude-asm-audit 343e53e17b bench: baseline under the fixed rq-mpmc default (job 6e71e9b1)
20-core box (taskset 0-19), 5 sets, at 0438f12. Replaces the rq-mutex
baseline (df41ff3) so sweep.py regress compares like with like. The
multi-thread rows show what the mutex contention was costing:
yield_many 20T 156.7ms -> 44.1ms (-72%), spawn_storm_busy 20T -72%,
chained_spawn/catch_unwind 20T -20%. Sweep ran clean (0 panics) —
the finding-13 fix holding under the full suite.
Note: mpsc_contention 1T is ~2x the mutex-default row (3155 -> 6123);
backend characteristic to investigate, not a fix regression (the
shootout shows the fix itself is noise-neutral).
2026-08-21 12:22:46 +00:00
claude-asm-audit 9b215573de fix(run_queue): make the Vyukov rings preemption-tolerant (finding 13)
A consumer OS-preempted between its dequeue_pos claim and its seq
release freezes one cell; once traffic laps the ring (~cap ops ≈ 1-2ms
at yield-storm throughput ≈ one scheduling quantum under load),
try_push reads the stale seq and the original algorithm's 'lap behind
=> full' inference misfires. The old push assert then converted that
liveness stall into an abort blaming a double enqueue that never
happened (soak: occupancy 174-181 of cap 16384 at every failure), and
the dead scheduler threads stranded actors => the observed hangs.
Pristine 5504ef3 failed 14/14 under an 8-spinner soak on the 20-core
box; a retry prototype passed 13/14, its one failure being a stall
that outlived a fixed 1M-spin bound — pure spinning starves the
descheduled consumer, so the wait must yield.

Fix, following crossbeam ArrayQueue's shape (fence + opposite-counter
check; cells hand off, COUNTERS give verdicts):
- push: on a lap-behind cell, fence(SeqCst) + occupancy check.
  occ < cap => transient stall => spin-then-yield backoff and retry;
  occ >= cap => the REAL at-most-once-enqueued violation => panic with
  a truthful message and the counters. enqueue_pos is loaded before
  dequeue_pos so racing pops only underestimate occupancy (no spurious
  panic).
- pop: symmetric counter check before an empty verdict; on a
  mid-publish producer, bounded wait then None — deliberate deviation
  from crossbeam's unbounded retry (spurious None is benign: RFC 018's
  enqueue-wake self-heals; StripedRing's probe must not hang on one
  stripe).
- StripedRing::push probe: yield-escalating backoff after a full
  refused lap (was a bare spin_loop).
- Hand-rolled Backoff (spin 2^n to 64, then yield_now); under loom
  every wait is a yield so models explore the stalled peer's progress.

New loom model mpmc_lap_onto_stalled_consumer_completes reproduces the
old panic in the first explored interleavings (verified FAILED against
5504ef3) and passes with the fix. Lib 60 + integration 41 + all 4 ring
loom models pass. rq-striped inherits the fix (stripes are MpmcRings).
2026-08-21 12:22:46 +00:00
claude-asm-audit bf55cef3e3 perf(run_queue): flip default backend rq-mutex -> rq-mpmc
Evidence (history.md session 5, findings 10-11, 20-core jobrunner run):
rq-mutex collapses with thread count on queue-heavy load (yield-storm
6.9x slower than striped at 20T; attributed cause of the baseline's
multi-thread regression on yield_many/chained_spawn), while rq-mpmc is
best-or-close everywhere: best 1-thread, best slot-on ping-pong
(1237us vs mutex 6754us at 20T), -23%/-41 cyc per yield roundtrip on
real hardware (single-core switch_cost, interleaved). rq-striped
remains the churn-heavy many-core option, selectable per build.

Docs + compile_error hints in run_queue.rs updated to name rq-mpmc as
the default. slot_state.rs untouched, no loom-relevant changes; lib
(60) + scheduler/channel/supervisor/park_wake/wake_slot/preempt (41)
pass under the new default.
2026-08-21 12:22:46 +00:00
claude-asm-audit 2f88264426 bench: refresh baseline.json — 20-core box, fe85197, rq-mutex
Re-measured via jobrunner (taskset -c 0-19 of 24, rust:1.97-slim,
sweep.py run --save-baseline, 5 sets). Supersedes the old 24-thread
baseline; multi-thread labels are now 'smarm 20-thread'. Taken under
the rq-mutex default *before* the backend flip — the yield_many /
chained_spawn multi-thread regression reproduces here and is attributed
to rq-mutex contention (history.md session 5, finding 10).
2026-08-21 12:22:45 +00:00
claude-asm-audit 6172f4231d perf(runtime): skip take_closure's locked swap after first resume
Every resume paid an unconditional AtomicPtr::swap (lock xchg, full
barrier) to check for a first-resume closure that is null on all
resumes after the first. A Relaxed null-load fast path is sound:
store_closure runs only before publish_queued, whose Release pairing
with try_claim's Acquire orders it before this call, so no writer can
race the load within an occupancy.

Measured on switch_cost (1-core sandbox, rq-mutex, cycles): mean
roundtrip 350-355 -> 324-328, ~7.5%. All lib + scheduler/channel/
supervisor tests pass.
2026-08-21 12:22:45 +00:00
claude-asm-audit d7082eb266 perf(context): pass actor sp through registers, not TLS
switch_to_actor takes the target sp in rdi and returns the actor's
next saved sp in rax (handed over by switch_to_scheduler's shim).
Deletes the ACTOR_SP thread-local and halves the helper calls per
one-way switch (2 -> 1); the scheduler loop also drops its
set_actor_sp/get_actor_sp TLS round-trips. SCHEDULER_SP stays: a
yielding actor at arbitrary call depth has no argument channel back.

asm before/after in outputs/history.md session 1-2. Tests: 354 pass.
2026-08-21 12:22:45 +00:00
Claude (sandbox) 741c10337b release: v0.7.0
Bump crate version to 0.7.0.
v0.7.0
2026-08-21 13:10:25 +02:00
Claude (sandbox) e570138da5 docs(roadmap): supervisor start order is not start readiness
Filed from the urus v0.3 endpoint work. start_child spawns and moves on,
so a later sibling can whereis an earlier named child before that child's
actor has run. Notes why blocking spawn is not the fix ('has begun
executing' != 'has bound its name', plus a per-accept round-trip tax and
every spawn becoming a context-switch point), that OTP has the same async
spawn and synchronises one level up in gen_server:start_link, the
readiness-ack shape if scheduled, and the structural workaround urus uses
today (registrar spawns its own consumers).
2026-08-20 13:20:42 +00:00
Claude (sandbox) 415effb2e9 feat(gen_server,gen_statem): lifetime is the actor's — refs are addresses; inline named run
Root cause behind the "pin the endpoint" gotcha and the trapping-wrapper
pattern: a gen_server had two lifetime authorities — its refs (last one
dropped → inbox closes → exit) and, when supervised, its supervisor. OTP has
one: a process lives until it stops, is shut down, or is killed; a pid is an
address. Root exit now shutting down every forest root removes the reason the
ref-governed idiom existed (a forgotten server no longer hangs the run), so
adopt the one rule:

- The server/machine loop holds one inbox sender for its life; the inbox
  never closes. GenServerRef / GenStatemRef are addresses. Explicit close is
  `shutdown()`; a forgotten one is swept at root exit.
- `NamedGenServerBuilder::run()` / `gen_statem::run_named(name, m)` run the
  loop inline as the current actor: a server is a direct ChildSpec child,
  gets the supervisor's shutdown as handle_shutdown / a shutdown row, re-binds
  its name on restart, and is addressed by name. The wrapper in
  examples/graceful_shutdown.rs is gone.
- gen_statem gains GenStatemName + whereis_machine/send/call/shutdown by name
  (parity with gen_server); the macro gets `Sm::new`.
- Root-exit sweep records `Event::RootSweep { target, trapping }` under
  smarm-trace ("root_sweep shutdown|stopped"): unsupervised leftovers are
  visible rather than silently owned-by-refs.
- Named start() name-clash path stops the spawned actor instead of relying
  on ref drop.

Tests: tests/gen_server_lifetime.rs, tests/gen_statem_lifetime.rs,
tests/root_sweep_trace.rs (feature-gated); three existing tests that used
drop-closes-inbox now use shutdown(). Docs/README/ROADMAP/Deep Dive updated.
2026-08-19 17:47:31 +00:00
Claude (sandbox) 849a424c8e docs,examples: graceful shutdown — new examples/graceful_shutdown.rs, README 'Stopping actors', named_genserver uses shutdown(), Deep Dive terminate note, ROADMAP open items 2026-08-19 16:28:42 +00:00
Claude (sandbox) 6ceb138f5f feat(gen_statem): graceful-shutdown parity — trap_exit, shutdown/exit rows, stop, terminate
Mirrors the gen_server surface in gen_statem's event model:
- Cx::trap_exit() (in the initial enter): a shutdown request then arrives
  as the Shutdown event, routed by state through `shutdown` rows (default
  for a state with no row: stop); linked-peer deaths as `exit <pat>` rows
  (default: drop). Non-trapping machines are stopped outright, as before.
- Cx::stop(): normal self-exit after the current event; `stop` tail keyword
  is sugar for { cx.stop(); prev }.
- Machine::terminate (optional `terminate { … }` macro block), run from a
  Drop guard on every exit path; the guard also drains armed timers.
- Machine::shutdown_ev / exit_ev (defaults None) so hand-written machines
  keep compiling; GenStatemRef::shutdown() is graceful and waits.
- Loop selects exits > timers > inbox.

Tests: tests/gen_statem_shutdown.rs.
2026-08-19 16:25:15 +00:00
Claude (sandbox) 250f31265b feat(runtime): root exit is graceful shutdown of the forest roots
The RFC 014 root-exit sweep hard-stopped every live slot once nothing was
runnable. That deferral privileged queued work over parked-with-a-pending-
wake work (a sleeper was killed, a queued cast was drained) and any attempt
to widen the notion of pending wake (timers, fd readiness) re-wedges the
run on the periodic-timer daemon the sweep exists to end.

Root exit now means "the program is done": finalize_actor delivers
request_shutdown to every forest root — each live actor whose parent is
the run (ROOT_PID) or is dead — synchronously, before the live-count
decrement. Supervisors cascade per child Shutdown policy; trapping actors
may Continue/drain with working timers and end the run when they stop
themselves; non-trapping actors are stopped outright. No forcing sweep.

Removes root_exited/root_swept, Pop::RootDrain and the idle-verdict
condition; adds tests/root_exit.rs.
2026-08-19 07:11:46 +00:00
claude 9c8f59ca53 feat(scheduler,supervisor,gen_server): graceful shutdown — request_shutdown, child Shutdown policy, handle_shutdown
Lift OTP's `exit(Pid, shutdown)` + child-spec `shutdown` wholesale.

scheduler / runtime
- `request_shutdown(pid)`: the polite stop. A target trapping exits gets an
  `ExitSignal { reason: DownReason::Shutdown }` on its trap inbox and keeps
  running; a non-trapping target is stopped as by `request_stop`, which is
  now documented as the hard stop (`exit(Pid, kill)`). Dead pid: no-op.
- `RuntimeHandle::request_shutdown` for the off-runtime (signal thread) path;
  `from == ROOT_PID` there.
- `DownReason::Shutdown` — appears only in ExitSignal, never in Down (a
  complying target exits *normally*).

supervisor
- `ChildSpec::shutdown(Shutdown::{BrutalKill, Timeout(d), Infinity})`,
  default Timeout(5s). Every supervisor-initiated stop (ordered shutdown and
  OneForAll/RestForOne sibling cycling) is: request_shutdown → await the
  child's Signal up to the grace → request_stop → await. Sequential, reverse
  start order.
- The supervisor traps exits; a Shutdown ExitSignal runs the ordered
  shutdown and `run()` returns normally, so `request_shutdown(root_sup)`
  tears a whole tree down top-down with each child's grace period.
- FIX: a hard `request_stop` on a supervisor previously orphaned its
  children (the ordered shutdown lived after the loop, and the unwind
  skipped it). `Live` (the by_pid map) now carries a drop guard that
  fire-and-forget hard-stops live children when unwinding.

gen_server
- `GenServerCtx::trap_exit()` opt-in in `init`; the trap inbox becomes arm 0
  of the loop's select. Shutdown ExitSignal → `handle_shutdown() ->
  ShutdownAction::{Exit, Continue}` (default Exit: loop breaks, `terminate`
  runs on the normal path and may block). Other ExitSignals →
  `handle_exit(sig)`.
- `GenServerCtx::stop_handle() -> StopHandle`, `stop()` ends the server
  after the current message with a *normal* exit — the missing
  `{stop, normal, State}`; `request_stop(self_pid())` was the only self-exit
  and it is abnormal (Transient restarts it).
- `GenServerRef::shutdown()` / `gen_server::shutdown(name)` now go through
  `request_shutdown`.

Tests: tests/shutdown.rs, tests/supervisor_shutdown.rs,
tests/gen_server_shutdown.rs. Full suite green; fmt + clippy --lib clean.
2026-08-19 06:34:48 +00:00
Claude (sandbox) 1002777ef3 feat(channel,runtime): off-runtime cross-thread wake for parked receivers and stop
Wakes issued from a non-scheduler OS thread were silent no-ops. Every
off-runtime wake primitive (unpark, unpark_at, request_stop) reaches the
runtime through the RUNTIME thread-local, which is unset on any foreign
thread — so a send from a plain std::thread enqueued its message but never
woke the parked receiver, and there was no way to drive a stop into a
runtime from an application thread (e.g. an OS-signal handler). The former
strands a parked recv forever; the latter is why a downstream server must
poll a shutdown flag instead of parking on it.

Generalize RFC 018's rule — a producer reaches the runtime through a Weak it
holds — from the IO backend to channel senders and to a new handle:

- A receiver captures a Weak<RuntimeInner> (provably live at that moment)
  alongside its (pid, epoch) when it parks. send() and the last-sender drop
  wake through scheduler::unpark_at_via, which takes the thread-local path
  when on a scheduler thread (preemption-gated, slot-eligible) and the
  captured Weak otherwise — the same cross-context wake the IO threads do.
- Runtime::handle() returns a Send + Sync RuntimeHandle carrying that Weak;
  RuntimeHandle::request_stop drives a cooperative stop from any thread and
  is a no-op once the runtime is dropped.

The in-runtime wake paths (recv/select timers) are unchanged; only the sites
reachable from a foreign thread route through the Weak. RuntimeHandle exposes
request_stop only: send-wake needs no user-facing handle, and off-runtime
unpark is covered because request_stop drives unpark on the upgraded inner.

tests/cross_thread_wake.rs: a foreign-thread send wakes a parked receiver; a
foreign-thread request_stop wakes and stops a parked actor; a RuntimeHandle
held across and beyond run never blocks all-done and degrades to a no-op.
2026-08-19 05:50:54 +00:00
smarm 8f2d513940 README: rewrite overview, add limitations, roadmap, and contribution notes
Expands the intro into an overview/limitations structure, documents
preemption, tracing/causal profiling, URUS, and adds a 'Coming up' and
'A note on open source' section.
2026-08-18 00:25:25 +02:00
Claude (sandbox)andClaude (sandbox) ca1c98336e feat(scheduler,runtime): non-panicking try_spawn for at-capacity load shedding
allocate_slot() panics on a full slab; for a load-shedding caller (an
accept loop spawning one actor per connection) that panic lands in the
spawning actor, which then crash-loops under Restart::Transient into the
still-full slab until its restart budget is spent — and the service stops
accepting entirely. Observed live (urus slowloris scaling, 2026-08-10).
A full slab is a routine overload condition for such callers, not an
invariant violation.

- RuntimeInner::try_allocate_slot() -> Option<u32>: the non-panicking
  core; a single pop under the free-list lock, so the claim is atomic
  (claim-or-report — no check-then-spawn TOCTOU, no headroom margin).
  allocate_slot() is now a thin panicking wrapper over it.
- scheduler::try_spawn / try_spawn_under_with -> Result<JoinHandle,
  SpawnError>: parity with spawn/spawn_under_with except a full slab
  returns Err(SpawnError::AtCapacity) instead of panicking. Minimal
  surface per the agreed strategy; the remaining _with/_addr mirrors are
  trivial wrappers if ever needed.
- Slot-first ordering on the try path (reverse of spawn's stack-first):
  under overload Err is the hot path, and a rejection costs one mutex
  pop — no mmap/pool-pop + init + recycle per shed unit of work. A
  drop-guard returns the claimed slot if stack allocation panics in the
  claim-to-install window (would otherwise leak and trip run()'s
  teardown slot-leak debug_assert).
- SpawnError: non_exhaustive, Display + std::error::Error.
- spawn and every existing call site untouched: the panic remains the
  correct loud invariant check at internal/bounded spawn sites.

tests/try_spawn.rs: parity when slots free; exact slab accounting at
capacity (Err, no panic, repeatable); custom-shape try refuses before
stack allocation; self-heal after slots free; plain spawn still panics
(surfaced via JoinError payload); 4-thread race for the last slots
claims exactly the free count; SpawnError impl checks.

Design doc: smarm-suggestion-try-spawn.md. Downstream consumer change
(canned 503 on AtCapacity in urus's accept loop) is urus scope, not
smarm.

(cherry picked from commit 36de4b36aeaa72b2a5f9f3797b9854652656dcf6)
v0.6.1
2026-08-13 15:03:16 +02:00
smarm-agent 95306c7f60 style: cargo fmt sweep under rustc 1.97.1 (toolchain reformat, no semantic change) 2026-08-13 05:56:49 +00:00
smarm-agent 1262cc30e3 monitor: widen stamp eligibility to watchable = named ∪ exported (soak sig 5)
The terminal record existed for watches that raced their target's death, but
e43c673 scoped its stamp to named tenancies — and the pid-identity watch
surface (§4 Slice 3) targets arbitrary actors, including anonymous ones whose
pids cross the boundary in contract replies. The first wild pid-face hit
(width-20 soak, pid_watch_test.exs:47, 1/600 full-suite: a monitor installed
while the child was alive delivered :noproc instead of {:smarm_exit, :panic})
is exactly the residual a0ba9be's commit body deferred.

ever_named becomes `watchable`, with a second set-site: mark_watchable(pid),
which the bridge calls wherever a smarm pid is encoded across the boundary —
BEAM can only watch pids it holds, and can only hold pids that crossed.
Anonymous never-exported churn (holder threads, egress tasks) stays
ineligible, preserving e43c673's LIFO-eviction protection unchanged.

mark_watchable takes the cold lock before the liveness screen: finalize
publishes Done and reads the bit under the same lock, so the mark either
lands before the death stamps or observes the tenancy dead and no-ops —
no lost-stamp window, and marking a corpse cannot invent history (pinned
in the test alongside the mark-while-alive stamp).
2026-08-13 05:56:19 +00:00
smarm-agent 461fe4b768 fix(runtime): only named tenancies stamp the terminal record — anonymous churn must not evict it
Discovered wiring the bridge consult: with an unconditional stamp, the record
for the very death being raced was the shortest-lived data in the runtime.
Every green thread is a slot tenant, the free list is LIFO — so the slot a
named server's death frees is the first one recycled, and the next throwaway
exit (monitor holders, chain-runner work, anything) overwrote the record
before a raced watch could consult it. Deterministic bridge repro: the
corpse resolved fine, terminal_reason read None every time.

register_with now flags the tenancy (ever_named, reset at reclaim) before
the binding lands — set outside the registry lock, so no successfully
registered actor can die unflagged and a failed register's overshoot is
harmless — and finalize stamps only flagged tenancies. Watchable identities
are exactly the named ones (the bridge's pid-identity path deliberately
keeps Erlang's raw :noproc), so nothing consultable is lost.

Contract test updated: the three death modes now self-register; a new
anonymous control pins that unregistered deaths neither stamp nor evict.
2026-08-13 05:56:19 +00:00
smarm-agent b937f1f50f monitor/registry: terminal-outcome record — a raced watch can recover the real down reason (soak sig 4)
A watch installed after its target's death has, until now, only NoProc to
report — but the bridge's proxies install their native watch asynchronously
after acquire returns, so a link established before a crash (from the BEAM's
view) could still lose the panic's translated reason to that blanket NoProc
(width-20 soak signature 4: link_test.exs:26, 1/600 full-suite, 3/2000
link-only, all whereis-miss; deterministic repro in the bridge suite).

Two primitives, no change to monitor()'s own Erlang-faithful stale-pid
semantics — the upgrade is the caller's deliberate act:

- finalize_actor stamps the slot with (generation, DownReason) under the same
  cold-lock block that publishes the outcome. The record survives reclaim,
  registry pruning, and the next tenant's install; only the slot's next death
  overwrites it. terminal_reason(pid) reads it generation-matched.
- resolve_name(name) is whereis with the corpse kept: the dead-holder arm
  returns the stored pid it prunes (NameResolution::Corpse) instead of
  discarding the only evidence of who died — whereis itself prunes on the way
  out, so a whereis-then-lookup consumer would find the evidence already
  destroyed. Live/Unbound match whereis's Some/None; the name heals exactly
  as before.

Contract pinned in tests/terminal_outcome_after_death.rs: one record per way
of dying (Exit/Panic/Stopped), no record while live, corpse capture + heal on
resolve_name, record independence from registry pruning, survival across slot
re-tenancy, overwrite at the next tenancy's death.
2026-08-13 05:56:19 +00:00
Claude (sandbox) 301e3463e3 chore(release): v0.6.0 — RFC 019: actor stack reserve & shrink
Per-actor stack shapes on every spawn surface (SpawnOpts stack_reserve/
guard_size, Config defaults, pool rule: only default-shaped recycle);
sampled stack high-water + MADV_FREE shrink at actor-park (THRESHOLD
256 KiB, COOLDOWN 64 parks, redzone 1 page); pool-recycle MADV_DONTNEED
above the retained 64 KiB entry end; SIGSEGV overflow diagnostics
(two-tier: in-guard definitive / 1 MiB overshoot 'stepped over', prior
handler chained for foreign faults) with per-scheduler sigaltstack; and
the per-actor introspection surface (ActorInfo.stack: reserve, guard,
sampled depth, parks_since_shrink, shrinks).

Amendments ratified during implementation, for the RFC changelog:
- DEFAULT_STACK_GUARD 64 KiB -> 1 MiB, following the kernel's post-Stack-
  Clash stack_guard_gap convention; PROT_NONE width is VA-only and free.
- §7's motivating segfault was a cargo-vendored gz build, not SQLite as
  the RFC text says (cc-built C lacks -fstack-clash-protection; distro
  libraries have it — the risky class is vendored builds).
- §4 hibernate() deferred to the jar (bolt-on: force-flag on the §3
  shrink path, ~10 lines when wanted).

Gates (jobrunner box, 2026-08-08): reclaim gate PASS at c3 and again at
tip (3.0 MiB LazyFree -> kernel reclaim -> Rss to one live page ->
re-spike bit-identical, live data intact; MADV_PAGEOUT stands in for
memcg — cgroup2 is RO in the job container — driving the same reclaim
path). E1 interleaved A/B vs v0.5.0: every ka cell (the E1 subject)
within +0.3..+2.9% at tip; close-mode control cells within noise except
t8-c4 close, which is bistable (~40-44k vs ~46-49k modes for BOTH
variants, base self-disagrees by 11% across rounds); 6 rounds across two
runs are inconclusive there and a 10-round focused run is noted in the
handoff as deferred follow-up, accepted for this release.

No breaking API changes since v0.5.0: SpawnOpts fields and ActorInfo
gained members (exhaustive-construction downstream will need the new
ActorInfo.stack field; urus does not construct it).
v0.6.0
2026-08-08 19:44:55 +00:00
Claude (sandbox) 410ba33d82 feat(introspect,runtime): per-actor stack surface on ActorInfo (RFC 019 §8)
- introspect::StackInfo { reserve, guard, depth_high_water,
  parks_since_shrink, shrinks } as ActorInfo.stack; re-exported at crate
  root beside ActorInfo.
- All reads lock-free: geometry from the c6 diag slot atomics, depth =
  top - hwm (the §2 sampled high-water; doc spells out sampled-not-exact
  and that 0 means never-descheduled-at-depth), counters straight off the
  §3 atomics. Coherence for the incarnation rides read_slot's existing
  generation check, same as overruns/messages_received.
- Slot::stack_introspect(): one pub(crate) tuple accessor beside the other
  counter accessors.
- Exact RSS deliberately absent per RFC (mincore = debug tooling only,
  never a runtime path); stack_shape(pid) untouched (cold-lock exact
  variant from c2).
- tests/introspect.rs: defaults surface (64 KiB reserve / 1 MiB guard /
  sampled ~32 KiB depth / gate park counted / zero shrinks) + live shrink
  counters (spike visible pre-shrink; shrinks>=1, cooldown counter reset,
  hwm reset after crossing COOLDOWN) read mid-run -- post-join the slot
  reclaim correctly hides the incarnation, which the first draft of the
  test learned the hard way.

FLAGGED (Claude-solo calls):
- Nested StackInfo struct over five flat ActorInfo fields (grain break;
  the five fields are one concern and ActorInfo is already 12 fields).
- Field names reserve/guard/shrinks (RFC says stack_reserve/stack_guard/
  shrink count; the stack_ prefix is redundant inside StackInfo).
2026-08-08 19:12:46 +00:00
Claude (sandbox) 5fd8aecf55 feat(signal,runtime,stack): SIGSEGV overflow diagnostics + 1 MiB guard default (RFC 019 §7)
- src/signal.rs: process-global SA_SIGINFO|SA_ONSTACK handler installed once
  at runtime::init (before any scheduler thread -> unracing PRIOR save);
  per-scheduler-thread 64 KiB sigaltstack registered at schedule_loop entry
  (a guard hit leaves no stack to handle on). Async-signal-safe throughout:
  classification is plain loads (const-init TLS Cell + slot atomics), print
  is fixed-buffer itoa + one write(2), death is SIG_DFL + refault at the
  same instruction (core-dumpable, correct wait status).
- Two-tier classification (agreed): in-guard = definitive; OVERSHOOT window
  below the guard = 'unprobed (FFI?) frame stepped over it' probable
  attribution -- the RFC's motivating incident (cargo-vendored gz, not
  SQLite as the RFC text says) faults there under a small guard. Pure
  classify() fn, 5 adversarial units incl. saturation at low addresses.
- DEFAULT_STACK_GUARD 64 KiB -> 1 MiB (agreed): kernel stack_guard_gap
  anchor post-Stack-Clash; PROT_NONE is VA-only (no RSS, no page tables,
  no overcommit charge) so width is free at any actor count.
- Unclassified faults reinstate the PRIOR sigaction and refault (agreed):
  std's own OS-thread overflow diagnostics survive our presence.
- Slot: diag_{stack_top,stack_reserve,stack_guard,pid} atomics written in
  install_actor pre-publish; readable without the cold lock (Stack lives
  under it); only consulted while CURRENT_SLOT points at the slot, so
  never stale where read. preempt::current_slot_ptr ungated from
  smarm-causal (now also the classifier's anchor).
- build.rs + cc (agreed Q3): canary/canary.c, 96 KiB local touched low-end
  first, -fno-stack-clash-protection pinned so hardened toolchains don't
  probe the canary into uselessness.
- tests/stack_diag.rs: subprocess x4 -- Rust recursion tier-1; FFI canary
  tier-1 at defaults (1 MiB guard catches the jump); tier-2 at guard=4 KiB
  ('stepped over', reproduces the incident); clean at reserve=256 KiB
  (the §1 knob is the fix, same frame).

FLAGGED (Claude-solo calls):
- OVERSHOOT_SLOP = 1 MiB (matches guard default/kernel gap; beyond it
  attribution would be dishonest).
- Altstack 64 KiB, mmap'd once per OS thread, never freed (bounded by
  thread count; reused across run()s via TLS flag).
- Foreign-fault reinstate permanently deregisters our handler; accepted --
  the process is dying either way.
- Diag geometry as 4 slot atomics (install-time cost only) over a per-switch
  TLS snapshot (hot-path stores).
2026-08-08 18:58:30 +00:00
Claude (sandbox) 7d8b9e0310 feat(stack,runtime): pool-recycle DONTNEED above the retained entry end (RFC 019 §6)
- stack::retain_range: pure checked span fn (retain page-up = zap less;
  None when retain covers the reserve, so the 64 KiB default config never
  pays a syscall) + 6 adversarial units mirroring shrink_range's.
- Stack::recycle_zap: advisory MADV_DONTNEED of [usable_base, top-RETAIN);
  stack is unowned at the call site, synchronous eager zap races nothing.
- recycle_stack: zap OFF-LOCK before pool admission (acquire_stack's
  no-syscall-under-the-pool-lock invariant); rare cap-overflow pays a
  wasted zap ahead of munmap, accepted over a second lock round-trip.
- pub const RECYCLE_RETAIN = 64 KiB beside the shrink knobs, ratified-as-
  constant rationale in doc.
- tests/stack_recycle.rs: mincore-based exact-zero-resident assert over
  the zap span. smaps was tried first and over-counts: a neighboring rw
  anon VMA can merge flush against the stack top (observed once under the
  full-suite run); the PROT_NONE guard pins the usable base exactly.

FLAGGED (Claude-solo calls):
- RFC §6 'above the bottom RETAIN' is direction-ambiguous in address
  terms; implemented as retain the ENTRY end (highest addresses, the
  pages the next actor faults first), zap the cold deep span below.
- Const named RECYCLE_RETAIN (RFC says RETAIN) to sit beside SHRINK_*.
2026-08-08 16:13:53 +00:00
Claude (sandbox) 8225716b11 feat(runtime,stack): sampled stack high-water + MADV_FREE shrink at actor-park (RFC 019 §§2–3)
hwm: AtomicUsize lands beside sp on the slot: the single context-save
site min-updates it (one branch + at most one Relaxed store into the
line the sp store just dirtied), install resets it to the fresh top.
Advisory by construction — correctness never depends on it. The mod-doc
ordering chain gains a line: hwm piggybacks the existing
Relaxed-store-before-Release pattern and adds no edges.

Shrink hook in the YieldIntent::Park arm only, before the park_return
Release transition — the owned window (obligation 1's assert-comment at
the site): after the sp store, before Parked is published, scheduler on
its own stack, actor saved and unstealable. It runs on both arms of the
park_return race (a consumed unpark flag means one wasted-but-harmless
madvise). The preempt/yield path deliberately never checks: §4's
bounded, self-healing leak under saturation, when syscalls are least
affordable.

SHRINK_THRESHOLD = 256 KiB and SHRINK_COOLDOWN = 64 parks are pub
constants with the ratified doc rationale, not Config fields. The freed
span is shrink_range(hwm, sp, page): whole pages of [hwm, sp − 1-page
redzone), rounded inward, checked arithmetic — adversarial inputs
collapse to None (obligation 2). MADV_FREE marks lazily; the kernel's
reclaim-under-pressure IS the hysteresis, cancel-on-write is the safety
net. parks_since_shrink + shrink_count ride the slot for the cooldown
and the future introspect surface.

Tests: 7 adversarial shrink_range units (inverted/empty spans, redzone
underflow, unaligned ends, sp-crossing sweep); integration — 8 MiB
reserve, ~3 MiB spike sampled via yield-at-depth, parks gated on
introspected Parked state past the cooldown, then ≥ 2 MiB LazyFree
asserted inside the stack's smaps range with live data intact; and the
inverse guard — a shallow never-spiking actor ends at exactly 0
LazyFree (also proves the parser isn't vacuously zero via the first
test).
2026-08-08 14:30:32 +00:00
Claude (sandbox) 3cb64eefc2 feat(scheduler,gen_server,gen_statem,introspect): SpawnOpts — per-actor stack shape on every spawn surface (RFC 019 §1)
SpawnOpts { stack_reserve, guard_size } with Option<usize> fields, None
resolving to the Config defaults at spawn time — a deliberate deviation
from the RFC's plain-usize struct so struct-update syntax works without
a runtime handle in scope. Threaded across the five surfaces:
spawn_with, spawn_under_with, spawn_addr_with,
GenServerBuilder::stack_opts (mirrored on NamedGenServerBuilder), and
gen_statem::spawn_with (gen_statem has no builder, so the opts ride a
_with variant — Claude-solo surface call, flagged for review). Existing
spawns forward defaults; no call-site churn.

introspect::stack_shape(pid) pulled forward (agreed) as the first slice
of the RFC 019 introspection surface, giving tests an observable.

Tests (tests/spawn_opts.rs): override/partial-override/rounding on each
surface; obligation 4 from the outside — a dead custom stack is never
handed to the next default spawn (LIFO pool would expose it), and the
reverse (default stacks ARE recycled); 8 MiB reserve behaviorally
permits ~1 MiB recursion. Also: silence unused-Result in the c1
runtime test (join now unwrapped).
2026-08-08 14:22:38 +00:00
Claude (sandbox) 0fe052bc7e feat(stack,runtime): per-shape actor stacks — Stack::new(reserve, guard), Config knobs, pool rule (RFC 019 §1)
Stack takes an explicit (reserve, guard) shape, both page-rounded and
stored; usable_base derives from the stored guard. Guard default raised
4 KiB -> 64 KiB (DEFAULT_STACK_GUARD): probestack makes one page enough
for Rust frames, but an unprobed C frame can leap a page in one sub rsp
— the motivating SQLite segfault. Reserve default stays 64 KiB
(DEFAULT_STACK_RESERVE); ACTOR_STACK_SIZE retired.

Config::{stack_reserve, stack_guard} thread the runtime defaults into
RuntimeInner pre-rounded. All acquisition/recycling now goes through
acquire_stack/recycle_stack carrying the pool rule: only default-shaped
stacks are pooled (pooled ⇒ default-shaped by induction); custom shapes
mmap fresh and munmap at death. Pool lock still dropped before any mmap.

No public spawn API change (SpawnOpts is the next commit).

Tests: shape rounding + accessors, wide-guard faults at both ends
(subprocess), Config::stack_reserve permits >64 KiB recursion that
previously could only segfault.
2026-08-08 14:18:27 +00:00
smarm a03a7ca01e chore(release): v0.5.0
Breaking API rename since v0.4.0: gen_server's ServerRef/ServerBuilder/
ServerCtx -> GenServerRef/GenServerBuilder/GenServerCtx, Watcher<G> is
now generic over its GenServer, and GenServer gained a required
associated Timer type for timer-fire payloads (arm_after/handle_timer).
Downstream consumers (urus) have been ported.
v0.5.0
2026-08-08 11:44:35 +02:00
smarm d4839f1d81 feat(runtime,io): driver-enqueues + park/wake idle path — retire the wake pipe
The swap (RFC 018). Schedulers no longer sleep on a shared level-triggered
wake pipe — the herd source that made the default 8-thread config 7x
slower than 2 threads (E1). They park on per-thread futex parkers via the
coordination layer; IO backends become producers behind a two-call
contract (make runnable, then the enqueue tail wakes exactly one parked
scheduler).

Deleted: the drain lock and the one-winner phase-1 drain; the shared
completions VecDeque; the wake pipe fds, poll_wake, drain_wake_pipe,
wake_scheduler, the FdReady/Blocking Completion enum; the 100us idle nap;
the per-pop io.lock liveness read; io.rs's as_millis timeout truncation.

Added:
- enqueue wake tail (fixes the silent enqueue): wake_one_if_idle, a fence
  + one Relaxed mask load when everyone is busy — the pure-compute hot
  path pays almost nothing.
- driver-enqueues: the pool thread stashes its result in the slot,
  decrements io_outstanding, unparks; the epoll thread removes+DELs the
  waiter under the waiters lock and unparks. Both reach the runtime via a
  Weak (no Arc cycle). The waiters map moves behind its own Arc<Mutex> so
  the epoll thread never takes the runtime io lock (teardown holds it
  while joining that thread).
- io_outstanding / io_fd_waiters atomics: the termination verdict reads
  two atomics instead of taking io.lock on every pop.
- timekeeper idle path: at most one parked scheduler holds the timer
  deadline (an expiry wakes one, not a herd); everyone else parks
  indefinitely and is woken by the enqueue tail.
- busy-path timer due-check (ratified design point (a)): under saturation
  nobody parks and no timekeeper exists, yet due timers must still fire —
  one Relaxed load of the earliest-deadline snapshot per loop, clock read
  only when a timer is armed. Maintained under the timers mutex.
- chain rule: a scheduler that pops with more work queued and a sibling
  parked wakes one, so surplus runs in parallel rather than behind it.

tests/park_wake.rs pins the two new observable properties: timers fire
under full scheduler saturation, and sub-ms sleeps are prompt (the
as_millis truncation regression). Full suite + all loom models green;
clippy --lib clean.
2026-07-24 09:12:12 +02:00
smarm 2854b560d6 feat(park): fenced producer fast path + earliest-deadline snapshot
Two integration-driven amendments ahead of the runtime swap:

wake_one_if_idle() realizes RFC 018's "empty-mask fast path is one
relaxed load" soundly: a bare relaxed load is a lost-wake in the Dekker
shape for the lock-free ring queues, so the producer publishes work,
fences (SeqCst), then reads the mask Relaxed — paired with a matching
fence between the consumer's bit-publish and its re-check in park().
The pure-compute hot path (mask 0) never takes the shared mask line
exclusive; the RMW read stays on the rare chain-rule path only. Loom
models 1/2 now drive the fenced pattern end to end.

next_deadline is the earliest KNOWN timer deadline, independent of
whether anyone is parked — which tk_armed cannot give: under saturation
nobody parks, nobody arms, yet due timers must still fire (ratified
design point (a): the busy-path due-check). Maintained under the timers
mutex (note_deadline on insert — which also carries the timekeeper
re-arm wake — refresh_deadline after pop/clear); read lock-free.
deadline_due() costs one Relaxed load and a branch when no timer exists;
the clock is read only when one does.
2026-07-24 09:12:12 +02:00
smarm 7b026cfe56 feat(park): scheduler coordination layer — parkers, idle mask, wake protocol (RFC 018)
Schedulers get an IO-agnostic sleep/wake primitive of their own: one
futex Parker per scheduler thread (permit semantics, std::thread::park
shaped — closes the check-then-park race), an AtomicU64 idle mask with a
set-bit → re-check → wait park protocol, wake_one (highest-bit LIFO,
CAS-clear before unpark: exactly one wakeup per call by construction),
wake_all for the terminal path, and the timekeeper role — at most one
parked scheduler holds the timer deadline, with an atomic armed-deadline
snapshot for the busy-path due-check and an insert-side re-arm wake.

Deadlines travel as nanosecond timespecs end to end; the wake pipe's
as_millis truncation is unrepresentable here. The Dekker publish/re-check
shape is resolved by the same-location-RMW handshake (AcqRel), not SeqCst
loads; loom verifies exactly this in four models (no-lost-wake, chain
propagation, timekeeper handoff, termination), run with
LOOM_MAX_PREEMPTIONS=3 — unbounded exploration is impractical for the
looped models. Loom/non-Linux builds park on a Mutex+Condvar via
sync_shim.

Standalone until the runtime swap (next commit): nothing outside tests
constructs a Coordinator yet, hence the temporary dead_code allow in
lib.rs.
2026-07-24 09:12:12 +02:00
smarm 006a3283e7 chore(hooks): clippy gate falls back to a nix-shell toolchain
Desktop migration: the home-manager rust here ships without the clippy
component. Prefer an installed cargo-clippy; otherwise run clippy from an
ephemeral nix-shell with a separate target dir (mixed-compiler artifacts
are an E0514 hard error). MSRV keeps the shell's older toolchain a
legitimate gate.
2026-07-24 09:12:12 +02:00
Markk116 8c764e9169 docs(monitor): user-facing rewrite of process monitors
Lead with the user's problem (learn when another actor dies without
it knowing you're watching), explain one-directional/one-shot
semantics and contrast briefly with link without assuming link.rs has
been read. Add a compiling doctest. Drop em-dashes. Correctness facts
about registration/death races and demonitor-after-fire safety kept,
reworded in plain terms and separated from the public item docs.
2026-07-24 08:44:56 +02:00
Markk116 41b9d6d056 docs(introspect): user-facing rewrite of runtime introspection
Was the worst offender for external-context references (RFC 016
Chunk 1/4, DECISION D1/D2, RFC 003/011), all removed. Lead with the
practical use cases (debugging, health checks, test assertions,
dashboards) for snapshot()/actor_info()/tree(), and explain the
per-actor-reads-not-a-world-freeze consistency model in plain terms
instead of citing a decision log. Add a compiling doctest.
2026-07-24 08:44:56 +02:00
Markk116 dd845f22fe docs(mutex): user-facing rewrite of the actor-blocking mutex
Every public item was previously undocumented. Lead with why Mutex<T>
exists (a channel/gen_server is overkill for plain shared state) and
how it differs from std::sync::Mutex (parks the actor not the OS
thread, every lock is timeout-bounded by default). Add a compiling
doctest. Document new/lock/lock_timeout/try_lock/set_default_timeout/
MutexGuard/LockTimeout/DEFAULT_TIMEOUT. Drop em-dashes; keep wake
protocol mechanics as contributor-facing comments on private internals.
2026-07-24 08:44:56 +02:00
Markk116 36a0a9832d docs(registry): user-facing rewrite of the name registry
Replace the 'what changed' diff-against-a-prior-design framing with a
plain explanation of what the registry is for (naming an actor so
others can find and message it by name) and a compiling doctest
(register/whereis/send/unregister). Cut all RFC/decision-number/bug-id
references and em-dashes; move type-erasure and locking-discipline
detail into an Implementation notes section for contributors.
2026-07-24 08:40:26 +02:00
Markk116 feda6517e5 docs(channel): user-facing rewrite of the MPSC channel primitive
Lead with what a channel is and how to use it (compiling doctest for
channel()/send/recv/close), before any internal rationale. Document
every previously-undocumented public item (channel(), Sender, Receiver,
SendError, RecvError). Move the RawMutex-vs-std::sync::Mutex rationale
and lock-class discipline into an Implementation notes section. Drop
em-dashes throughout.
2026-07-24 08:40:26 +02:00
Markk116 8625ae4c35 docs(scheduler): user-facing rewrite of the actor/spawn/run entry point
Lead with what an actor is and how to start one with run()/spawn(),
following gen_server.rs's example-first style. Add a compiling module
doctest. Drop RFC references and em-dashes; keep internal mechanics
(preemption gating, thread-local borrow rules) as plain contributor
comments rather than public-facing doc prose.
2026-07-24 08:40:26 +02:00
Claude (sandbox) d9addeba5e test(causal): controller test no longer races the sweep's site snapshot
run_experiments snapshots the site registry once at entry. stage-x is
registered lazily (first causal_site! execution in the worker), so on a
1-core box the snapshot deterministically wins whenever a sibling test
has already paid the tsc_hz calibration — stage-x missed the sweep and
the summary assert tripped (silently, pre-propagation; the earlier
cascade attribution was incomplete). The worker now signals after its
first site entry and the test waits on it before starting the sweep.
Jarred separately: the lazy-registration trap is a lib UX hazard worth a
doc note or warm-up guidance.
2026-07-18 21:59:21 +00:00
Claude (sandbox) d5a3ba1934 fix(causal): born current — slot reuse booked phantom park forgiveness
A fresh/reused slot started with causal_delay=0 and causal_parked=true:
the first resume then 'forgave' the entire monotone global backlog, once
per spawn, booked as park forgiveness. Under close-mode conn churn (~95k
spawns/s) that is millions of phantom forgiven ms per 700ms window, even
in 0% cells — the books could not balance under actor churn while
ka-mode stayed plausible (few spawns). Impacts were unaffected: the
first resume always precedes the first check, so nobody ever spun.

reset_counters now installs Coz's new-thread rule: causal_delay starts
at the current global ledger and causal_parked starts false — a newborn
neither owes nor is forgiven the process's history, and delay injected
while it sits spawn-queued (runnable, not blocked) is owed and paid at
its first check, the semantics audit_zero_pct_window_absorbs_leftover_
debt pins. That test (born failing, unmasked by the run() propagation
fix) and the new actor_churn_between_experiments_forgives_nothing
regression test both go green.
2026-07-18 21:55:56 +00:00
Claude (sandbox) 7eae56a296 fix(runtime): a root actor panic escapes run()
The trampoline caught the root's panic, recorded it as Outcome::Panic on
the slot, and run() dropped the initial handle without reading it — every
assert inside run(), the standard test-suite pattern, was silently
vacuous (found live: a failing-first test passed). run() now reads the
root outcome before the handle drop and resume_unwinds the payload after
full teardown, so a caller's catch_unwind leaves the Runtime reusable;
Exit and Stopped return normally. The payload message is printed before
re-raising (the throw-site hook output was suppressed in-actor).

Correct the two tests this unmasked, both born failing and never run:
the select loser-arm test kept a closed arm in the set (the documented
closed-arm rule: a closed arm reports ready forever — observe the
disconnect and drop it); the send_after-to-dead test expected Ok(None)
from a closed+empty channel (documented: Err(RecvError), which proves
nothing-delivered even more strongly).
2026-07-18 21:52:50 +00:00
Claude (sandbox) 527f045e17 feat(causal): offcpu column + closed-books eff in the attrib probe
The probe now prints eff+offcpu next to eff — attributed plus the
counted runnable off-CPU gaps over ground-truth in-site time — the
per-window check that the located mechanism accounts for the whole
residual (~1.00 = books closed, no remaining silent loss). Audit line
gains the offcpu column, same delta-terms convention as the lib
renderer. Header doc rewritten from hypothesis to resolution.
2026-07-13 12:46:58 +00:00
Claude (sandbox) 0ee3fe7330 feat(causal): offcpu audit bucket — the @50 deficit located (RFC 007)
GPU sweep decomposed the ~23ms/700ms @50 injection deficit: every
ledger bucket is ~zero (drop park 0, discards 0, drop yield ~0.3ms),
books balance at absorbed+forgiven = 4x injected in all 48 cells, and
the 0%-cell contamination signature is absent. The probe pins the
residual: eff 0.933-0.943, a constant 22-27µs missing per site entry
= ~4.9 slice-expiry yields/entry x ~5.6µs runqueue wait. The
"deficit" is runnable off-CPU time inside the site — wall time the
probe's ground truth counts but on-CPU attribution correctly skips
(Coz model: speeding the site's code does not shrink queue-wait).

Measure-only reclassification, no behaviour change: a yield in the
target site stashes (tsc, experiment epoch) on the slot; the next
on_resume counts the gap into OFFCPU_IN_SITE_{CYCLES,N} (would-be
delta terms, MAX_SAMPLE_CYCLES-capped) iff the epoch still matches
and the word is live — a gap straddling end()/a same-word begin()
(live in the probe's 50,50 schedule) is dropped, never a leaked
cooldown. Parks excluded: blocked time is forgiveness territory.
New offcpu column in render_ledger_audit; LedgerCounters and
ExperimentResult grow the two fields. Fidelity footer now states
the on-CPU basis (deliberate wording change to the pinned summary;
the substring pin test still holds). +2 tests (counted gap; epoch
straddle) + render assert.
2026-07-13 12:46:15 +00:00
Claude (sandbox) 9bfeb2c6a2 feat(causal): ledger-audit output in the pipeline demo and attrib probe
- causal_pipeline: SMARM_CAUSAL_AUDIT=1 appends render_ledger_audit()
  after the summary; pinned summary format untouched.
- causal_attrib_probe: per-pct audit line (absorbed/forgiven/drops/
  discards) under the existing eff line, so in-site-vs-attributed and
  the loss buckets land in one place for the sweep.

1-core smoke (work + wide): books balance — absorbed = 2x injected and
forgiven = 2x injected, i.e. owed = injected x (N-1) with N=5 actors,
zero outstanding. drop park = 0 even in wide mode (the bottleneck's
queue is never empty, so its in-guard recv never parks); drop yield is
~1500 events but ~0.1ms per window, confirming slice-expiry yields
sample at their own checkpoint. Deficit decomposition needs the
parallel box.
2026-07-13 12:11:24 +00:00
Claude (sandbox) a3be8f0977 feat(causal): ledger audit — decompose the @50 injection deficit (RFC 007)
Measure-only counters for the deficit hunt (~23ms short per 700ms window
at 50% on the bottleneck site; superlinear vs 25%). Nothing here changes
injection or absorption; the sweep decides the fix.

Buckets, windowed per cell into new ExperimentResult fields (audit
snapshot taken at end() — spin/attribution freeze there, forgiveness
does not):
- spin_absorbed / park_forgiven: where owed delay actually went. Spin
  during a 0% cell is the baseline-contamination signature — leftover
  debt from a prior window being paid in a later one (checks gate on
  the experiment word, so cooldowns pay nothing and debt carries over).
- drop_park / drop_yield (+counts): the deschedule path flushes no
  sample tail — an in-target-site park or yield silently loses
  [last sample -> now]; on_resume re-arms before the actor runs again.
  New on_deschedule hook in all three intent arms (real park; explicit/
  slice-expiry yield; requeued park counts as yield — it never blocked).
  Slice-expiry yields sample at the descheduling checkpoint, so a fat
  yield bucket points at explicit yield_now or requeued parks.
- discard_overmax (+count, in would-be delta terms so columns compare
  against injected_cycles) / discard_unarmed: the attribute() clamps,
  previously silent.

LedgerCounters + ledger_counters() expose cumulative totals (tests,
run-level prints); render_ledger_audit() is the per-cell companion to
render_summary, which stays byte-identical (pinned). ExperimentResult
now derives Default so literals survive future audit-field growth.

Tests: +7 (spin counted, forgiveness counted, in-site park drop, in-site
yield drop, overmax discard, 0%-window leftover absorption — synthesized
deterministically via inject_delay_cycles_for_test with no experiment
active — and the audit render). 22/22 causal.
2026-07-13 12:07:19 +00:00