Commit Graph
213 Commits
Author SHA1 Message Date
claude-asm-audit 69a52a5578 Merge branch 'rfc010-cluster': RFC 010 clustering (c1–c16 + Phase 6) onto the v0.7.0 + perf-audit tree
Both lines branched from ca1c983 (v0.6.1 + try_spawn). master carried
graceful shutdown, gen_server/gen_statem lifetime, v0.7.0 and the 16
perf-audit commits; rfc010-cluster carried clustering behind
`--features cluster`. Two resolutions beyond the automatic merge:

- tests/channel.rs: both sides deflaked the spawn-then-monitor race in
  channel_ops_interleaved_with_monitor_churn_multi_thread. Kept master's
  `spawn_monitor` (monitor registered before publish) over the cluster
  side's `go`-gated spawn; same intent, API-level fix.
- src/cluster/envelope.rs: master added `DownReason::Shutdown`, which
  made the `Frame::Down` reason codec non-exhaustive. New wire tag
  DR_SHUTDOWN = 6, encoded and decoded symmetrically. `Shutdown` never
  rides in a `Down` by contract (a target that honours the request
  exits normally); the tag exists so the codec stays total. Tags 1–5
  unchanged.

Gates on the merged tree (rustc 1.98.1): default 405/0, cluster 492/0
(0 ignored beyond the 11 pre-existing `ignore` doctests), clippy --lib
-D warnings on both configs, doctests. cluster_disconnect still ~1.8s
(SMARM_FAST_TIMING plumbing intact).
2026-09-12 05:25:50 +00:00
claude-asm-audit 8fd724a6b6 bench: regenerate baseline.json on the reconciled tree (job 3a1af71f)
Rebased onto upstream v0.7.0, with the Weak-per-park fix (93bc83a) and the
root-sweep lock-free guard (3b0e06e) both in. 20-core box, rq-mpmc default,
wake slot on. Key order kept, so the diff is values only.

Not a flat baseline, and the regress pass in outputs/f21_job_3a1af71f.log
records why. Against 09ad759 the short 1-thread sections carry ~30-45 µs of
residual fixed cost (the root-exit sweep still walks all 16_384 slot words even
though it no longer locks them), and ping_pong_steady is +25% 1T / +23% 20T with
only ~30 µs of that explained — ~28 ns/roundtrip on the channel park chain that
nothing in the upstream diff accounts for. Several 20-thread sections moved
10-18% alongside it. That is finding 20c, open, and regenerating here means
`regress` can no longer see it: the comparison lives in history.md and the job
log instead.

switch_cost 126-128 cyc/roundtrip, unchanged, so the yield path is not involved.
2026-08-21 14:04:19 +00:00
claude-asm-audit 3b0e06ef13 perf(runtime): reject non-Live slots lock-free in the root-exit sweep
`shutdown_forest_roots` locked every slot's `cold` to discover `actor == None`.
The slab is `max_actors` entries (16_384 by default) and is almost entirely
vacant at root exit, so the sweep cost one uncontended mutex round-trip per
slot: a fixed ~152 µs per run on the 5900X, visible as deep_recursion 1T
289 -> 441 µs and the same +152 on chained_spawn 1T, fan_out_compute 1T and
deep_recursion 20T. `general.rs` times `init` and `run` together, so it landed
in every smarm bench number while no tokio section paid it.

`is_live_for` is a lock-free snapshot of the slot word; a non-Live slot has no
actor to shut down, and the ones that survive the check take the lock below as
before. The sweep is not weakened: the snapshot can go Live just after we pass
it, but that was already true of a spawn landing after the scan finished, and
there is no second sweep either way.

Suite green under reltest, including upstream's root_exit / shutdown /
supervisor_shutdown / root_sweep_trace (the last also under --features
smarm-trace, which is what asserts the sweep actually visits its targets).
2026-08-21 13:53:17 +00:00
claude-asm-audit 93bc83a5b3 perf(channel): capture the receiver's runtime Weak once per channel, not once per park
Upstream's off-runtime wake (1002777) captured a Weak<RuntimeInner> in the
parked_receiver tuple on every park, so every park/unpark round-trip paid an
Arc::downgrade plus the matching Weak drop — a locked RMW pair on one globally
shared counter, on the hot path of every channel workload.

The receiver never migrates between runtimes, so one capture is enough: the
Weak moves into Inner<T>, filled in on first park (a branch on an Option
thereafter), and unpark_at_via takes a closure so the fallback only re-locks
and clones off a scheduler thread, where the in-runtime fast path has already
declined. runtime_weak() returns Option and no longer panics off-runtime.

Sandbox 1T channel steady round-trip: 290 -> 270 ns (upstream's shape measured
427 -> 445 on its own base). Full suite green under reltest, including
upstream's tests/cross_thread_wake.rs.
2026-08-21 12:30:18 +00:00
claude-asm-audit a9341e2d82 fix(tls): actor-side thread-local accessors are #[inline(never)] + fence — LLVM caches TLS addresses across context switches (finding 19)
Without LTO (any downstream `cargo build --release`), LLVM keeps `%fs:0` in a
callee-saved register across `switch_to_scheduler`; after the actor migrates,
`ACTOR_DONE`/`PREEMPTION_ENABLED`/`CURRENT_PID` etc. hit the OLD thread's TLS.
8 multi-thread tests aborted with "scheduler resumed a done actor" (error 5).
Thin LTO only worked because it picked the local-exec model.

- context.rs: module doc "Thread-locals and migration" (the rule), tls_fence()
- every TLS accessor reachable from an actor stack is out of line:
  actor::{current_pid, publish_outcome (new)}, preempt::{maybe_preempt,
  check_cancelled, current_slot_ptr, note_*, preemption_swap/enabled (new;
  replaces direct PREEMPTION_ENABLED.with in NoPreempt/RawMutex/with_runtime/
  trace/debug asserts)}, context::{get,set}_scheduler_sp, runtime::{sched_slot,
  slot_push, set_yield_intent}, causal, trace
- raw_mutex order checks: out of line only under debug_assertions
- Cargo.toml: [profile.reltest] = release without LTO; 41 test bins link in
  ~1.5 s each instead of ~12 s (suite 11 min -> 1.5 min) AND it is the
  regression oracle for this bug: `cargo test --profile reltest`
Both profiles: 41/41 + doctests green.
2026-08-21 12:23:46 +00:00
claude-asm-audit d300a9d536 docs: fix 10 rustdoc link warnings, deny broken/private/redundant intra-doc links, doctests deny(warnings) (3 doctests had unused vars/imports) 2026-08-21 12:23:46 +00:00
claude-asm-audit bc0a5e8656 scheduler: spawn_monitor / spawn_monitor_with — monitor registered on the child's slot before publish, so the Down always carries the real reason (spawn-then-monitor could race to NoProc); tests; channel test uses it 2026-08-21 12:23:46 +00:00
claude-asm-audit ea2b222cff runtime: wake_slot default ON (RFC 005 accepted, finding 18); supervisor stops survivors sequentially so reverse-order teardown is guaranteed; deflake gen_server/channel tests; baseline.json regenerated slot-on; ROADMAP notes 2026-08-21 12:23:33 +00:00
claude-asm-audit 7001f04b65 diag(runtime): wake-path counters in SchedulerStats + RuntimeStats::wake_diag; RQDIAG/DIAG lines in rq_runtime and general (target 5, finding 17) 2026-08-21 12:22:46 +00:00
claude-asm-audit 77938cd31d bench: split ping_pong_oneshot — spawn_pair_control + ping_pong_steady; refresh baseline (job e982b5b5)
general.rs sections 5/6: spawn_pair_control is ping_pong_oneshot with the
messages removed (2 spawn + 2 join per round); ping_pong_steady is one
persistent pair × 10k roundtrips over unbounded MPSC (smarm::channel vs
tokio::sync::mpsc::unbounded_channel). Box (5900X, 3729fff, rq-mpmc):
oneshot 804/control 633 → 79% spawn+join; steady 142 ns vs tokio 128
(0.90×) at 1T, 1.8 µs/roundtrip at 20T. history.md finding 16.

baseline.json regenerated from the same run (20T labels, 14 benches).
2026-08-21 12:22:46 +00:00
claude-asm-audit 5ddd122711 perf(preempt): rdtsc unserialised by default; causal attribution opts into lfence
reset_timeslice paid an lfence pipeline drain on every resume via the
shared rdtsc() helper. Nothing in preempt.rs needs it: the timeslice
arm/expiry compare against a ~1e5-cycle slice, and an early stamp only
makes the slice look more used. The consumer that does need it — causal
site attribution, where a speculative early read misattributes a site's
tail — now calls rdtsc_serialising() explicitly (cold_check sample,
SiteGuard enter/exit).

Measured vs 31dc26a (baseline commit), 24-core box, rq-mpmc default:
- switch_cost, taskset -c 2, 9 interleaved old/new pairs: mean_cyc
  149 -> 117 per roundtrip (-32, -21%); mean_ns 43.2 -> 34.9.
- sweep.py regress + run, 20T, two runs agree:
  yield_in_hot_loop 1T 40171 -> 30277/30497 us (-24%)
  yield_many        1T 12513 ->  9965/10013 us (-20%)
  ping_pong_oneshot unchanged (RFC 005 wake-slot handoffs skip
  reset_timeslice, so that path never paid the lfence).
  All other rows within the box's ~+-15% noise floor.
- 1-core sandbox: 218 -> 206 cyc (understates by ~3x).
- bench binary lfence count 24 -> 4 (survivors = the bench's own
  rdtscp;lfence bracket). Tests green.
Baseline (benches/baseline.json) not re-saved in this commit.
2026-08-21 12:22:46 +00:00
claude-asm-audit 343e53e17b bench: baseline under the fixed rq-mpmc default (job 6e71e9b1)
20-core box (taskset 0-19), 5 sets, at 0438f12. Replaces the rq-mutex
baseline (df41ff3) so sweep.py regress compares like with like. The
multi-thread rows show what the mutex contention was costing:
yield_many 20T 156.7ms -> 44.1ms (-72%), spawn_storm_busy 20T -72%,
chained_spawn/catch_unwind 20T -20%. Sweep ran clean (0 panics) —
the finding-13 fix holding under the full suite.
Note: mpsc_contention 1T is ~2x the mutex-default row (3155 -> 6123);
backend characteristic to investigate, not a fix regression (the
shootout shows the fix itself is noise-neutral).
2026-08-21 12:22:46 +00:00
claude-asm-audit 9b215573de fix(run_queue): make the Vyukov rings preemption-tolerant (finding 13)
A consumer OS-preempted between its dequeue_pos claim and its seq
release freezes one cell; once traffic laps the ring (~cap ops ≈ 1-2ms
at yield-storm throughput ≈ one scheduling quantum under load),
try_push reads the stale seq and the original algorithm's 'lap behind
=> full' inference misfires. The old push assert then converted that
liveness stall into an abort blaming a double enqueue that never
happened (soak: occupancy 174-181 of cap 16384 at every failure), and
the dead scheduler threads stranded actors => the observed hangs.
Pristine 5504ef3 failed 14/14 under an 8-spinner soak on the 20-core
box; a retry prototype passed 13/14, its one failure being a stall
that outlived a fixed 1M-spin bound — pure spinning starves the
descheduled consumer, so the wait must yield.

Fix, following crossbeam ArrayQueue's shape (fence + opposite-counter
check; cells hand off, COUNTERS give verdicts):
- push: on a lap-behind cell, fence(SeqCst) + occupancy check.
  occ < cap => transient stall => spin-then-yield backoff and retry;
  occ >= cap => the REAL at-most-once-enqueued violation => panic with
  a truthful message and the counters. enqueue_pos is loaded before
  dequeue_pos so racing pops only underestimate occupancy (no spurious
  panic).
- pop: symmetric counter check before an empty verdict; on a
  mid-publish producer, bounded wait then None — deliberate deviation
  from crossbeam's unbounded retry (spurious None is benign: RFC 018's
  enqueue-wake self-heals; StripedRing's probe must not hang on one
  stripe).
- StripedRing::push probe: yield-escalating backoff after a full
  refused lap (was a bare spin_loop).
- Hand-rolled Backoff (spin 2^n to 64, then yield_now); under loom
  every wait is a yield so models explore the stalled peer's progress.

New loom model mpmc_lap_onto_stalled_consumer_completes reproduces the
old panic in the first explored interleavings (verified FAILED against
5504ef3) and passes with the fix. Lib 60 + integration 41 + all 4 ring
loom models pass. rq-striped inherits the fix (stripes are MpmcRings).
2026-08-21 12:22:46 +00:00
claude-asm-audit bf55cef3e3 perf(run_queue): flip default backend rq-mutex -> rq-mpmc
Evidence (history.md session 5, findings 10-11, 20-core jobrunner run):
rq-mutex collapses with thread count on queue-heavy load (yield-storm
6.9x slower than striped at 20T; attributed cause of the baseline's
multi-thread regression on yield_many/chained_spawn), while rq-mpmc is
best-or-close everywhere: best 1-thread, best slot-on ping-pong
(1237us vs mutex 6754us at 20T), -23%/-41 cyc per yield roundtrip on
real hardware (single-core switch_cost, interleaved). rq-striped
remains the churn-heavy many-core option, selectable per build.

Docs + compile_error hints in run_queue.rs updated to name rq-mpmc as
the default. slot_state.rs untouched, no loom-relevant changes; lib
(60) + scheduler/channel/supervisor/park_wake/wake_slot/preempt (41)
pass under the new default.
2026-08-21 12:22:46 +00:00
claude-asm-audit 2f88264426 bench: refresh baseline.json — 20-core box, fe85197, rq-mutex
Re-measured via jobrunner (taskset -c 0-19 of 24, rust:1.97-slim,
sweep.py run --save-baseline, 5 sets). Supersedes the old 24-thread
baseline; multi-thread labels are now 'smarm 20-thread'. Taken under
the rq-mutex default *before* the backend flip — the yield_many /
chained_spawn multi-thread regression reproduces here and is attributed
to rq-mutex contention (history.md session 5, finding 10).
2026-08-21 12:22:45 +00:00
claude-asm-audit 6172f4231d perf(runtime): skip take_closure's locked swap after first resume
Every resume paid an unconditional AtomicPtr::swap (lock xchg, full
barrier) to check for a first-resume closure that is null on all
resumes after the first. A Relaxed null-load fast path is sound:
store_closure runs only before publish_queued, whose Release pairing
with try_claim's Acquire orders it before this call, so no writer can
race the load within an occupancy.

Measured on switch_cost (1-core sandbox, rq-mutex, cycles): mean
roundtrip 350-355 -> 324-328, ~7.5%. All lib + scheduler/channel/
supervisor tests pass.
2026-08-21 12:22:45 +00:00
claude-asm-audit d7082eb266 perf(context): pass actor sp through registers, not TLS
switch_to_actor takes the target sp in rdi and returns the actor's
next saved sp in rax (handed over by switch_to_scheduler's shim).
Deletes the ACTOR_SP thread-local and halves the helper calls per
one-way switch (2 -> 1); the scheduler loop also drops its
set_actor_sp/get_actor_sp TLS round-trips. SCHEDULER_SP stays: a
yielding actor at arbitrary call depth has no argument channel back.

asm before/after in outputs/history.md session 1-2. Tests: 354 pass.
2026-08-21 12:22:45 +00:00
Claude (sandbox) 741c10337b release: v0.7.0
Bump crate version to 0.7.0.
v0.7.0
2026-08-21 13:10:25 +02:00
Claude (sandbox) e570138da5 docs(roadmap): supervisor start order is not start readiness
Filed from the urus v0.3 endpoint work. start_child spawns and moves on,
so a later sibling can whereis an earlier named child before that child's
actor has run. Notes why blocking spawn is not the fix ('has begun
executing' != 'has bound its name', plus a per-accept round-trip tax and
every spawn becoming a context-switch point), that OTP has the same async
spawn and synchronises one level up in gen_server:start_link, the
readiness-ack shape if scheduled, and the structural workaround urus uses
today (registrar spawns its own consumers).
2026-08-20 13:20:42 +00:00
Claude (sandbox) 415effb2e9 feat(gen_server,gen_statem): lifetime is the actor's — refs are addresses; inline named run
Root cause behind the "pin the endpoint" gotcha and the trapping-wrapper
pattern: a gen_server had two lifetime authorities — its refs (last one
dropped → inbox closes → exit) and, when supervised, its supervisor. OTP has
one: a process lives until it stops, is shut down, or is killed; a pid is an
address. Root exit now shutting down every forest root removes the reason the
ref-governed idiom existed (a forgotten server no longer hangs the run), so
adopt the one rule:

- The server/machine loop holds one inbox sender for its life; the inbox
  never closes. GenServerRef / GenStatemRef are addresses. Explicit close is
  `shutdown()`; a forgotten one is swept at root exit.
- `NamedGenServerBuilder::run()` / `gen_statem::run_named(name, m)` run the
  loop inline as the current actor: a server is a direct ChildSpec child,
  gets the supervisor's shutdown as handle_shutdown / a shutdown row, re-binds
  its name on restart, and is addressed by name. The wrapper in
  examples/graceful_shutdown.rs is gone.
- gen_statem gains GenStatemName + whereis_machine/send/call/shutdown by name
  (parity with gen_server); the macro gets `Sm::new`.
- Root-exit sweep records `Event::RootSweep { target, trapping }` under
  smarm-trace ("root_sweep shutdown|stopped"): unsupervised leftovers are
  visible rather than silently owned-by-refs.
- Named start() name-clash path stops the spawned actor instead of relying
  on ref drop.

Tests: tests/gen_server_lifetime.rs, tests/gen_statem_lifetime.rs,
tests/root_sweep_trace.rs (feature-gated); three existing tests that used
drop-closes-inbox now use shutdown(). Docs/README/ROADMAP/Deep Dive updated.
2026-08-19 17:47:31 +00:00
Claude (sandbox) 849a424c8e docs,examples: graceful shutdown — new examples/graceful_shutdown.rs, README 'Stopping actors', named_genserver uses shutdown(), Deep Dive terminate note, ROADMAP open items 2026-08-19 16:28:42 +00:00
Claude (sandbox) 6ceb138f5f feat(gen_statem): graceful-shutdown parity — trap_exit, shutdown/exit rows, stop, terminate
Mirrors the gen_server surface in gen_statem's event model:
- Cx::trap_exit() (in the initial enter): a shutdown request then arrives
  as the Shutdown event, routed by state through `shutdown` rows (default
  for a state with no row: stop); linked-peer deaths as `exit <pat>` rows
  (default: drop). Non-trapping machines are stopped outright, as before.
- Cx::stop(): normal self-exit after the current event; `stop` tail keyword
  is sugar for { cx.stop(); prev }.
- Machine::terminate (optional `terminate { … }` macro block), run from a
  Drop guard on every exit path; the guard also drains armed timers.
- Machine::shutdown_ev / exit_ev (defaults None) so hand-written machines
  keep compiling; GenStatemRef::shutdown() is graceful and waits.
- Loop selects exits > timers > inbox.

Tests: tests/gen_statem_shutdown.rs.
2026-08-19 16:25:15 +00:00
Claude (sandbox) 250f31265b feat(runtime): root exit is graceful shutdown of the forest roots
The RFC 014 root-exit sweep hard-stopped every live slot once nothing was
runnable. That deferral privileged queued work over parked-with-a-pending-
wake work (a sleeper was killed, a queued cast was drained) and any attempt
to widen the notion of pending wake (timers, fd readiness) re-wedges the
run on the periodic-timer daemon the sweep exists to end.

Root exit now means "the program is done": finalize_actor delivers
request_shutdown to every forest root — each live actor whose parent is
the run (ROOT_PID) or is dead — synchronously, before the live-count
decrement. Supervisors cascade per child Shutdown policy; trapping actors
may Continue/drain with working timers and end the run when they stop
themselves; non-trapping actors are stopped outright. No forcing sweep.

Removes root_exited/root_swept, Pop::RootDrain and the idle-verdict
condition; adds tests/root_exit.rs.
2026-08-19 07:11:46 +00:00
claude 9c8f59ca53 feat(scheduler,supervisor,gen_server): graceful shutdown — request_shutdown, child Shutdown policy, handle_shutdown
Lift OTP's `exit(Pid, shutdown)` + child-spec `shutdown` wholesale.

scheduler / runtime
- `request_shutdown(pid)`: the polite stop. A target trapping exits gets an
  `ExitSignal { reason: DownReason::Shutdown }` on its trap inbox and keeps
  running; a non-trapping target is stopped as by `request_stop`, which is
  now documented as the hard stop (`exit(Pid, kill)`). Dead pid: no-op.
- `RuntimeHandle::request_shutdown` for the off-runtime (signal thread) path;
  `from == ROOT_PID` there.
- `DownReason::Shutdown` — appears only in ExitSignal, never in Down (a
  complying target exits *normally*).

supervisor
- `ChildSpec::shutdown(Shutdown::{BrutalKill, Timeout(d), Infinity})`,
  default Timeout(5s). Every supervisor-initiated stop (ordered shutdown and
  OneForAll/RestForOne sibling cycling) is: request_shutdown → await the
  child's Signal up to the grace → request_stop → await. Sequential, reverse
  start order.
- The supervisor traps exits; a Shutdown ExitSignal runs the ordered
  shutdown and `run()` returns normally, so `request_shutdown(root_sup)`
  tears a whole tree down top-down with each child's grace period.
- FIX: a hard `request_stop` on a supervisor previously orphaned its
  children (the ordered shutdown lived after the loop, and the unwind
  skipped it). `Live` (the by_pid map) now carries a drop guard that
  fire-and-forget hard-stops live children when unwinding.

gen_server
- `GenServerCtx::trap_exit()` opt-in in `init`; the trap inbox becomes arm 0
  of the loop's select. Shutdown ExitSignal → `handle_shutdown() ->
  ShutdownAction::{Exit, Continue}` (default Exit: loop breaks, `terminate`
  runs on the normal path and may block). Other ExitSignals →
  `handle_exit(sig)`.
- `GenServerCtx::stop_handle() -> StopHandle`, `stop()` ends the server
  after the current message with a *normal* exit — the missing
  `{stop, normal, State}`; `request_stop(self_pid())` was the only self-exit
  and it is abnormal (Transient restarts it).
- `GenServerRef::shutdown()` / `gen_server::shutdown(name)` now go through
  `request_shutdown`.

Tests: tests/shutdown.rs, tests/supervisor_shutdown.rs,
tests/gen_server_shutdown.rs. Full suite green; fmt + clippy --lib clean.
2026-08-19 06:34:48 +00:00
Claude (sandbox) 1002777ef3 feat(channel,runtime): off-runtime cross-thread wake for parked receivers and stop
Wakes issued from a non-scheduler OS thread were silent no-ops. Every
off-runtime wake primitive (unpark, unpark_at, request_stop) reaches the
runtime through the RUNTIME thread-local, which is unset on any foreign
thread — so a send from a plain std::thread enqueued its message but never
woke the parked receiver, and there was no way to drive a stop into a
runtime from an application thread (e.g. an OS-signal handler). The former
strands a parked recv forever; the latter is why a downstream server must
poll a shutdown flag instead of parking on it.

Generalize RFC 018's rule — a producer reaches the runtime through a Weak it
holds — from the IO backend to channel senders and to a new handle:

- A receiver captures a Weak<RuntimeInner> (provably live at that moment)
  alongside its (pid, epoch) when it parks. send() and the last-sender drop
  wake through scheduler::unpark_at_via, which takes the thread-local path
  when on a scheduler thread (preemption-gated, slot-eligible) and the
  captured Weak otherwise — the same cross-context wake the IO threads do.
- Runtime::handle() returns a Send + Sync RuntimeHandle carrying that Weak;
  RuntimeHandle::request_stop drives a cooperative stop from any thread and
  is a no-op once the runtime is dropped.

The in-runtime wake paths (recv/select timers) are unchanged; only the sites
reachable from a foreign thread route through the Weak. RuntimeHandle exposes
request_stop only: send-wake needs no user-facing handle, and off-runtime
unpark is covered because request_stop drives unpark on the upgraded inner.

tests/cross_thread_wake.rs: a foreign-thread send wakes a parked receiver; a
foreign-thread request_stop wakes and stops a parked actor; a RuntimeHandle
held across and beyond run never blocks all-done and degrades to a no-op.
2026-08-19 05:50:54 +00:00
claude-asm-audit 4c0e42152f feat(cluster): RFC 010 c10–c16, follow-ups and Phase 6 (squash of 16ef583..d9c62a8)
Tree snapshot of d9c62a8 (2026-08-18). The 20 source commits between
16ef583 (c9) and d9c62a8 were never pushed and the clone that held them
was lost; this commit carries their combined tree verbatim so the build
history stays auditable from the c1–c9 commits below it. Original
hashes as recorded in the session handoff:

  c10  f03e94d  pid targeting + auto-serialization (RemotePid, D14 name
                on the wire); Phase 3 gate
  c11  7ef4bad  DownReason::Disconnected, wire tag 5
  c12  d124162  remote monitors (Monitor/Demonitor/Down frames)
  c13  9de967b  connection-loss synthesis (A+B: Monitors::teardown +
                unread-command Disconnected); Phase 4 gate
  c14  7e822b7  eager pg eviction (reaper actor, ReaperInboxes)
       dbe1a22  InboundVerdict::label(), trace::Event::ClusterInbound
       31a9877  tests/channel.rs monitor-churn target gated on `go`
       653559e  Discovery::Withdrawn{name, addr}
  c15  b41d76e  distributed pg: Sync on NodeUp, Join/Leave broadcast,
                NodeDown sweep, members_all; PgMsg wire type
  c16  fafa881  pick_any / dispatch_any; Phase 5 complete
  Phase 6 Tier A:
       195c73e  p4  NodeEvent::NodeDown(NodeInfo)
       48fd766  p1  connector Candidate{name, addr, state}
       ce8cf99  p2+p7 conn.rs select arms as Vec<Arm>; Outbound::Drained
       bf24988  p6  RemotePid::from_local -> Option
       9ae0380  p3  PeerStanding{Free, Claimed, Dialing}
       c7d62a1  p11 cluster::Timing knobs, threaded by value
       46f171d  p11 cluster_disconnect un-ignored on SMARM_FAST_TIMING
  Phase 6 Tier B:
       391a9ae  p5  cluster::RemoteDownReason{Local, Disconnected};
                    DownReason::Disconnected removed from core
       7ddd908  p9  pg ctl channel unconditional, one cfg seam at spawn
       d9c62a8      PeerNameMismatch parks the candidate; ClusterDial trace

Verified at d9c62a8: default 361/0, cluster 448/0, clippy --lib on
default / cluster / cluster+smarm-trace, fmt, 10x flake on
cluster_dial_mismatch, 5x on cluster_pg.
2026-08-18 16:00:00 +02:00
smarm 8f2d513940 README: rewrite overview, add limitations, roadmap, and contribution notes
Expands the intro into an overview/limitations structure, documents
preemption, tracing/causal profiling, URUS, and adds a 'Coming up' and
'A note on open source' section.
2026-08-18 00:25:25 +02:00
Claude 16ef583455 feat(cluster): RFC 010 c9 — remote Name sends: outbound table + the one inbound seam
Outbound (D13, ratified 2026-08-15): a module-private table node → dedicated
Sender<Frame>, populated/torn down by the manager inside the same serialized
handlers that own connection lifetime (Register / Disconnect / reap /
terminate), on RuntimeInner beside the exposure state. remote::send is one
leaf-lock lookup + one channel send: no gen_server on the data plane (rejected
send-through-manager — worse than the c8 argument, it is the data plane), no
published channel a leaked conn pid could inject raw frames into (rejected
send_dyn-at-conn-pid — module privacy cannot fence a published channel).
Ok(()) = handed to the connection's inbox, local knowledge only (RFC §3);
missing entry / closed channel = NotConnected. The outbound sender is a
SEPARATE channel from cmd_tx on purpose: a clone would let the table hold
lifetime authority (D9 violation); its closure is not a stop signal.
Buffering toward a slow peer is unbounded (BEAM busy_dist_port shape);
backpressure out of scope, documented not silently absent.

Inbound: remote::deliver_named is THE resolution seam (RFC v2) — exposed-set
check (unexposed = unreachable, the safety) → hash check against what the
name was exposed with → registry whereis → c8's decode_deliver. Verdicts are
local-only (InboundVerdict, discarded by the conn actor for now; a trace
hook is the place). The conn actor gains a third select arm (outbound
inbox → wire) and interprets SendNamed; Send/Monitor/Down are consumed for
liveness and ignored until c10/c11.

Public surface: RemoteName<M> (node, Name<M>), RemoteSendError, NotConnected,
send, send_remote_raw (untyped escape hatch so tests can put deliberately
wrong frames on the wire — the typed API cannot express a hash mismatch).

CORE (found the hard way this chunk): publishing a second channel of the
same message type on one live actor silently replaced and CLOSED the first,
so a select/recv on it returned 'closed' immediately forever — a hot loop
starving the single-threaded scheduler. publish_channel now asserts when an
existing same-type channel's receiver is still alive AND the new sender is
not a clone of it (Sender::same_channel via Arc::ptr_eq); replacing a
dead-receiver channel stays silent (an actor re-registering after dropping
its inbox is legit). Message names the sanctioned shapes. Three registry
tests pin panic / cloned-sender-ok / dead-receiver-ok. Full default suite
clean.

Harness: wait_line drains stderr before dumping on timeout/EOF (eprintln!
diagnostics no longer vanish); Node::transcript() accessor for
ordering-proof assertions.

tests/cluster_remote_send.rs 1/0, 10/10 flake runs, subprocess harness:
cross-node name-send delivers; unexposed name unreachable (registered
locally, never delivered); wrong hash never misroutes (unknown-hash and
known-hash-wrong-channel flavours); send to an unconnected node =
NotConnected locally; every frame-bearing send is Ok — 'handed to
transport', asserted and documented at the test. Negatives are proven by
STREAM ORDERING (they precede the positive on one in-order connection),
not by sleeping. All cluster suites regression-clean (envelope 15,
handshake 11, transport 11, lifecycle 1, liveness 3, connect 9, two_node 3,
membership 4, mesh 2, expose 5); clippy --lib green both configs; fmt clean;
default build compiles.
2026-08-15 21:22:56 +00:00
Claude a7f98f8d48 feat(cluster): RFC 010 c8 — exposure registry + fixed-seed type hashing
Nothing local is remotely reachable by default (RFC §4). expose(Name<M>)
marks a name remotely addressable and registers M's decoder under
type_hash::<M>(); expose_type::<M>() registers only the decoder (the
reply-to path). exposed_names() is the auditable remote surface.

D3's watchable fold, resolved against the code as it stands (stated in the
module docs): register() ALREADY stamps every named holder watchable ('no
successfully-registered actor can die unflagged', registry.rs), so an
exposed name's holder needs no extra mark — and re-registration after a
holder's death re-stamps the new holder for free, which a per-tenancy mark
taken at expose time could not do. The cluster's own mark_watchable
set-site is therefore the pid crossing the wire (frame serialization, c10)
— the exact analog of the membrane crossing. c8 adds only the name/type
state neither the registry nor slot bits can carry. No new pid registry;
RFC §4 honored.

One-viable calls, flagged:
- State lives on RuntimeInner (the pg pattern: leaf RawMutex field,
  cfg-gated behind cluster, zero-cost-when-off per c1) — c9's inbound
  decode consults it per frame; manager-held state would serialize every
  remote delivery through one gen_server.
- type_hash = FNV-1a 64 (fixed seed: the offset basis) over TypeId: a
  constant of the binary — stable across runs of the same build (the scope
  the build-hash handshake reduces the mesh to), deliberately not across
  builds. Collisions degrade to decode error / refused channel, never a
  misroute (the NoChannel guarantee, RFC §3).
- Decoder = decode-and-deliver-to-pid Arc closure capturing M (the one
  typed site): decode_payload then send_dyn. Wire-name → pid resolution
  stays OUTSIDE — that is c9's single seam, which calls decode_deliver.
  Arc so the call happens with the exposure lock RELEASED: send_dyn takes
  the registry lock, a mutual Leaf (the runtime asserts on nesting — caught
  live by the first test run).
- expose is a name-level fact, valid for an unregistered name (names
  late-bind; c9 resolves per delivery).

tests/cluster_expose.rs 5/0 stable x5, purely local per roadmap:
exposed/unexposed lookup + audit listing; decoder registration and the
delivery contract (happy path into a registered String channel; unknown
hash; corrupt bytes; wrong channel refused — never misrouted); distinct
types distinct hashes; expose/bridge-crossing agreement via the shared
watchable observable (terminal_reason after holder death); hash stability
across runs in the same binary via a c4-harness re-exec. Payload types are
std types — the crate's serde is derive-less by design, user crates bring
their own derive.

All cluster suites regression-clean (envelope 15, handshake 11, transport
11, lifecycle 1, liveness 3, connect 9, two_node 3, membership 4, mesh 2);
clippy --lib green both configs; fmt clean; default build compiles.
2026-08-15 07:25:42 +00:00
Claude 1282c3a08d feat(cluster): RFC 010 c7b — discovery Strategy, static seeds, connector dial loop
Phase 2 gate: 3-node mesh under the subprocess harness, repeatable (10/10).

Strategy (ratified): push-based, spawned as its own actor by the connector —
it emits Discovery events into a channel whenever it learns something and
may run forever; the connector owns all retry/backoff state. StaticSeeds
announces its list once and exits. Discovery is #[non_exhaustive] and
additive-only (candidates announced, never withdrawn) so expiry can land
later without breaking strategies.

One-viable correction to the ratified Discovery shape, flagged: a candidate
is a (name, addr) PAIR, not a bare address. The dial path and the D7
tie-break are keyed by peer name (the dial intent must be registered before
connecting so a crossing inbound Hello sees it), so an anonymous dial would
reintroduce exactly the simultaneous-connect flap D7 exists to prevent.
Discovery mechanisms know names — that is what they discover.

Connector: plain select-loop actor (the c6 shape) folding cmd inbox,
discovery stream, membership stream, and the earliest retry deadline into
one wait. It tracks who is up by SUBSCRIBING TO MEMBERSHIP like any
consumer — first consumer of c7a's snapshot-then-stream surface, no
privileged channel into the manager. Backoff: 250ms doubling to a 5s cap
(the c6c class of one-viable constants), reset on node_up; node_down
schedules a prompt redial with a fresh sequence. A candidate bearing the
local name is parked (that seed is us); every other failure retries — in
particular NameTaken can be our own ghost at the peer, not yet reaped by
its liveness timer, so it must not park. Dials run inline in the loop, the
acceptor's deliberate serialization (each attempt bounded by the connect +
handshake deadlines).

cluster::start(Config {node_name, meta, listen_addr, strategy}) is now the
integrated node start: supervised manager + acceptor + connector. It
completes the node identity: build_hash = cluster::BUILD_HASH (first
consumer, closing the c6d loose end) and incarnation = self_incarnation()
— unix-epoch MILLIS truncated to u32, not seconds: a supervised
crash-and-restart inside one second is routine, and seconds would collide
the ghost with its successor. Cluster handle: local_addr()/local()/
shutdown(); drop stops acceptor+connector loops, manager subtree detaches
(same split as AcceptorHandle alone).

Roadmap-binding, asserted in review: no consumer touches the connection
table — Manager.conns and ConnEntry stay private; the only exposures are
Call::Peers (sorted names, pre-existing) and the membership surface.

tests/cluster_mesh.rs 2/0, 10/10 flake runs: (1) 3-node mesh forms; kill
one (SIGKILL via Drop, per the retractable-state trap: roles park forever)
=> node_down at both survivors; restart same name => new incarnation at
every observer, distinguishable from the ghost; (2) seed unreachable at
start (pre-reserved closed port; accepted micro steal-window, documented)
then arriving later => edge forms via the retry path. All cluster suites
regression-clean (envelope 15, handshake 11, transport 11, lifecycle 1,
liveness 3, connect 9, two_node 3, membership 4); clippy --lib green both
configs; fmt clean; default build compiles.
2026-08-15 07:14:13 +00:00
Claude 160967939b feat(cluster): RFC 010 c7a — membership events + view at the manager
node_up/node_down are derived facts of the manager's own register/remove
events, so the membership state lives in the manager (no cross-actor race
between 'connection exists' and 'node is up'); src/cluster/membership.rs is
the consumer surface: NodeEvent/NodeInfo, subscribe(), view(). The conn
table stays private — no consumer touches it (roadmap-binding).

Ratified semantics: subscribe() is snapshot-then-stream — one NodeUp per
live peer is queued before the subscription joins the list, exact because
gen_server handlers are serialized. Dropped subscribers are pruned on the
next emit (closed channel), no monitor needed.

One-viable call, flagged: NodeId is memoized per (name, incarnation) — a
compact local alias for the wire identity, per pg.rs's framing. A reconnect
blip at the same incarnation keeps its id; a restart (new incarnation) gets
a fresh one, so a ghost and its successor are always distinguishable.
Allocation starts at 1; NodeId(0) stays pg::DEFAULT_NODE_ID (self).

Call::Register now carries the whole handshake Peer (the path already has
it; node_up needs incarnation + meta).

tests/cluster_membership.rs 4/0 stable over 5 runs (live up/down over
localhost TCP with commanded and EOF teardown; late-subscriber snapshot +
view agreement; restart-vs-blip id identity; dead-subscriber pruning). All
cluster suites regression-clean; clippy --lib green both configs; fmt
clean; default build compiles.
2026-08-15 07:07:14 +00:00
Claude 112d6b2e65 feat(cluster): RFC 010 c6d — build_hash derivation
cluster::BUILD_HASH, the derived value for LocalNode::build_hash: a
compile-time const, FNV-1a 64 over the build script's input string
(rustc -V + sorted enabled CARGO_FEATURE_* set) with PROTO_VERSION folded
in as a continuation of the same state — a proto bump moves the hash even
on an identical toolchain. build.rs emits only the raw inputs
(SMARM_BUILD_HASH_INPUTS); the hashing lives in src/cluster.rs next to
PROTO_VERSION rather than build.rs parsing it out of a source file. The
env var is emitted unconditionally — one string costs the default build
nothing, and the const itself is behind the cluster feature with the rest
of the module.

Domain is the flagged lean (one-viable-answer, veto at diff review):
toolchain + declared features + proto version. Tightenable later without a
wire change — it is just a u64 on the Hello.

Unit tests: FNV-1a 64 published vectors (empty/'a'/'foobar'); the proto
fold is a continuation of the same FNV state and moves the hash;
BUILD_HASH is const-evaluable and nonzero. All cluster suites
regression-clean (7 files); clippy --lib green both configs; fmt clean;
default build compiles.
2026-08-15 06:37:57 +00:00
Claude ad4958421f feat(cluster): RFC 010 c6c — heartbeat send + fixed-timeout liveness + teardown
The timeout arm of the connection actor's select: HEARTBEAT_INTERVAL (1s)
paces outbound Frame::Heartbeat (first at spawn, so the peer's window
starts fed) and LIVENESS_TIMEOUT (4s = 4 intervals) declares the peer dead
when no inbound frame arrives inside it — any frame resets the window, so
heartbeats keep an idle connection alive and real traffic (c8+) counts for
free. Fire => close + exit; the manager's monitor reaps the table entry as
on every other exit path. Fixed timeout per RFC v2 §5 (control connection,
heartbeats can't queue behind bulk). Intervals are the one-viable-answer
call flagged for veto at diff review.

The pump was made non-blocking to keep the deadlines honest: a plain recv()
blocks into the socket while the buffer holds a partial frame, parking the
actor past its timers. Two additive FramedConn methods (read_once,
next_buffered): exactly one socket read per level-triggered readable wake
(cannot block, cannot strand — leftovers re-signal), then drain every
complete buffered frame. Liveness resets only on complete frames.

No-fd transports (loopback) still get the command-only loop: no readiness
means no timers, same caveat as recv_deadline.

tests/cluster_conn_liveness.rs 3/0, stable over 5 runs (raw far end over
localhost TCP: heartbeats appear unprompted; mute peer still up at half
the window, gone after it; heartbeat-only peer survives 1.5x the window,
then reaped once silenced). Cluster suites regression-clean; clippy --lib
green both configs; fmt clean.
2026-08-15 06:36:07 +00:00
Claude c8ed858e4c feat(cluster): RFC 010 c6b — handshake on the accept/connect path
Drive the c5 machines as straight-line code on the path (D8): dial_handshake
and accept_handshake do the IO on a shared FramedConn, and a connection actor
is spawned only after a successful handshake. Rejects, tie-break losses (D7),
protocol faults and timeouts are all resolved on the path by closing, so no
actor ever exists for a connection that did not establish. The whole
FramedConn travels into spawn_established, carrying any read-ahead past the
handshake frames.

Handshake deadlines land here rather than in c6c: FramedConn::recv_deadline
enforces them between reads via the connection's fd arm, so a peer that
connects and goes silent cannot wedge the acceptor.

Connection lifetime moves to the manager (pulled forward from c7). The path
registers each established connection and hands over its ConnHandle; the
manager owns it, monitors the actor, and tears the connection down on
Disconnect, on peer close, or at manager shutdown. spawn_established returns
a Pid, so a connection neither outlives nor dies with whichever actor
established it — the ownership that made two-node teardown unorderable.

The manager also tracks in-flight dial intents, monitored so a panicking
dial cannot wedge the tie-break, and answers HelloCtx for the accept path.
2026-08-14 21:09:08 +00:00
Claude bbaaa062e3 feat(cluster): RFC 010 c6a — connection actor + manager subtree
Per-peer connection actor as a single select-loop plain actor owning the
whole FramedConn: one select folds its command inbox and the transport's
readable arm, so reads and control share one execution context — no reader
thread, no read/write split. The handshake is bypassed here (c6b wires it);
the actor is spawned already-established and self-registers with the manager.

Manager gen_server: the peer-name -> conn-pid registry and the uniqueness
source the handshake's NameTaken depends on. It monitors each connection, so
the table self-heals on any exit path. Explicit supervision subtree keeps the
manager up; connections are dynamic and monitored, never restarted (c7
re-dials).

Transport gains an additive Conn::readable_arm -> Option<FdArm> (default None;
TCP returns its fd's arm, loopback stays None). Existing c3 transport tests
unchanged.

Lifecycle test over localhost TCP: up reflected in the table, commanded
shutdown reaps exactly one, peer EOF reaps the other.
2026-08-14 19:23:39 +00:00
claude 9e49038474 feat(cluster): RFC 010 c5 — handshake as a pure state machine
Frames in, actions out — no IO, no clocks, no actors; the c6 connection
actor will drive it. Initiator (dial: emit Hello, interpret the single
response) and Responder (accept: judge the first frame) as consuming-self
machines; check order proto -> hash -> name -> tie-break. Driver-supplied
HelloCtx carries the two facts the pure machine cannot know (name claimed,
own dial in flight). Tie-break ratified as a wire fact: the smaller name's
dial survives; the losing inbound closes silently (both ends compute the
same verdict, no reject frame needed). Peer's own name offered => NameTaken.
build_hash is config-supplied; derivation lands with c6.
2026-08-14 17:17:14 +00:00
Claude 8a9e2b81b1 test(cluster): RFC 010 c4 — subprocess two-node harness
The runtime is a process singleton, so multi-node tests mean multiple
processes. tests/common/mod.rs is the reusable harness (precedent: RFC
019 c6 / tests/stack_diag.rs self-re-exec, extended to live tailing):
re-execs the current test binary as named roles, tails stdout/stderr on
reader threads, waits on protocol-visible lines with bounded timeouts
(panic dumps carry the full transcript), and Drop SIGKILLs+reaps so a
panicking test leaves no orphan or zombie. Children run with
--test-threads=1 --quiet --nocapture; the last flag is load-bearing —
libtest's capture would otherwise swallow role output.

Port assignment is race-free by construction: children bind port 0 and
announce the concrete address (LISTENING <addr>); the parent never
pre-picks.

Smoke suite per roadmap: two real nodes, handshake-less TCP connect
through the real framed codec (one Heartbeat across, clean close seen on
both sides, both exit 0), plus reap-on-drop proven via ESRCH and
nonzero-exit surfacing. Flake budget stated in the module doc: 10 s
bound per wait, 10/10 clean at authoring, >1/100 failures = regression.
2026-08-14 14:46:28 +00:00
Claude 39ab92871e feat(cluster): RFC 010 c3 — transport trait, framed codec, TCP + loopback impls
The control-connection abstraction (RFC v2 §5): object-safe Transport/
Listener/Conn over opaque pre-resolved addresses (resolution stays the c9
seam), with FramedConn as the single shared byte->Frame codec feeding
Frame::decode's incremental contract. Nothing forecloses additional
per-peer connections for the jarred bulk plane; the membrane is not a
transport (D2).

TCP parks the calling actor via scheduler fd readiness (MSG_NOSIGNAL
writes, EINPROGRESS dial resolved through SO_ERROR). Loopback is the
shipped in-memory test transport: OS-thread-blocking condvar pipes with
TCP-shaped close semantics, per-instance address registry.

Conformance suite runs the same codec over both impls: roundtrips both
directions, framing across split writes, coalesced frames, peer-close
mid-frame as TruncatedByPeer (not EOF), clean close as Ok(None). Plus
impl-specific establishment/error cases and a 4 MiB cross-buffer TCP
frame under real backpressure.
2026-08-14 14:31:39 +00:00
Claude 3850f6099b feat(cluster): RFC 010 c2 — owned wire envelope
Frame enum per the RFC inventory; hand-rolled encode/decode with u32 LE
length prefix + u8 tag; strings u16-prefixed, payload blobs u32-prefixed;
MAX_FRAME_LEN cap (control plane never carries bulk, §5). Streaming decode:
Ok(None) = need more bytes, every Err = corruption. postcard confined to
encode_payload/decode_payload — the single codec seam (§2); postcard gains
the alloc feature for to_allocvec (still no_std-aligned, no default
features). DownReason/RejectReason travel as single tag bytes; Down tag 5
is reserved for c11's Disconnected.

Tests: per-frame roundtrip, back-to-back frames, golden heartbeat bytes,
zero-length payload, every-prefix incomplete, unknown frame/enum tags,
length prefix lying long (with and without bytes present) and short,
truncation mid-string, adversarial lengths (u32::MAX, cap+1, zero), UTF-8
corruption, payload seam roundtrip through a real Send frame.
2026-08-14 13:23:50 +00:00
Claude 58a2fe3046 feat(cluster): RFC 010 c1 — cluster feature flag + optional serde/postcard deps
Off by default; the default build stays libc-only. serde (payload contract)
and postcard (payload codec) are optional, no default features. src/cluster.rs
is an intentionally empty cfg-gated stub so the flag's default-build
invariance is reviewable in isolation.
2026-08-14 13:20:20 +00:00
Claude (sandbox)andClaude (sandbox) ca1c98336e feat(scheduler,runtime): non-panicking try_spawn for at-capacity load shedding
allocate_slot() panics on a full slab; for a load-shedding caller (an
accept loop spawning one actor per connection) that panic lands in the
spawning actor, which then crash-loops under Restart::Transient into the
still-full slab until its restart budget is spent — and the service stops
accepting entirely. Observed live (urus slowloris scaling, 2026-08-10).
A full slab is a routine overload condition for such callers, not an
invariant violation.

- RuntimeInner::try_allocate_slot() -> Option<u32>: the non-panicking
  core; a single pop under the free-list lock, so the claim is atomic
  (claim-or-report — no check-then-spawn TOCTOU, no headroom margin).
  allocate_slot() is now a thin panicking wrapper over it.
- scheduler::try_spawn / try_spawn_under_with -> Result<JoinHandle,
  SpawnError>: parity with spawn/spawn_under_with except a full slab
  returns Err(SpawnError::AtCapacity) instead of panicking. Minimal
  surface per the agreed strategy; the remaining _with/_addr mirrors are
  trivial wrappers if ever needed.
- Slot-first ordering on the try path (reverse of spawn's stack-first):
  under overload Err is the hot path, and a rejection costs one mutex
  pop — no mmap/pool-pop + init + recycle per shed unit of work. A
  drop-guard returns the claimed slot if stack allocation panics in the
  claim-to-install window (would otherwise leak and trip run()'s
  teardown slot-leak debug_assert).
- SpawnError: non_exhaustive, Display + std::error::Error.
- spawn and every existing call site untouched: the panic remains the
  correct loud invariant check at internal/bounded spawn sites.

tests/try_spawn.rs: parity when slots free; exact slab accounting at
capacity (Err, no panic, repeatable); custom-shape try refuses before
stack allocation; self-heal after slots free; plain spawn still panics
(surfaced via JoinError payload); 4-thread race for the last slots
claims exactly the free count; SpawnError impl checks.

Design doc: smarm-suggestion-try-spawn.md. Downstream consumer change
(canned 503 on AtCapacity in urus's accept loop) is urus scope, not
smarm.

(cherry picked from commit 36de4b36aeaa72b2a5f9f3797b9854652656dcf6)
v0.6.1
2026-08-13 15:03:16 +02:00
smarm-agent 95306c7f60 style: cargo fmt sweep under rustc 1.97.1 (toolchain reformat, no semantic change) 2026-08-13 05:56:49 +00:00
smarm-agent 1262cc30e3 monitor: widen stamp eligibility to watchable = named ∪ exported (soak sig 5)
The terminal record existed for watches that raced their target's death, but
e43c673 scoped its stamp to named tenancies — and the pid-identity watch
surface (§4 Slice 3) targets arbitrary actors, including anonymous ones whose
pids cross the boundary in contract replies. The first wild pid-face hit
(width-20 soak, pid_watch_test.exs:47, 1/600 full-suite: a monitor installed
while the child was alive delivered :noproc instead of {:smarm_exit, :panic})
is exactly the residual a0ba9be's commit body deferred.

ever_named becomes `watchable`, with a second set-site: mark_watchable(pid),
which the bridge calls wherever a smarm pid is encoded across the boundary —
BEAM can only watch pids it holds, and can only hold pids that crossed.
Anonymous never-exported churn (holder threads, egress tasks) stays
ineligible, preserving e43c673's LIFO-eviction protection unchanged.

mark_watchable takes the cold lock before the liveness screen: finalize
publishes Done and reads the bit under the same lock, so the mark either
lands before the death stamps or observes the tenancy dead and no-ops —
no lost-stamp window, and marking a corpse cannot invent history (pinned
in the test alongside the mark-while-alive stamp).
2026-08-13 05:56:19 +00:00
smarm-agent 461fe4b768 fix(runtime): only named tenancies stamp the terminal record — anonymous churn must not evict it
Discovered wiring the bridge consult: with an unconditional stamp, the record
for the very death being raced was the shortest-lived data in the runtime.
Every green thread is a slot tenant, the free list is LIFO — so the slot a
named server's death frees is the first one recycled, and the next throwaway
exit (monitor holders, chain-runner work, anything) overwrote the record
before a raced watch could consult it. Deterministic bridge repro: the
corpse resolved fine, terminal_reason read None every time.

register_with now flags the tenancy (ever_named, reset at reclaim) before
the binding lands — set outside the registry lock, so no successfully
registered actor can die unflagged and a failed register's overshoot is
harmless — and finalize stamps only flagged tenancies. Watchable identities
are exactly the named ones (the bridge's pid-identity path deliberately
keeps Erlang's raw :noproc), so nothing consultable is lost.

Contract test updated: the three death modes now self-register; a new
anonymous control pins that unregistered deaths neither stamp nor evict.
2026-08-13 05:56:19 +00:00
smarm-agent b937f1f50f monitor/registry: terminal-outcome record — a raced watch can recover the real down reason (soak sig 4)
A watch installed after its target's death has, until now, only NoProc to
report — but the bridge's proxies install their native watch asynchronously
after acquire returns, so a link established before a crash (from the BEAM's
view) could still lose the panic's translated reason to that blanket NoProc
(width-20 soak signature 4: link_test.exs:26, 1/600 full-suite, 3/2000
link-only, all whereis-miss; deterministic repro in the bridge suite).

Two primitives, no change to monitor()'s own Erlang-faithful stale-pid
semantics — the upgrade is the caller's deliberate act:

- finalize_actor stamps the slot with (generation, DownReason) under the same
  cold-lock block that publishes the outcome. The record survives reclaim,
  registry pruning, and the next tenant's install; only the slot's next death
  overwrites it. terminal_reason(pid) reads it generation-matched.
- resolve_name(name) is whereis with the corpse kept: the dead-holder arm
  returns the stored pid it prunes (NameResolution::Corpse) instead of
  discarding the only evidence of who died — whereis itself prunes on the way
  out, so a whereis-then-lookup consumer would find the evidence already
  destroyed. Live/Unbound match whereis's Some/None; the name heals exactly
  as before.

Contract pinned in tests/terminal_outcome_after_death.rs: one record per way
of dying (Exit/Panic/Stopped), no record while live, corpse capture + heal on
resolve_name, record independence from registry pruning, survival across slot
re-tenancy, overwrite at the next tenancy's death.
2026-08-13 05:56:19 +00:00
Claude (sandbox) 301e3463e3 chore(release): v0.6.0 — RFC 019: actor stack reserve & shrink
Per-actor stack shapes on every spawn surface (SpawnOpts stack_reserve/
guard_size, Config defaults, pool rule: only default-shaped recycle);
sampled stack high-water + MADV_FREE shrink at actor-park (THRESHOLD
256 KiB, COOLDOWN 64 parks, redzone 1 page); pool-recycle MADV_DONTNEED
above the retained 64 KiB entry end; SIGSEGV overflow diagnostics
(two-tier: in-guard definitive / 1 MiB overshoot 'stepped over', prior
handler chained for foreign faults) with per-scheduler sigaltstack; and
the per-actor introspection surface (ActorInfo.stack: reserve, guard,
sampled depth, parks_since_shrink, shrinks).

Amendments ratified during implementation, for the RFC changelog:
- DEFAULT_STACK_GUARD 64 KiB -> 1 MiB, following the kernel's post-Stack-
  Clash stack_guard_gap convention; PROT_NONE width is VA-only and free.
- §7's motivating segfault was a cargo-vendored gz build, not SQLite as
  the RFC text says (cc-built C lacks -fstack-clash-protection; distro
  libraries have it — the risky class is vendored builds).
- §4 hibernate() deferred to the jar (bolt-on: force-flag on the §3
  shrink path, ~10 lines when wanted).

Gates (jobrunner box, 2026-08-08): reclaim gate PASS at c3 and again at
tip (3.0 MiB LazyFree -> kernel reclaim -> Rss to one live page ->
re-spike bit-identical, live data intact; MADV_PAGEOUT stands in for
memcg — cgroup2 is RO in the job container — driving the same reclaim
path). E1 interleaved A/B vs v0.5.0: every ka cell (the E1 subject)
within +0.3..+2.9% at tip; close-mode control cells within noise except
t8-c4 close, which is bistable (~40-44k vs ~46-49k modes for BOTH
variants, base self-disagrees by 11% across rounds); 6 rounds across two
runs are inconclusive there and a 10-round focused run is noted in the
handoff as deferred follow-up, accepted for this release.

No breaking API changes since v0.5.0: SpawnOpts fields and ActorInfo
gained members (exhaustive-construction downstream will need the new
ActorInfo.stack field; urus does not construct it).
v0.6.0
2026-08-08 19:44:55 +00:00
Claude (sandbox) 410ba33d82 feat(introspect,runtime): per-actor stack surface on ActorInfo (RFC 019 §8)
- introspect::StackInfo { reserve, guard, depth_high_water,
  parks_since_shrink, shrinks } as ActorInfo.stack; re-exported at crate
  root beside ActorInfo.
- All reads lock-free: geometry from the c6 diag slot atomics, depth =
  top - hwm (the §2 sampled high-water; doc spells out sampled-not-exact
  and that 0 means never-descheduled-at-depth), counters straight off the
  §3 atomics. Coherence for the incarnation rides read_slot's existing
  generation check, same as overruns/messages_received.
- Slot::stack_introspect(): one pub(crate) tuple accessor beside the other
  counter accessors.
- Exact RSS deliberately absent per RFC (mincore = debug tooling only,
  never a runtime path); stack_shape(pid) untouched (cold-lock exact
  variant from c2).
- tests/introspect.rs: defaults surface (64 KiB reserve / 1 MiB guard /
  sampled ~32 KiB depth / gate park counted / zero shrinks) + live shrink
  counters (spike visible pre-shrink; shrinks>=1, cooldown counter reset,
  hwm reset after crossing COOLDOWN) read mid-run -- post-join the slot
  reclaim correctly hides the incarnation, which the first draft of the
  test learned the hard way.

FLAGGED (Claude-solo calls):
- Nested StackInfo struct over five flat ActorInfo fields (grain break;
  the five fields are one concern and ActorInfo is already 12 fields).
- Field names reserve/guard/shrinks (RFC says stack_reserve/stack_guard/
  shrink count; the stack_ prefix is redundant inside StackInfo).
2026-08-08 19:12:46 +00:00
Claude (sandbox) 5fd8aecf55 feat(signal,runtime,stack): SIGSEGV overflow diagnostics + 1 MiB guard default (RFC 019 §7)
- src/signal.rs: process-global SA_SIGINFO|SA_ONSTACK handler installed once
  at runtime::init (before any scheduler thread -> unracing PRIOR save);
  per-scheduler-thread 64 KiB sigaltstack registered at schedule_loop entry
  (a guard hit leaves no stack to handle on). Async-signal-safe throughout:
  classification is plain loads (const-init TLS Cell + slot atomics), print
  is fixed-buffer itoa + one write(2), death is SIG_DFL + refault at the
  same instruction (core-dumpable, correct wait status).
- Two-tier classification (agreed): in-guard = definitive; OVERSHOOT window
  below the guard = 'unprobed (FFI?) frame stepped over it' probable
  attribution -- the RFC's motivating incident (cargo-vendored gz, not
  SQLite as the RFC text says) faults there under a small guard. Pure
  classify() fn, 5 adversarial units incl. saturation at low addresses.
- DEFAULT_STACK_GUARD 64 KiB -> 1 MiB (agreed): kernel stack_guard_gap
  anchor post-Stack-Clash; PROT_NONE is VA-only (no RSS, no page tables,
  no overcommit charge) so width is free at any actor count.
- Unclassified faults reinstate the PRIOR sigaction and refault (agreed):
  std's own OS-thread overflow diagnostics survive our presence.
- Slot: diag_{stack_top,stack_reserve,stack_guard,pid} atomics written in
  install_actor pre-publish; readable without the cold lock (Stack lives
  under it); only consulted while CURRENT_SLOT points at the slot, so
  never stale where read. preempt::current_slot_ptr ungated from
  smarm-causal (now also the classifier's anchor).
- build.rs + cc (agreed Q3): canary/canary.c, 96 KiB local touched low-end
  first, -fno-stack-clash-protection pinned so hardened toolchains don't
  probe the canary into uselessness.
- tests/stack_diag.rs: subprocess x4 -- Rust recursion tier-1; FFI canary
  tier-1 at defaults (1 MiB guard catches the jump); tier-2 at guard=4 KiB
  ('stepped over', reproduces the incident); clean at reserve=256 KiB
  (the §1 knob is the fix, same frame).

FLAGGED (Claude-solo calls):
- OVERSHOOT_SLOP = 1 MiB (matches guard default/kernel gap; beyond it
  attribution would be dishonest).
- Altstack 64 KiB, mmap'd once per OS thread, never freed (bounded by
  thread count; reused across run()s via TLS flag).
- Foreign-fault reinstate permanently deregisters our handler; accepted --
  the process is dying either way.
- Diag geometry as 4 slot atomics (install-time cost only) over a per-switch
  TLS snapshot (hot-path stores).
2026-08-08 18:58:30 +00:00
Claude (sandbox) 7d8b9e0310 feat(stack,runtime): pool-recycle DONTNEED above the retained entry end (RFC 019 §6)
- stack::retain_range: pure checked span fn (retain page-up = zap less;
  None when retain covers the reserve, so the 64 KiB default config never
  pays a syscall) + 6 adversarial units mirroring shrink_range's.
- Stack::recycle_zap: advisory MADV_DONTNEED of [usable_base, top-RETAIN);
  stack is unowned at the call site, synchronous eager zap races nothing.
- recycle_stack: zap OFF-LOCK before pool admission (acquire_stack's
  no-syscall-under-the-pool-lock invariant); rare cap-overflow pays a
  wasted zap ahead of munmap, accepted over a second lock round-trip.
- pub const RECYCLE_RETAIN = 64 KiB beside the shrink knobs, ratified-as-
  constant rationale in doc.
- tests/stack_recycle.rs: mincore-based exact-zero-resident assert over
  the zap span. smaps was tried first and over-counts: a neighboring rw
  anon VMA can merge flush against the stack top (observed once under the
  full-suite run); the PROT_NONE guard pins the usable base exactly.

FLAGGED (Claude-solo calls):
- RFC §6 'above the bottom RETAIN' is direction-ambiguous in address
  terms; implemented as retain the ENTRY end (highest addresses, the
  pages the next actor faults first), zap the cold deep span below.
- Const named RECYCLE_RETAIN (RFC says RETAIN) to sit beside SHRINK_*.
2026-08-08 16:13:53 +00:00
Claude (sandbox) 8225716b11 feat(runtime,stack): sampled stack high-water + MADV_FREE shrink at actor-park (RFC 019 §§2–3)
hwm: AtomicUsize lands beside sp on the slot: the single context-save
site min-updates it (one branch + at most one Relaxed store into the
line the sp store just dirtied), install resets it to the fresh top.
Advisory by construction — correctness never depends on it. The mod-doc
ordering chain gains a line: hwm piggybacks the existing
Relaxed-store-before-Release pattern and adds no edges.

Shrink hook in the YieldIntent::Park arm only, before the park_return
Release transition — the owned window (obligation 1's assert-comment at
the site): after the sp store, before Parked is published, scheduler on
its own stack, actor saved and unstealable. It runs on both arms of the
park_return race (a consumed unpark flag means one wasted-but-harmless
madvise). The preempt/yield path deliberately never checks: §4's
bounded, self-healing leak under saturation, when syscalls are least
affordable.

SHRINK_THRESHOLD = 256 KiB and SHRINK_COOLDOWN = 64 parks are pub
constants with the ratified doc rationale, not Config fields. The freed
span is shrink_range(hwm, sp, page): whole pages of [hwm, sp − 1-page
redzone), rounded inward, checked arithmetic — adversarial inputs
collapse to None (obligation 2). MADV_FREE marks lazily; the kernel's
reclaim-under-pressure IS the hysteresis, cancel-on-write is the safety
net. parks_since_shrink + shrink_count ride the slot for the cooldown
and the future introspect surface.

Tests: 7 adversarial shrink_range units (inverted/empty spans, redzone
underflow, unaligned ends, sp-crossing sweep); integration — 8 MiB
reserve, ~3 MiB spike sampled via yield-at-depth, parks gated on
introspected Parked state past the cooldown, then ≥ 2 MiB LazyFree
asserted inside the stack's smaps range with live data intact; and the
inverse guard — a shallow never-spiking actor ends at exactly 0
LazyFree (also proves the parser isn't vacuously zero via the first
test).
2026-08-08 14:30:32 +00:00