Commit Graph
12 Commits
Author SHA1 Message Date
claude-asm-audit d300a9d536 docs: fix 10 rustdoc link warnings, deny broken/private/redundant intra-doc links, doctests deny(warnings) (3 doctests had unused vars/imports) 2026-08-21 12:23:46 +00:00
claude-asm-audit bc0a5e8656 scheduler: spawn_monitor / spawn_monitor_with — monitor registered on the child's slot before publish, so the Down always carries the real reason (spawn-then-monitor could race to NoProc); tests; channel test uses it 2026-08-21 12:23:46 +00:00
claude-asm-audit ea2b222cff runtime: wake_slot default ON (RFC 005 accepted, finding 18); supervisor stops survivors sequentially so reverse-order teardown is guaranteed; deflake gen_server/channel tests; baseline.json regenerated slot-on; ROADMAP notes 2026-08-21 12:23:33 +00:00
claude-asm-audit 7001f04b65 diag(runtime): wake-path counters in SchedulerStats + RuntimeStats::wake_diag; RQDIAG/DIAG lines in rq_runtime and general (target 5, finding 17) 2026-08-21 12:22:46 +00:00
claude-asm-audit 77938cd31d bench: split ping_pong_oneshot — spawn_pair_control + ping_pong_steady; refresh baseline (job e982b5b5)
general.rs sections 5/6: spawn_pair_control is ping_pong_oneshot with the
messages removed (2 spawn + 2 join per round); ping_pong_steady is one
persistent pair × 10k roundtrips over unbounded MPSC (smarm::channel vs
tokio::sync::mpsc::unbounded_channel). Box (5900X, 3729fff, rq-mpmc):
oneshot 804/control 633 → 79% spawn+join; steady 142 ns vs tokio 128
(0.90×) at 1T, 1.8 µs/roundtrip at 20T. history.md finding 16.

baseline.json regenerated from the same run (20T labels, 14 benches).
2026-08-21 12:22:46 +00:00
claude-asm-audit 5ddd122711 perf(preempt): rdtsc unserialised by default; causal attribution opts into lfence
reset_timeslice paid an lfence pipeline drain on every resume via the
shared rdtsc() helper. Nothing in preempt.rs needs it: the timeslice
arm/expiry compare against a ~1e5-cycle slice, and an early stamp only
makes the slice look more used. The consumer that does need it — causal
site attribution, where a speculative early read misattributes a site's
tail — now calls rdtsc_serialising() explicitly (cold_check sample,
SiteGuard enter/exit).

Measured vs 31dc26a (baseline commit), 24-core box, rq-mpmc default:
- switch_cost, taskset -c 2, 9 interleaved old/new pairs: mean_cyc
  149 -> 117 per roundtrip (-32, -21%); mean_ns 43.2 -> 34.9.
- sweep.py regress + run, 20T, two runs agree:
  yield_in_hot_loop 1T 40171 -> 30277/30497 us (-24%)
  yield_many        1T 12513 ->  9965/10013 us (-20%)
  ping_pong_oneshot unchanged (RFC 005 wake-slot handoffs skip
  reset_timeslice, so that path never paid the lfence).
  All other rows within the box's ~+-15% noise floor.
- 1-core sandbox: 218 -> 206 cyc (understates by ~3x).
- bench binary lfence count 24 -> 4 (survivors = the bench's own
  rdtscp;lfence bracket). Tests green.
Baseline (benches/baseline.json) not re-saved in this commit.
2026-08-21 12:22:46 +00:00
claude-asm-audit 343e53e17b bench: baseline under the fixed rq-mpmc default (job 6e71e9b1)
20-core box (taskset 0-19), 5 sets, at 0438f12. Replaces the rq-mutex
baseline (df41ff3) so sweep.py regress compares like with like. The
multi-thread rows show what the mutex contention was costing:
yield_many 20T 156.7ms -> 44.1ms (-72%), spawn_storm_busy 20T -72%,
chained_spawn/catch_unwind 20T -20%. Sweep ran clean (0 panics) —
the finding-13 fix holding under the full suite.
Note: mpsc_contention 1T is ~2x the mutex-default row (3155 -> 6123);
backend characteristic to investigate, not a fix regression (the
shootout shows the fix itself is noise-neutral).
2026-08-21 12:22:46 +00:00
claude-asm-audit 9b215573de fix(run_queue): make the Vyukov rings preemption-tolerant (finding 13)
A consumer OS-preempted between its dequeue_pos claim and its seq
release freezes one cell; once traffic laps the ring (~cap ops ≈ 1-2ms
at yield-storm throughput ≈ one scheduling quantum under load),
try_push reads the stale seq and the original algorithm's 'lap behind
=> full' inference misfires. The old push assert then converted that
liveness stall into an abort blaming a double enqueue that never
happened (soak: occupancy 174-181 of cap 16384 at every failure), and
the dead scheduler threads stranded actors => the observed hangs.
Pristine 5504ef3 failed 14/14 under an 8-spinner soak on the 20-core
box; a retry prototype passed 13/14, its one failure being a stall
that outlived a fixed 1M-spin bound — pure spinning starves the
descheduled consumer, so the wait must yield.

Fix, following crossbeam ArrayQueue's shape (fence + opposite-counter
check; cells hand off, COUNTERS give verdicts):
- push: on a lap-behind cell, fence(SeqCst) + occupancy check.
  occ < cap => transient stall => spin-then-yield backoff and retry;
  occ >= cap => the REAL at-most-once-enqueued violation => panic with
  a truthful message and the counters. enqueue_pos is loaded before
  dequeue_pos so racing pops only underestimate occupancy (no spurious
  panic).
- pop: symmetric counter check before an empty verdict; on a
  mid-publish producer, bounded wait then None — deliberate deviation
  from crossbeam's unbounded retry (spurious None is benign: RFC 018's
  enqueue-wake self-heals; StripedRing's probe must not hang on one
  stripe).
- StripedRing::push probe: yield-escalating backoff after a full
  refused lap (was a bare spin_loop).
- Hand-rolled Backoff (spin 2^n to 64, then yield_now); under loom
  every wait is a yield so models explore the stalled peer's progress.

New loom model mpmc_lap_onto_stalled_consumer_completes reproduces the
old panic in the first explored interleavings (verified FAILED against
5504ef3) and passes with the fix. Lib 60 + integration 41 + all 4 ring
loom models pass. rq-striped inherits the fix (stripes are MpmcRings).
2026-08-21 12:22:46 +00:00
claude-asm-audit bf55cef3e3 perf(run_queue): flip default backend rq-mutex -> rq-mpmc
Evidence (history.md session 5, findings 10-11, 20-core jobrunner run):
rq-mutex collapses with thread count on queue-heavy load (yield-storm
6.9x slower than striped at 20T; attributed cause of the baseline's
multi-thread regression on yield_many/chained_spawn), while rq-mpmc is
best-or-close everywhere: best 1-thread, best slot-on ping-pong
(1237us vs mutex 6754us at 20T), -23%/-41 cyc per yield roundtrip on
real hardware (single-core switch_cost, interleaved). rq-striped
remains the churn-heavy many-core option, selectable per build.

Docs + compile_error hints in run_queue.rs updated to name rq-mpmc as
the default. slot_state.rs untouched, no loom-relevant changes; lib
(60) + scheduler/channel/supervisor/park_wake/wake_slot/preempt (41)
pass under the new default.
2026-08-21 12:22:46 +00:00
claude-asm-audit 2f88264426 bench: refresh baseline.json — 20-core box, fe85197, rq-mutex
Re-measured via jobrunner (taskset -c 0-19 of 24, rust:1.97-slim,
sweep.py run --save-baseline, 5 sets). Supersedes the old 24-thread
baseline; multi-thread labels are now 'smarm 20-thread'. Taken under
the rq-mutex default *before* the backend flip — the yield_many /
chained_spawn multi-thread regression reproduces here and is attributed
to rq-mutex contention (history.md session 5, finding 10).
2026-08-21 12:22:45 +00:00
claude-asm-audit 6172f4231d perf(runtime): skip take_closure's locked swap after first resume
Every resume paid an unconditional AtomicPtr::swap (lock xchg, full
barrier) to check for a first-resume closure that is null on all
resumes after the first. A Relaxed null-load fast path is sound:
store_closure runs only before publish_queued, whose Release pairing
with try_claim's Acquire orders it before this call, so no writer can
race the load within an occupancy.

Measured on switch_cost (1-core sandbox, rq-mutex, cycles): mean
roundtrip 350-355 -> 324-328, ~7.5%. All lib + scheduler/channel/
supervisor tests pass.
2026-08-21 12:22:45 +00:00
claude-asm-audit d7082eb266 perf(context): pass actor sp through registers, not TLS
switch_to_actor takes the target sp in rdi and returns the actor's
next saved sp in rax (handed over by switch_to_scheduler's shim).
Deletes the ACTOR_SP thread-local and halves the helper calls per
one-way switch (2 -> 1); the scheduler loop also drops its
set_actor_sp/get_actor_sp TLS round-trips. SCHEDULER_SP stays: a
yielding actor at arbitrary call depth has no argument channel back.

asm before/after in outputs/history.md session 1-2. Tests: 354 pass.
2026-08-21 12:22:45 +00:00