Commit Graph
5 Commits
Author SHA1 Message Date
claude-asm-audit 9b215573de fix(run_queue): make the Vyukov rings preemption-tolerant (finding 13)
A consumer OS-preempted between its dequeue_pos claim and its seq
release freezes one cell; once traffic laps the ring (~cap ops ≈ 1-2ms
at yield-storm throughput ≈ one scheduling quantum under load),
try_push reads the stale seq and the original algorithm's 'lap behind
=> full' inference misfires. The old push assert then converted that
liveness stall into an abort blaming a double enqueue that never
happened (soak: occupancy 174-181 of cap 16384 at every failure), and
the dead scheduler threads stranded actors => the observed hangs.
Pristine 5504ef3 failed 14/14 under an 8-spinner soak on the 20-core
box; a retry prototype passed 13/14, its one failure being a stall
that outlived a fixed 1M-spin bound — pure spinning starves the
descheduled consumer, so the wait must yield.

Fix, following crossbeam ArrayQueue's shape (fence + opposite-counter
check; cells hand off, COUNTERS give verdicts):
- push: on a lap-behind cell, fence(SeqCst) + occupancy check.
  occ < cap => transient stall => spin-then-yield backoff and retry;
  occ >= cap => the REAL at-most-once-enqueued violation => panic with
  a truthful message and the counters. enqueue_pos is loaded before
  dequeue_pos so racing pops only underestimate occupancy (no spurious
  panic).
- pop: symmetric counter check before an empty verdict; on a
  mid-publish producer, bounded wait then None — deliberate deviation
  from crossbeam's unbounded retry (spurious None is benign: RFC 018's
  enqueue-wake self-heals; StripedRing's probe must not hang on one
  stripe).
- StripedRing::push probe: yield-escalating backoff after a full
  refused lap (was a bare spin_loop).
- Hand-rolled Backoff (spin 2^n to 64, then yield_now); under loom
  every wait is a yield so models explore the stalled peer's progress.

New loom model mpmc_lap_onto_stalled_consumer_completes reproduces the
old panic in the first explored interleavings (verified FAILED against
5504ef3) and passes with the fix. Lib 60 + integration 41 + all 4 ring
loom models pass. rq-striped inherits the fix (stripes are MpmcRings).
2026-08-21 12:22:46 +00:00
claude-asm-audit bf55cef3e3 perf(run_queue): flip default backend rq-mutex -> rq-mpmc
Evidence (history.md session 5, findings 10-11, 20-core jobrunner run):
rq-mutex collapses with thread count on queue-heavy load (yield-storm
6.9x slower than striped at 20T; attributed cause of the baseline's
multi-thread regression on yield_many/chained_spawn), while rq-mpmc is
best-or-close everywhere: best 1-thread, best slot-on ping-pong
(1237us vs mutex 6754us at 20T), -23%/-41 cyc per yield roundtrip on
real hardware (single-core switch_cost, interleaved). rq-striped
remains the churn-heavy many-core option, selectable per build.

Docs + compile_error hints in run_queue.rs updated to name rq-mpmc as
the default. slot_state.rs untouched, no loom-relevant changes; lib
(60) + scheduler/channel/supervisor/park_wake/wake_slot/preempt (41)
pass under the new default.
2026-08-21 12:22:46 +00:00
claude-asm-audit 2f88264426 bench: refresh baseline.json — 20-core box, fe85197, rq-mutex
Re-measured via jobrunner (taskset -c 0-19 of 24, rust:1.97-slim,
sweep.py run --save-baseline, 5 sets). Supersedes the old 24-thread
baseline; multi-thread labels are now 'smarm 20-thread'. Taken under
the rq-mutex default *before* the backend flip — the yield_many /
chained_spawn multi-thread regression reproduces here and is attributed
to rq-mutex contention (history.md session 5, finding 10).
2026-08-21 12:22:45 +00:00
claude-asm-audit 6172f4231d perf(runtime): skip take_closure's locked swap after first resume
Every resume paid an unconditional AtomicPtr::swap (lock xchg, full
barrier) to check for a first-resume closure that is null on all
resumes after the first. A Relaxed null-load fast path is sound:
store_closure runs only before publish_queued, whose Release pairing
with try_claim's Acquire orders it before this call, so no writer can
race the load within an occupancy.

Measured on switch_cost (1-core sandbox, rq-mutex, cycles): mean
roundtrip 350-355 -> 324-328, ~7.5%. All lib + scheduler/channel/
supervisor tests pass.
2026-08-21 12:22:45 +00:00
claude-asm-audit d7082eb266 perf(context): pass actor sp through registers, not TLS
switch_to_actor takes the target sp in rdi and returns the actor's
next saved sp in rax (handed over by switch_to_scheduler's shim).
Deletes the ACTOR_SP thread-local and halves the helper calls per
one-way switch (2 -> 1); the scheduler loop also drops its
set_actor_sp/get_actor_sp TLS round-trips. SCHEDULER_SP stays: a
yielding actor at arbitrary call depth has no argument channel back.

asm before/after in outputs/history.md session 1-2. Tests: 354 pass.
2026-08-21 12:22:45 +00:00