Commit Graph
14 Commits
Author SHA1 Message Date
claude-asm-audit 93bc83a5b3 perf(channel): capture the receiver's runtime Weak once per channel, not once per park
Upstream's off-runtime wake (1002777) captured a Weak<RuntimeInner> in the
parked_receiver tuple on every park, so every park/unpark round-trip paid an
Arc::downgrade plus the matching Weak drop — a locked RMW pair on one globally
shared counter, on the hot path of every channel workload.

The receiver never migrates between runtimes, so one capture is enough: the
Weak moves into Inner<T>, filled in on first park (a branch on an Option
thereafter), and unpark_at_via takes a closure so the fallback only re-locks
and clones off a scheduler thread, where the in-runtime fast path has already
declined. runtime_weak() returns Option and no longer panics off-runtime.

Sandbox 1T channel steady round-trip: 290 -> 270 ns (upstream's shape measured
427 -> 445 on its own base). Full suite green under reltest, including
upstream's tests/cross_thread_wake.rs.
2026-08-21 12:30:18 +00:00
claude-asm-audit a9341e2d82 fix(tls): actor-side thread-local accessors are #[inline(never)] + fence — LLVM caches TLS addresses across context switches (finding 19)
Without LTO (any downstream `cargo build --release`), LLVM keeps `%fs:0` in a
callee-saved register across `switch_to_scheduler`; after the actor migrates,
`ACTOR_DONE`/`PREEMPTION_ENABLED`/`CURRENT_PID` etc. hit the OLD thread's TLS.
8 multi-thread tests aborted with "scheduler resumed a done actor" (error 5).
Thin LTO only worked because it picked the local-exec model.

- context.rs: module doc "Thread-locals and migration" (the rule), tls_fence()
- every TLS accessor reachable from an actor stack is out of line:
  actor::{current_pid, publish_outcome (new)}, preempt::{maybe_preempt,
  check_cancelled, current_slot_ptr, note_*, preemption_swap/enabled (new;
  replaces direct PREEMPTION_ENABLED.with in NoPreempt/RawMutex/with_runtime/
  trace/debug asserts)}, context::{get,set}_scheduler_sp, runtime::{sched_slot,
  slot_push, set_yield_intent}, causal, trace
- raw_mutex order checks: out of line only under debug_assertions
- Cargo.toml: [profile.reltest] = release without LTO; 41 test bins link in
  ~1.5 s each instead of ~12 s (suite 11 min -> 1.5 min) AND it is the
  regression oracle for this bug: `cargo test --profile reltest`
Both profiles: 41/41 + doctests green.
2026-08-21 12:23:46 +00:00
claude-asm-audit d300a9d536 docs: fix 10 rustdoc link warnings, deny broken/private/redundant intra-doc links, doctests deny(warnings) (3 doctests had unused vars/imports) 2026-08-21 12:23:46 +00:00
claude-asm-audit bc0a5e8656 scheduler: spawn_monitor / spawn_monitor_with — monitor registered on the child's slot before publish, so the Down always carries the real reason (spawn-then-monitor could race to NoProc); tests; channel test uses it 2026-08-21 12:23:46 +00:00
claude-asm-audit ea2b222cff runtime: wake_slot default ON (RFC 005 accepted, finding 18); supervisor stops survivors sequentially so reverse-order teardown is guaranteed; deflake gen_server/channel tests; baseline.json regenerated slot-on; ROADMAP notes 2026-08-21 12:23:33 +00:00
claude-asm-audit 7001f04b65 diag(runtime): wake-path counters in SchedulerStats + RuntimeStats::wake_diag; RQDIAG/DIAG lines in rq_runtime and general (target 5, finding 17) 2026-08-21 12:22:46 +00:00
claude-asm-audit 77938cd31d bench: split ping_pong_oneshot — spawn_pair_control + ping_pong_steady; refresh baseline (job e982b5b5)
general.rs sections 5/6: spawn_pair_control is ping_pong_oneshot with the
messages removed (2 spawn + 2 join per round); ping_pong_steady is one
persistent pair × 10k roundtrips over unbounded MPSC (smarm::channel vs
tokio::sync::mpsc::unbounded_channel). Box (5900X, 3729fff, rq-mpmc):
oneshot 804/control 633 → 79% spawn+join; steady 142 ns vs tokio 128
(0.90×) at 1T, 1.8 µs/roundtrip at 20T. history.md finding 16.

baseline.json regenerated from the same run (20T labels, 14 benches).
2026-08-21 12:22:46 +00:00
claude-asm-audit 5ddd122711 perf(preempt): rdtsc unserialised by default; causal attribution opts into lfence
reset_timeslice paid an lfence pipeline drain on every resume via the
shared rdtsc() helper. Nothing in preempt.rs needs it: the timeslice
arm/expiry compare against a ~1e5-cycle slice, and an early stamp only
makes the slice look more used. The consumer that does need it — causal
site attribution, where a speculative early read misattributes a site's
tail — now calls rdtsc_serialising() explicitly (cold_check sample,
SiteGuard enter/exit).

Measured vs 31dc26a (baseline commit), 24-core box, rq-mpmc default:
- switch_cost, taskset -c 2, 9 interleaved old/new pairs: mean_cyc
  149 -> 117 per roundtrip (-32, -21%); mean_ns 43.2 -> 34.9.
- sweep.py regress + run, 20T, two runs agree:
  yield_in_hot_loop 1T 40171 -> 30277/30497 us (-24%)
  yield_many        1T 12513 ->  9965/10013 us (-20%)
  ping_pong_oneshot unchanged (RFC 005 wake-slot handoffs skip
  reset_timeslice, so that path never paid the lfence).
  All other rows within the box's ~+-15% noise floor.
- 1-core sandbox: 218 -> 206 cyc (understates by ~3x).
- bench binary lfence count 24 -> 4 (survivors = the bench's own
  rdtscp;lfence bracket). Tests green.
Baseline (benches/baseline.json) not re-saved in this commit.
2026-08-21 12:22:46 +00:00
claude-asm-audit 343e53e17b bench: baseline under the fixed rq-mpmc default (job 6e71e9b1)
20-core box (taskset 0-19), 5 sets, at 0438f12. Replaces the rq-mutex
baseline (df41ff3) so sweep.py regress compares like with like. The
multi-thread rows show what the mutex contention was costing:
yield_many 20T 156.7ms -> 44.1ms (-72%), spawn_storm_busy 20T -72%,
chained_spawn/catch_unwind 20T -20%. Sweep ran clean (0 panics) —
the finding-13 fix holding under the full suite.
Note: mpsc_contention 1T is ~2x the mutex-default row (3155 -> 6123);
backend characteristic to investigate, not a fix regression (the
shootout shows the fix itself is noise-neutral).
2026-08-21 12:22:46 +00:00
claude-asm-audit 9b215573de fix(run_queue): make the Vyukov rings preemption-tolerant (finding 13)
A consumer OS-preempted between its dequeue_pos claim and its seq
release freezes one cell; once traffic laps the ring (~cap ops ≈ 1-2ms
at yield-storm throughput ≈ one scheduling quantum under load),
try_push reads the stale seq and the original algorithm's 'lap behind
=> full' inference misfires. The old push assert then converted that
liveness stall into an abort blaming a double enqueue that never
happened (soak: occupancy 174-181 of cap 16384 at every failure), and
the dead scheduler threads stranded actors => the observed hangs.
Pristine 5504ef3 failed 14/14 under an 8-spinner soak on the 20-core
box; a retry prototype passed 13/14, its one failure being a stall
that outlived a fixed 1M-spin bound — pure spinning starves the
descheduled consumer, so the wait must yield.

Fix, following crossbeam ArrayQueue's shape (fence + opposite-counter
check; cells hand off, COUNTERS give verdicts):
- push: on a lap-behind cell, fence(SeqCst) + occupancy check.
  occ < cap => transient stall => spin-then-yield backoff and retry;
  occ >= cap => the REAL at-most-once-enqueued violation => panic with
  a truthful message and the counters. enqueue_pos is loaded before
  dequeue_pos so racing pops only underestimate occupancy (no spurious
  panic).
- pop: symmetric counter check before an empty verdict; on a
  mid-publish producer, bounded wait then None — deliberate deviation
  from crossbeam's unbounded retry (spurious None is benign: RFC 018's
  enqueue-wake self-heals; StripedRing's probe must not hang on one
  stripe).
- StripedRing::push probe: yield-escalating backoff after a full
  refused lap (was a bare spin_loop).
- Hand-rolled Backoff (spin 2^n to 64, then yield_now); under loom
  every wait is a yield so models explore the stalled peer's progress.

New loom model mpmc_lap_onto_stalled_consumer_completes reproduces the
old panic in the first explored interleavings (verified FAILED against
5504ef3) and passes with the fix. Lib 60 + integration 41 + all 4 ring
loom models pass. rq-striped inherits the fix (stripes are MpmcRings).
2026-08-21 12:22:46 +00:00
claude-asm-audit bf55cef3e3 perf(run_queue): flip default backend rq-mutex -> rq-mpmc
Evidence (history.md session 5, findings 10-11, 20-core jobrunner run):
rq-mutex collapses with thread count on queue-heavy load (yield-storm
6.9x slower than striped at 20T; attributed cause of the baseline's
multi-thread regression on yield_many/chained_spawn), while rq-mpmc is
best-or-close everywhere: best 1-thread, best slot-on ping-pong
(1237us vs mutex 6754us at 20T), -23%/-41 cyc per yield roundtrip on
real hardware (single-core switch_cost, interleaved). rq-striped
remains the churn-heavy many-core option, selectable per build.

Docs + compile_error hints in run_queue.rs updated to name rq-mpmc as
the default. slot_state.rs untouched, no loom-relevant changes; lib
(60) + scheduler/channel/supervisor/park_wake/wake_slot/preempt (41)
pass under the new default.
2026-08-21 12:22:46 +00:00
claude-asm-audit 2f88264426 bench: refresh baseline.json — 20-core box, fe85197, rq-mutex
Re-measured via jobrunner (taskset -c 0-19 of 24, rust:1.97-slim,
sweep.py run --save-baseline, 5 sets). Supersedes the old 24-thread
baseline; multi-thread labels are now 'smarm 20-thread'. Taken under
the rq-mutex default *before* the backend flip — the yield_many /
chained_spawn multi-thread regression reproduces here and is attributed
to rq-mutex contention (history.md session 5, finding 10).
2026-08-21 12:22:45 +00:00
claude-asm-audit 6172f4231d perf(runtime): skip take_closure's locked swap after first resume
Every resume paid an unconditional AtomicPtr::swap (lock xchg, full
barrier) to check for a first-resume closure that is null on all
resumes after the first. A Relaxed null-load fast path is sound:
store_closure runs only before publish_queued, whose Release pairing
with try_claim's Acquire orders it before this call, so no writer can
race the load within an occupancy.

Measured on switch_cost (1-core sandbox, rq-mutex, cycles): mean
roundtrip 350-355 -> 324-328, ~7.5%. All lib + scheduler/channel/
supervisor tests pass.
2026-08-21 12:22:45 +00:00
claude-asm-audit d7082eb266 perf(context): pass actor sp through registers, not TLS
switch_to_actor takes the target sp in rdi and returns the actor's
next saved sp in rax (handed over by switch_to_scheduler's shim).
Deletes the ACTOR_SP thread-local and halves the helper calls per
one-way switch (2 -> 1); the scheduler loop also drops its
set_actor_sp/get_actor_sp TLS round-trips. SCHEDULER_SP stays: a
yielding actor at arbitrary call depth has no argument channel back.

asm before/after in outputs/history.md session 1-2. Tests: 354 pass.
2026-08-21 12:22:45 +00:00