- benches/rq_micro.rs: raw-structure microbench, threads x producer:consumer ratio sweep. Benches all three queue types in one binary (they compile in every build; only the runtime alias is feature-selected), so no rebuild dance. Queues sized to the op count so the occupancy contract is met. - benches/rq_runtime.rs: whole-runtime benches with the selected variant: yield-storm (pure queue churn), ping-pong-pairs (park/unpark latency), spawn-storm (slab + free list + queue under churn). Scheduler-count sweep. - scripts/bench_rq.sh: rebuilds rq_runtime per rq-* feature, runs rq_micro once, aggregates RQCSV lines into bench_results/summary.csv. - All knobs via SMARM_BENCH_* env vars; house table format + machine lines. - run_queue module is now #[doc(hidden)] pub (types + push/pop/len + MpmcRing::with_capacity) solely so the external bench binary can drive the raw structures. docs(roadmap): phase 4 ticked (harness done; numbers from the 20-core box). New fast-follow per review: assert the invariants we lean on — debug_assert! on hot paths, loud assert!/panic on cold ones, at the point of reliance; sweep existing code during the phase-5 audit, adopt as house style. Validated end-to-end at smoke scale on the 1-core sandbox: full driver run, 24-row summary.csv across micro (3 structures x sweeps) and runtime (3 variants x 3 benches x thread sweep).
8.6 KiB
8.6 KiB
smarm v0.5 — Runtime decomposition & run-path scaling
Goal: dismantle the single Mutex<SharedState> along the run path so the
runtime scales to dozens of cores, with crash isolation hardened so a torn
write or a stray cancellation can never poison shared runtime state.
Guiding rules established this cycle:
- Lock count that matters is hot-path locks. Cold lifecycle paths (spawn/join/monitor/link/finalize) may keep a small per-slot lock; the run path (yield/park/unpark/pop/resume) targets zero locks.
- Lock ordering: io-before-shared, and slot locks are leaves (never hold two slot locks at once). Any new lock states its position in this order.
- No unwind inside a runtime critical section. The stop sentinel only fires at lock-free observation points.
- Never hold a thread-local guard across a possible switch point. An
actor can be preempted at any allocation and resume on a different OS
thread; a live
RefCellborrow (or stdMutexGuard) then pairs its acquire/release across two threads' locals — count underflow / UB.with_runtime,with_shared, and everyRawMutexguard disable preemption for exactly this reason (phase 2 found this the hard way).
Phase 1 — State decomposition: easy peel-offs ✅ DONE
Split independent concerns out of SharedState so the global lock guards less.
next_monitor_id→AtomicU64onRuntimeInner(lock-free id minting).timers→ ownMutex<Timers>(only drain-winner + blocking prims touch it).io→ ownMutex<Option<IoThread>>(completion queue already self-locked).pending_closures: Vec→ folded intoSlot::pending_closure(per-actor data).- Termination check split: read io liveness before
shared; ordering argument documented inschedule_loop. - Poison fix: gate
check_cancelled()inmaybe_preemptbehindPREEMPTION_ENABLED, so a cancellation sentinel can never unwind while astd::sync::Mutex(shared/channel) is held. Regression test:tests/poison_stop.rs.
SharedState now holds only: slots, free_list, run_queue, root_pid.
Phase 2 — Slot table split ✅ DONE
Make slot lookup lock-free and per-slot state independently mutable.
- Fixed slab:
Box<[Slot]>, slots never move → stable addresses, lock-free index. Assert on exhaustion with a panic message namingConfig::max_actors(n)as the fix. Defaultmax_actors = 16_384. (Mental note: must become unbounded or configurable-bounded later — see "Deferred" below. Do not let the fixed cap calcify into an assumption.) - Generation packed INTO the state word (better than the planned separate
AtomicU32): oneAtomicU64 = (gen << 32) | state, so the gen check is atomic with every transition — no ABA, no spurious unparks, by CAS. - Per-slot CAS state machine
Vacant/Queued/Running/RunningNotified/ Parked/Donereplacingstate+pending_unpark. Also fixed a latent lost wakeup in the Blocking-IO completion path (result set for a still-Running actor without a flag). sp→ relaxedAtomicUsize; stop flag + first-resume closure asAtomicPtrs — resume path fully lock-free.- Per-slot raw non-poisoning futex mutex (
src/raw_mutex.rs) for the cold collections. Guard entersNoPreempt. - Free list →
RawMutex<Vec<u32>>(a leaf lock). Treiber/striped only if the phase-4 spawn-storm bench shows contention — revisit then. finalize_actorlink cascade locks peers one at a time; acyclicity argument written at the site.link()registers target-first with the race argument at the site.live_actorsatomic; termination = io_out == 0 (read pre-queue-lock) && queue empty && live == 0; decrement-last ordering documented as the correctness crux at the site.
Phase 3 — Pluggable run queue ✅ DONE (shootout = phase 4 harness)
RunQueuealias insrc/run_queue.rs, compile-time selected,compile_error!guards for zero / >1 features. No runtime dispatch. All variants compile in every build (unit tests always run); the feature only picks the alias. Non-default variants:--no-default-features --features rq-…. -rq-mutex—Mutex<VecDeque>, the control/baseline. -rq-mpmc— single hand-rolled Vyukov bounded MPMC ring (per-cell seq numbers). Strict FIFO; "one hot cache line". -rq-striped— M Vyukov rings, fetch-add ticket distribution. Relaxed FIFO, reordering bounded by ~M. Predicted winner @20c.- Bounded rings sound via slab cap + at-most-once-enqueued (occupancy ≤
max_actors); mpmc capacity = next_pow2(max_actors), striped Σcapacity ≈ 2×max_actors with probe-from-home-stripe. A full ring panics as an invariant violation (double enqueue), never spins. - All hand-rolled, dependency-free.
- Two contracts surfaced by the extraction, documented + debug-asserted
in
run_queue.rs: queue ops require preemption disabled (a producer suspended mid-publish stalls every consumer behind its cell — livelock); and pop-None is a snapshot, not a fence — termination is counter-first (live == 0alone implies the queue holds nothing actionable; argument rewritten at theschedule_loopsite).
Phase 4 — Bench harness ✅ DONE (harness; real numbers from the 20-core box)
benches/rq_micro.rs: raw structures, threads × p:c ratio sweep. One binary covers all three structures (types compile in every build).benches/rq_runtime.rs: yield-storm, ping-pong-pairs, spawn-storm, sweeping scheduler count; variant baked in by feature.scripts/bench_rq.sh: rebuilds perrq-*feature, aggregates RQCSV lines intobench_results/summary.csv. Knobs via SMARM_BENCH_* env; e.g.SMARM_BENCH_THREADS="1 2 4 8 16 20" ./scripts/bench_rq.sh.- Harness validated end-to-end at smoke scale on the 1-core sandbox; contention curves and the actual variant decision come from the 20-core box.
Phase 5 — Safety hardening & model checking
loom(dev-dependency only, x86 Linux) model tests for the slot state machine and each ring variant.NoPreempt/ no-unwind-under-lock audit across all internal critical sections; assert the invariant in debug builds where feasible.
Fast follow (post-v0.5, written down so it isn't lost)
- Assert the invariants we lean on. This cycle accumulated load-bearing
invariants: at-most-once-enqueued, queue-ops-under-NoPreempt, never two
cold locks, cold-path generation re-verify under the lock, finalize's
decrement-last, pushes-pair-with-Queued-transitions, thread-local guards
never crossing a switch point. Whenever code RELIES on one and a cheap
check exists, assert it at the point of reliance —
debug_assert!on hot paths, fullassert!/loud panic on cold ones — so a violation fails at the breakage site, not three modules downstream (the slab-overflow panic and the queue-op preemption debug_assert are the pattern). Sweep the existing code for missed spots; new code adopts it as house style. The phase-5 audit is the natural vehicle for the sweep. - Channel mutex migration.
channel::Inner<T>isArc<Mutex<_>>of the same poison class as the old shared lock;recv_matcheven runs a user predicate under it. The Phase-1check_cancelledgating already removes the unwind source globally, so channels are poison-safe today — but phase 2 surfaced a second, sharper reason to migrate: the stdMutexGuardis held with preemption ENABLED, so a timeslice switch inside a channel critical section migrates the actor and unlocks the pthread mutex from a different OS thread — technically UB (Linux futexes happen to tolerate it). TheRawMutexguard disables preemption and is cross-thread-release sound by construction. Migrate as the first post-v0.5 change.
Deferred / explicitly out of scope for v0.5
- Unbounded / configurable-bounded actor count. v0.5 ships a fixed slab
with a loud assert. Revisit with a segmented slab (array of
AtomicPtr<Segment>, doubling segment sizes, append-only) once we actually hit the cap or a workload demands it. - Idle-wakeup eventcount. Idle scheduler threads keep the current 100µs poll-sleep. A futex-based eventcount is a later optimization if benches show idle latency matters.
- User-facing safe data structures (ArcSwap-style cells, structurally- shared persistent maps). Context for the runtime work, not in scope here.
- IO fd hygiene on actor death (pre-existing v0.2 TODO in
io.rs).