- next_monitor_id -> AtomicU64 on RuntimeInner - timers -> own Mutex<Timers>; io -> own Mutex<Option<IoThread>> (lock order: io-before-shared) - pending_closures Vec folded into Slot::pending_closure - termination check reads io liveness before shared; ordering argument documented - poison fix: check_cancelled() no longer fires while preemption is disabled, so a cancellation unwind can never poison a runtime/channel mutex - regression test: tests/poison_stop.rs - ROADMAP_v0.5.md added
6.4 KiB
6.4 KiB
smarm v0.5 — Runtime decomposition & run-path scaling
Goal: dismantle the single Mutex<SharedState> along the run path so the
runtime scales to dozens of cores, with crash isolation hardened so a torn
write or a stray cancellation can never poison shared runtime state.
Guiding rules established this cycle:
- Lock count that matters is hot-path locks. Cold lifecycle paths (spawn/join/monitor/link/finalize) may keep a small per-slot lock; the run path (yield/park/unpark/pop/resume) targets zero locks.
- Lock ordering: io-before-shared, and slot locks are leaves (never hold two slot locks at once). Any new lock states its position in this order.
- No unwind inside a runtime critical section. The stop sentinel only fires at lock-free observation points.
Phase 1 — State decomposition: easy peel-offs ✅ DONE
Split independent concerns out of SharedState so the global lock guards less.
next_monitor_id→AtomicU64onRuntimeInner(lock-free id minting).timers→ ownMutex<Timers>(only drain-winner + blocking prims touch it).io→ ownMutex<Option<IoThread>>(completion queue already self-locked).pending_closures: Vec→ folded intoSlot::pending_closure(per-actor data).- Termination check split: read io liveness before
shared; ordering argument documented inschedule_loop. - Poison fix: gate
check_cancelled()inmaybe_preemptbehindPREEMPTION_ENABLED, so a cancellation sentinel can never unwind while astd::sync::Mutex(shared/channel) is held. Regression test:tests/poison_stop.rs.
SharedState now holds only: slots, free_list, run_queue, root_pid.
Phase 2 — Slot table split (NEXT, the big correctness step)
Make slot lookup lock-free and per-slot state independently mutable.
- Fixed slab:
Box<[SlotCell; max_actors]>, slots never move → stable addresses, lock-free index. Assert on exhaustion with a panic message namingConfig::max_actors(n)as the fix. Defaultmax_actors = 16_384. (Mental note: must become unbounded or configurable-bounded later — see "Deferred" below. Do not let the fixed cap calcify into an assumption.) - Per-slot
generation: AtomicU32outside the lock for cheap stale-PID rejection. - Per-slot CAS state machine replacing
state+pending_unpark:Parked / Queued / Running / RunningNotified / Donein oneAtomicU8.unpark= CAS Parked→Queued (then push) or Running→RunningNotified; park-return = CAS Running→Parked, else re-queue. Eliminates the lost-wakeup bool by turning the race window into a state. sp→ relaxedAtomicUsize(happens-before carried by queue release-push / acquire-pop).- Per-slot raw non-poisoning mutex (spin-then-futex, Linux; embedded-
friendlier than
std::sync::Mutex) for the cold collections:waiters,monitors,links,supervisor_channel,outcome,pending_io_result,outstanding_handles. Guard entersNoPreempt. - Free list → its own structure (lock-free Treiber stack or striped).
finalize_actorlink cascade: lock peers one at a time (leaf rule); write the acyclicity argument next to it.live_actors/io_outstanding→ atomics; termination = both zero + queue empty. Finalize decrementslivestrictly after all wakeups it produced are enqueued — document this ordering as the correctness crux.
Phase 3 — Pluggable run queue + bench shootout
- Extract
RunQueuebehind atype RunQueue = ...alias selected at compile time by mutually-exclusive features, with acompile_error!guard if zero or >1 is set. No runtime dispatch. -rq-mutex—Mutex<VecDeque>, the control/baseline. -rq-mpmc— single hand-rolled Vyukov bounded MPMC ring (per-cell seq numbers). Strict FIFO; "one hot cache line". -rq-striped— M Vyukov rings, fetch-add ticket distribution. Relaxed FIFO, reordering bounded by ~M. Predicted winner @20c. - Bounded rings are sound because the slab cap + at-most-once-enqueued
invariant bound occupancy ≤
max_actors; size Σcapacity >max_actorsso push is infallible (probe next stripe when full; terminates). - All hand-rolled, dependency-free (portability incl. possible embedded).
Phase 4 — Bench harness
- Raw-structure microbench: N threads, push/pop throughput vs thread count, sweeping producer:consumer ratios. Isolates the data structure.
- Runtime-level: yield-storm, ping-pong-pairs, spawn-storm, all sweeping
scheduler count. Reuses existing
benches/harness style. - Driver script rebuilding per
rq-*feature to compare variants in one go. - Sandbox validates correctness via oversubscribed
Config::exact(N)on 1 core; real contention numbers come from the 20-core box.
Phase 5 — Safety hardening & model checking
loom(dev-dependency only, x86 Linux) model tests for the slot state machine and each ring variant.NoPreempt/ no-unwind-under-lock audit across all internal critical sections; assert the invariant in debug builds where feasible.
Fast follow (post-v0.5, written down so it isn't lost)
- Channel mutex migration.
channel::Inner<T>isArc<Mutex<_>>of the same poison class as the old shared lock;recv_matcheven runs a user predicate under it. The Phase-1check_cancelledgating already removes the unwind source globally, so channels are safe today — but migrating the channel to the raw non-poisoning mutex (Phase 2) is a clean follow-up rather than bundled into the run-path rework.
Deferred / explicitly out of scope for v0.5
- Unbounded / configurable-bounded actor count. v0.5 ships a fixed slab
with a loud assert. Revisit with a segmented slab (array of
AtomicPtr<Segment>, doubling segment sizes, append-only) once we actually hit the cap or a workload demands it. - Idle-wakeup eventcount. Idle scheduler threads keep the current 100µs poll-sleep. A futex-based eventcount is a later optimization if benches show idle latency matters.
- User-facing safe data structures (ArcSwap-style cells, structurally- shared persistent maps). Context for the runtime work, not in scope here.
- IO fd hygiene on actor death (pre-existing v0.2 TODO in
io.rs).