Files
smarm/ROADMAP_v0.5.md
T
Claude 6d9f3698d4 feat(bench): phase 4 — run-queue bench harness + shootout driver
- benches/rq_micro.rs: raw-structure microbench, threads x producer:consumer
  ratio sweep. Benches all three queue types in one binary (they compile in
  every build; only the runtime alias is feature-selected), so no rebuild
  dance. Queues sized to the op count so the occupancy contract is met.
- benches/rq_runtime.rs: whole-runtime benches with the selected variant:
  yield-storm (pure queue churn), ping-pong-pairs (park/unpark latency),
  spawn-storm (slab + free list + queue under churn). Scheduler-count sweep.
- scripts/bench_rq.sh: rebuilds rq_runtime per rq-* feature, runs rq_micro
  once, aggregates RQCSV lines into bench_results/summary.csv.
- All knobs via SMARM_BENCH_* env vars; house table format + machine lines.
- run_queue module is now #[doc(hidden)] pub (types + push/pop/len +
  MpmcRing::with_capacity) solely so the external bench binary can drive the
  raw structures.

docs(roadmap): phase 4 ticked (harness done; numbers from the 20-core box).
New fast-follow per review: assert the invariants we lean on — debug_assert!
on hot paths, loud assert!/panic on cold ones, at the point of reliance;
sweep existing code during the phase-5 audit, adopt as house style.

Validated end-to-end at smoke scale on the 1-core sandbox: full driver run,
24-row summary.csv across micro (3 structures x sweeps) and runtime
(3 variants x 3 benches x thread sweep).
2026-06-09 20:44:10 +00:00

143 lines
8.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# smarm v0.5 — Runtime decomposition & run-path scaling
Goal: dismantle the single `Mutex<SharedState>` along the run path so the
runtime scales to dozens of cores, with crash isolation hardened so a torn
write or a stray cancellation can never poison shared runtime state.
Guiding rules established this cycle:
- **Lock count that matters is hot-path locks.** Cold lifecycle paths
(spawn/join/monitor/link/finalize) may keep a small per-slot lock; the run
path (yield/park/unpark/pop/resume) targets zero locks.
- **Lock ordering: io-before-shared, and slot locks are leaves** (never hold
two slot locks at once). Any new lock states its position in this order.
- **No unwind inside a runtime critical section.** The stop sentinel only
fires at lock-free observation points.
- **Never hold a thread-local guard across a possible switch point.** An
actor can be preempted at any allocation and resume on a different OS
thread; a live `RefCell` borrow (or std `MutexGuard`) then pairs its
acquire/release across two threads' locals — count underflow / UB.
`with_runtime`, `with_shared`, and every `RawMutex` guard disable
preemption for exactly this reason (phase 2 found this the hard way).
---
## Phase 1 — State decomposition: easy peel-offs ✅ DONE
Split independent concerns out of `SharedState` so the global lock guards less.
- [x] `next_monitor_id``AtomicU64` on `RuntimeInner` (lock-free id minting).
- [x] `timers` → own `Mutex<Timers>` (only drain-winner + blocking prims touch it).
- [x] `io` → own `Mutex<Option<IoThread>>` (completion queue already self-locked).
- [x] `pending_closures: Vec` → folded into `Slot::pending_closure` (per-actor data).
- [x] Termination check split: read io liveness *before* `shared`; ordering
argument documented in `schedule_loop`.
- [x] **Poison fix:** gate `check_cancelled()` in `maybe_preempt` behind
`PREEMPTION_ENABLED`, so a cancellation sentinel can never unwind while a
`std::sync::Mutex` (shared/channel) is held. Regression test:
`tests/poison_stop.rs`.
`SharedState` now holds only: `slots`, `free_list`, `run_queue`, `root_pid`.
## Phase 2 — Slot table split ✅ DONE
Make slot lookup lock-free and per-slot state independently mutable.
- [x] Fixed slab: `Box<[Slot]>`, slots never move → stable
addresses, lock-free index. **Assert on exhaustion** with a panic message
naming `Config::max_actors(n)` as the fix. Default `max_actors = 16_384`.
(Mental note: must become unbounded or configurable-bounded later — see
"Deferred" below. Do not let the fixed cap calcify into an assumption.)
- [x] Generation packed INTO the state word (better than the planned separate
`AtomicU32`): one `AtomicU64 = (gen << 32) | state`, so the gen check is
atomic with every transition — no ABA, no spurious unparks, by CAS.
- [x] Per-slot CAS state machine `Vacant/Queued/Running/RunningNotified/
Parked/Done` replacing `state` + `pending_unpark`. Also fixed a latent
lost wakeup in the Blocking-IO completion path (result set for a
still-Running actor without a flag).
- [x] `sp` → relaxed `AtomicUsize`; stop flag + first-resume closure as
`AtomicPtr`s — resume path fully lock-free.
- [x] Per-slot raw non-poisoning futex mutex (`src/raw_mutex.rs`) for the
cold collections. Guard enters `NoPreempt`.
- [x] Free list → `RawMutex<Vec<u32>>` (a leaf lock). Treiber/striped only if
the phase-4 spawn-storm bench shows contention — revisit then.
- [x] `finalize_actor` link cascade locks peers one at a time; acyclicity
argument written at the site. `link()` registers target-first with the
race argument at the site.
- [x] `live_actors` atomic; termination = io_out == 0 (read pre-queue-lock)
&& queue empty && live == 0; decrement-last ordering documented as the
correctness crux at the site.
## Phase 3 — Pluggable run queue ✅ DONE (shootout = phase 4 harness)
- [x] `RunQueue` alias in `src/run_queue.rs`, compile-time selected,
`compile_error!` guards for zero / >1 features. No runtime dispatch.
All variants compile in every build (unit tests always run); the
feature only picks the alias. Non-default variants:
`--no-default-features --features rq-…`.
- `rq-mutex` — `Mutex<VecDeque>`, the control/baseline.
- `rq-mpmc` — single hand-rolled Vyukov bounded MPMC ring
(per-cell seq numbers). Strict FIFO; "one hot cache line".
- `rq-striped` — M Vyukov rings, fetch-add ticket distribution. Relaxed
FIFO, reordering bounded by ~M. Predicted winner @20c.
- [x] Bounded rings sound via slab cap + at-most-once-enqueued (occupancy ≤
`max_actors`); mpmc capacity = next_pow2(max_actors), striped
Σcapacity ≈ 2×max_actors with probe-from-home-stripe. A full ring
panics as an invariant violation (double enqueue), never spins.
- [x] All hand-rolled, dependency-free.
- [x] Two contracts surfaced by the extraction, documented + debug-asserted
in `run_queue.rs`: queue ops require preemption disabled (a producer
suspended mid-publish stalls every consumer behind its cell — livelock);
and pop-None is a snapshot, not a fence — termination is counter-first
(`live == 0` alone implies the queue holds nothing actionable; argument
rewritten at the `schedule_loop` site).
## Phase 4 — Bench harness ✅ DONE (harness; real numbers from the 20-core box)
- [x] `benches/rq_micro.rs`: raw structures, threads × p:c ratio sweep. One
binary covers all three structures (types compile in every build).
- [x] `benches/rq_runtime.rs`: yield-storm, ping-pong-pairs, spawn-storm,
sweeping scheduler count; variant baked in by feature.
- [x] `scripts/bench_rq.sh`: rebuilds per `rq-*` feature, aggregates RQCSV
lines into `bench_results/summary.csv`. Knobs via SMARM_BENCH_* env;
e.g. `SMARM_BENCH_THREADS="1 2 4 8 16 20" ./scripts/bench_rq.sh`.
- [x] Harness validated end-to-end at smoke scale on the 1-core sandbox;
contention curves and the actual variant decision come from the
20-core box.
## Phase 5 — Safety hardening & model checking
- [ ] `loom` (dev-dependency only, x86 Linux) model tests for the slot state
machine and each ring variant.
- [ ] `NoPreempt` / no-unwind-under-lock audit across all internal critical
sections; assert the invariant in debug builds where feasible.
---
## Fast follow (post-v0.5, written down so it isn't lost)
- **Assert the invariants we lean on.** This cycle accumulated load-bearing
invariants: at-most-once-enqueued, queue-ops-under-NoPreempt, never two
cold locks, cold-path generation re-verify under the lock, finalize's
decrement-last, pushes-pair-with-Queued-transitions, thread-local guards
never crossing a switch point. Whenever code RELIES on one and a cheap
check exists, assert it at the point of reliance — `debug_assert!` on hot
paths, full `assert!`/loud panic on cold ones — so a violation fails at
the breakage site, not three modules downstream (the slab-overflow panic
and the queue-op preemption debug_assert are the pattern). Sweep the
existing code for missed spots; new code adopts it as house style. The
phase-5 audit is the natural vehicle for the sweep.
- **Channel mutex migration.** `channel::Inner<T>` is `Arc<Mutex<_>>` of the
same poison class as the old shared lock; `recv_match` even runs a user
predicate under it. The Phase-1 `check_cancelled` gating already removes the
unwind *source* globally, so channels are poison-safe today — but phase 2
surfaced a second, sharper reason to migrate: the std `MutexGuard` is held
with preemption ENABLED, so a timeslice switch inside a channel critical
section migrates the actor and unlocks the pthread mutex from a different
OS thread — technically UB (Linux futexes happen to tolerate it). The
`RawMutex` guard disables preemption and is cross-thread-release sound by
construction. Migrate as the first post-v0.5 change.
## Deferred / explicitly out of scope for v0.5
- **Unbounded / configurable-bounded actor count.** v0.5 ships a fixed slab
with a loud assert. Revisit with a segmented slab (array of
`AtomicPtr<Segment>`, doubling segment sizes, append-only) once we actually
hit the cap or a workload demands it.
- **Idle-wakeup eventcount.** Idle scheduler threads keep the current 100µs
poll-sleep. A futex-based eventcount is a later optimization if benches show
idle latency matters.
- **User-facing safe data structures** (ArcSwap-style cells, structurally-
shared persistent maps). Context for the runtime work, not in scope here.
- **IO fd hygiene on actor death** (pre-existing v0.2 TODO in `io.rs`).