# smarm — Roadmap ## Shipped (compacted — full cycle plans and deviation records live in git history) Cycles before v0.7 (v0.4 actor primitives, v0.5 runtime decomposition & pluggable run queue, v0.6 actor ergonomics): see `git log ROADMAP.md`. ### v0.7 — select, on epoch-stamped consuming wakes ✅ A 24-bit park-epoch packed into the slot word gives every wait an identity: registrations carry `(pid, epoch)`, every successful wake **consumes** the epoch, stale wakes die at one failed CAS. Subsumed the per-primitive wait seqs (channel, mutex, timer); the only wildcard wake left is `request_stop` (terminal). On top: `select`/`select_timeout` — ready-index wait over many receivers, priority order, no cancellation pass. Loom theorems re-proved on the new word. Notable deviations: `retire_wait` for the no-park exit path (plan missed it); the Drop-guard deregistration was superseded by relaxed single-receiver asserts; a closed arm is ready *forever* (documented gotcha). Commits `4913835`…`00128f3`, `a0a93b6`. ### v0.8 — gen_server: handle_info / handle_down + io fd hygiene ✅ Spent `select` on the server loop: static info arms (`type Info`, `ServerBuilder::with_info`) and dynamic monitor forwarding (`ServerCtx`/`Watcher` + a control arm), priority downs → control → infos → inbox. Closed the v0.2 fd hole: a drop guard in `wait_fd` DELs the kernel registration on unwind (the leak was worse than documented — a stale waiters entry permanently poisoned the fd). Deviation: the plain-inbox fast path narrowed; servers holding a `Watcher` select forever. Commits `e5d1b3b`, `24b95c9`, `f6969e5`. --- ## Decision record — queue topology 🔒 CLOSED (2026-06-10) The run-queue shootout (harness `6d9f369`, 24-core sweep 1–24 schedulers, report `bench_report_rq_shootout.html`) landed in **RFC 005's World 3**: the three queue variants are within 10–15% of each other in `rq_runtime` at every scheduler count ≥ 4, on all three workloads. Striped's real micro-bench dominance (24.8 M ops/s at 24t, 4.3× over mutex) does not propagate to the runtime because the queue is not the hot path in any measured workload. USL fits put σ at 5.5–7.6 (super-serial) for ping-pong: the ceiling is per-wake latency through the park/unpark protocol, not queue push/pop. Consequences: - **`rq-mutex` stays the default** — simplest correct, no capacity constraints, locking model already integrated with the preemption-disabled queue-op invariant. - **Feature plumbing stays as is.** All three variants keep compiling in every build; `rq-mpmc`/`rq-striped` remain selectable for benching. - **Reopening is benchmark-driven only.** The report documents the conditional upgrade paths if a future workload qualifies: mpmc for message-passing-dominant loads at N ≤ 8 (1.3× lower σ, 0.16 µs ping-pong at 1t); striped for high-contention balanced push/pop at N ≥ 16 — neither is a scheduler workload as measured. - **Effort redirects to the wake path**: RFC 005 (billed as a latency patch, per its own World 3 framing), RFC 004, and eventually per-switch cost. --- ## v0.9 — Wake-path latency Goal: attack per-wake latency — the measured ceiling — with the two specified mechanisms, each benched against a clean baseline before the next lands. Order: RFC 005 first (the slot is measured against the just-benched idle policy), then RFC 004 (measured against the slot-accepted baseline), then an interaction pass. Full specs: `rfc_005-wake-slot.md`, `rfc_004-tunable-scheduler-idle-policy.md` (artefacts.kalsbeek.dev). ### 1. Wake slot (RFC 005) Per-scheduler, thread-local, capacity-one wake cache, checked before the shared queue. Runtime-selected via `Config { wake_slot: bool }`, default off until accepted (one binary benches both arms). - **Push policy:** slot-eligible iff the wake originates from actor context (`current_pid().is_some()`). Scheduler-context wakes (timer/IO drain) and spawns always go shared. - **Displacement:** newest wake takes the slot, occupant pushed shared (Go semantics); the branch swap is the fallback if benches look pathological. - **Pop order:** slot, then shared. Slot-popped actors inherit the waker's remaining timeslice — a handoff chain is bounded by one slice, so the shared queue is consulted at least once per slice per scheduler (the starvation bound, zero new counters). - **Invariant care:** the slot push replaces `run_queue.push` at the tail of the `Parked → Queued` CAS; at-most-once-enqueued holds verbatim (pid in slot ⊕ shared queue). Stall blast radius grows by exactly one actor. - Counters: `slot_hits`, `slot_displacements`. ### 2. Slot shootout Extend `rq_runtime`/`bench_rq.sh` with the slot on/off dimension (cheap — it's a Config knob, not a feature rebuild). Same sweep as the rq shootout. - **ping-pong-pairs** — target metric; expect the win, single- and multi-scheduler. - **yield-storm** — regression guard; yields never touch the slot, any delta is pop-path overhead. - **spawn-storm** — neutrality check; spawns bypass the slot by policy. Acceptance flips the default on and re-baselines for RFC 004. ### 3. Spinning workers (RFC 004) Bounded spin-before-park for idle schedulers, killing the ~100 µs `thread::sleep` worst case on cross-thread handoff. Two Config knobs: `spin_budget_cycles` (0 recovers today's behaviour) and `max_spinners` (default N/2); one new atomic `n_spinning`. No run-queue changes — lands on the frozen `rq-mutex` substrate, independent of the slot mechanically. ### 4. RFC 004 bench + interaction pass Re-run the sweep with spinning enabled, slot on and off. The known interaction: once idle pickup latency ≪ slice, the slot's latency trade (occupant waits out the waker's slice while other schedulers idle) may stop paying — RFC 005's documented escape hatch is Go's "spinners exist → push shared instead" check. In scope only if the data demands it. --- ## Later ### Unwakeable idle sleep when io is absent (terminal-wake residual) The `(Some(deadline), None)` idle branch — timers pending, io subsystem never initialized — blocks in `thread::sleep` with no wake mechanism at all. The terminal wake (writes the wake pipe at AllDone) cannot reach it: no io, no pipe. Same stall as the fixed bug, in any no-io runtime: a sibling that blocked on an orphaned deadline sleeps it out in full after everything else finished. Candidates, mutually exclusive: (a) clamp the sleep (cheap, but turns idle into periodic wakeups), or (b) park the branch on a condvar/futex the AllDone path signals — and at that point consider making the condvar the idle primitive for the no-io runtime generally (a cross-thread unpark could signal it too, see below). Decide before any no-io deployment. ### Cross-thread unpark `RuntimeInner::enqueue` does not wake idle sibling schedulers — only io completions write the wake pipe. Mid-flight this is masked (the enqueuing thread is awake and eats the work itself), but it costs parallelism: work enqueued by a busy thread waits until the sibling's idle poll times out. Candidate from the urus chunk-2 session; needs bench evidence (does the shared-queue handoff latency actually show up?) before a mechanism is picked. ### Per-switch cost (context shims, epoch protocol) The shootout's residual: per-wake latency is 0.16–0.18 µs at N=1 and 0.8–1.2 µs at N=8+, dominated by the context-switch shims and the epoch protocol, not the queue. On current evidence this is the larger constant — "the whole game" alongside the v0.9 work — but there is no spec yet. Needs a profiling spike (where do the cycles actually go per park/unpark round-trip) and then an RFC before it can be scheduled. ### Unbounded / configurable-bounded actor count Fixed slab with a loud assert (`Config::max_actors(n)`, default 16 384). Revisit with a segmented slab (array of `AtomicPtr`, doubling segment sizes, append-only) once the cap is actually hit. Do not let it calcify. ### arm-port validation & merge `arm-port` branch carries an AAPCS64 context-switch backend, never run on hardware. Build + run full test suite on an aarch64 device; check `chained_spawn` / `yield_many` bench medians; merge and update README. --- ## Invariants & gotchas (respect these across all cycles) - **Shared mutex is non-reentrant.** `Sender::send` can call `unpark` → `with_shared`. Never send on a channel while holding the shared lock. Pattern: `mem::take` data under the lock, send after releasing. See `finalize_actor`. - **`finalize_actor` order:** take stack/waiters/monitors under lock + set Done/outcome → recycle stack → deliver supervisor Signal + monitor Downs → unpark joiners → reclaim slot if `outstanding_handles==0`. Death notifications always precede reclamation. - **Slot lifecycle reset in THREE places:** `Slot::vacant()`, `reclaim_slot()` (runtime.rs), slot-init block in `spawn_under` (scheduler.rs). Any new `Slot` field must be reset in all three. - **Pid = (index, generation).** Stale handles caught by generation mismatch in `slot()/slot_mut()`. The monitor `NoProc` path relies on this. - **The only wildcard wake is `request_stop`, and it is terminal.** Every registration-based waker (channel sends, mutex grants, wait-timers, io completions, joiner wakes, `select` arms) carries the wait's park-epoch and wakes through `unpark_at`; every successful wake consumes the epoch. Wakes are therefore *meaningful*: one-shot park sites interpret them without loops, and `select` needs no cancellation pass. When adding a new waker, decide which form it is — if its registration handle can outlive the wait it was created for, it MUST be epoch-stamped; a wait that can exit without parking MUST `retire_wait` first (see slot_state.rs). - **`select` exists; a unified per-process mailbox still does not.** The supervisor keeps its single `supervisor_channel` funnel; `recv_match` stays per-channel. `select` composes channels at the wait, not into one queue — gen_server's `handle_info`/`handle_down` (v0.8) are built on exactly that composition, with documented arm priority (downs → control → infos → inbox) instead of mailbox FIFO. A hot higher-priority arm starves lower ones by design; that's the contract. - **Cooperative-only.** Preemption and cancellation both depend on the actor reaching `check!()`/yield/alloc/blocking points. - **Lock order is Leaf → Channel, one of each at most** (debug-asserted in `raw_mutex.rs`). Leaf = cold locks / free list / stack pool / registry, mutual leaves. A channel lock may be taken under a Leaf (finalize/monitor clone senders living in slots); nothing may be locked under a channel lock. - **Queue ops require preemption disabled.** A producer suspended mid-publish stalls every consumer — livelock. `with_runtime`, `with_shared`, and `RawMutex` guards all disable preemption for their span. - **`run()` is single-thread** (`Config::exact(1)`); tests rely on deterministic single-thread ordering. Multi-thread via `runtime::init(Config…)`.