Files
smarm/ROADMAP.md
T
Claude 4e4b617559 docs(roadmap): restructure around the rq shootout verdict
Shootout (6d9f369 harness, 24-core sweep) landed in RFC 005's World 3:
queue variants indistinguishable in rq_runtime; per-wake latency is the
ceiling. Record the queue-topology closure (rq-mutex frozen as default,
benchmark-driven reopening only), define v0.9 as the wake-path latency
cycle (RFC 005 slot -> bench -> RFC 004 spinning workers -> bench +
interaction pass), and add per-switch cost to Later pending a profiling
spike + RFC.

Completed cycles compacted to the two trailing minors (v0.7, v0.8);
earlier history lives in git.
2026-06-10 21:58:40 +00:00

9.6 KiB
Raw Blame History

smarm — Roadmap

Shipped (compacted — full cycle plans and deviation records live in git history)

Cycles before v0.7 (v0.4 actor primitives, v0.5 runtime decomposition & pluggable run queue, v0.6 actor ergonomics): see git log ROADMAP.md.

v0.7 — select, on epoch-stamped consuming wakes

A 24-bit park-epoch packed into the slot word gives every wait an identity: registrations carry (pid, epoch), every successful wake consumes the epoch, stale wakes die at one failed CAS. Subsumed the per-primitive wait seqs (channel, mutex, timer); the only wildcard wake left is request_stop (terminal). On top: select/select_timeout — ready-index wait over many receivers, priority order, no cancellation pass. Loom theorems re-proved on the new word. Notable deviations: retire_wait for the no-park exit path (plan missed it); the Drop-guard deregistration was superseded by relaxed single-receiver asserts; a closed arm is ready forever (documented gotcha). Commits 491383500128f3, a0a93b6.

v0.8 — gen_server: handle_info / handle_down + io fd hygiene

Spent select on the server loop: static info arms (type Info, ServerBuilder::with_info) and dynamic monitor forwarding (ServerCtx/Watcher + a control arm), priority downs → control → infos → inbox. Closed the v0.2 fd hole: a drop guard in wait_fd DELs the kernel registration on unwind (the leak was worse than documented — a stale waiters entry permanently poisoned the fd). Deviation: the plain-inbox fast path narrowed; servers holding a Watcher select forever. Commits e5d1b3b, 24b95c9, f6969e5.


Decision record — queue topology 🔒 CLOSED (2026-06-10)

The run-queue shootout (harness 6d9f369, 24-core sweep 124 schedulers, report bench_report_rq_shootout.html) landed in RFC 005's World 3: the three queue variants are within 1015% of each other in rq_runtime at every scheduler count ≥ 4, on all three workloads. Striped's real micro-bench dominance (24.8 M ops/s at 24t, 4.3× over mutex) does not propagate to the runtime because the queue is not the hot path in any measured workload. USL fits put σ at 5.57.6 (super-serial) for ping-pong: the ceiling is per-wake latency through the park/unpark protocol, not queue push/pop.

Consequences:

  • rq-mutex stays the default — simplest correct, no capacity constraints, locking model already integrated with the preemption-disabled queue-op invariant.
  • Feature plumbing stays as is. All three variants keep compiling in every build; rq-mpmc/rq-striped remain selectable for benching.
  • Reopening is benchmark-driven only. The report documents the conditional upgrade paths if a future workload qualifies: mpmc for message-passing-dominant loads at N ≤ 8 (1.3× lower σ, 0.16 µs ping-pong at 1t); striped for high-contention balanced push/pop at N ≥ 16 — neither is a scheduler workload as measured.
  • Effort redirects to the wake path: RFC 005 (billed as a latency patch, per its own World 3 framing), RFC 004, and eventually per-switch cost.

v0.9 — Wake-path latency

Goal: attack per-wake latency — the measured ceiling — with the two specified mechanisms, each benched against a clean baseline before the next lands. Order: RFC 005 first (the slot is measured against the just-benched idle policy), then RFC 004 (measured against the slot-accepted baseline), then an interaction pass.

Full specs: rfc_005-wake-slot.md, rfc_004-tunable-scheduler-idle-policy.md (artefacts.kalsbeek.dev).

1. Wake slot (RFC 005)

Per-scheduler, thread-local, capacity-one wake cache, checked before the shared queue. Runtime-selected via Config { wake_slot: bool }, default off until accepted (one binary benches both arms).

  • Push policy: slot-eligible iff the wake originates from actor context (current_pid().is_some()). Scheduler-context wakes (timer/IO drain) and spawns always go shared.
  • Displacement: newest wake takes the slot, occupant pushed shared (Go semantics); the branch swap is the fallback if benches look pathological.
  • Pop order: slot, then shared. Slot-popped actors inherit the waker's remaining timeslice — a handoff chain is bounded by one slice, so the shared queue is consulted at least once per slice per scheduler (the starvation bound, zero new counters).
  • Invariant care: the slot push replaces run_queue.push at the tail of the Parked → Queued CAS; at-most-once-enqueued holds verbatim (pid in slot ⊕ shared queue). Stall blast radius grows by exactly one actor.
  • Counters: slot_hits, slot_displacements.

2. Slot shootout

Extend rq_runtime/bench_rq.sh with the slot on/off dimension (cheap — it's a Config knob, not a feature rebuild). Same sweep as the rq shootout.

  • ping-pong-pairs — target metric; expect the win, single- and multi-scheduler.
  • yield-storm — regression guard; yields never touch the slot, any delta is pop-path overhead.
  • spawn-storm — neutrality check; spawns bypass the slot by policy. Acceptance flips the default on and re-baselines for RFC 004.

3. Spinning workers (RFC 004)

Bounded spin-before-park for idle schedulers, killing the ~100 µs thread::sleep worst case on cross-thread handoff. Two Config knobs: spin_budget_cycles (0 recovers today's behaviour) and max_spinners (default N/2); one new atomic n_spinning. No run-queue changes — lands on the frozen rq-mutex substrate, independent of the slot mechanically.

4. RFC 004 bench + interaction pass

Re-run the sweep with spinning enabled, slot on and off. The known interaction: once idle pickup latency ≪ slice, the slot's latency trade (occupant waits out the waker's slice while other schedulers idle) may stop paying — RFC 005's documented escape hatch is Go's "spinners exist → push shared instead" check. In scope only if the data demands it.


Later

Per-switch cost (context shims, epoch protocol)

The shootout's residual: per-wake latency is 0.160.18 µs at N=1 and 0.81.2 µs at N=8+, dominated by the context-switch shims and the epoch protocol, not the queue. On current evidence this is the larger constant — "the whole game" alongside the v0.9 work — but there is no spec yet. Needs a profiling spike (where do the cycles actually go per park/unpark round-trip) and then an RFC before it can be scheduled.

Unbounded / configurable-bounded actor count

Fixed slab with a loud assert (Config::max_actors(n), default 16 384). Revisit with a segmented slab (array of AtomicPtr<Segment>, doubling segment sizes, append-only) once the cap is actually hit. Do not let it calcify.

arm-port validation & merge

arm-port branch carries an AAPCS64 context-switch backend, never run on hardware. Build + run full test suite on an aarch64 device; check chained_spawn / yield_many bench medians; merge and update README.


Invariants & gotchas (respect these across all cycles)

  • Shared mutex is non-reentrant. Sender::send can call unparkwith_shared. Never send on a channel while holding the shared lock. Pattern: mem::take data under the lock, send after releasing. See finalize_actor.
  • finalize_actor order: take stack/waiters/monitors under lock + set Done/outcome → recycle stack → deliver supervisor Signal + monitor Downs → unpark joiners → reclaim slot if outstanding_handles==0. Death notifications always precede reclamation.
  • Slot lifecycle reset in THREE places: Slot::vacant(), reclaim_slot() (runtime.rs), slot-init block in spawn_under (scheduler.rs). Any new Slot field must be reset in all three.
  • Pid = (index, generation). Stale handles caught by generation mismatch in slot()/slot_mut(). The monitor NoProc path relies on this.
  • The only wildcard wake is request_stop, and it is terminal. Every registration-based waker (channel sends, mutex grants, wait-timers, io completions, joiner wakes, select arms) carries the wait's park-epoch and wakes through unpark_at; every successful wake consumes the epoch. Wakes are therefore meaningful: one-shot park sites interpret them without loops, and select needs no cancellation pass. When adding a new waker, decide which form it is — if its registration handle can outlive the wait it was created for, it MUST be epoch-stamped; a wait that can exit without parking MUST retire_wait first (see slot_state.rs).
  • select exists; a unified per-process mailbox still does not. The supervisor keeps its single supervisor_channel funnel; recv_match stays per-channel. select composes channels at the wait, not into one queue — gen_server's handle_info/handle_down (v0.8) are built on exactly that composition, with documented arm priority (downs → control → infos → inbox) instead of mailbox FIFO. A hot higher-priority arm starves lower ones by design; that's the contract.
  • Cooperative-only. Preemption and cancellation both depend on the actor reaching check!()/yield/alloc/blocking points.
  • Lock order is Leaf → Channel, one of each at most (debug-asserted in raw_mutex.rs). Leaf = cold locks / free list / stack pool / registry, mutual leaves. A channel lock may be taken under a Leaf (finalize/monitor clone senders living in slots); nothing may be locked under a channel lock.
  • Queue ops require preemption disabled. A producer suspended mid-publish stalls every consumer — livelock. with_runtime, with_shared, and RawMutex guards all disable preemption for their span.
  • run() is single-thread (Config::exact(1)); tests rely on deterministic single-thread ordering. Multi-thread via runtime::init(Config…).