Shootout (6d9f369 harness, 24-core sweep) landed in RFC 005's World 3:
queue variants indistinguishable in rq_runtime; per-wake latency is the
ceiling. Record the queue-topology closure (rq-mutex frozen as default,
benchmark-driven reopening only), define v0.9 as the wake-path latency
cycle (RFC 005 slot -> bench -> RFC 004 spinning workers -> bench +
interaction pass), and add per-switch cost to Later pending a profiling
spike + RFC.
Completed cycles compacted to the two trailing minors (v0.7, v0.8);
earlier history lives in git.
177 lines
9.6 KiB
Markdown
177 lines
9.6 KiB
Markdown
# smarm — Roadmap
|
||
|
||
## Shipped (compacted — full cycle plans and deviation records live in git history)
|
||
|
||
Cycles before v0.7 (v0.4 actor primitives, v0.5 runtime decomposition &
|
||
pluggable run queue, v0.6 actor ergonomics): see `git log ROADMAP.md`.
|
||
|
||
### v0.7 — select, on epoch-stamped consuming wakes ✅
|
||
A 24-bit park-epoch packed into the slot word gives every wait an identity:
|
||
registrations carry `(pid, epoch)`, every successful wake **consumes** the
|
||
epoch, stale wakes die at one failed CAS. Subsumed the per-primitive wait
|
||
seqs (channel, mutex, timer); the only wildcard wake left is `request_stop`
|
||
(terminal). On top: `select`/`select_timeout` — ready-index wait over many
|
||
receivers, priority order, no cancellation pass. Loom theorems re-proved on
|
||
the new word. Notable deviations: `retire_wait` for the no-park exit path
|
||
(plan missed it); the Drop-guard deregistration was superseded by relaxed
|
||
single-receiver asserts; a closed arm is ready *forever* (documented gotcha).
|
||
Commits `4913835`…`00128f3`, `a0a93b6`.
|
||
|
||
### v0.8 — gen_server: handle_info / handle_down + io fd hygiene ✅
|
||
Spent `select` on the server loop: static info arms (`type Info`,
|
||
`ServerBuilder::with_info`) and dynamic monitor forwarding
|
||
(`ServerCtx`/`Watcher` + a control arm), priority downs → control → infos →
|
||
inbox. Closed the v0.2 fd hole: a drop guard in `wait_fd` DELs the kernel
|
||
registration on unwind (the leak was worse than documented — a stale waiters
|
||
entry permanently poisoned the fd). Deviation: the plain-inbox fast path
|
||
narrowed; servers holding a `Watcher` select forever.
|
||
Commits `e5d1b3b`, `24b95c9`, `f6969e5`.
|
||
|
||
---
|
||
|
||
## Decision record — queue topology 🔒 CLOSED (2026-06-10)
|
||
|
||
The run-queue shootout (harness `6d9f369`, 24-core sweep 1–24 schedulers,
|
||
report `bench_report_rq_shootout.html`) landed in **RFC 005's World 3**: the
|
||
three queue variants are within 10–15% of each other in `rq_runtime` at every
|
||
scheduler count ≥ 4, on all three workloads. Striped's real micro-bench
|
||
dominance (24.8 M ops/s at 24t, 4.3× over mutex) does not propagate to the
|
||
runtime because the queue is not the hot path in any measured workload. USL
|
||
fits put σ at 5.5–7.6 (super-serial) for ping-pong: the ceiling is per-wake
|
||
latency through the park/unpark protocol, not queue push/pop.
|
||
|
||
Consequences:
|
||
- **`rq-mutex` stays the default** — simplest correct, no capacity
|
||
constraints, locking model already integrated with the preemption-disabled
|
||
queue-op invariant.
|
||
- **Feature plumbing stays as is.** All three variants keep compiling in
|
||
every build; `rq-mpmc`/`rq-striped` remain selectable for benching.
|
||
- **Reopening is benchmark-driven only.** The report documents the
|
||
conditional upgrade paths if a future workload qualifies: mpmc for
|
||
message-passing-dominant loads at N ≤ 8 (1.3× lower σ, 0.16 µs ping-pong
|
||
at 1t); striped for high-contention balanced push/pop at N ≥ 16 — neither
|
||
is a scheduler workload as measured.
|
||
- **Effort redirects to the wake path**: RFC 005 (billed as a latency patch,
|
||
per its own World 3 framing), RFC 004, and eventually per-switch cost.
|
||
|
||
---
|
||
|
||
## v0.9 — Wake-path latency
|
||
|
||
Goal: attack per-wake latency — the measured ceiling — with the two specified
|
||
mechanisms, each benched against a clean baseline before the next lands.
|
||
Order: RFC 005 first (the slot is measured against the just-benched idle
|
||
policy), then RFC 004 (measured against the slot-accepted baseline), then an
|
||
interaction pass.
|
||
|
||
Full specs: `rfc_005-wake-slot.md`, `rfc_004-tunable-scheduler-idle-policy.md`
|
||
(artefacts.kalsbeek.dev).
|
||
|
||
### 1. Wake slot (RFC 005)
|
||
Per-scheduler, thread-local, capacity-one wake cache, checked before the
|
||
shared queue. Runtime-selected via `Config { wake_slot: bool }`, default off
|
||
until accepted (one binary benches both arms).
|
||
- **Push policy:** slot-eligible iff the wake originates from actor context
|
||
(`current_pid().is_some()`). Scheduler-context wakes (timer/IO drain) and
|
||
spawns always go shared.
|
||
- **Displacement:** newest wake takes the slot, occupant pushed shared (Go
|
||
semantics); the branch swap is the fallback if benches look pathological.
|
||
- **Pop order:** slot, then shared. Slot-popped actors inherit the waker's
|
||
remaining timeslice — a handoff chain is bounded by one slice, so the
|
||
shared queue is consulted at least once per slice per scheduler (the
|
||
starvation bound, zero new counters).
|
||
- **Invariant care:** the slot push replaces `run_queue.push` at the tail of
|
||
the `Parked → Queued` CAS; at-most-once-enqueued holds verbatim (pid in
|
||
slot ⊕ shared queue). Stall blast radius grows by exactly one actor.
|
||
- Counters: `slot_hits`, `slot_displacements`.
|
||
|
||
### 2. Slot shootout
|
||
Extend `rq_runtime`/`bench_rq.sh` with the slot on/off dimension (cheap —
|
||
it's a Config knob, not a feature rebuild). Same sweep as the rq shootout.
|
||
- **ping-pong-pairs** — target metric; expect the win, single- and
|
||
multi-scheduler.
|
||
- **yield-storm** — regression guard; yields never touch the slot, any delta
|
||
is pop-path overhead.
|
||
- **spawn-storm** — neutrality check; spawns bypass the slot by policy.
|
||
Acceptance flips the default on and re-baselines for RFC 004.
|
||
|
||
### 3. Spinning workers (RFC 004)
|
||
Bounded spin-before-park for idle schedulers, killing the ~100 µs
|
||
`thread::sleep` worst case on cross-thread handoff. Two Config knobs:
|
||
`spin_budget_cycles` (0 recovers today's behaviour) and `max_spinners`
|
||
(default N/2); one new atomic `n_spinning`. No run-queue changes — lands on
|
||
the frozen `rq-mutex` substrate, independent of the slot mechanically.
|
||
|
||
### 4. RFC 004 bench + interaction pass
|
||
Re-run the sweep with spinning enabled, slot on and off. The known
|
||
interaction: once idle pickup latency ≪ slice, the slot's latency trade
|
||
(occupant waits out the waker's slice while other schedulers idle) may stop
|
||
paying — RFC 005's documented escape hatch is Go's "spinners exist → push
|
||
shared instead" check. In scope only if the data demands it.
|
||
|
||
---
|
||
|
||
## Later
|
||
|
||
### Per-switch cost (context shims, epoch protocol)
|
||
The shootout's residual: per-wake latency is 0.16–0.18 µs at N=1 and
|
||
0.8–1.2 µs at N=8+, dominated by the context-switch shims and the epoch
|
||
protocol, not the queue. On current evidence this is the larger constant —
|
||
"the whole game" alongside the v0.9 work — but there is no spec yet. Needs a
|
||
profiling spike (where do the cycles actually go per park/unpark round-trip)
|
||
and then an RFC before it can be scheduled.
|
||
|
||
### Unbounded / configurable-bounded actor count
|
||
Fixed slab with a loud assert (`Config::max_actors(n)`, default 16 384).
|
||
Revisit with a segmented slab (array of `AtomicPtr<Segment>`, doubling segment
|
||
sizes, append-only) once the cap is actually hit. Do not let it calcify.
|
||
|
||
### arm-port validation & merge
|
||
`arm-port` branch carries an AAPCS64 context-switch backend, never run on
|
||
hardware. Build + run full test suite on an aarch64 device; check
|
||
`chained_spawn` / `yield_many` bench medians; merge and update README.
|
||
|
||
---
|
||
|
||
## Invariants & gotchas (respect these across all cycles)
|
||
|
||
- **Shared mutex is non-reentrant.** `Sender::send` can call `unpark` →
|
||
`with_shared`. Never send on a channel while holding the shared lock. Pattern:
|
||
`mem::take` data under the lock, send after releasing. See `finalize_actor`.
|
||
- **`finalize_actor` order:** take stack/waiters/monitors under lock + set
|
||
Done/outcome → recycle stack → deliver supervisor Signal + monitor Downs →
|
||
unpark joiners → reclaim slot if `outstanding_handles==0`. Death notifications
|
||
always precede reclamation.
|
||
- **Slot lifecycle reset in THREE places:** `Slot::vacant()`, `reclaim_slot()`
|
||
(runtime.rs), slot-init block in `spawn_under` (scheduler.rs). Any new `Slot`
|
||
field must be reset in all three.
|
||
- **Pid = (index, generation).** Stale handles caught by generation mismatch in
|
||
`slot()/slot_mut()`. The monitor `NoProc` path relies on this.
|
||
- **The only wildcard wake is `request_stop`, and it is terminal.** Every
|
||
registration-based waker (channel sends, mutex grants, wait-timers, io
|
||
completions, joiner wakes, `select` arms) carries the wait's park-epoch
|
||
and wakes through `unpark_at`; every successful wake consumes the epoch.
|
||
Wakes are therefore *meaningful*: one-shot park sites interpret them
|
||
without loops, and `select` needs no cancellation pass. When adding a new
|
||
waker, decide which form it is — if its registration handle can outlive
|
||
the wait it was created for, it MUST be epoch-stamped; a wait that can
|
||
exit without parking MUST `retire_wait` first (see slot_state.rs).
|
||
- **`select` exists; a unified per-process mailbox still does not.** The
|
||
supervisor keeps its single `supervisor_channel` funnel; `recv_match`
|
||
stays per-channel. `select` composes channels at the wait, not into one
|
||
queue — gen_server's `handle_info`/`handle_down` (v0.8) are built on
|
||
exactly that composition, with documented arm priority (downs → control
|
||
→ infos → inbox) instead of mailbox FIFO. A hot higher-priority arm
|
||
starves lower ones by design; that's the contract.
|
||
- **Cooperative-only.** Preemption and cancellation both depend on the actor
|
||
reaching `check!()`/yield/alloc/blocking points.
|
||
- **Lock order is Leaf → Channel, one of each at most** (debug-asserted in
|
||
`raw_mutex.rs`). Leaf = cold locks / free list / stack pool / registry,
|
||
mutual leaves. A channel lock may be taken under a Leaf (finalize/monitor
|
||
clone senders living in slots); nothing may be locked under a channel lock.
|
||
- **Queue ops require preemption disabled.** A producer suspended mid-publish
|
||
stalls every consumer — livelock. `with_runtime`, `with_shared`, and
|
||
`RawMutex` guards all disable preemption for their span.
|
||
- **`run()` is single-thread** (`Config::exact(1)`); tests rely on deterministic
|
||
single-thread ordering. Multi-thread via `runtime::init(Config…)`.
|