Files
smarm/ROADMAP.md
T
smarm-agent 4d4a2a6c9b roadmap: ship v0.9, promote Process groups to up-next
- Compact v0.9 (Wake-path latency) into Shipped; record the RFC 004 spinning
  excision as a deviation (preserved on branch rfc-004-spinning).
- Drop the v0.7 entry, folding its reference into the Shipped history pointer.
- Promote Process groups from Later into the up-next slot v0.9 vacated;
  Per-switch cost now leads Later/Highest priority.
2026-06-15 12:39:30 +00:00

12 KiB
Raw Blame History

smarm — Roadmap

Shipped (compacted — full cycle plans and deviation records live in git history)

Cycles before v0.8 (v0.4 actor primitives, v0.5 runtime decomposition & pluggable run queue, v0.6 actor ergonomics, v0.7 select on epoch-stamped consuming wakes): see git log ROADMAP.md.

v0.8 — gen_server: handle_info / handle_down + io fd hygiene

Spent select on the server loop: static info arms (type Info, ServerBuilder::with_info) and dynamic monitor forwarding (ServerCtx/Watcher + a control arm), priority downs → control → infos → inbox. Closed the v0.2 fd hole: a drop guard in wait_fd DELs the kernel registration on unwind (the leak was worse than documented — a stale waiters entry permanently poisoned the fd). Deviation: the plain-inbox fast path narrowed; servers holding a Watcher select forever. Commits e5d1b3b, 24b95c9, f6969e5.

v0.9 — Wake-path latency

Attacked per-wake latency with the RFC 005 wake slot: a per-scheduler, thread-local, capacity-one wake cache checked before the shared queue, pushed only from actor context, slot-then-shared pop with the waker's residual slice as the starvation bound (slot_hits/slot_displacements). Benched via the slot on/off dimension of rq_runtime — ping-pong-pairs (win), yield-storm (regression guard), spawn-storm (neutrality); results annotated in RFC 005. The RFC 004 spinning-workers experiment, originally scoped here, was evaluated and excised (not worth the code cost; preserved on branch rfc-004-spinning). Also a false-sharing fix (align(64) on SchedulerStats) and a termination wake for idle siblings. Commits 2708042, 37d9319, eddf3fe.


Decision record — queue topology 🔒 CLOSED (2026-06-10)

The run-queue shootout (harness 6d9f369, 24-core sweep 124 schedulers, report bench_report_rq_shootout.html) landed in RFC 005's World 3: the three queue variants are within 1015% of each other in rq_runtime at every scheduler count ≥ 4, on all three workloads. Consequences:

  • rq-mutex stays the default — simplest correct, no capacity constraints, locking model already integrated.
  • Feature plumbing stays as is. All three variants keep compiling in every build; rq-mpmc/rq-striped remain selectable for benching.
  • Reopening is benchmark-driven only. The report documents the conditional upgrade paths if a future workload qualifies: mpmc for message-passing-dominant loads at N ≤ 8; striped for high-contention balanced push/pop at N ≥ 16. Neither is a scheduler workload as measured.
  • Effort redirects to the wake path: RFC 005 (billed as a latency patch, per its own World 3 framing), RFC 004, and eventually per-switch cost.

Up next — Process groups: the primitive pubsub & channels should have sat on

Context: urus is a webserver written on top of smarm to provide a testing target. A named pid→multiset map with monitor-backed removal: registry.rs generalised from name↔pid bimap to name→multiset, the death hook reused verbatim. Local first. The point is that urus pubsub collapses into a pg consumer (subscribe = join, broadcast = send-to-members) instead of being a bespoke mechanism, and the same group set reads two ways — fan-out (all members) vs discovery/pool (one member), with different netsplit consequences. Urus shipped pubsub/channels predate this and want reframing on top of it. Foundational, so early in the post-v0.9 stack. Needs an RFC.


Later

Highest priority

Per-switch cost (context shims, epoch protocol)

The shootout's residual: per-wake latency is 0.160.18 µs at N=1 and 0.81.2 µs at N=8+, dominated by the context-switch shims and the epoch protocol, not the queue. On current evidence this is the larger constant — "the whole game" alongside the v0.9 work — but there is no spec yet. Needs a profiling spike (where do the cycles actually go per park/unpark round-trip) and then an RFC before it can be scheduled.

send_after / cancel_timer

Message-delivery timer on the existing min-heap (timer.rs): deliver a value to a channel at a deadline, cancellable. Unlocks the gen_server idioms with no clean expression today — heartbeat, debounce, retry backoff, session expiry. Small;

Introspection — process_info / get_state / tree dump

trace.rs is the seed. What an actor is parked on, queue depth, stack size; a gen_server state snapshot; a supervision-tree walk. The native edge over tokio — actors already carry pid, name, parent where tokio tasks are anonymous — so it costs little and differentiates a lot. Needs an RFC.

Worker pool behaviour

Supervised, interchangeable workers with restart semantics over a shared inbox (poolboy / NimblePool shape) — distinct from connection pools (bb8/deadpool), which pool resources, not supervised processes. Sits on supervisor.rs + gen_server.rs. Needs an RFC.

Medium Priority

Demand-driven pipelines — GenStage / Broadway shape

Supervised producer/consumer stages where consumers signal demand upstream, with batching, ack, partitioning. The clearest thing hex has and crates.io lacks (stream combinators and bounded channels are not a supervised demand-contract stage graph), and the natural fit for ingestion-shaped workloads. Builds on channels + gen_server

  • supervisor. Needs an RFC.

Unwakeable idle sleep when io is absent (terminal-wake residual)

The (Some(deadline), None) idle branch — timers pending, io subsystem never initialized — blocks in thread::sleep with no wake mechanism at all. The terminal wake (writes the wake pipe at AllDone) cannot reach it: no io, no pipe. Same stall as the fixed bug, in any no-io runtime: a sibling that blocked on an orphaned deadline sleeps it out in full after everything else finished. Candidates, mutually exclusive: (a) clamp the sleep (cheap, but turns idle into periodic wakeups), or (b) park the branch on a condvar/futex the AllDone path signals — and at that point consider making the condvar the idle primitive for the no-io runtime generally (a cross-thread unpark could signal it too, see below). Decide before any no-io deployment.

Cross-thread unpark

RuntimeInner::enqueue does not wake idle sibling schedulers — only io completions write the wake pipe. Mid-flight this is masked (the enqueuing thread is awake and eats the work itself), but it costs parallelism: work enqueued by a busy thread waits until the sibling's idle poll times out. Needs bench evidence (does the shared-queue handoff latency actually show up?) before a mechanism is picked.

Unbounded / configurable-bounded actor count

Fixed slab with a loud assert (Config::max_actors(n), default 16 384). Revisit with a segmented slab (array of AtomicPtr<Segment>, doubling segment sizes, append-only) once the cap is actually hit. Do not let it calcify.

arm-port validation & merge

arm-port branch carries an AAPCS64 context-switch backend, never run on hardware. Build + run full test suite on an aarch64 device; check chained_spawn / yield_many bench medians; merge and update README.

Low priority

gen_statem — postponement + state timeouts only

A thin layer over gen_server, not a new behaviour. The (state, event) dispatch matrix is free from the type system and not worth porting. The two mechanisms that are: event postponement (defer events in the wrong state, replay on transition — selective receive, codified) and state timeouts (auto-cancel on state change). Device-connection FSMs are the canonical use. Wants send_after underneath. Needs an RFC.

Clustering — distribution epic

Sequenced deliberately after v0.9 and the per-switch-cost spike. A fat stack of RFCs, not one. Spine settled in discussion; decisions still open:

  • Explicit remote boundary, never transparency. Serialization colours edges (channel types), not functions — local edges stay zero-copy Send, only remote edges take a RemoteRef<T: Serialize + DeserializeOwned>. No hidden latency when a peer migrates; the refactor is visible by construction.
  • One binary, role as runtime config (ROLE=… REGION=… SEEDS=…); a build-hash handshake enforces same-binary type identity and sidesteps cross-version type agreement. Roles select which supervision subtree mounts.
  • Distributed pg falls out of local pg + a membership/gossip layer, and distributed pubsub falls out of that for free; per-member metadata (region, load) enables fly-style nearest-member routing.
  • Migratable gen_servers as a sub-layer: only behaviours migrate (a raw actor's stack is opaque; a gen_server between callbacks is just its State), gated by Serialize bounds + an on_arrive reacquire hook, addressed by name not pid. The BEAM can't do this — leaning on the behaviour layer is what buys it. Requires State to be serializable, so should probaby spec a trait MigratableGenServer: GenServer where Self::State: Migratable , or something to that extent, so we can lean on the type system to make sure we don't accidentally make state that cannot be serialised.
  • CRDT presence is the high-value, genuinely-hard layer above distributed pg, kept out of the pg primitive (the pg2 strong-consistency lesson). Furthest out. Can maybe defer to rust ecosystem

Look into

app actors block AllDone; no external stop path

Agent working on urus (see same git server as smarm) reported a lazily spawned actor never returning, blocking program shutdown. Maybe we should do something about it. Agent worked around it by giving the actor an atomic bool to spin on. See urus example crud for exact impl.


Invariants & gotchas (respect these across all cycles)

  • Shared mutex is non-reentrant. Sender::send can call unparkwith_shared. Never send on a channel while holding the shared lock. Pattern: mem::take data under the lock, send after releasing. See finalize_actor.
  • finalize_actor order: take stack/waiters/monitors under lock + set Done/outcome → recycle stack → deliver supervisor Signal + monitor Downs → unpark joiners → reclaim slot if outstanding_handles==0. Death notifications always precede reclamation.
  • Slot lifecycle reset in THREE places: Slot::vacant(), reclaim_slot() (runtime.rs), slot-init block in spawn_under (scheduler.rs). Any new Slot field must be reset in all three.
  • Pid = (index, generation). Stale handles caught by generation mismatch in slot()/slot_mut(). The monitor NoProc path relies on this.
  • The only wildcard wake is request_stop, and it is terminal. Every registration-based waker (channel sends, mutex grants, wait-timers, io completions, joiner wakes, select arms) carries the wait's park-epoch and wakes through unpark_at; every successful wake consumes the epoch. Wakes are therefore meaningful: one-shot park sites interpret them without loops, and select needs no cancellation pass. When adding a new waker, decide which form it is — if its registration handle can outlive the wait it was created for, it MUST be epoch-stamped; a wait that can exit without parking MUST retire_wait first (see slot_state.rs).
  • select exists; a unified per-process mailbox still does not. The supervisor keeps its single supervisor_channel funnel; recv_match stays per-channel. select composes channels at the wait, not into one queue — gen_server's handle_info/handle_down (v0.8) are built on exactly that composition, with documented arm priority (downs → control → infos → inbox) instead of mailbox FIFO. A hot higher-priority arm starves lower ones by design; that's the contract.
  • Cooperative-only. Preemption and cancellation both depend on the actor reaching check!()/yield/alloc/blocking points.
  • Lock order is Leaf → Channel, one of each at most (debug-asserted in raw_mutex.rs). Leaf = cold locks / free list / stack pool / registry, mutual leaves. A channel lock may be taken under a Leaf (finalize/monitor clone senders living in slots); nothing may be locked under a channel lock.
  • Queue ops require preemption disabled. A producer suspended mid-publish stalls every consumer — livelock. with_runtime, with_shared, and RawMutex guards all disable preemption for their span.
  • run() is single-thread (Config::exact(1)); tests rely on deterministic single-thread ordering. Multi-thread via runtime::init(Config…).