26 Commits
Author SHA1 Message Date
Claude (sandbox) 741c10337b release: v0.7.0
Bump crate version to 0.7.0.
2026-08-21 13:10:25 +02:00
Claude (sandbox) e570138da5 docs(roadmap): supervisor start order is not start readiness
Filed from the urus v0.3 endpoint work. start_child spawns and moves on,
so a later sibling can whereis an earlier named child before that child's
actor has run. Notes why blocking spawn is not the fix ('has begun
executing' != 'has bound its name', plus a per-accept round-trip tax and
every spawn becoming a context-switch point), that OTP has the same async
spawn and synchronises one level up in gen_server:start_link, the
readiness-ack shape if scheduled, and the structural workaround urus uses
today (registrar spawns its own consumers).
2026-08-20 13:20:42 +00:00
Claude (sandbox) 415effb2e9 feat(gen_server,gen_statem): lifetime is the actor's — refs are addresses; inline named run
Root cause behind the "pin the endpoint" gotcha and the trapping-wrapper
pattern: a gen_server had two lifetime authorities — its refs (last one
dropped → inbox closes → exit) and, when supervised, its supervisor. OTP has
one: a process lives until it stops, is shut down, or is killed; a pid is an
address. Root exit now shutting down every forest root removes the reason the
ref-governed idiom existed (a forgotten server no longer hangs the run), so
adopt the one rule:

- The server/machine loop holds one inbox sender for its life; the inbox
  never closes. GenServerRef / GenStatemRef are addresses. Explicit close is
  `shutdown()`; a forgotten one is swept at root exit.
- `NamedGenServerBuilder::run()` / `gen_statem::run_named(name, m)` run the
  loop inline as the current actor: a server is a direct ChildSpec child,
  gets the supervisor's shutdown as handle_shutdown / a shutdown row, re-binds
  its name on restart, and is addressed by name. The wrapper in
  examples/graceful_shutdown.rs is gone.
- gen_statem gains GenStatemName + whereis_machine/send/call/shutdown by name
  (parity with gen_server); the macro gets `Sm::new`.
- Root-exit sweep records `Event::RootSweep { target, trapping }` under
  smarm-trace ("root_sweep shutdown|stopped"): unsupervised leftovers are
  visible rather than silently owned-by-refs.
- Named start() name-clash path stops the spawned actor instead of relying
  on ref drop.

Tests: tests/gen_server_lifetime.rs, tests/gen_statem_lifetime.rs,
tests/root_sweep_trace.rs (feature-gated); three existing tests that used
drop-closes-inbox now use shutdown(). Docs/README/ROADMAP/Deep Dive updated.
2026-08-19 17:47:31 +00:00
Claude (sandbox) 849a424c8e docs,examples: graceful shutdown — new examples/graceful_shutdown.rs, README 'Stopping actors', named_genserver uses shutdown(), Deep Dive terminate note, ROADMAP open items 2026-08-19 16:28:42 +00:00
Claude (sandbox) 6ceb138f5f feat(gen_statem): graceful-shutdown parity — trap_exit, shutdown/exit rows, stop, terminate
Mirrors the gen_server surface in gen_statem's event model:
- Cx::trap_exit() (in the initial enter): a shutdown request then arrives
  as the Shutdown event, routed by state through `shutdown` rows (default
  for a state with no row: stop); linked-peer deaths as `exit <pat>` rows
  (default: drop). Non-trapping machines are stopped outright, as before.
- Cx::stop(): normal self-exit after the current event; `stop` tail keyword
  is sugar for { cx.stop(); prev }.
- Machine::terminate (optional `terminate { … }` macro block), run from a
  Drop guard on every exit path; the guard also drains armed timers.
- Machine::shutdown_ev / exit_ev (defaults None) so hand-written machines
  keep compiling; GenStatemRef::shutdown() is graceful and waits.
- Loop selects exits > timers > inbox.

Tests: tests/gen_statem_shutdown.rs.
2026-08-19 16:25:15 +00:00
Claude (sandbox) 250f31265b feat(runtime): root exit is graceful shutdown of the forest roots
The RFC 014 root-exit sweep hard-stopped every live slot once nothing was
runnable. That deferral privileged queued work over parked-with-a-pending-
wake work (a sleeper was killed, a queued cast was drained) and any attempt
to widen the notion of pending wake (timers, fd readiness) re-wedges the
run on the periodic-timer daemon the sweep exists to end.

Root exit now means "the program is done": finalize_actor delivers
request_shutdown to every forest root — each live actor whose parent is
the run (ROOT_PID) or is dead — synchronously, before the live-count
decrement. Supervisors cascade per child Shutdown policy; trapping actors
may Continue/drain with working timers and end the run when they stop
themselves; non-trapping actors are stopped outright. No forcing sweep.

Removes root_exited/root_swept, Pop::RootDrain and the idle-verdict
condition; adds tests/root_exit.rs.
2026-08-19 07:11:46 +00:00
claude 9c8f59ca53 feat(scheduler,supervisor,gen_server): graceful shutdown — request_shutdown, child Shutdown policy, handle_shutdown
Lift OTP's `exit(Pid, shutdown)` + child-spec `shutdown` wholesale.

scheduler / runtime
- `request_shutdown(pid)`: the polite stop. A target trapping exits gets an
  `ExitSignal { reason: DownReason::Shutdown }` on its trap inbox and keeps
  running; a non-trapping target is stopped as by `request_stop`, which is
  now documented as the hard stop (`exit(Pid, kill)`). Dead pid: no-op.
- `RuntimeHandle::request_shutdown` for the off-runtime (signal thread) path;
  `from == ROOT_PID` there.
- `DownReason::Shutdown` — appears only in ExitSignal, never in Down (a
  complying target exits *normally*).

supervisor
- `ChildSpec::shutdown(Shutdown::{BrutalKill, Timeout(d), Infinity})`,
  default Timeout(5s). Every supervisor-initiated stop (ordered shutdown and
  OneForAll/RestForOne sibling cycling) is: request_shutdown → await the
  child's Signal up to the grace → request_stop → await. Sequential, reverse
  start order.
- The supervisor traps exits; a Shutdown ExitSignal runs the ordered
  shutdown and `run()` returns normally, so `request_shutdown(root_sup)`
  tears a whole tree down top-down with each child's grace period.
- FIX: a hard `request_stop` on a supervisor previously orphaned its
  children (the ordered shutdown lived after the loop, and the unwind
  skipped it). `Live` (the by_pid map) now carries a drop guard that
  fire-and-forget hard-stops live children when unwinding.

gen_server
- `GenServerCtx::trap_exit()` opt-in in `init`; the trap inbox becomes arm 0
  of the loop's select. Shutdown ExitSignal → `handle_shutdown() ->
  ShutdownAction::{Exit, Continue}` (default Exit: loop breaks, `terminate`
  runs on the normal path and may block). Other ExitSignals →
  `handle_exit(sig)`.
- `GenServerCtx::stop_handle() -> StopHandle`, `stop()` ends the server
  after the current message with a *normal* exit — the missing
  `{stop, normal, State}`; `request_stop(self_pid())` was the only self-exit
  and it is abnormal (Transient restarts it).
- `GenServerRef::shutdown()` / `gen_server::shutdown(name)` now go through
  `request_shutdown`.

Tests: tests/shutdown.rs, tests/supervisor_shutdown.rs,
tests/gen_server_shutdown.rs. Full suite green; fmt + clippy --lib clean.
2026-08-19 06:34:48 +00:00
Claude (sandbox) 1002777ef3 feat(channel,runtime): off-runtime cross-thread wake for parked receivers and stop
Wakes issued from a non-scheduler OS thread were silent no-ops. Every
off-runtime wake primitive (unpark, unpark_at, request_stop) reaches the
runtime through the RUNTIME thread-local, which is unset on any foreign
thread — so a send from a plain std::thread enqueued its message but never
woke the parked receiver, and there was no way to drive a stop into a
runtime from an application thread (e.g. an OS-signal handler). The former
strands a parked recv forever; the latter is why a downstream server must
poll a shutdown flag instead of parking on it.

Generalize RFC 018's rule — a producer reaches the runtime through a Weak it
holds — from the IO backend to channel senders and to a new handle:

- A receiver captures a Weak<RuntimeInner> (provably live at that moment)
  alongside its (pid, epoch) when it parks. send() and the last-sender drop
  wake through scheduler::unpark_at_via, which takes the thread-local path
  when on a scheduler thread (preemption-gated, slot-eligible) and the
  captured Weak otherwise — the same cross-context wake the IO threads do.
- Runtime::handle() returns a Send + Sync RuntimeHandle carrying that Weak;
  RuntimeHandle::request_stop drives a cooperative stop from any thread and
  is a no-op once the runtime is dropped.

The in-runtime wake paths (recv/select timers) are unchanged; only the sites
reachable from a foreign thread route through the Weak. RuntimeHandle exposes
request_stop only: send-wake needs no user-facing handle, and off-runtime
unpark is covered because request_stop drives unpark on the upgraded inner.

tests/cross_thread_wake.rs: a foreign-thread send wakes a parked receiver; a
foreign-thread request_stop wakes and stops a parked actor; a RuntimeHandle
held across and beyond run never blocks all-done and degrades to a no-op.
2026-08-19 05:50:54 +00:00
smarm 8f2d513940 README: rewrite overview, add limitations, roadmap, and contribution notes
Expands the intro into an overview/limitations structure, documents
preemption, tracing/causal profiling, URUS, and adds a 'Coming up' and
'A note on open source' section.
2026-08-18 00:25:25 +02:00
Claude (sandbox)andClaude (sandbox) ca1c98336e feat(scheduler,runtime): non-panicking try_spawn for at-capacity load shedding
allocate_slot() panics on a full slab; for a load-shedding caller (an
accept loop spawning one actor per connection) that panic lands in the
spawning actor, which then crash-loops under Restart::Transient into the
still-full slab until its restart budget is spent — and the service stops
accepting entirely. Observed live (urus slowloris scaling, 2026-08-10).
A full slab is a routine overload condition for such callers, not an
invariant violation.

- RuntimeInner::try_allocate_slot() -> Option<u32>: the non-panicking
  core; a single pop under the free-list lock, so the claim is atomic
  (claim-or-report — no check-then-spawn TOCTOU, no headroom margin).
  allocate_slot() is now a thin panicking wrapper over it.
- scheduler::try_spawn / try_spawn_under_with -> Result<JoinHandle,
  SpawnError>: parity with spawn/spawn_under_with except a full slab
  returns Err(SpawnError::AtCapacity) instead of panicking. Minimal
  surface per the agreed strategy; the remaining _with/_addr mirrors are
  trivial wrappers if ever needed.
- Slot-first ordering on the try path (reverse of spawn's stack-first):
  under overload Err is the hot path, and a rejection costs one mutex
  pop — no mmap/pool-pop + init + recycle per shed unit of work. A
  drop-guard returns the claimed slot if stack allocation panics in the
  claim-to-install window (would otherwise leak and trip run()'s
  teardown slot-leak debug_assert).
- SpawnError: non_exhaustive, Display + std::error::Error.
- spawn and every existing call site untouched: the panic remains the
  correct loud invariant check at internal/bounded spawn sites.

tests/try_spawn.rs: parity when slots free; exact slab accounting at
capacity (Err, no panic, repeatable); custom-shape try refuses before
stack allocation; self-heal after slots free; plain spawn still panics
(surfaced via JoinError payload); 4-thread race for the last slots
claims exactly the free count; SpawnError impl checks.

Design doc: smarm-suggestion-try-spawn.md. Downstream consumer change
(canned 503 on AtCapacity in urus's accept loop) is urus scope, not
smarm.

(cherry picked from commit 36de4b36aeaa72b2a5f9f3797b9854652656dcf6)
2026-08-13 15:03:16 +02:00
smarm-agent 95306c7f60 style: cargo fmt sweep under rustc 1.97.1 (toolchain reformat, no semantic change) 2026-08-13 05:56:49 +00:00
smarm-agent 1262cc30e3 monitor: widen stamp eligibility to watchable = named ∪ exported (soak sig 5)
The terminal record existed for watches that raced their target's death, but
e43c673 scoped its stamp to named tenancies — and the pid-identity watch
surface (§4 Slice 3) targets arbitrary actors, including anonymous ones whose
pids cross the boundary in contract replies. The first wild pid-face hit
(width-20 soak, pid_watch_test.exs:47, 1/600 full-suite: a monitor installed
while the child was alive delivered :noproc instead of {:smarm_exit, :panic})
is exactly the residual a0ba9be's commit body deferred.

ever_named becomes `watchable`, with a second set-site: mark_watchable(pid),
which the bridge calls wherever a smarm pid is encoded across the boundary —
BEAM can only watch pids it holds, and can only hold pids that crossed.
Anonymous never-exported churn (holder threads, egress tasks) stays
ineligible, preserving e43c673's LIFO-eviction protection unchanged.

mark_watchable takes the cold lock before the liveness screen: finalize
publishes Done and reads the bit under the same lock, so the mark either
lands before the death stamps or observes the tenancy dead and no-ops —
no lost-stamp window, and marking a corpse cannot invent history (pinned
in the test alongside the mark-while-alive stamp).
2026-08-13 05:56:19 +00:00
smarm-agent 461fe4b768 fix(runtime): only named tenancies stamp the terminal record — anonymous churn must not evict it
Discovered wiring the bridge consult: with an unconditional stamp, the record
for the very death being raced was the shortest-lived data in the runtime.
Every green thread is a slot tenant, the free list is LIFO — so the slot a
named server's death frees is the first one recycled, and the next throwaway
exit (monitor holders, chain-runner work, anything) overwrote the record
before a raced watch could consult it. Deterministic bridge repro: the
corpse resolved fine, terminal_reason read None every time.

register_with now flags the tenancy (ever_named, reset at reclaim) before
the binding lands — set outside the registry lock, so no successfully
registered actor can die unflagged and a failed register's overshoot is
harmless — and finalize stamps only flagged tenancies. Watchable identities
are exactly the named ones (the bridge's pid-identity path deliberately
keeps Erlang's raw :noproc), so nothing consultable is lost.

Contract test updated: the three death modes now self-register; a new
anonymous control pins that unregistered deaths neither stamp nor evict.
2026-08-13 05:56:19 +00:00
smarm-agent b937f1f50f monitor/registry: terminal-outcome record — a raced watch can recover the real down reason (soak sig 4)
A watch installed after its target's death has, until now, only NoProc to
report — but the bridge's proxies install their native watch asynchronously
after acquire returns, so a link established before a crash (from the BEAM's
view) could still lose the panic's translated reason to that blanket NoProc
(width-20 soak signature 4: link_test.exs:26, 1/600 full-suite, 3/2000
link-only, all whereis-miss; deterministic repro in the bridge suite).

Two primitives, no change to monitor()'s own Erlang-faithful stale-pid
semantics — the upgrade is the caller's deliberate act:

- finalize_actor stamps the slot with (generation, DownReason) under the same
  cold-lock block that publishes the outcome. The record survives reclaim,
  registry pruning, and the next tenant's install; only the slot's next death
  overwrites it. terminal_reason(pid) reads it generation-matched.
- resolve_name(name) is whereis with the corpse kept: the dead-holder arm
  returns the stored pid it prunes (NameResolution::Corpse) instead of
  discarding the only evidence of who died — whereis itself prunes on the way
  out, so a whereis-then-lookup consumer would find the evidence already
  destroyed. Live/Unbound match whereis's Some/None; the name heals exactly
  as before.

Contract pinned in tests/terminal_outcome_after_death.rs: one record per way
of dying (Exit/Panic/Stopped), no record while live, corpse capture + heal on
resolve_name, record independence from registry pruning, survival across slot
re-tenancy, overwrite at the next tenancy's death.
2026-08-13 05:56:19 +00:00
Claude (sandbox) 301e3463e3 chore(release): v0.6.0 — RFC 019: actor stack reserve & shrink
Per-actor stack shapes on every spawn surface (SpawnOpts stack_reserve/
guard_size, Config defaults, pool rule: only default-shaped recycle);
sampled stack high-water + MADV_FREE shrink at actor-park (THRESHOLD
256 KiB, COOLDOWN 64 parks, redzone 1 page); pool-recycle MADV_DONTNEED
above the retained 64 KiB entry end; SIGSEGV overflow diagnostics
(two-tier: in-guard definitive / 1 MiB overshoot 'stepped over', prior
handler chained for foreign faults) with per-scheduler sigaltstack; and
the per-actor introspection surface (ActorInfo.stack: reserve, guard,
sampled depth, parks_since_shrink, shrinks).

Amendments ratified during implementation, for the RFC changelog:
- DEFAULT_STACK_GUARD 64 KiB -> 1 MiB, following the kernel's post-Stack-
  Clash stack_guard_gap convention; PROT_NONE width is VA-only and free.
- §7's motivating segfault was a cargo-vendored gz build, not SQLite as
  the RFC text says (cc-built C lacks -fstack-clash-protection; distro
  libraries have it — the risky class is vendored builds).
- §4 hibernate() deferred to the jar (bolt-on: force-flag on the §3
  shrink path, ~10 lines when wanted).

Gates (jobrunner box, 2026-08-08): reclaim gate PASS at c3 and again at
tip (3.0 MiB LazyFree -> kernel reclaim -> Rss to one live page ->
re-spike bit-identical, live data intact; MADV_PAGEOUT stands in for
memcg — cgroup2 is RO in the job container — driving the same reclaim
path). E1 interleaved A/B vs v0.5.0: every ka cell (the E1 subject)
within +0.3..+2.9% at tip; close-mode control cells within noise except
t8-c4 close, which is bistable (~40-44k vs ~46-49k modes for BOTH
variants, base self-disagrees by 11% across rounds); 6 rounds across two
runs are inconclusive there and a 10-round focused run is noted in the
handoff as deferred follow-up, accepted for this release.

No breaking API changes since v0.5.0: SpawnOpts fields and ActorInfo
gained members (exhaustive-construction downstream will need the new
ActorInfo.stack field; urus does not construct it).
2026-08-08 19:44:55 +00:00
Claude (sandbox) 410ba33d82 feat(introspect,runtime): per-actor stack surface on ActorInfo (RFC 019 §8)
- introspect::StackInfo { reserve, guard, depth_high_water,
  parks_since_shrink, shrinks } as ActorInfo.stack; re-exported at crate
  root beside ActorInfo.
- All reads lock-free: geometry from the c6 diag slot atomics, depth =
  top - hwm (the §2 sampled high-water; doc spells out sampled-not-exact
  and that 0 means never-descheduled-at-depth), counters straight off the
  §3 atomics. Coherence for the incarnation rides read_slot's existing
  generation check, same as overruns/messages_received.
- Slot::stack_introspect(): one pub(crate) tuple accessor beside the other
  counter accessors.
- Exact RSS deliberately absent per RFC (mincore = debug tooling only,
  never a runtime path); stack_shape(pid) untouched (cold-lock exact
  variant from c2).
- tests/introspect.rs: defaults surface (64 KiB reserve / 1 MiB guard /
  sampled ~32 KiB depth / gate park counted / zero shrinks) + live shrink
  counters (spike visible pre-shrink; shrinks>=1, cooldown counter reset,
  hwm reset after crossing COOLDOWN) read mid-run -- post-join the slot
  reclaim correctly hides the incarnation, which the first draft of the
  test learned the hard way.

FLAGGED (Claude-solo calls):
- Nested StackInfo struct over five flat ActorInfo fields (grain break;
  the five fields are one concern and ActorInfo is already 12 fields).
- Field names reserve/guard/shrinks (RFC says stack_reserve/stack_guard/
  shrink count; the stack_ prefix is redundant inside StackInfo).
2026-08-08 19:12:46 +00:00
Claude (sandbox) 5fd8aecf55 feat(signal,runtime,stack): SIGSEGV overflow diagnostics + 1 MiB guard default (RFC 019 §7)
- src/signal.rs: process-global SA_SIGINFO|SA_ONSTACK handler installed once
  at runtime::init (before any scheduler thread -> unracing PRIOR save);
  per-scheduler-thread 64 KiB sigaltstack registered at schedule_loop entry
  (a guard hit leaves no stack to handle on). Async-signal-safe throughout:
  classification is plain loads (const-init TLS Cell + slot atomics), print
  is fixed-buffer itoa + one write(2), death is SIG_DFL + refault at the
  same instruction (core-dumpable, correct wait status).
- Two-tier classification (agreed): in-guard = definitive; OVERSHOOT window
  below the guard = 'unprobed (FFI?) frame stepped over it' probable
  attribution -- the RFC's motivating incident (cargo-vendored gz, not
  SQLite as the RFC text says) faults there under a small guard. Pure
  classify() fn, 5 adversarial units incl. saturation at low addresses.
- DEFAULT_STACK_GUARD 64 KiB -> 1 MiB (agreed): kernel stack_guard_gap
  anchor post-Stack-Clash; PROT_NONE is VA-only (no RSS, no page tables,
  no overcommit charge) so width is free at any actor count.
- Unclassified faults reinstate the PRIOR sigaction and refault (agreed):
  std's own OS-thread overflow diagnostics survive our presence.
- Slot: diag_{stack_top,stack_reserve,stack_guard,pid} atomics written in
  install_actor pre-publish; readable without the cold lock (Stack lives
  under it); only consulted while CURRENT_SLOT points at the slot, so
  never stale where read. preempt::current_slot_ptr ungated from
  smarm-causal (now also the classifier's anchor).
- build.rs + cc (agreed Q3): canary/canary.c, 96 KiB local touched low-end
  first, -fno-stack-clash-protection pinned so hardened toolchains don't
  probe the canary into uselessness.
- tests/stack_diag.rs: subprocess x4 -- Rust recursion tier-1; FFI canary
  tier-1 at defaults (1 MiB guard catches the jump); tier-2 at guard=4 KiB
  ('stepped over', reproduces the incident); clean at reserve=256 KiB
  (the §1 knob is the fix, same frame).

FLAGGED (Claude-solo calls):
- OVERSHOOT_SLOP = 1 MiB (matches guard default/kernel gap; beyond it
  attribution would be dishonest).
- Altstack 64 KiB, mmap'd once per OS thread, never freed (bounded by
  thread count; reused across run()s via TLS flag).
- Foreign-fault reinstate permanently deregisters our handler; accepted --
  the process is dying either way.
- Diag geometry as 4 slot atomics (install-time cost only) over a per-switch
  TLS snapshot (hot-path stores).
2026-08-08 18:58:30 +00:00
Claude (sandbox) 7d8b9e0310 feat(stack,runtime): pool-recycle DONTNEED above the retained entry end (RFC 019 §6)
- stack::retain_range: pure checked span fn (retain page-up = zap less;
  None when retain covers the reserve, so the 64 KiB default config never
  pays a syscall) + 6 adversarial units mirroring shrink_range's.
- Stack::recycle_zap: advisory MADV_DONTNEED of [usable_base, top-RETAIN);
  stack is unowned at the call site, synchronous eager zap races nothing.
- recycle_stack: zap OFF-LOCK before pool admission (acquire_stack's
  no-syscall-under-the-pool-lock invariant); rare cap-overflow pays a
  wasted zap ahead of munmap, accepted over a second lock round-trip.
- pub const RECYCLE_RETAIN = 64 KiB beside the shrink knobs, ratified-as-
  constant rationale in doc.
- tests/stack_recycle.rs: mincore-based exact-zero-resident assert over
  the zap span. smaps was tried first and over-counts: a neighboring rw
  anon VMA can merge flush against the stack top (observed once under the
  full-suite run); the PROT_NONE guard pins the usable base exactly.

FLAGGED (Claude-solo calls):
- RFC §6 'above the bottom RETAIN' is direction-ambiguous in address
  terms; implemented as retain the ENTRY end (highest addresses, the
  pages the next actor faults first), zap the cold deep span below.
- Const named RECYCLE_RETAIN (RFC says RETAIN) to sit beside SHRINK_*.
2026-08-08 16:13:53 +00:00
Claude (sandbox) 8225716b11 feat(runtime,stack): sampled stack high-water + MADV_FREE shrink at actor-park (RFC 019 §§2–3)
hwm: AtomicUsize lands beside sp on the slot: the single context-save
site min-updates it (one branch + at most one Relaxed store into the
line the sp store just dirtied), install resets it to the fresh top.
Advisory by construction — correctness never depends on it. The mod-doc
ordering chain gains a line: hwm piggybacks the existing
Relaxed-store-before-Release pattern and adds no edges.

Shrink hook in the YieldIntent::Park arm only, before the park_return
Release transition — the owned window (obligation 1's assert-comment at
the site): after the sp store, before Parked is published, scheduler on
its own stack, actor saved and unstealable. It runs on both arms of the
park_return race (a consumed unpark flag means one wasted-but-harmless
madvise). The preempt/yield path deliberately never checks: §4's
bounded, self-healing leak under saturation, when syscalls are least
affordable.

SHRINK_THRESHOLD = 256 KiB and SHRINK_COOLDOWN = 64 parks are pub
constants with the ratified doc rationale, not Config fields. The freed
span is shrink_range(hwm, sp, page): whole pages of [hwm, sp − 1-page
redzone), rounded inward, checked arithmetic — adversarial inputs
collapse to None (obligation 2). MADV_FREE marks lazily; the kernel's
reclaim-under-pressure IS the hysteresis, cancel-on-write is the safety
net. parks_since_shrink + shrink_count ride the slot for the cooldown
and the future introspect surface.

Tests: 7 adversarial shrink_range units (inverted/empty spans, redzone
underflow, unaligned ends, sp-crossing sweep); integration — 8 MiB
reserve, ~3 MiB spike sampled via yield-at-depth, parks gated on
introspected Parked state past the cooldown, then ≥ 2 MiB LazyFree
asserted inside the stack's smaps range with live data intact; and the
inverse guard — a shallow never-spiking actor ends at exactly 0
LazyFree (also proves the parser isn't vacuously zero via the first
test).
2026-08-08 14:30:32 +00:00
Claude (sandbox) 3cb64eefc2 feat(scheduler,gen_server,gen_statem,introspect): SpawnOpts — per-actor stack shape on every spawn surface (RFC 019 §1)
SpawnOpts { stack_reserve, guard_size } with Option<usize> fields, None
resolving to the Config defaults at spawn time — a deliberate deviation
from the RFC's plain-usize struct so struct-update syntax works without
a runtime handle in scope. Threaded across the five surfaces:
spawn_with, spawn_under_with, spawn_addr_with,
GenServerBuilder::stack_opts (mirrored on NamedGenServerBuilder), and
gen_statem::spawn_with (gen_statem has no builder, so the opts ride a
_with variant — Claude-solo surface call, flagged for review). Existing
spawns forward defaults; no call-site churn.

introspect::stack_shape(pid) pulled forward (agreed) as the first slice
of the RFC 019 introspection surface, giving tests an observable.

Tests (tests/spawn_opts.rs): override/partial-override/rounding on each
surface; obligation 4 from the outside — a dead custom stack is never
handed to the next default spawn (LIFO pool would expose it), and the
reverse (default stacks ARE recycled); 8 MiB reserve behaviorally
permits ~1 MiB recursion. Also: silence unused-Result in the c1
runtime test (join now unwrapped).
2026-08-08 14:22:38 +00:00
Claude (sandbox) 0fe052bc7e feat(stack,runtime): per-shape actor stacks — Stack::new(reserve, guard), Config knobs, pool rule (RFC 019 §1)
Stack takes an explicit (reserve, guard) shape, both page-rounded and
stored; usable_base derives from the stored guard. Guard default raised
4 KiB -> 64 KiB (DEFAULT_STACK_GUARD): probestack makes one page enough
for Rust frames, but an unprobed C frame can leap a page in one sub rsp
— the motivating SQLite segfault. Reserve default stays 64 KiB
(DEFAULT_STACK_RESERVE); ACTOR_STACK_SIZE retired.

Config::{stack_reserve, stack_guard} thread the runtime defaults into
RuntimeInner pre-rounded. All acquisition/recycling now goes through
acquire_stack/recycle_stack carrying the pool rule: only default-shaped
stacks are pooled (pooled ⇒ default-shaped by induction); custom shapes
mmap fresh and munmap at death. Pool lock still dropped before any mmap.

No public spawn API change (SpawnOpts is the next commit).

Tests: shape rounding + accessors, wide-guard faults at both ends
(subprocess), Config::stack_reserve permits >64 KiB recursion that
previously could only segfault.
2026-08-08 14:18:27 +00:00
smarm a03a7ca01e chore(release): v0.5.0
Breaking API rename since v0.4.0: gen_server's ServerRef/ServerBuilder/
ServerCtx -> GenServerRef/GenServerBuilder/GenServerCtx, Watcher<G> is
now generic over its GenServer, and GenServer gained a required
associated Timer type for timer-fire payloads (arm_after/handle_timer).
Downstream consumers (urus) have been ported.
2026-08-08 11:44:35 +02:00
smarm d4839f1d81 feat(runtime,io): driver-enqueues + park/wake idle path — retire the wake pipe
The swap (RFC 018). Schedulers no longer sleep on a shared level-triggered
wake pipe — the herd source that made the default 8-thread config 7x
slower than 2 threads (E1). They park on per-thread futex parkers via the
coordination layer; IO backends become producers behind a two-call
contract (make runnable, then the enqueue tail wakes exactly one parked
scheduler).

Deleted: the drain lock and the one-winner phase-1 drain; the shared
completions VecDeque; the wake pipe fds, poll_wake, drain_wake_pipe,
wake_scheduler, the FdReady/Blocking Completion enum; the 100us idle nap;
the per-pop io.lock liveness read; io.rs's as_millis timeout truncation.

Added:
- enqueue wake tail (fixes the silent enqueue): wake_one_if_idle, a fence
  + one Relaxed mask load when everyone is busy — the pure-compute hot
  path pays almost nothing.
- driver-enqueues: the pool thread stashes its result in the slot,
  decrements io_outstanding, unparks; the epoll thread removes+DELs the
  waiter under the waiters lock and unparks. Both reach the runtime via a
  Weak (no Arc cycle). The waiters map moves behind its own Arc<Mutex> so
  the epoll thread never takes the runtime io lock (teardown holds it
  while joining that thread).
- io_outstanding / io_fd_waiters atomics: the termination verdict reads
  two atomics instead of taking io.lock on every pop.
- timekeeper idle path: at most one parked scheduler holds the timer
  deadline (an expiry wakes one, not a herd); everyone else parks
  indefinitely and is woken by the enqueue tail.
- busy-path timer due-check (ratified design point (a)): under saturation
  nobody parks and no timekeeper exists, yet due timers must still fire —
  one Relaxed load of the earliest-deadline snapshot per loop, clock read
  only when a timer is armed. Maintained under the timers mutex.
- chain rule: a scheduler that pops with more work queued and a sibling
  parked wakes one, so surplus runs in parallel rather than behind it.

tests/park_wake.rs pins the two new observable properties: timers fire
under full scheduler saturation, and sub-ms sleeps are prompt (the
as_millis truncation regression). Full suite + all loom models green;
clippy --lib clean.
2026-07-24 09:12:12 +02:00
smarm 2854b560d6 feat(park): fenced producer fast path + earliest-deadline snapshot
Two integration-driven amendments ahead of the runtime swap:

wake_one_if_idle() realizes RFC 018's "empty-mask fast path is one
relaxed load" soundly: a bare relaxed load is a lost-wake in the Dekker
shape for the lock-free ring queues, so the producer publishes work,
fences (SeqCst), then reads the mask Relaxed — paired with a matching
fence between the consumer's bit-publish and its re-check in park().
The pure-compute hot path (mask 0) never takes the shared mask line
exclusive; the RMW read stays on the rare chain-rule path only. Loom
models 1/2 now drive the fenced pattern end to end.

next_deadline is the earliest KNOWN timer deadline, independent of
whether anyone is parked — which tk_armed cannot give: under saturation
nobody parks, nobody arms, yet due timers must still fire (ratified
design point (a): the busy-path due-check). Maintained under the timers
mutex (note_deadline on insert — which also carries the timekeeper
re-arm wake — refresh_deadline after pop/clear); read lock-free.
deadline_due() costs one Relaxed load and a branch when no timer exists;
the clock is read only when one does.
2026-07-24 09:12:12 +02:00
smarm 7b026cfe56 feat(park): scheduler coordination layer — parkers, idle mask, wake protocol (RFC 018)
Schedulers get an IO-agnostic sleep/wake primitive of their own: one
futex Parker per scheduler thread (permit semantics, std::thread::park
shaped — closes the check-then-park race), an AtomicU64 idle mask with a
set-bit → re-check → wait park protocol, wake_one (highest-bit LIFO,
CAS-clear before unpark: exactly one wakeup per call by construction),
wake_all for the terminal path, and the timekeeper role — at most one
parked scheduler holds the timer deadline, with an atomic armed-deadline
snapshot for the busy-path due-check and an insert-side re-arm wake.

Deadlines travel as nanosecond timespecs end to end; the wake pipe's
as_millis truncation is unrepresentable here. The Dekker publish/re-check
shape is resolved by the same-location-RMW handshake (AcqRel), not SeqCst
loads; loom verifies exactly this in four models (no-lost-wake, chain
propagation, timekeeper handoff, termination), run with
LOOM_MAX_PREEMPTIONS=3 — unbounded exploration is impractical for the
looped models. Loom/non-Linux builds park on a Mutex+Condvar via
sync_shim.

Standalone until the runtime swap (next commit): nothing outside tests
constructs a Coordinator yet, hence the temporary dead_code allow in
lib.rs.
2026-07-24 09:12:12 +02:00
smarm 006a3283e7 chore(hooks): clippy gate falls back to a nix-shell toolchain
Desktop migration: the home-manager rust here ships without the clippy
component. Prefer an installed cargo-clippy; otherwise run clippy from an
ephemeral nix-shell with a separate target dir (mixed-compiler artifacts
are an E0514 hard error). MSRV keeps the shell's older toolchain a
legitimate gate.
2026-07-24 09:12:12 +02:00
94 changed files with 9366 additions and 1494 deletions
+16
View File
@@ -2,7 +2,23 @@
# smarm pre-commit gate: clippy the library (src/) with warnings as errors. # smarm pre-commit gate: clippy the library (src/) with warnings as errors.
# unwrap_used / expect_used are denied (Cargo.toml [lints.clippy]): library # unwrap_used / expect_used are denied (Cargo.toml [lints.clippy]): library
# code must not hide a panic behind unwrap/expect. Tests/examples are not gated. # code must not hide a panic behind unwrap/expect. Tests/examples are not gated.
#
# Toolchain resolution: prefer an installed cargo-clippy; on machines whose
# rust comes without the clippy component (e.g. NixOS home-manager), fall
# back to an ephemeral nix-shell toolchain. The fallback uses its own target
# dir (target/clippy) because the shell's rustc version may differ from the
# default toolchain's — mixed-compiler artifacts in one target dir are an
# E0514 hard error. MSRV (Cargo.toml rust-version) keeps the older shell
# toolchain a legitimate gate.
set -eu set -eu
[ -f "$HOME/.cargo/env" ] && . "$HOME/.cargo/env" [ -f "$HOME/.cargo/env" ] && . "$HOME/.cargo/env"
cd "$(git rev-parse --show-toplevel)" cd "$(git rev-parse --show-toplevel)"
if cargo clippy --version >/dev/null 2>&1; then
cargo clippy --lib -- -D warnings cargo clippy --lib -- -D warnings
elif command -v nix-shell >/dev/null 2>&1; then
nix-shell -p clippy -p cargo -p rustc \
--run 'CARGO_TARGET_DIR=target/clippy cargo clippy --lib -- -D warnings'
else
echo "pre-commit: cargo clippy unavailable and no nix-shell fallback" >&2
exit 1
fi
+4 -1
View File
@@ -1,6 +1,6 @@
[package] [package]
name = "smarm" name = "smarm"
version = "0.4.0" version = "0.7.0"
edition = "2021" edition = "2021"
rust-version = "1.95" rust-version = "1.95"
@@ -39,6 +39,9 @@ rq-mutex = []
rq-mpmc = [] rq-mpmc = []
rq-striped = [] rq-striped = []
[build-dependencies]
cc = "1"
[dependencies] [dependencies]
libc = "0.2" libc = "0.2"
+69 -43
View File
@@ -1,35 +1,38 @@
# smarm # smarm
> SMARM — Smarm, Marks Actor Runtime Machinery. A proof-of-concept green-thread actor runtime for Rust. > SMARM: Smarm, Marks Actor Runtime Machinery. A proof-of-concept green-thread actor runtime for Rust.
Implements the core ideas in [`Achitecture.md`](.docs/Architecture.md): green-thread actors on a SMARM is my attempt to implement the erlang/OTP philosophy in the Rust programming language. This has yielded a fault-tolerant, fast, and scalable runtime. This runtime allows the creation of asynchronous applications in Rust without the function coloring associated with the async/await system. It encourages the writing of simple, synchronous code, and largely elides the need for lifetime annotations.
shared heap, scheduled cooperatively, communicating only by `Send` messages.
Erlang's isolation model without Erlang's copying GC, Rust's zero-copy
ownership transfers without async's function colouring.
The scheduler is multi-threaded — one OS thread per available CPU, all drawing
from a shared run queue. The single-threaded `run()` entry point is kept as a
convenience wrapper around `runtime::init(Config::exact(1)).run(f)`.
## What's here ## Overview
SMARM implements green-thread actors on a shared heap, communicating only by `Send` messages. By sharing the heap, SMARM avoids the copying overhead of Erlang, which is safe to do due to Rust's borrow checker.
On top of the core runtime mechanics, SMARM also provides a library of primitives for making applications closely inspired by erlang/OTP. This includes generic servers (gen_servers), generic state machines (gen_statem), and supervision trees.
Supervision trees are the core primitive to allow your application to survive an unexpected panic. Supervisors are processes dedicated to monitoring other processes, which can restart these should they fail. This means that when set up properly an application may 'self-heal' when encountering unforeseen circumstances.
SMARM is not cooperatively scheduled; it uses preemption. This means a heavy task will not starve out other lighter tasks. Everything will make steady progress, which translates to very beneficial behaviour under (over)load: average latency goes up, but tail latency does not blow up.
To help diagnose these unforeseen circumstances, smarm may be compiled with its `tracing` feature, which emits a full trace using [Perfetto](https://perfetto.dev/).
Should you want to optimize your application, SMARM is unusually well poised to help. As the runtime functionally controls time, SMARM comes with a built in causal profiler, under the `causal` feature.
I also built a Phoenix-Framework inspired HTTP 1.1 library on top of SMARM called [URUS](https://git.kalsbeek.dev/Markk116/urus), which implements Pub/Sub, Channels, and basic amenities like Websockets and Server Sent Events.
## Limitations
This runtime requires naked assembly to function, and has thus far only been implemented for x86-64 assembly. It expects an operating system that supports virtual address space, and is therefore not (yet) suited for embedded targets. The IO implementation is currently based around the Linux kernel's `epoll` mechanism, meaning it requires a (GNU+)Linux distribution to run.
The preemption mechanism works by wrapping the memory allocator and checking how many CPU cycles you have used compared to your timeslice budget. This allows preemption to fire in most normal code, but tight zero-allocation loops do not get caught and require manual insertion of `check!()` if you want preemption to function.
This library is still in its early stages, and while I try my best with loom and tests, stable operation cannot be guaranteed. Therefore it is not (yet) recommended for production use.
At this moment, stack memory for each green thread is capped. Uncapping this may lead to performance benefits for deeply recursive algorithms that in a traditional async runtime might require pointer-chases through the heap. This is as yet unrealised.
At this stage, the codebase is largely LLM-generated, which is obvious if you start to read through the internals. While I did the design, and I keep the LLM under tight rein, the codebase is not in a state that I am very happy with. This also goes for the documentation.
| Module | What it does |
|--------------|------------------------------------------------------------------------|
| `stack` | `mmap`'d growable stack with guard page; SIGSEGV on overflow |
| `context` | `#[naked]` x86-64 context-switch shims, callee-saved regs only |
| `preempt` | Allocator-driven preemption; `check!()` macro for no-alloc loops |
| `pid` | `(index, generation)` PIDs; stale handles are detectable, not silent |
| `actor` | Trampoline + `catch_unwind` boundary at the actor entry point |
| `scheduler` | Run queue, slot table, spawn/join, parking, idle path |
| `channel` | Unbounded MPSC channel; `recv` parks the actor; `recv_timeout` bounds it; `select`/`select_timeout` park on many receivers at once (ready-index, priority order) |
| `mutex` | `Mutex<T>` with mandatory timeout; FIFO waiters; parks the green thread |
| `timer` | Min-heap of `(deadline, reason)`; `Sleep` and `WaitTimeout` reasons |
| `io` | `block_on_io` for blocking work; `wait_readable`/`wait_writable` + `read`/`write` via epoll |
| `supervisor` | `Signal::Exit`/`Panic`/`Stopped` funnelled to a parent; `OneForOne`/`OneForAll`/`RestForOne` strategies + restart-intensity cap |
| `monitor` | `monitor(pid)` → `Monitor { id, target, rx }`; one-shot `Down` via `rx`; `demonitor(&m)` tears one registration down; unidirectional death notice |
| `link` | bidirectional `link`/`unlink`; abnormal death propagates (cooperative stop, or an `ExitSignal` message under `trap_exit`) |
| `gen_server` | `call`/`call_timeout` (sync request-reply) / `cast` (async) over one inbox; `handle_info` over static info arms + `handle_down` via `Watcher`-fed monitors, selected ahead of the inbox; `ServerRef`/`ServerBuilder` + `init`/`terminate` hooks; server-down via channel closure |
| `registry` | `register`/`whereis`/`name_of`: name ↔ pid bimap; lazy generation-checked cleanup |
## Quick taste ## Quick taste
@@ -51,6 +54,29 @@ run(|| {
}); });
``` ```
## Stopping actors
Two strengths, as in OTP. `request_stop(pid)` is `exit(Pid, kill)`: a cooperative
hard stop, unwinding at the actor's next observation point. `request_shutdown(pid)`
is `exit(Pid, shutdown)`: an actor that traps exits (`trap_exit()`, or
`ctx.trap_exit()` in a gen_server / `cx.trap_exit()` in a gen_statem) receives it
as a signal — `handle_shutdown` / a `shutdown` row — and may drain before stopping
itself; one that does not trap is stopped outright. Supervisors trap:
`request_shutdown(sup)` tears the tree down top-down, each child per its
`ChildSpec` `Shutdown` policy (`Timeout(d)`, `Infinity`, `BrutalKill`). The run's
root actor returning means "the program is done": every top-level actor gets a
`request_shutdown`, and `run()` returns when they are gone. From outside the
runtime (a signal thread), `Runtime::handle().request_shutdown(pid)` does the
same. `examples/graceful_shutdown.rs` shows all of it.
A gen_server or gen_statem lives until it stops, is shut down, or is killed; its
refs are addresses — dropping them never ends it (a forgotten one is swept at
root exit; with `--features smarm-trace` each such sweep is a `root_sweep`
trace line). The supervised shape is `GenServerBuilder::named(N).run()` /
`gen_statem::run_named(N, m)`: the server runs inline as the `ChildSpec` child
itself, so the supervisor's shutdown reaches it directly, a restart re-binds the
name, and the program addresses it by name.
## Layout ## Layout
``` ```
@@ -67,12 +93,9 @@ benches/
## Building and running ## Building and running
Standard Cargo. Requires Rust 1.95 or newer (the `#[naked]` attribute went stable Standard Cargo. Requires Rust 1.95 or newer (the `#[naked]` attribute went stable in 1.88; we use a few unrelated post-1.88 features). I have worked hard to keep this library as dependency-free as possible. `master` is x86-64 Linux only. An experimental, **untested** aarch64 context-switch backend lives on the `arm-port` branch (extracted into a `target_arch`-gated `src/arch/`); it has not been validated on hardware yet. macOS remains on the deferred list because of the epoll dependency.
in 1.88; we use a few unrelated post-1.88 features). `master` is x86-64 Linux
only. An experimental, **untested** aarch64 context-switch backend lives on the
`arm-port` branch (extracted into a `target_arch`-gated `src/arch/`); it has not
been validated on hardware yet. macOS remains on the deferred list because of the
epoll dependency.
```sh ```sh
cargo test # all tests cargo test # all tests
@@ -80,26 +103,29 @@ cargo test --test mutex # one module
cargo bench # primes benchmark vs tokio cargo bench # primes benchmark vs tokio
``` ```
## What's not here
See the **Defer** section of `Architecture.md`.
`join!` for handle groups, stack growth via remap,
hierarchical timer wheel, fd-wait timeouts, `Signal::Timeout`. Each is
mechanism we know how to add; none belongs in this iteration.
## Docs ## Docs
| Document | What it covers | | Document | What it covers |
|---|---| |---|---|
| [`Architecture.md`](./docs/Architecture.md) | Design intent, runtime model, and deferred work | | [`Architecture.md`](./docs/Architecture.md) | Design intent, runtime model, and deferred work |
| [`smarm - Deep Dive.html`](./docs/smarm%20-%20Deep%20Dive.html) | Generated walkthrough of the system; good starting point | | [`smarm - Deep Dive.html`](./docs/smarm%20-%20Deep%20Dive.html) | Generated walkthrough of the system; good starting point if you want to learn about the internals |
| [`BENCHMARKS_AND_TUNING.md`](./docs/BENCHMARKS_AND_TUNING.md) | Where smarm wins and loses vs tokio, preemption knob recommendations | | [`BENCHMARKS_AND_TUNING.md`](./docs/BENCHMARKS_AND_TUNING.md) | Where smarm wins and loses vs tokio, preemption knob recommendations |
| [`benchmarks.md`](./docs/benchmarks.md) | Raw benchmark results, methodology, and tuning experiment log | | [`benchmarks.md`](./docs/benchmarks.md) | Raw benchmark results, methodology, and tuning experiment log |
## Coming up
Clustering: clustering multiple SMARM nodes together is in the pipeline.
SMARM-BEAM Interop: Running SMARM as a supervised node under the BEAM via a Rustler NIF works, including message passing and supervision trees that span the runtimes. However, this library is still too unstable to release.
SMARM is an interesting platform for implementing a 'dataflow' library, but work on this has not yet started.
## Contributing ## Contributing
This is a personal proof-of-concept. There's no PR workflow. If you fork it and do something interesting, just send me an email. If it's nice, I'll upstream the changes. This started as a personal proof-of-concept, but it is starting to outgrow that name. If you want to contribute, please get in contact to discuss what you want to work on. Code without prior communication is not welcome.
## A note on open source
An open source project is a gift, and by giving it, it is no longer mine. I highly enourage you to fork it, to make it your own. This repository, however, is still mine.
--- ---
+48 -4
View File
@@ -77,10 +77,10 @@ Delivered surface:
`ServerBuilder::start` untouched; free `call` / `cast` / `whereis_server`; `ServerBuilder::start` untouched; free `call` / `cast` / `whereis_server`;
`ServerRef::shutdown` + free `shutdown` as the sys-style synchronous stop. `ServerRef::shutdown` + free `shutdown` as the sys-style synchronous stop.
- **Root-exit teardown** (final phase): the run's initial actor is the root; - **Root-exit teardown** (final phase): the run's initial actor is the root;
when it exits, the scheduler's idle verdict stops the parked-forever remainder when it exits the run winds down. *(Reworked with the graceful-shutdown work:
(deferred past the queue drain, so actors with in-flight work finish rather root exit now delivers `request_shutdown` to every forest root — see
than unwinding on the stop). Closes the "app actor blocks AllDone" stall — see "Root exit" below and `tests/root_exit.rs`.)* Closes the "app actor blocks
Look into, below. AllDone" stall — see Look into, below.
Extends — does not retire — the "select exists; a unified per-process mailbox Extends — does not retire — the "select exists; a unified per-process mailbox
still does not" invariant: 014 adds addressable *delivery*, not a unified inbox; still does not" invariant: 014 adds addressable *delivery*, not a unified inbox;
@@ -260,8 +260,52 @@ path the atomic-bool workaround stood in for. Re-check the urus crud repro to
confirm the workaround can be retired (the teardown is cooperative — an actor in confirm the workaround can be retired (the teardown is cooperative — an actor in
a tight loop with no observation point still can't be stopped). a tight loop with no observation point still can't be stopped).
**Update (graceful shutdown):** the RFC 014 sweep was a hard `request_stop` of
every live slot, deferred until nothing was runnable — which killed a sleeping
actor (timer pending) but drained a queued one, for no principled reason. It is
now the OTP semantics: root exit = "the program is done" = `request_shutdown`
to every **forest root** (live actor whose parent is the run or is dead), run
synchronously on the root's finalize path. Supervisors cascade with their child
`Shutdown` policies; trapping actors may `Continue`/drain (timers keep working)
and end the run when they stop themselves; non-trapping actors are stopped
outright — `join` what you need finished. No forcing sweep follows.
--- ---
### Open items from the graceful-shutdown work (not scheduled)
- ~~gen_server / gen_statem as a direct supervised child.~~ Done: lifetime is
the actor's (refs are addresses, the loop holds an inbox sender);
`NamedGenServerBuilder::run` / `gen_statem::run_named` run the loop inline as
the `ChildSpec` child; the root-exit sweep traces each leftover as
`root_sweep` under `smarm-trace`.
- Supervisor `Live` drop-guard sweep is `request_stop` (kill propagates as
kill); OTP would deliver a trappable `killed`. Chosen for boundedness.
- A root-exit shutdown reaches only actors live *at that instant*; a
non-trapping forest root that spawns before it unwinds leaves that spawn
to itself (Erlang: an unlinked spawn is nobody's child).
- **Supervisor start *order* is not start *readiness*.** `start_child`
spawns and moves straight on, so an earlier child is merely *scheduled*,
not initialised, when a later sibling starts. A later child that resolves
an earlier one by name (`whereis_server`) can therefore miss it — the
classic "named registry sibling, then its consumers" tree. Ordered
`OneForOne`/`RestForOne` shutdown is unaffected (reverse order is honoured
and each stop *is* awaited); this is a start-side gap only.
Making `spawn` itself block does NOT fix it — it would only shrink the
window to "child has begun executing", while the property callers need is
"child has bound its name / opened its socket", which only the child can
declare. It would also tax the hot path (one round-trip per accepted
connection) and turn every spawn into a context-switch point. OTP has the
same async `spawn` and puts the synchronisation one level up:
`gen_server:start_link` blocks the caller until `init/1` returns.
Fix shape when scheduled: a readiness ack in the supervisor's child-start
path (`ChildSpec` variant whose factory receives a ready-signal;
`NamedGenServerBuilder::run` acks after its name bind, gen_server default
acks after `init`; plain closures ack at spawn as today, i.e. opt-in with
no cost to existing children). Until then the workaround is structural:
have the registrar spawn its own consumers so the ordering is program
order inside one actor, not a cross-actor guarantee (urus v0.3 endpoint
does exactly this).
## Invariants & gotchas (respect these across all cycles) ## Invariants & gotchas (respect these across all cycles)
- **Shared mutex is non-reentrant.** `Sender::send` can call `unpark` → - **Shared mutex is non-reentrant.** `Sender::send` can call `unpark` →
+229
View File
@@ -0,0 +1,229 @@
# urus / smarm handoff — updated 2026-08-19 (session 3)
## TL;DR for the next session
**smarm is done for now** (5 unpushed commits on local `master`, see below).
**Next = urus v0.3 endpoint refactor.** You should NOT need to read smarm
scheduler internals; the contract you build on is fully described here and in
`smarm_full/examples/graceful_shutdown.rs` (read that file first — it is the
exact shape urus's tree will take) plus `smarm_full/tests/root_exit.rs`.
### The smarm contract urus builds on (all on local master, verified by tests)
- `request_stop(pid)` = kill (cooperative hard stop). `request_shutdown(pid)` =
polite: trapping target gets `ExitSignal{reason: Shutdown}`, non-trapping is
stopped outright. `RuntimeHandle::{request_stop,request_shutdown}` do the same
from any OS thread (signal handler); grab `rt.handle()` before `rt.run`.
- Supervisor traps; `request_shutdown(sup)` = ordered reverse-start shutdown,
per-child `ChildSpec::shutdown(Shutdown::{Timeout(d)|Infinity|BrutalKill})`
(default Timeout(5s)); sup then returns normally. `request_stop(sup)`
hard-stops children too (no orphans).
- gen_server: `ctx.trap_exit()` in init; `handle_shutdown() -> Exit|Continue`;
`handle_exit(sig)`; `ctx.stop_handle().stop()` = normal self-exit;
`terminate()` may block only on the graceful path (Exit / stop / inbox close).
`GenServerRef::shutdown()` is graceful and waits.
- gen_statem: same in event clothes — `cx.trap_exit()` in initial enter,
`shutdown` rows (default `stop`), `exit sig` rows, `cx.stop()` / `stop` tail,
optional `terminate { }` block. `GenStatemRef::shutdown()`.
- **Root exit = program done**: when the root actor returns, the runtime
`request_shutdown`s every *forest root* (live actor whose parent is the run
or dead). Supervisors cascade; trapping actors may drain (timers keep
working) and end the run when they stop; non-trapping are stopped; **no
forcing sweep** (`join` what must finish). The old "wait until nothing
runnable then kill all" deferral is gone.
- **Gotcha for urus:** a gen_server's lifetime is governed by its refs — drop
the last `GenServerRef` and the inbox closes → clean exit *even mid-drain*.
The endpoint must be pinned (named, or its ref held by the supervisor
wrapper) or it will terminate the moment the root drops its ref.
- **Known gap (ROADMAP open item):** a gen_server can't be a direct `ChildSpec`
child; use the trapping wrapper pattern in `examples/graceful_shutdown.rs::
drainer_child` (starts `under(self_pid())`, forwards shutdown, waits). Doing
an inline `GenServerBuilder::run()` first may be worth a short smarm detour —
decide with Markk.
### smarm commits this session (local master, NOT pushed, NOT tagged)
`250f312` root-exit = graceful shutdown of forest roots (tests/root_exit.rs)
`6ceb138` gen_statem shutdown parity (tests/gen_statem_shutdown.rs)
`849a424` docs + examples/graceful_shutdown.rs + README "Stopping actors"
On top of `1002777` (cross-thread wake) and `9c8f59c` (graceful shutdown).
Cargo.toml still `0.6.1`. Release cut (push, tag — v0.7 is justified by the
API surface — version bump) is Markk's. Full suite, doc tests, examples,
`cargo fmt`, `cargo clippy --lib` all clean. (`clippy --tests` has pre-existing
unwrap lints in tests/fd_select.rs, untouched.)
### Decisions taken this session (Markk)
- Root exit means "program done" (Go/tokio/OTP), not "wait for pending work";
the previously agreed "sleep(50ms) must finish" test was dropped as encoding
the wrong contract (a timer-wheel gate would re-wedge periodic-timer daemons).
- No behaviour-preserving deferral, no forcing second sweep.
- Examples/docs done in the same session; urus next session.
---
# Previous handoff (still accurate where not superseded above)
## Next-session goal
Phase 1 is **done and committed**; Phase 2 is next:
1. **smarm v0.6.2** — cross-thread wake root fix. **DONE**, committed on `master`
as `1002777`. Not yet tagged, not yet version-bumped (Cargo.toml still reads
`0.6.1`), and **not yet pushed to origin** — it exists only in the delivered
snapshot zip and the local sandbox clone. Cutting the release (push + tag
`v0.6.2` + bump `0.6.1`→`0.6.2`) is Markk's step.
2. **urus v0.3** — endpoint refactor. Working against `smarm = { path = "../smarm_full" }`
with `git update-index --skip-worktree Cargo.toml` (Markk approved); release commit
swaps back to the tag once Markk cuts it. Phase 2 plan below is STALE where it says
drain-in-terminate; the endpoint is a trapping GenServer: `handle_shutdown` →
`Continue`, enter Draining, `StopHandle::stop()` when the conn set empties.
Decisions below are locked unless marked *(confirm)*.
## Reconstruction (the sandbox resets between sessions)
A fresh sandbox has an empty home and **no Rust toolchain**. To restore:
- Install rustup/cargo. smarm reformats under **rustc 1.97.1**; urus `rust-version`
is 1.95. Use 1.97.1.
- urus: `git clone https://git.kalsbeek.dev/Markk116/urus` — `origin` is registered
and public-read. master `8bdec97` = the v0.2.x line. (Zips in outputs are stale;
prefer the remote now.)
- smarm: `git clone https://git.kalsbeek.dev/Markk116/smarm`. Latest tag **v0.6.1**
(`ca1c983`). The cross-thread wake fix is committed as `1002777` on top of the
post-v0.6.1 README commit `8f2d513` (= origin/master). **It is NOT on origin
yet** — a fresh clone won't have it until Markk pushes. Restore it from the
snapshot zip if working before the push. **v0.6.2 is not yet tagged.**
- urus pins smarm by git **tag** in `Cargo.toml` (currently `v0.6.0`). A trivial
first commit bumps it to `v0.6.1` (also picks up `try_spawn` + monitor
terminal-outcome fixes).
## Why (context — the finding that drives the plan)
urus's shutdown machinery (the `AtomicBool` listener flag + the `SHUTDOWN_POLL`
loop in `serve.rs`) is scaffolding around two smarm properties. Their statuses
differ, which is the whole point:
- **Issue A — lossy stop vs a QUEUED actor: ALREADY FIXED in smarm.** Commit
`7bab4d2` added an entry-side `check_cancelled()` in `park_current`. A
`request_stop` against a listener parked in `wait_readable_timeout` now unwinds
cleanly (it parks via `try_select_timeout → park_current`). urus's flag + its
stale "smarm's lossy stop-while-QUEUED window" comment can be deleted.
- **Issue B — foreign-thread wake is a no-op: FIXED in `1002777` (was present
through v0.6.1).** The gap: `unpark`/`unpark_at`/`request_stop` all route through
`try_with_runtime`, which reads a thread-local that is `None` on any non-scheduler
thread, so no cross-thread wake worked — a signal handler / OS thread could not
wake *or* stop a parked actor, which is why `serve.rs` polls the shutdown signal
instead of parking on it. Now closed (see Phase 1 below): urus's `SHUTDOWN_POLL`
loop can be deleted and its `Handle::shutdown` can park on a handle-driven stop.
## Phase 1 — smarm cross-thread wake (root fix) — DONE (`1002777`)
Shipped as one commit generalizing RFC 018 (a producer reaches the runtime through
a `Weak` it holds) from the IO backend to channel senders and a new handle:
- **`Runtime::handle() -> RuntimeHandle`** (`Send + Sync`), holding a
`Weak<RuntimeInner>`. Grab it before `rt.run` and hand it to the signal thread.
- **`RuntimeHandle::request_stop<A>(Pid<A>)`** — upgrades the Weak and calls
`request_stop_inner` on the inner; no-op if the runtime is gone. This is the
signal-handler-drives-shutdown path; it cascades the ordered stop down the tree
exactly like an in-runtime `request_stop`.
- **Send-wake:** the receiver captures `scheduler::runtime_weak()` into its
`parked_receiver` tuple **at park time** (not at channel creation — the resolved
sub-decision; a parked receiver is a live actor so the Weak is provably upgradable,
and it scopes the capture to when a wake is possible). `send()` and last-sender
`drop` wake via `scheduler::unpark_at_via(pid, epoch, &weak)`: thread-local path
when on a scheduler thread (preempt-gated, slot-eligible), captured Weak otherwise.
In-runtime timer wakes (recv/select) were left on `scheduler::unpark_at`.
**API scope decision (signed off):** `RuntimeHandle` exposes **`request_stop` only**.
No public `unpark`/`unpark_at` on the handle — send-wake needs no user-facing handle,
and "unpark off-runtime" is covered because `request_stop` drives `unpark` on the
upgraded inner. No `is_alive()`. Both are one-line additions if a consumer appears.
**No RFC written** — pattern was already established (RFC 018), agreed not needed.
Tests: `tests/cross_thread_wake.rs` (foreign-thread send wakes a parked receiver;
foreign-thread `request_stop` wakes+stops a parked actor; a lingering handle never
blocks all-done and degrades to a no-op once the runtime drops). Full suite green;
`cargo fmt` + `cargo clippy --lib` clean.
**Remaining release step (Markk):** push `master`, tag `v0.6.2`, bump Cargo.toml
`0.6.1`→`0.6.2`. Left paired with the tag as the release cut, not done in `1002777`.
## Phase 1b — smarm graceful shutdown (OTP lift) — DONE (`9c8f59c`, on top of `1002777`)
Decided this session (Markk): B — fix at the smarm level rather than a two-stop
split in urus. No RFC (Markk: "just implement it"). Shipped, tested, committed on
the local `master`, **not pushed, not tagged**. It should ship as the same
release as 1002777 (v0.6.2, or v0.7 given the API surface — Markk's call).
- `request_shutdown(pid)` / `RuntimeHandle::request_shutdown` = `exit(Pid, shutdown)`;
`request_stop` = `exit(Pid, kill)`. Trapping target gets `ExitSignal{reason:
DownReason::Shutdown}`; non-trapping is stopped outright.
- `ChildSpec::shutdown(Shutdown::{BrutalKill, Timeout(d), Infinity})`, default 5s.
Supervisor traps exits; `request_shutdown(sup)` = ordered top-down shutdown,
returns normally. **Also fixed**: `request_stop(sup)` used to ORPHAN children
(probe-verified; the handoff's "cascade" claim was wrong) — `Live` drop guard now
hard-stops them.
- gen_server: `ctx.trap_exit()`, `handle_shutdown() -> ShutdownAction::{Exit,
Continue}`, `handle_exit(ExitSignal)`, `ctx.stop_handle().stop()` = normal
self-exit (`{stop, normal}`; previously impossible — only abnormal `Stopped`).
`GenServerRef::shutdown()` is graceful now.
- Root finding that forced this: gen_server `terminate()` runs from a Drop guard,
mid-unwind on the stop path; any park in it = double panic = abort. So
"drain-in-terminate()" (the old Phase 2 plan) was never viable.
### ~~Next-session smarm work~~ DONE this session (see TL;DR)
1. **Root-exit sweep**: make `Pop::RootDrain` also require an empty timer wheel
(and it already requires nothing runnable; io_out is only checked for AllDone —
check whether it should gate RootDrain too). TDD: an actor in `sleep(50ms)` when
the root returns must finish, not be swept. Then a `Reservoir`-style test that a
*truly* parked-forever daemon still gets swept.
2. **gen_statem parity**: `ctx.trap_exit()`, `handle_shutdown -> ShutdownAction`,
`handle_exit`, stop handle. Mirror gen_server; mechanical.
3. **Examples review**: `examples/*.rs` predate all of this. Rework where they show
shutdown/teardown to use `request_shutdown`, `Shutdown` policies, and
`StopHandle`; `named_genserver.rs` first (uses `shutdown`). Also
`docs/smarm - Deep Dive.html` says terminate() must be non-blocking — now only
true on the unwind paths; and README could use a "Stopping actors" paragraph
(request_stop = kill, request_shutdown = shutdown, Shutdown policy).
### Open smarm items found on the way (noted, not scheduled)
- (root-exit sweep and gen_statem parity moved up to the scheduled list.)
- Sweep in the supervisor `Live` drop guard is `request_stop` (kill propagates as
kill); OTP would deliver a trappable `killed`. Chosen for boundedness.
## Phase 2 — urus v0.3: endpoint refactor (after v0.6.2 is tagged)
Target = the spec's original shape (`urus-spec.md` §2.1/§6: `listener_sup` under the
**user's** root supervisor). Deviation to unwind: `serve` owning `rt.run`.
- App owns the runtime: `smarm::init(cfg).run(|| root_sup.run())`, root e.g.
`RestForOne[ app actors…, urus::endpoint(config, pipeline) ]`. This is what kills
the `Arc<OnceLock>` idiom for the right reason (app state born in-runtime as a
supervised, ordered child).
- `urus::endpoint` = one GenServer child owning the registry + an **internal**
listener sub-supervisor + drain-in-`terminate()`. Listeners stay internal, not
app-visible peers.
- Shutdown = `request_stop` the root supervisor (or via the runtime handle from a
signal thread) → cascades down → `endpoint.terminate()` runs the drain
(`drain_timeout`, force-stop sweep).
- **DELETE:** the `AtomicBool` listener flag (A fixed) and the `SHUTDOWN_POLL` loop
+ its apologetic comment (B fixed → park, don't poll).
- Nuance: `request_stop` → `Signal::Stopped` is *abnormal* → `Transient` restarts.
Stop-without-restart = stop the **supervisor**, not the children.
- *(confirm)* Keep `serve`/`serve_with`/`serve_with_shutdown` as thin wrappers that
build the one-child tree internally, so the simple case stays one line.
- *(confirm)* Keep `Handle`/`ShutdownSignal`? Now that cross-thread wake works,
`Handle::shutdown` can map to handle-driven `request_stop` on the endpoint.
- Breaking → cut **urus v0.3**; bump the smarm pin to `v0.6.2` here.
## Working norms
- Every bash call: `export PATH=$HOME/.cargo/bin:$PATH` (once the toolchain's in).
- **TDD**: failing test first, then implement; keep suites green.
- **Hammer ritual** for ANY connection-lifecycle change (the urus #2 shutdown work
qualifies): 35× subset (`shutdown timeout reaped slowloris streaming chunked sse
stalled ws_ channels session`) + 3× full + 1× trace. `scripts/hammer.sh` does NOT
pass feature flags — loop manually with `--features phoenix`. Subset filter must
NOT use `--test integration` (session tests live in lib).
- Example smoke tests: hold the server's stdin open (`mkfifo` + `sleep > fifo`) or
the Enter-to-shutdown thread fires on EOF instantly.
- Background procs are reaped BETWEEN bash calls; `pkill -f` matches your own shell.
- Artefact store (specs): `curl -H "Authorization: Bearer sk-llmingest-2e45d80c63db24c6781e761eb2a9a58e83d9f48ef77a42185bad311d07c80e68" https://artefacts.kalsbeek.dev/artifacts/<name>`
— `urus-spec.md`, `urus-bench-spec.md`, `rfc_008-implementation-notes.md`, …
- smarm feature flags: `smarm-trace`, `smarm-causal` (urus re-exports both).
## Local cross-repo testing — KEEP OUT OF COMMITS
To test urus #2 against un-tagged smarm 0.6.2, point urus's `Cargo.toml` smarm dep
at a local path (`smarm = { path = "../smarm" }`) instead of the git tag.
- Must NOT land in commits. Guard: `git update-index --skip-worktree Cargo.toml`
after editing (undo with `--no-skip-worktree`), or stash before committing.
- The committed `Cargo.toml` stays pinned to the git tag; restore the tag (bumped to
`v0.6.2`) for the release commit.
+67 -26
View File
@@ -26,7 +26,9 @@ use std::time::Instant;
const ITERS: u32 = 15; const ITERS: u32 = 15;
fn available_threads() -> usize { fn available_threads() -> usize {
std::thread::available_parallelism().map(|n| n.get()).unwrap_or(1) std::thread::available_parallelism()
.map(|n| n.get())
.unwrap_or(1)
} }
fn env_sets() -> u32 { fn env_sets() -> u32 {
@@ -108,17 +110,15 @@ fn bench_chained_smarm(threads: usize) -> (u64, u128) {
fn bench_chained_tokio_current() -> (u64, u128) { fn bench_chained_tokio_current() -> (u64, u128) {
let counter = Arc::new(AtomicU64::new(0)); let counter = Arc::new(AtomicU64::new(0));
let c2 = counter.clone(); let c2 = counter.clone();
let rt = tokio::runtime::Builder::new_current_thread().build().unwrap(); let rt = tokio::runtime::Builder::new_current_thread()
.build()
.unwrap();
let start = Instant::now(); let start = Instant::now();
let local = tokio::task::LocalSet::new(); let local = tokio::task::LocalSet::new();
local.block_on(&rt, async move { local.block_on(&rt, async move {
// Use a oneshot done channel like tokio's own chained_spawn bench. // Use a oneshot done channel like tokio's own chained_spawn bench.
let (done_tx, done_rx) = tokio::sync::oneshot::channel(); let (done_tx, done_rx) = tokio::sync::oneshot::channel();
fn iter( fn iter(c: Arc<AtomicU64>, done: tokio::sync::oneshot::Sender<()>, n: u64) {
c: Arc<AtomicU64>,
done: tokio::sync::oneshot::Sender<()>,
n: u64,
) {
if n == 0 { if n == 0 {
let _ = done.send(()); let _ = done.send(());
} else { } else {
@@ -186,7 +186,9 @@ fn bench_yield_smarm(threads: usize) -> (u64, u128) {
} }
fn bench_yield_tokio_current() -> (u64, u128) { fn bench_yield_tokio_current() -> (u64, u128) {
let rt = tokio::runtime::Builder::new_current_thread().build().unwrap(); let rt = tokio::runtime::Builder::new_current_thread()
.build()
.unwrap();
let start = Instant::now(); let start = Instant::now();
let local = tokio::task::LocalSet::new(); let local = tokio::task::LocalSet::new();
local.block_on(&rt, async move { local.block_on(&rt, async move {
@@ -235,11 +237,22 @@ const PRIME_N: u64 = 400_000;
const PRIME_WORKERS: u64 = 64; const PRIME_WORKERS: u64 = 64;
fn is_prime(n: u64) -> bool { fn is_prime(n: u64) -> bool {
if n < 2 { return false; } if n < 2 {
if n < 4 { return true; } return false;
if n % 2 == 0 { return false; } }
if n < 4 {
return true;
}
if n % 2 == 0 {
return false;
}
let mut i = 3u64; let mut i = 3u64;
while i * i <= n { if n % i == 0 { return false; } i += 2; } while i * i <= n {
if n % i == 0 {
return false;
}
i += 2;
}
true true
} }
@@ -250,7 +263,11 @@ fn count_primes(lo: u64, hi: u64) -> u64 {
fn primes_slice(w: u64) -> (u64, u64) { fn primes_slice(w: u64) -> (u64, u64) {
let per = PRIME_N / PRIME_WORKERS; let per = PRIME_N / PRIME_WORKERS;
let lo = w * per; let lo = w * per;
let hi = if w + 1 == PRIME_WORKERS { PRIME_N } else { lo + per }; let hi = if w + 1 == PRIME_WORKERS {
PRIME_N
} else {
lo + per
};
(lo, hi) (lo, hi)
} }
@@ -267,7 +284,9 @@ fn bench_primes_smarm(threads: usize) -> (u64, u128) {
tc.fetch_add(count_primes(lo, hi), Ordering::Relaxed); tc.fetch_add(count_primes(lo, hi), Ordering::Relaxed);
})); }));
} }
for h in handles { h.join().unwrap(); } for h in handles {
h.join().unwrap();
}
}); });
(total.load(Ordering::Relaxed), start.elapsed().as_micros()) (total.load(Ordering::Relaxed), start.elapsed().as_micros())
} }
@@ -275,7 +294,9 @@ fn bench_primes_smarm(threads: usize) -> (u64, u128) {
fn bench_primes_tokio_current() -> (u64, u128) { fn bench_primes_tokio_current() -> (u64, u128) {
let total = Arc::new(AtomicU64::new(0)); let total = Arc::new(AtomicU64::new(0));
let t2 = total.clone(); let t2 = total.clone();
let rt = tokio::runtime::Builder::new_current_thread().build().unwrap(); let rt = tokio::runtime::Builder::new_current_thread()
.build()
.unwrap();
let start = Instant::now(); let start = Instant::now();
let local = tokio::task::LocalSet::new(); let local = tokio::task::LocalSet::new();
local.block_on(&rt, async move { local.block_on(&rt, async move {
@@ -287,7 +308,9 @@ fn bench_primes_tokio_current() -> (u64, u128) {
tc.fetch_add(count_primes(lo, hi), Ordering::Relaxed); tc.fetch_add(count_primes(lo, hi), Ordering::Relaxed);
})); }));
} }
for h in handles { let _ = h.await; } for h in handles {
let _ = h.await;
}
}); });
(total.load(Ordering::Relaxed), start.elapsed().as_micros()) (total.load(Ordering::Relaxed), start.elapsed().as_micros())
} }
@@ -309,7 +332,9 @@ fn bench_primes_tokio_multi() -> (u64, u128) {
tc.fetch_add(count_primes(lo, hi), Ordering::Relaxed); tc.fetch_add(count_primes(lo, hi), Ordering::Relaxed);
})); }));
} }
for h in handles { let _ = h.await; } for h in handles {
let _ = h.await;
}
}); });
(total.load(Ordering::Relaxed), start.elapsed().as_micros()) (total.load(Ordering::Relaxed), start.elapsed().as_micros())
} }
@@ -344,7 +369,9 @@ fn bench_pp_smarm(threads: usize) -> (u64, u128) {
} }
fn bench_pp_tokio_current() -> (u64, u128) { fn bench_pp_tokio_current() -> (u64, u128) {
let rt = tokio::runtime::Builder::new_current_thread().build().unwrap(); let rt = tokio::runtime::Builder::new_current_thread()
.build()
.unwrap();
let start = Instant::now(); let start = Instant::now();
let local = tokio::task::LocalSet::new(); let local = tokio::task::LocalSet::new();
local.block_on(&rt, async move { local.block_on(&rt, async move {
@@ -395,7 +422,6 @@ fn bench_pp_tokio_multi() -> (u64, u128) {
// main // main
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// Knob helper — reads SMARM_ALLOC_INTERVAL / SMARM_TIMESLICE_CYCLES env vars // Knob helper — reads SMARM_ALLOC_INTERVAL / SMARM_TIMESLICE_CYCLES env vars
// so the sweep script can override the preemption knobs without recompiling. // so the sweep script can override the preemption knobs without recompiling.
@@ -404,10 +430,14 @@ fn bench_pp_tokio_multi() -> (u64, u128) {
fn bench_cfg(threads: usize) -> smarm::runtime::Config { fn bench_cfg(threads: usize) -> smarm::runtime::Config {
let mut cfg = smarm::runtime::Config::exact(threads); let mut cfg = smarm::runtime::Config::exact(threads);
if let Ok(v) = std::env::var("SMARM_ALLOC_INTERVAL") { if let Ok(v) = std::env::var("SMARM_ALLOC_INTERVAL") {
if let Ok(n) = v.parse::<u32>() { cfg = cfg.alloc_interval(n); } if let Ok(n) = v.parse::<u32>() {
cfg = cfg.alloc_interval(n);
}
} }
if let Ok(v) = std::env::var("SMARM_TIMESLICE_CYCLES") { if let Ok(v) = std::env::var("SMARM_TIMESLICE_CYCLES") {
if let Ok(n) = v.parse::<u64>() { cfg = cfg.timeslice_cycles(n); } if let Ok(n) = v.parse::<u64>() {
cfg = cfg.timeslice_cycles(n);
}
} }
cfg cfg
} }
@@ -417,7 +447,10 @@ fn main() {
println!("smarm general benchmarks"); println!("smarm general benchmarks");
println!("available parallelism: {n} threads"); println!("available parallelism: {n} threads");
let sets = env_sets(); let sets = env_sets();
println!("ITERS={ITERS}×{sets} sets = {} samples (+1 warmup, discarded)", ITERS * sets); println!(
"ITERS={ITERS}×{sets} sets = {} samples (+1 warmup, discarded)",
ITERS * sets
);
println!( println!(
"CHAIN_DEPTH={CHAIN_DEPTH}, YIELD_TASKS={YIELD_TASKS}×{YIELD_ROUNDS}, \ "CHAIN_DEPTH={CHAIN_DEPTH}, YIELD_TASKS={YIELD_TASKS}×{YIELD_ROUNDS}, \
PRIME_N={PRIME_N}/{PRIME_WORKERS} workers, PP_ROUNDS={PP_ROUNDS}" PRIME_N={PRIME_N}/{PRIME_WORKERS} workers, PP_ROUNDS={PP_ROUNDS}"
@@ -426,21 +459,29 @@ fn main() {
// ---- 1. chained_spawn ---- // ---- 1. chained_spawn ----
print_header(&format!("chained_spawn: depth {CHAIN_DEPTH}")); print_header(&format!("chained_spawn: depth {CHAIN_DEPTH}"));
run_n("smarm 1-thread", ITERS, || bench_chained_smarm(1)); run_n("smarm 1-thread", ITERS, || bench_chained_smarm(1));
run_n(&format!("smarm {n}-thread"), ITERS, || bench_chained_smarm(n)); run_n(&format!("smarm {n}-thread"), ITERS, || {
bench_chained_smarm(n)
});
run_n("tokio current_thread", ITERS, bench_chained_tokio_current); run_n("tokio current_thread", ITERS, bench_chained_tokio_current);
run_n("tokio multi-thread", ITERS, bench_chained_tokio_multi); run_n("tokio multi-thread", ITERS, bench_chained_tokio_multi);
// ---- 2. yield_many ---- // ---- 2. yield_many ----
print_header(&format!("yield_many: {YIELD_TASKS} tasks × {YIELD_ROUNDS} yields")); print_header(&format!(
"yield_many: {YIELD_TASKS} tasks × {YIELD_ROUNDS} yields"
));
run_n("smarm 1-thread", ITERS, || bench_yield_smarm(1)); run_n("smarm 1-thread", ITERS, || bench_yield_smarm(1));
run_n(&format!("smarm {n}-thread"), ITERS, || bench_yield_smarm(n)); run_n(&format!("smarm {n}-thread"), ITERS, || bench_yield_smarm(n));
run_n("tokio current_thread", ITERS, bench_yield_tokio_current); run_n("tokio current_thread", ITERS, bench_yield_tokio_current);
run_n("tokio multi-thread", ITERS, bench_yield_tokio_multi); run_n("tokio multi-thread", ITERS, bench_yield_tokio_multi);
// ---- 3. fan_out_compute ---- // ---- 3. fan_out_compute ----
print_header(&format!("fan_out_compute: primes in [2, {PRIME_N}) across {PRIME_WORKERS}")); print_header(&format!(
"fan_out_compute: primes in [2, {PRIME_N}) across {PRIME_WORKERS}"
));
run_n("smarm 1-thread", ITERS, || bench_primes_smarm(1)); run_n("smarm 1-thread", ITERS, || bench_primes_smarm(1));
run_n(&format!("smarm {n}-thread"), ITERS, || bench_primes_smarm(n)); run_n(&format!("smarm {n}-thread"), ITERS, || {
bench_primes_smarm(n)
});
run_n("tokio current_thread", ITERS, bench_primes_tokio_current); run_n("tokio current_thread", ITERS, bench_primes_tokio_current);
run_n("tokio multi-thread", ITERS, bench_primes_tokio_multi); run_n("tokio multi-thread", ITERS, bench_primes_tokio_multi);
+76 -27
View File
@@ -64,11 +64,22 @@ const PRIME_N: u64 = 400_000;
const WORKERS: u64 = 64; const WORKERS: u64 = 64;
fn is_prime(n: u64) -> bool { fn is_prime(n: u64) -> bool {
if n < 2 { return false; } if n < 2 {
if n < 4 { return true; } return false;
if n % 2 == 0 { return false; } }
if n < 4 {
return true;
}
if n % 2 == 0 {
return false;
}
let mut i = 3u64; let mut i = 3u64;
while i * i <= n { if n % i == 0 { return false; } i += 2; } while i * i <= n {
if n % i == 0 {
return false;
}
i += 2;
}
true true
} }
@@ -96,7 +107,9 @@ fn bench_primes_smarm(threads: usize) -> (u64, u128) {
tc.fetch_add(count_primes(lo, hi), Ordering::Relaxed); tc.fetch_add(count_primes(lo, hi), Ordering::Relaxed);
})); }));
} }
for h in handles { h.join().unwrap(); } for h in handles {
h.join().unwrap();
}
}); });
(total.load(Ordering::Relaxed), start.elapsed().as_micros()) (total.load(Ordering::Relaxed), start.elapsed().as_micros())
} }
@@ -104,7 +117,9 @@ fn bench_primes_smarm(threads: usize) -> (u64, u128) {
fn bench_primes_tokio_current() -> (u64, u128) { fn bench_primes_tokio_current() -> (u64, u128) {
let total = Arc::new(AtomicU64::new(0)); let total = Arc::new(AtomicU64::new(0));
let t2 = total.clone(); let t2 = total.clone();
let rt = tokio::runtime::Builder::new_current_thread().build().unwrap(); let rt = tokio::runtime::Builder::new_current_thread()
.build()
.unwrap();
let start = Instant::now(); let start = Instant::now();
let local = tokio::task::LocalSet::new(); let local = tokio::task::LocalSet::new();
local.block_on(&rt, async move { local.block_on(&rt, async move {
@@ -116,7 +131,9 @@ fn bench_primes_tokio_current() -> (u64, u128) {
tc.fetch_add(count_primes(lo, hi), Ordering::Relaxed); tc.fetch_add(count_primes(lo, hi), Ordering::Relaxed);
})); }));
} }
for h in handles { let _ = h.await; } for h in handles {
let _ = h.await;
}
}); });
(total.load(Ordering::Relaxed), start.elapsed().as_micros()) (total.load(Ordering::Relaxed), start.elapsed().as_micros())
} }
@@ -138,17 +155,21 @@ fn bench_primes_tokio_multi() -> (u64, u128) {
tc.fetch_add(count_primes(lo, hi), Ordering::Relaxed); tc.fetch_add(count_primes(lo, hi), Ordering::Relaxed);
})); }));
} }
for h in handles { let _ = h.await; } for h in handles {
let _ = h.await;
}
}); });
(total.load(Ordering::Relaxed), start.elapsed().as_micros()) (total.load(Ordering::Relaxed), start.elapsed().as_micros())
} }
fn bench_primes_baseline() -> (u64, u128) { fn bench_primes_baseline() -> (u64, u128) {
let start = Instant::now(); let start = Instant::now();
let total: u64 = (0..WORKERS).map(|w| { let total: u64 = (0..WORKERS)
.map(|w| {
let (lo, hi) = primes_slice(w); let (lo, hi) = primes_slice(w);
count_primes(lo, hi) count_primes(lo, hi)
}).sum(); })
.sum();
(total, start.elapsed().as_micros()) (total, start.elapsed().as_micros())
} }
@@ -167,15 +188,17 @@ fn bench_pingpong_smarm(threads: usize) -> (u64, u128) {
tx_a.send(0).unwrap(); tx_a.send(0).unwrap();
loop { loop {
let v = rx_b.recv().unwrap(); let v = rx_b.recv().unwrap();
if v >= PING_ROUNDS { break; } if v >= PING_ROUNDS {
break;
}
tx_a.send(v + 1).unwrap(); tx_a.send(v + 1).unwrap();
} }
}); });
let hb = smarm::spawn(move || { let hb = smarm::spawn(move || loop {
loop {
let v = rx_a.recv().unwrap(); let v = rx_a.recv().unwrap();
tx_b.send(v + 1).unwrap(); tx_b.send(v + 1).unwrap();
if v + 1 >= PING_ROUNDS { break; } if v + 1 >= PING_ROUNDS {
break;
} }
}); });
ha.join().unwrap(); ha.join().unwrap();
@@ -198,7 +221,9 @@ fn bench_pingpong_tokio_current() -> (u64, u128) {
tx_a.send(0).unwrap(); tx_a.send(0).unwrap();
loop { loop {
let v = rx_b.recv().await.unwrap(); let v = rx_b.recv().await.unwrap();
if v >= PING_ROUNDS { break; } if v >= PING_ROUNDS {
break;
}
tx_a.send(v + 1).unwrap(); tx_a.send(v + 1).unwrap();
} }
}); });
@@ -206,7 +231,9 @@ fn bench_pingpong_tokio_current() -> (u64, u128) {
loop { loop {
let v = rx_a.recv().await.unwrap(); let v = rx_a.recv().await.unwrap();
tx_b.send(v + 1).unwrap(); tx_b.send(v + 1).unwrap();
if v + 1 >= PING_ROUNDS { break; } if v + 1 >= PING_ROUNDS {
break;
}
} }
}); });
let _ = ha.await; let _ = ha.await;
@@ -229,7 +256,9 @@ fn bench_pingpong_tokio_multi() -> (u64, u128) {
tx_a.send(0).unwrap(); tx_a.send(0).unwrap();
loop { loop {
let v = rx_b.recv().await.unwrap(); let v = rx_b.recv().await.unwrap();
if v >= PING_ROUNDS { break; } if v >= PING_ROUNDS {
break;
}
tx_a.send(v + 1).unwrap(); tx_a.send(v + 1).unwrap();
} }
}); });
@@ -237,7 +266,9 @@ fn bench_pingpong_tokio_multi() -> (u64, u128) {
loop { loop {
let v = rx_a.recv().await.unwrap(); let v = rx_a.recv().await.unwrap();
tx_b.send(v + 1).unwrap(); tx_b.send(v + 1).unwrap();
if v + 1 >= PING_ROUNDS { break; } if v + 1 >= PING_ROUNDS {
break;
}
} }
}); });
let _ = ha.await; let _ = ha.await;
@@ -264,7 +295,9 @@ fn bench_spawn_smarm(threads: usize) -> (u64, u128) {
cc.fetch_add(1, Ordering::Relaxed); cc.fetch_add(1, Ordering::Relaxed);
})); }));
} }
for h in handles { h.join().unwrap(); } for h in handles {
h.join().unwrap();
}
}); });
(counter.load(Ordering::Relaxed), start.elapsed().as_micros()) (counter.load(Ordering::Relaxed), start.elapsed().as_micros())
} }
@@ -272,7 +305,9 @@ fn bench_spawn_smarm(threads: usize) -> (u64, u128) {
fn bench_spawn_tokio_current() -> (u64, u128) { fn bench_spawn_tokio_current() -> (u64, u128) {
let counter = Arc::new(AtomicU64::new(0)); let counter = Arc::new(AtomicU64::new(0));
let c = counter.clone(); let c = counter.clone();
let rt = tokio::runtime::Builder::new_current_thread().build().unwrap(); let rt = tokio::runtime::Builder::new_current_thread()
.build()
.unwrap();
let start = Instant::now(); let start = Instant::now();
let local = tokio::task::LocalSet::new(); let local = tokio::task::LocalSet::new();
local.block_on(&rt, async move { local.block_on(&rt, async move {
@@ -283,7 +318,9 @@ fn bench_spawn_tokio_current() -> (u64, u128) {
cc.fetch_add(1, Ordering::Relaxed); cc.fetch_add(1, Ordering::Relaxed);
})); }));
} }
for h in handles { let _ = h.await; } for h in handles {
let _ = h.await;
}
}); });
(counter.load(Ordering::Relaxed), start.elapsed().as_micros()) (counter.load(Ordering::Relaxed), start.elapsed().as_micros())
} }
@@ -304,7 +341,9 @@ fn bench_spawn_tokio_multi() -> (u64, u128) {
cc.fetch_add(1, Ordering::Relaxed); cc.fetch_add(1, Ordering::Relaxed);
})); }));
} }
for h in handles { let _ = h.await; } for h in handles {
let _ = h.await;
}
}); });
(counter.load(Ordering::Relaxed), start.elapsed().as_micros()) (counter.load(Ordering::Relaxed), start.elapsed().as_micros())
} }
@@ -320,22 +359,32 @@ fn main() {
println!("PRIME_N={PRIME_N}, WORKERS={WORKERS}, PING_ROUNDS={PING_ROUNDS}, SPAWN_COUNT={SPAWN_COUNT}"); println!("PRIME_N={PRIME_N}, WORKERS={WORKERS}, PING_ROUNDS={PING_ROUNDS}, SPAWN_COUNT={SPAWN_COUNT}");
// ---- Primes ---- // ---- Primes ----
print_header(&format!("Fan-out/fan-in: count primes in [2, {PRIME_N}) across {WORKERS} workers")); print_header(&format!(
"Fan-out/fan-in: count primes in [2, {PRIME_N}) across {WORKERS} workers"
));
run_n("baseline (serial)", ITERS, bench_primes_baseline); run_n("baseline (serial)", ITERS, bench_primes_baseline);
run_n("smarm single-thread", ITERS, || bench_primes_smarm(1)); run_n("smarm single-thread", ITERS, || bench_primes_smarm(1));
run_n(&format!("smarm {n}-thread"), ITERS, || bench_primes_smarm(n)); run_n(&format!("smarm {n}-thread"), ITERS, || {
bench_primes_smarm(n)
});
run_n("tokio current_thread", ITERS, bench_primes_tokio_current); run_n("tokio current_thread", ITERS, bench_primes_tokio_current);
run_n("tokio multi-thread", ITERS, bench_primes_tokio_multi); run_n("tokio multi-thread", ITERS, bench_primes_tokio_multi);
// ---- Ping-pong ---- // ---- Ping-pong ----
print_header(&format!("Ping-pong: {PING_ROUNDS} round-trips between two actors")); print_header(&format!(
"Ping-pong: {PING_ROUNDS} round-trips between two actors"
));
run_n("smarm single-thread", ITERS, || bench_pingpong_smarm(1)); run_n("smarm single-thread", ITERS, || bench_pingpong_smarm(1));
run_n(&format!("smarm {n}-thread"), ITERS, || bench_pingpong_smarm(n)); run_n(&format!("smarm {n}-thread"), ITERS, || {
bench_pingpong_smarm(n)
});
run_n("tokio current_thread", ITERS, bench_pingpong_tokio_current); run_n("tokio current_thread", ITERS, bench_pingpong_tokio_current);
run_n("tokio multi-thread", ITERS, bench_pingpong_tokio_multi); run_n("tokio multi-thread", ITERS, bench_pingpong_tokio_multi);
// ---- Spawn throughput ---- // ---- Spawn throughput ----
print_header(&format!("Spawn throughput: {SPAWN_COUNT} actors spawned and joined")); print_header(&format!(
"Spawn throughput: {SPAWN_COUNT} actors spawned and joined"
));
run_n("smarm single-thread", ITERS, || bench_spawn_smarm(1)); run_n("smarm single-thread", ITERS, || bench_spawn_smarm(1));
run_n(&format!("smarm {n}-thread"), ITERS, || bench_spawn_smarm(n)); run_n(&format!("smarm {n}-thread"), ITERS, || bench_spawn_smarm(n));
run_n("tokio current_thread", ITERS, bench_spawn_tokio_current); run_n("tokio current_thread", ITERS, bench_spawn_tokio_current);
+24 -7
View File
@@ -16,12 +16,20 @@ const WORKERS: u64 = 16;
const ITERATIONS: u32 = 5; const ITERATIONS: u32 = 5;
fn is_prime(n: u64) -> bool { fn is_prime(n: u64) -> bool {
if n < 2 { return false; } if n < 2 {
if n < 4 { return true; } return false;
if n % 2 == 0 { return false; } }
if n < 4 {
return true;
}
if n % 2 == 0 {
return false;
}
let mut i = 3u64; let mut i = 3u64;
while i * i <= n { while i * i <= n {
if n % i == 0 { return false; } if n % i == 0 {
return false;
}
i += 2; i += 2;
} }
true true
@@ -30,7 +38,9 @@ fn is_prime(n: u64) -> bool {
fn count_primes_in(lo: u64, hi: u64) -> u64 { fn count_primes_in(lo: u64, hi: u64) -> u64 {
let mut count = 0u64; let mut count = 0u64;
for n in lo..hi { for n in lo..hi {
if is_prime(n) { count += 1; } if is_prime(n) {
count += 1;
}
} }
count count
} }
@@ -38,7 +48,11 @@ fn count_primes_in(lo: u64, hi: u64) -> u64 {
fn slice(worker: u64) -> (u64, u64) { fn slice(worker: u64) -> (u64, u64) {
let per = N / WORKERS; let per = N / WORKERS;
let lo = worker * per; let lo = worker * per;
let hi = if worker + 1 == WORKERS { N } else { (worker + 1) * per }; let hi = if worker + 1 == WORKERS {
N
} else {
(worker + 1) * per
};
(lo, hi) (lo, hi)
} }
@@ -125,7 +139,10 @@ fn main() {
"Counting primes in [2, {}) across {} workers, {} iterations each\n", "Counting primes in [2, {}) across {} workers, {} iterations each\n",
N, WORKERS, ITERATIONS N, WORKERS, ITERATIONS
); );
println!("{:>12} | {:>15} | {:>16} | {:>15} | {:>15}", "runtime", "primes found", "median", "min", "max"); println!(
"{:>12} | {:>15} | {:>16} | {:>15} | {:>15}",
"runtime", "primes found", "median", "min", "max"
);
println!("{}", "-".repeat(80)); println!("{}", "-".repeat(80));
run_n("baseline", ITERATIONS, bench_baseline); run_n("baseline", ITERATIONS, bench_baseline);
+44 -7
View File
@@ -27,12 +27,19 @@ use std::sync::Arc;
use std::time::Instant; use std::time::Instant;
fn env_usize(key: &str, default: usize) -> usize { fn env_usize(key: &str, default: usize) -> usize {
std::env::var(key).ok().and_then(|v| v.parse().ok()).unwrap_or(default) std::env::var(key)
.ok()
.and_then(|v| v.parse().ok())
.unwrap_or(default)
} }
fn env_threads() -> Vec<usize> { fn env_threads() -> Vec<usize> {
std::env::var("SMARM_BENCH_THREADS") std::env::var("SMARM_BENCH_THREADS")
.map(|v| v.split_whitespace().filter_map(|t| t.parse().ok()).collect()) .map(|v| {
v.split_whitespace()
.filter_map(|t| t.parse().ok())
.collect()
})
.unwrap_or_else(|_| vec![1, 2, 4]) .unwrap_or_else(|_| vec![1, 2, 4])
} }
@@ -53,7 +60,11 @@ fn drive<Q: Send + Sync + 'static>(
for p in 0..producers { for p in 0..producers {
let q = q.clone(); let q = q.clone();
// Give the last producer the remainder. // Give the last producer the remainder.
let n = if p == producers - 1 { items - per * (producers - 1) } else { per }; let n = if p == producers - 1 {
items - per * (producers - 1)
} else {
per
};
hs.push(std::thread::spawn(move || { hs.push(std::thread::spawn(move || {
let pid = Pid::new(p as u32, 0); let pid = Pid::new(p as u32, 0);
for _ in 0..n { for _ in 0..n {
@@ -132,7 +143,12 @@ fn main() {
for &t in &threads_sweep { for &t in &threads_sweep {
for (p, c) in ratios_for(t) { for (p, c) in ratios_for(t) {
for s in ["mutex", "mpmc", "striped"] { for s in ["mutex", "mpmc", "striped"] {
cases.push(Case { structure: s, threads: t, producers: p, consumers: c }); cases.push(Case {
structure: s,
threads: t,
producers: p,
consumers: c,
});
} }
} }
} }
@@ -147,7 +163,14 @@ fn main() {
if case.threads < 2 { if case.threads < 2 {
drive_single(&*q, MutexQueue::push, MutexQueue::pop, items) drive_single(&*q, MutexQueue::push, MutexQueue::pop, items)
} else { } else {
drive(q, MutexQueue::push, MutexQueue::pop, case.producers, case.consumers, items) drive(
q,
MutexQueue::push,
MutexQueue::pop,
case.producers,
case.consumers,
items,
)
} }
} }
"mpmc" => { "mpmc" => {
@@ -155,7 +178,14 @@ fn main() {
if case.threads < 2 { if case.threads < 2 {
drive_single(&*q, MpmcRing::push, MpmcRing::pop, items) drive_single(&*q, MpmcRing::push, MpmcRing::pop, items)
} else { } else {
drive(q, MpmcRing::push, MpmcRing::pop, case.producers, case.consumers, items) drive(
q,
MpmcRing::push,
MpmcRing::pop,
case.producers,
case.consumers,
items,
)
} }
} }
"striped" => { "striped" => {
@@ -163,7 +193,14 @@ fn main() {
if case.threads < 2 { if case.threads < 2 {
drive_single(&*q, StripedRing::push, StripedRing::pop, items) drive_single(&*q, StripedRing::push, StripedRing::pop, items)
} else { } else {
drive(q, StripedRing::push, StripedRing::pop, case.producers, case.consumers, items) drive(
q,
StripedRing::push,
StripedRing::pop,
case.producers,
case.consumers,
items,
)
} }
} }
_ => unreachable!(), _ => unreachable!(),
+21 -4
View File
@@ -54,12 +54,19 @@ fn variant() -> &'static str {
} }
fn env_usize(key: &str, default: usize) -> usize { fn env_usize(key: &str, default: usize) -> usize {
std::env::var(key).ok().and_then(|v| v.parse().ok()).unwrap_or(default) std::env::var(key)
.ok()
.and_then(|v| v.parse().ok())
.unwrap_or(default)
} }
fn env_threads() -> Vec<usize> { fn env_threads() -> Vec<usize> {
std::env::var("SMARM_BENCH_THREADS") std::env::var("SMARM_BENCH_THREADS")
.map(|v| v.split_whitespace().filter_map(|t| t.parse().ok()).collect()) .map(|v| {
v.split_whitespace()
.filter_map(|t| t.parse().ok())
.collect()
})
.unwrap_or_else(|_| vec![1, 2, 4]) .unwrap_or_else(|_| vec![1, 2, 4])
} }
@@ -238,12 +245,22 @@ fn main() {
); );
println!( println!(
"RQCSV,runtime,{},{},{},{},{},{},{}", "RQCSV,runtime,{},{},{},{},{},{},{}",
variant(), slot_str, name, t, work, mid.us, per_s variant(),
slot_str,
name,
t,
work,
mid.us,
per_s
); );
if slot { if slot {
println!( println!(
"RQSLOT,{},{},{},{},{}", "RQSLOT,{},{},{},{},{}",
variant(), name, t, mid.hits, mid.displacements variant(),
name,
t,
mid.hits,
mid.displacements
); );
} }
} }
+55 -19
View File
@@ -37,7 +37,9 @@ use std::time::Instant;
const ITERS: u32 = 15; const ITERS: u32 = 15;
fn available_threads() -> usize { fn available_threads() -> usize {
std::thread::available_parallelism().map(|n| n.get()).unwrap_or(1) std::thread::available_parallelism()
.map(|n| n.get())
.unwrap_or(1)
} }
fn env_sets() -> u32 { fn env_sets() -> u32 {
@@ -116,7 +118,9 @@ fn bench_recurse_smarm(threads: usize) -> (u64, u128) {
fn bench_recurse_tokio_current() -> (u64, u128) { fn bench_recurse_tokio_current() -> (u64, u128) {
let counter = Arc::new(AtomicU64::new(0)); let counter = Arc::new(AtomicU64::new(0));
let c2 = counter.clone(); let c2 = counter.clone();
let rt = tokio::runtime::Builder::new_current_thread().build().unwrap(); let rt = tokio::runtime::Builder::new_current_thread()
.build()
.unwrap();
let start = Instant::now(); let start = Instant::now();
let local = tokio::task::LocalSet::new(); let local = tokio::task::LocalSet::new();
local.block_on(&rt, async move { local.block_on(&rt, async move {
@@ -199,7 +203,9 @@ fn bench_hot_smarm() -> (u64, u128) {
} }
fn bench_hot_tokio_current() -> (u64, u128) { fn bench_hot_tokio_current() -> (u64, u128) {
let rt = tokio::runtime::Builder::new_current_thread().build().unwrap(); let rt = tokio::runtime::Builder::new_current_thread()
.build()
.unwrap();
let start = Instant::now(); let start = Instant::now();
let local = tokio::task::LocalSet::new(); let local = tokio::task::LocalSet::new();
local.block_on(&rt, async move { local.block_on(&rt, async move {
@@ -249,7 +255,9 @@ fn bench_unc_smarm() -> (u64, u128) {
} }
fn bench_unc_tokio_current() -> (u64, u128) { fn bench_unc_tokio_current() -> (u64, u128) {
let rt = tokio::runtime::Builder::new_current_thread().build().unwrap(); let rt = tokio::runtime::Builder::new_current_thread()
.build()
.unwrap();
let start = Instant::now(); let start = Instant::now();
let local = tokio::task::LocalSet::new(); let local = tokio::task::LocalSet::new();
local.block_on(&rt, async move { local.block_on(&rt, async move {
@@ -297,8 +305,12 @@ fn bench_panic_smarm(threads: usize) -> (u64, u128) {
} }
for h in handles { for h in handles {
match h.join() { match h.join() {
Ok(()) => { ok2.fetch_add(1, Ordering::Relaxed); } Ok(()) => {
Err(_) => { err2.fetch_add(1, Ordering::Relaxed); } ok2.fetch_add(1, Ordering::Relaxed);
}
Err(_) => {
err2.fetch_add(1, Ordering::Relaxed);
}
} }
} }
}); });
@@ -312,7 +324,9 @@ fn bench_panic_tokio_current() -> (u64, u128) {
let err = Arc::new(AtomicU64::new(0)); let err = Arc::new(AtomicU64::new(0));
let ok2 = ok.clone(); let ok2 = ok.clone();
let err2 = err.clone(); let err2 = err.clone();
let rt = tokio::runtime::Builder::new_current_thread().build().unwrap(); let rt = tokio::runtime::Builder::new_current_thread()
.build()
.unwrap();
let prev_hook = std::panic::take_hook(); let prev_hook = std::panic::take_hook();
std::panic::set_hook(Box::new(|_| {})); std::panic::set_hook(Box::new(|_| {}));
let start = Instant::now(); let start = Instant::now();
@@ -328,8 +342,12 @@ fn bench_panic_tokio_current() -> (u64, u128) {
} }
for h in handles { for h in handles {
match h.await { match h.await {
Ok(()) => { ok2.fetch_add(1, Ordering::Relaxed); } Ok(()) => {
Err(_) => { err2.fetch_add(1, Ordering::Relaxed); } ok2.fetch_add(1, Ordering::Relaxed);
}
Err(_) => {
err2.fetch_add(1, Ordering::Relaxed);
}
} }
} }
}); });
@@ -361,8 +379,12 @@ fn bench_panic_tokio_multi() -> (u64, u128) {
} }
for h in handles { for h in handles {
match h.await { match h.await {
Ok(()) => { ok2.fetch_add(1, Ordering::Relaxed); } Ok(()) => {
Err(_) => { err2.fetch_add(1, Ordering::Relaxed); } ok2.fetch_add(1, Ordering::Relaxed);
}
Err(_) => {
err2.fetch_add(1, Ordering::Relaxed);
}
} }
} }
}); });
@@ -375,7 +397,6 @@ fn bench_panic_tokio_multi() -> (u64, u128) {
// main // main
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// Knob helper — reads SMARM_ALLOC_INTERVAL / SMARM_TIMESLICE_CYCLES env vars // Knob helper — reads SMARM_ALLOC_INTERVAL / SMARM_TIMESLICE_CYCLES env vars
// so the sweep script can override the preemption knobs without recompiling. // so the sweep script can override the preemption knobs without recompiling.
@@ -384,10 +405,14 @@ fn bench_panic_tokio_multi() -> (u64, u128) {
fn bench_cfg(threads: usize) -> smarm::runtime::Config { fn bench_cfg(threads: usize) -> smarm::runtime::Config {
let mut cfg = smarm::runtime::Config::exact(threads); let mut cfg = smarm::runtime::Config::exact(threads);
if let Ok(v) = std::env::var("SMARM_ALLOC_INTERVAL") { if let Ok(v) = std::env::var("SMARM_ALLOC_INTERVAL") {
if let Ok(n) = v.parse::<u32>() { cfg = cfg.alloc_interval(n); } if let Ok(n) = v.parse::<u32>() {
cfg = cfg.alloc_interval(n);
}
} }
if let Ok(v) = std::env::var("SMARM_TIMESLICE_CYCLES") { if let Ok(v) = std::env::var("SMARM_TIMESLICE_CYCLES") {
if let Ok(n) = v.parse::<u64>() { cfg = cfg.timeslice_cycles(n); } if let Ok(n) = v.parse::<u64>() {
cfg = cfg.timeslice_cycles(n);
}
} }
cfg cfg
} }
@@ -397,7 +422,10 @@ fn main() {
println!("smarm smarm-favored benchmarks"); println!("smarm smarm-favored benchmarks");
println!("available parallelism: {n} threads"); println!("available parallelism: {n} threads");
let sets = env_sets(); let sets = env_sets();
println!("ITERS={ITERS}×{sets} sets = {} samples (+1 warmup, discarded)", ITERS * sets); println!(
"ITERS={ITERS}×{sets} sets = {} samples (+1 warmup, discarded)",
ITERS * sets
);
println!( println!(
"RECURSE_DEPTH={RECURSE_DEPTH}, HOT_YIELDS={HOT_YIELDS}×2, \ "RECURSE_DEPTH={RECURSE_DEPTH}, HOT_YIELDS={HOT_YIELDS}×2, \
UNCONT_MSGS={UNCONT_MSGS}, PANIC_TASKS={PANIC_TASKS}" UNCONT_MSGS={UNCONT_MSGS}, PANIC_TASKS={PANIC_TASKS}"
@@ -406,22 +434,30 @@ fn main() {
// ---- 9. deep_recursion ---- // ---- 9. deep_recursion ----
print_header(&format!("deep_recursion: depth {RECURSE_DEPTH}")); print_header(&format!("deep_recursion: depth {RECURSE_DEPTH}"));
run_n("smarm 1-thread", ITERS, || bench_recurse_smarm(1)); run_n("smarm 1-thread", ITERS, || bench_recurse_smarm(1));
run_n(&format!("smarm {n}-thread"), ITERS, || bench_recurse_smarm(n)); run_n(&format!("smarm {n}-thread"), ITERS, || {
bench_recurse_smarm(n)
});
run_n("tokio current_thread", ITERS, bench_recurse_tokio_current); run_n("tokio current_thread", ITERS, bench_recurse_tokio_current);
run_n("tokio multi-thread", ITERS, bench_recurse_tokio_multi); run_n("tokio multi-thread", ITERS, bench_recurse_tokio_multi);
// ---- 10. yield_in_hot_loop ---- // ---- 10. yield_in_hot_loop ----
print_header(&format!("yield_in_hot_loop: 2 actors × {HOT_YIELDS} yields (single thread)")); print_header(&format!(
"yield_in_hot_loop: 2 actors × {HOT_YIELDS} yields (single thread)"
));
run_n("smarm 1-thread", ITERS, bench_hot_smarm); run_n("smarm 1-thread", ITERS, bench_hot_smarm);
run_n("tokio current_thread", ITERS, bench_hot_tokio_current); run_n("tokio current_thread", ITERS, bench_hot_tokio_current);
// ---- 11. uncontended_channel ---- // ---- 11. uncontended_channel ----
print_header(&format!("uncontended_channel: 1→1, {UNCONT_MSGS} msgs (single thread)")); print_header(&format!(
"uncontended_channel: 1→1, {UNCONT_MSGS} msgs (single thread)"
));
run_n("smarm 1-thread", ITERS, bench_unc_smarm); run_n("smarm 1-thread", ITERS, bench_unc_smarm);
run_n("tokio current_thread", ITERS, bench_unc_tokio_current); run_n("tokio current_thread", ITERS, bench_unc_tokio_current);
// ---- 12. catch_unwind_panics ---- // ---- 12. catch_unwind_panics ----
print_header(&format!("catch_unwind_panics: {PANIC_TASKS} tasks, 50% panic")); print_header(&format!(
"catch_unwind_panics: {PANIC_TASKS} tasks, 50% panic"
));
run_n("smarm 1-thread", ITERS, || bench_panic_smarm(1)); run_n("smarm 1-thread", ITERS, || bench_panic_smarm(1));
run_n(&format!("smarm {n}-thread"), ITERS, || bench_panic_smarm(n)); run_n(&format!("smarm {n}-thread"), ITERS, || bench_panic_smarm(n));
run_n("tokio current_thread", ITERS, bench_panic_tokio_current); run_n("tokio current_thread", ITERS, bench_panic_tokio_current);
+30 -5
View File
@@ -73,7 +73,10 @@ fn variant() -> &'static str {
} }
fn env_usize(key: &str, default: usize) -> usize { fn env_usize(key: &str, default: usize) -> usize {
std::env::var(key).ok().and_then(|v| v.parse().ok()).unwrap_or(default) std::env::var(key)
.ok()
.and_then(|v| v.parse().ok())
.unwrap_or(default)
} }
// -------------------------------------------------------------------------- // --------------------------------------------------------------------------
@@ -226,7 +229,11 @@ fn main() {
let mean_cyc = pooled_cyc.iter().map(|&v| v as f64).sum::<f64>() / n.max(1) as f64; let mean_cyc = pooled_cyc.iter().map(|&v| v as f64).sum::<f64>() / n.max(1) as f64;
// Derived effective frequency: cycles per ns = GHz. Cross-checks the two // Derived effective frequency: cycles per ns = GHz. Cross-checks the two
// lenses against the box's known base clock. // lenses against the box's known base clock.
let derived_ghz = if mean_ns > 0.0 { mean_cyc / mean_ns } else { 0.0 }; let derived_ghz = if mean_ns > 0.0 {
mean_cyc / mean_ns
} else {
0.0
};
let p50 = pct(&pooled_ns, 50.0); let p50 = pct(&pooled_ns, 50.0);
let p90 = pct(&pooled_ns, 90.0); let p90 = pct(&pooled_ns, 90.0);
@@ -241,8 +248,14 @@ fn main() {
" rounds={} warmup={} runs={} (instrumentation floor: {} ns / {} cyc, subtracted)", " rounds={} warmup={} runs={} (instrumentation floor: {} ns / {} cyc, subtracted)",
rounds, warmup, runs, floor_ns, floor_cyc rounds, warmup, runs, floor_ns, floor_cyc
); );
println!(" {:<10} {:<10} {:<10} {:<10} {:<10}", "p50 ns", "p90 ns", "p99 ns", "min ns", "max ns"); println!(
println!(" {:<10} {:<10} {:<10} {:<10} {:<10}", p50, p90, p99, lo, hi); " {:<10} {:<10} {:<10} {:<10} {:<10}",
"p50 ns", "p90 ns", "p99 ns", "min ns", "max ns"
);
println!(
" {:<10} {:<10} {:<10} {:<10} {:<10}",
p50, p90, p99, lo, hi
);
println!( println!(
" mean {:.1} ns | mean {:.0} cyc | derived {:.3} GHz", " mean {:.1} ns | mean {:.0} cyc | derived {:.3} GHz",
mean_ns, mean_cyc, derived_ghz mean_ns, mean_cyc, derived_ghz
@@ -251,6 +264,18 @@ fn main() {
// Greppable line — same spirit as SPINCSV. // Greppable line — same spirit as SPINCSV.
println!( println!(
"SWITCHCSV,{},{},{},{},{},{},{},{},{},{},{:.1},{:.0},{:.3}", "SWITCHCSV,{},{},{},{},{},{},{},{},{},{},{:.1},{:.0},{:.3}",
variant(), mode, rounds, runs, n, p50, p90, p99, lo, hi, mean_ns, mean_cyc, derived_ghz variant(),
mode,
rounds,
runs,
n,
p50,
p90,
p99,
lo,
hi,
mean_ns,
mean_cyc,
derived_ghz
); );
} }
+105 -33
View File
@@ -36,7 +36,9 @@ use std::time::{Duration, Instant};
const ITERS: u32 = 15; const ITERS: u32 = 15;
fn available_threads() -> usize { fn available_threads() -> usize {
std::thread::available_parallelism().map(|n| n.get()).unwrap_or(1) std::thread::available_parallelism()
.map(|n| n.get())
.unwrap_or(1)
} }
fn env_sets() -> u32 { fn env_sets() -> u32 {
@@ -114,11 +116,15 @@ fn bench_storm_smarm(threads: usize) -> (u64, u128) {
cc.fetch_add(1, Ordering::Relaxed); cc.fetch_add(1, Ordering::Relaxed);
})); }));
} }
for h in handles { h.join().unwrap(); } for h in handles {
h.join().unwrap();
}
// Tear down background. // Tear down background.
s2.store(true, Ordering::Relaxed); s2.store(true, Ordering::Relaxed);
for h in bg_handles { h.join().unwrap(); } for h in bg_handles {
h.join().unwrap();
}
}); });
(counter.load(Ordering::Relaxed), start.elapsed().as_micros()) (counter.load(Ordering::Relaxed), start.elapsed().as_micros())
} }
@@ -129,7 +135,9 @@ fn bench_storm_tokio_current() -> (u64, u128) {
let c2 = counter.clone(); let c2 = counter.clone();
let s2 = stop.clone(); let s2 = stop.clone();
let rt = tokio::runtime::Builder::new_current_thread().build().unwrap(); let rt = tokio::runtime::Builder::new_current_thread()
.build()
.unwrap();
let start = Instant::now(); let start = Instant::now();
let local = tokio::task::LocalSet::new(); let local = tokio::task::LocalSet::new();
local.block_on(&rt, async move { local.block_on(&rt, async move {
@@ -149,9 +157,13 @@ fn bench_storm_tokio_current() -> (u64, u128) {
cc.fetch_add(1, Ordering::Relaxed); cc.fetch_add(1, Ordering::Relaxed);
})); }));
} }
for h in handles { let _ = h.await; } for h in handles {
let _ = h.await;
}
s2.store(true, Ordering::Relaxed); s2.store(true, Ordering::Relaxed);
for h in bg_handles { let _ = h.await; } for h in bg_handles {
let _ = h.await;
}
}); });
(counter.load(Ordering::Relaxed), start.elapsed().as_micros()) (counter.load(Ordering::Relaxed), start.elapsed().as_micros())
} }
@@ -184,9 +196,13 @@ fn bench_storm_tokio_multi() -> (u64, u128) {
cc.fetch_add(1, Ordering::Relaxed); cc.fetch_add(1, Ordering::Relaxed);
})); }));
} }
for h in handles { let _ = h.await; } for h in handles {
let _ = h.await;
}
s2.store(true, Ordering::Relaxed); s2.store(true, Ordering::Relaxed);
for h in bg_handles { let _ = h.await; } for h in bg_handles {
let _ = h.await;
}
}); });
(counter.load(Ordering::Relaxed), start.elapsed().as_micros()) (counter.load(Ordering::Relaxed), start.elapsed().as_micros())
} }
@@ -219,14 +235,21 @@ fn bench_mpsc_smarm(threads: usize) -> (u64, u128) {
} }
let _ = count; // discard; run() closure must return () let _ = count; // discard; run() closure must return ()
}); });
for h in prod_handles { h.join().unwrap(); } for h in prod_handles {
h.join().unwrap();
}
let _ = consumer.join().unwrap(); let _ = consumer.join().unwrap();
}); });
(MPSC_PRODUCERS * MPSC_PER_PRODUCER, start.elapsed().as_micros()) (
MPSC_PRODUCERS * MPSC_PER_PRODUCER,
start.elapsed().as_micros(),
)
} }
fn bench_mpsc_tokio_current() -> (u64, u128) { fn bench_mpsc_tokio_current() -> (u64, u128) {
let rt = tokio::runtime::Builder::new_current_thread().build().unwrap(); let rt = tokio::runtime::Builder::new_current_thread()
.build()
.unwrap();
let start = Instant::now(); let start = Instant::now();
let local = tokio::task::LocalSet::new(); let local = tokio::task::LocalSet::new();
local.block_on(&rt, async move { local.block_on(&rt, async move {
@@ -248,10 +271,15 @@ fn bench_mpsc_tokio_current() -> (u64, u128) {
} }
count count
}); });
for h in prod_handles { let _ = h.await; } for h in prod_handles {
let _ = h.await;
}
let _ = consumer.await; let _ = consumer.await;
}); });
(MPSC_PRODUCERS * MPSC_PER_PRODUCER, start.elapsed().as_micros()) (
MPSC_PRODUCERS * MPSC_PER_PRODUCER,
start.elapsed().as_micros(),
)
} }
fn bench_mpsc_tokio_multi() -> (u64, u128) { fn bench_mpsc_tokio_multi() -> (u64, u128) {
@@ -279,10 +307,15 @@ fn bench_mpsc_tokio_multi() -> (u64, u128) {
} }
count count
}); });
for h in prod_handles { let _ = h.await; } for h in prod_handles {
let _ = h.await;
}
let _ = consumer.await; let _ = consumer.await;
}); });
(MPSC_PRODUCERS * MPSC_PER_PRODUCER, start.elapsed().as_micros()) (
MPSC_PRODUCERS * MPSC_PER_PRODUCER,
start.elapsed().as_micros(),
)
} }
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -308,7 +341,9 @@ fn bench_timers_smarm(threads: usize) -> (u64, u128) {
smarm::sleep(Duration::from_millis(ms)); smarm::sleep(Duration::from_millis(ms));
})); }));
} }
for h in handles { h.join().unwrap(); } for h in handles {
h.join().unwrap();
}
}); });
(TIMER_ACTORS, start.elapsed().as_micros()) (TIMER_ACTORS, start.elapsed().as_micros())
} }
@@ -328,7 +363,9 @@ fn bench_timers_tokio_current() -> (u64, u128) {
tokio::time::sleep(Duration::from_millis(ms)).await; tokio::time::sleep(Duration::from_millis(ms)).await;
})); }));
} }
for h in handles { let _ = h.await; } for h in handles {
let _ = h.await;
}
}); });
(TIMER_ACTORS, start.elapsed().as_micros()) (TIMER_ACTORS, start.elapsed().as_micros())
} }
@@ -348,7 +385,9 @@ fn bench_timers_tokio_multi() -> (u64, u128) {
tokio::time::sleep(Duration::from_millis(ms)).await; tokio::time::sleep(Duration::from_millis(ms)).await;
})); }));
} }
for h in handles { let _ = h.await; } for h in handles {
let _ = h.await;
}
}); });
(TIMER_ACTORS, start.elapsed().as_micros()) (TIMER_ACTORS, start.elapsed().as_micros())
} }
@@ -361,11 +400,22 @@ const SCALING_N: u64 = 400_000;
const SCALING_WORKERS: u64 = 64; const SCALING_WORKERS: u64 = 64;
fn is_prime(n: u64) -> bool { fn is_prime(n: u64) -> bool {
if n < 2 { return false; } if n < 2 {
if n < 4 { return true; } return false;
if n % 2 == 0 { return false; } }
if n < 4 {
return true;
}
if n % 2 == 0 {
return false;
}
let mut i = 3u64; let mut i = 3u64;
while i * i <= n { if n % i == 0 { return false; } i += 2; } while i * i <= n {
if n % i == 0 {
return false;
}
i += 2;
}
true true
} }
@@ -376,7 +426,11 @@ fn count_primes(lo: u64, hi: u64) -> u64 {
fn scaling_slice(w: u64) -> (u64, u64) { fn scaling_slice(w: u64) -> (u64, u64) {
let per = SCALING_N / SCALING_WORKERS; let per = SCALING_N / SCALING_WORKERS;
let lo = w * per; let lo = w * per;
let hi = if w + 1 == SCALING_WORKERS { SCALING_N } else { lo + per }; let hi = if w + 1 == SCALING_WORKERS {
SCALING_N
} else {
lo + per
};
(lo, hi) (lo, hi)
} }
@@ -393,7 +447,9 @@ fn bench_scaling_smarm(threads: usize) -> (u64, u128) {
tc.fetch_add(count_primes(lo, hi), Ordering::Relaxed); tc.fetch_add(count_primes(lo, hi), Ordering::Relaxed);
})); }));
} }
for h in handles { h.join().unwrap(); } for h in handles {
h.join().unwrap();
}
}); });
(total.load(Ordering::Relaxed), start.elapsed().as_micros()) (total.load(Ordering::Relaxed), start.elapsed().as_micros())
} }
@@ -415,7 +471,9 @@ fn bench_scaling_tokio_multi(threads: usize) -> (u64, u128) {
tc.fetch_add(count_primes(lo, hi), Ordering::Relaxed); tc.fetch_add(count_primes(lo, hi), Ordering::Relaxed);
})); }));
} }
for h in handles { let _ = h.await; } for h in handles {
let _ = h.await;
}
}); });
(total.load(Ordering::Relaxed), start.elapsed().as_micros()) (total.load(Ordering::Relaxed), start.elapsed().as_micros())
} }
@@ -424,7 +482,6 @@ fn bench_scaling_tokio_multi(threads: usize) -> (u64, u128) {
// main // main
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// Knob helper — reads SMARM_ALLOC_INTERVAL / SMARM_TIMESLICE_CYCLES env vars // Knob helper — reads SMARM_ALLOC_INTERVAL / SMARM_TIMESLICE_CYCLES env vars
// so the sweep script can override the preemption knobs without recompiling. // so the sweep script can override the preemption knobs without recompiling.
@@ -433,10 +490,14 @@ fn bench_scaling_tokio_multi(threads: usize) -> (u64, u128) {
fn bench_cfg(threads: usize) -> smarm::runtime::Config { fn bench_cfg(threads: usize) -> smarm::runtime::Config {
let mut cfg = smarm::runtime::Config::exact(threads); let mut cfg = smarm::runtime::Config::exact(threads);
if let Ok(v) = std::env::var("SMARM_ALLOC_INTERVAL") { if let Ok(v) = std::env::var("SMARM_ALLOC_INTERVAL") {
if let Ok(n) = v.parse::<u32>() { cfg = cfg.alloc_interval(n); } if let Ok(n) = v.parse::<u32>() {
cfg = cfg.alloc_interval(n);
}
} }
if let Ok(v) = std::env::var("SMARM_TIMESLICE_CYCLES") { if let Ok(v) = std::env::var("SMARM_TIMESLICE_CYCLES") {
if let Ok(n) = v.parse::<u64>() { cfg = cfg.timeslice_cycles(n); } if let Ok(n) = v.parse::<u64>() {
cfg = cfg.timeslice_cycles(n);
}
} }
cfg cfg
} }
@@ -446,7 +507,10 @@ fn main() {
println!("smarm tokio-favored benchmarks"); println!("smarm tokio-favored benchmarks");
println!("available parallelism: {n} threads"); println!("available parallelism: {n} threads");
let sets = env_sets(); let sets = env_sets();
println!("ITERS={ITERS}×{sets} sets = {} samples (+1 warmup, discarded)", ITERS * sets); println!(
"ITERS={ITERS}×{sets} sets = {} samples (+1 warmup, discarded)",
ITERS * sets
);
println!( println!(
"STORM_BACKGROUND={STORM_BACKGROUND}, STORM_SPAWN={STORM_SPAWN}, \ "STORM_BACKGROUND={STORM_BACKGROUND}, STORM_SPAWN={STORM_SPAWN}, \
MPSC={MPSC_PRODUCERS}×{MPSC_PER_PRODUCER}, \ MPSC={MPSC_PRODUCERS}×{MPSC_PER_PRODUCER}, \
@@ -477,7 +541,9 @@ fn main() {
"many_timers: {TIMER_ACTORS} actors sleeping {TIMER_MIN_MS}–{TIMER_MAX_MS} ms" "many_timers: {TIMER_ACTORS} actors sleeping {TIMER_MIN_MS}–{TIMER_MAX_MS} ms"
)); ));
run_n("smarm 1-thread", ITERS, || bench_timers_smarm(1)); run_n("smarm 1-thread", ITERS, || bench_timers_smarm(1));
run_n(&format!("smarm {n}-thread"), ITERS, || bench_timers_smarm(n)); run_n(&format!("smarm {n}-thread"), ITERS, || {
bench_timers_smarm(n)
});
run_n("tokio current_thread", ITERS, bench_timers_tokio_current); run_n("tokio current_thread", ITERS, bench_timers_tokio_current);
run_n("tokio multi-thread", ITERS, bench_timers_tokio_multi); run_n("tokio multi-thread", ITERS, bench_timers_tokio_multi);
@@ -487,13 +553,19 @@ fn main() {
)); ));
let sweep: Vec<usize> = { let sweep: Vec<usize> = {
let mut v = vec![1usize, 2, 4]; let mut v = vec![1usize, 2, 4];
if n > 4 && !v.contains(&n) { v.push(n); } if n > 4 && !v.contains(&n) {
v.push(n);
}
v.into_iter().filter(|t| *t <= n).collect() v.into_iter().filter(|t| *t <= n).collect()
}; };
for t in &sweep { for t in &sweep {
run_n(&format!("smarm {t}-thread"), ITERS, || bench_scaling_smarm(*t)); run_n(&format!("smarm {t}-thread"), ITERS, || {
bench_scaling_smarm(*t)
});
} }
for t in &sweep { for t in &sweep {
run_n(&format!("tokio multi {t}-thread"), ITERS, || bench_scaling_tokio_multi(*t)); run_n(&format!("tokio multi {t}-thread"), ITERS, || {
bench_scaling_tokio_multi(*t)
});
} }
} }
+11
View File
@@ -0,0 +1,11 @@
fn main() {
// RFC 019 §7 test canary (agreed Q3): compiled without stack-clash
// protection so its 96 KiB local is a genuine one-displacement guard
// jumper; distro-hardened compilers would otherwise probe it page-wise
// and defeat the test's purpose.
cc::Build::new()
.file("canary/canary.c")
.flag_if_supported("-fno-stack-clash-protection")
.compile("smarm_canary");
println!("cargo:rerun-if-changed=canary/canary.c");
}
+14
View File
@@ -0,0 +1,14 @@
/* RFC 019 §7 FFI canary: an honest unprobed C frame with a 96 KiB local,
* touched from its LOW end first — the exact "one sub rsp steps over a small
* guard" pattern the RFC's motivating incident hit (a cargo-vendored gz
* build; cc-invoked builds do not enable -fstack-clash-protection, and this
* file pins that off explicitly so the canary stays a canary even on
* hardened-default toolchains). */
void smarm_canary_burn(void) {
volatile char buf[96 * 1024];
buf[0] = 1; /* deepest address first */
for (unsigned i = 0; i < sizeof buf; i += 4096) {
buf[i] = (char)i;
}
buf[sizeof buf - 1] = 1;
}
+1 -1
View File
@@ -75,7 +75,7 @@ genuine advantage over tokio's task abort model.
### Spawn-heavy workloads (19–70×) ### Spawn-heavy workloads (19–70×)
Every smarm actor `mmap`s a 64 KiB stack with a guard page. This is Every smarm actor `mmap`s a 64 KiB stack reserve with a 64 KiB PROT_NONE guard below (both per-actor configurable since RFC 019; the reserve is demand-paged). This is
a syscall. Tokio tasks are heap-allocated state machines — no stack, a syscall. Tokio tasks are heap-allocated state machines — no stack,
no syscall, ~100 bytes each. For workloads that spawn thousands of no syscall, ~100 bytes each. For workloads that spawn thousands of
short-lived actors per second, this is a structural disadvantage. short-lived actors per second, this is a structural disadvantage.
+4 -3
View File
@@ -1620,8 +1620,9 @@
wait: <code>select</code> priority is <strong>Down arms › Watcher arm › info channels (declaration order) › wait: <code>select</code> priority is <strong>Down arms › Watcher arm › info channels (declaration order) ›
inbox</strong>, rebuilt each turn. A hot inbox can't starve a death notice or a system message; inbox</strong>, rebuilt each turn. A hot inbox can't starve a death notice or a system message;
conversely a hot info channel <em>can</em> starve the inbox — deliberately. A closed info arm is conversely a hot info channel <em>can</em> starve the inbox — deliberately. A closed info arm is
silently dropped from the set; a closed <em>inbox</em> (every <code>ServerRef</code> gone) is graceful silently dropped from the set. The inbox never closes — the loop holds one sender for its whole
shutdown.</p> life, so a <code>GenServerRef</code> is an address, not an owner: the server ends only by
<code>StopHandle::stop</code>, a shutdown, a hard stop, or a panic.</p>
<h3>Death needs no monitor</h3> <h3>Death needs no monitor</h3>
<p>Server death detection falls out of channel closure. Already dead → the inbox is closed and <p>Server death detection falls out of channel closure. Already dead → the inbox is closed and
@@ -1700,7 +1701,7 @@
</div> </div>
<div class="module-card"> <div class="module-card">
<div class="module-name" style="color:var(--red)">Panics in <code>terminate()</code></div> <div class="module-name" style="color:var(--red)">Panics in <code>terminate()</code></div>
<p>gen_server's <code>terminate()</code> runs from a drop guard, possibly mid-unwind. A panic inside it during an unwind is a double panic → process abort, no supervision tree to save you. Keep it cheap, non-blocking, non-panicking.</p> <p>gen_server's <code>terminate()</code> runs from a drop guard, possibly mid-unwind. A panic inside it during an unwind is a double panic → process abort, no supervision tree to save you. On the panic and hard-stop paths keep it cheap, non-blocking, non-panicking. Only the graceful path (<code>handle_shutdown → Exit</code>, <code>StopHandle::stop</code>) runs it outside an unwind, where it may do real work.</p>
</div> </div>
<div class="module-card"> <div class="module-card">
<div class="module-name" style="color:var(--yellow)">Cold locks are leaf locks</div> <div class="module-name" style="color:var(--yellow)">Cold locks are leaf locks</div>
+3 -1
View File
@@ -67,7 +67,9 @@ fn main() {
println!("calibration: {per_us} work iters/µs"); println!("calibration: {per_us} work iters/µs");
let work_us = move |us: u64| work_iters(us * per_us); let work_us = move |us: u64| work_iters(us * per_us);
let cores = std::thread::available_parallelism().map(|n| n.get()).unwrap_or(1); let cores = std::thread::available_parallelism()
.map(|n| n.get())
.unwrap_or(1);
println!("cores: {cores}"); println!("cores: {cores}");
if cores < 4 { if cores < 4 {
println!("probe: SKIPPED (needs the stages in parallel)"); println!("probe: SKIPPED (needs the stages in parallel)");
+8 -8
View File
@@ -35,8 +35,8 @@
#![deny(dead_code, unreachable_patterns)] #![deny(dead_code, unreachable_patterns)]
use smarm::gen_statem::{spawn, Cx, GenStatemRef, Machine, Reply, Resolution, Step};
use smarm::run; use smarm::run;
use smarm::gen_statem::{spawn, Cx, Machine, Reply, Resolution, Step, GenStatemRef};
// === user types ============================================================ // === user types ============================================================
@@ -123,7 +123,11 @@ impl DoorSm {
fn start(init: Door) -> GenStatemRef<DoorSm> { fn start(init: Door) -> GenStatemRef<DoorSm> {
spawn(DoorSm { spawn(DoorSm {
state: init, state: init,
data: Data { enters: 0, pushes: 0, knocks: 0 }, data: Data {
enters: 0,
pushes: 0,
knocks: 0,
},
}) })
} }
@@ -193,16 +197,12 @@ impl Machine for DoorSm {
(Door::Closed, Ev::Cast(Cast::Push | Cast::Unlock(_))) => Resolution::Unhandled, (Door::Closed, Ev::Cast(Cast::Push | Cast::Unlock(_))) => Resolution::Unhandled,
// --- Locked (branching row: handler picks within UnlockOutcome) - // --- Locked (branching row: handler picks within UnlockOutcome) -
(Door::Locked, Ev::Cast(Cast::Unlock(key))) => { (Door::Locked, Ev::Cast(Cast::Unlock(key))) => Resolution::To(on_unlock(key).into()),
Resolution::To(on_unlock(key).into())
}
// Routed out in phase 1; listed only to keep this match total. // Routed out in phase 1; listed only to keep this match total.
(Door::Locked, Ev::Cast(Cast::Knock)) => { (Door::Locked, Ev::Cast(Cast::Knock)) => {
unreachable!("postponed event is replayed, not dispatched here") unreachable!("postponed event is replayed, not dispatched here")
} }
(Door::Locked, Ev::Cast(Cast::Push | Cast::Pull | Cast::Lock)) => { (Door::Locked, Ev::Cast(Cast::Push | Cast::Pull | Cast::Lock)) => Resolution::Unhandled,
Resolution::Unhandled
}
// --- state-independent queries (reply, then stay) --------------- // --- state-independent queries (reply, then stay) ---------------
(_, Ev::Call(Call::GetState(r))) => { (_, Ev::Call(Call::GetState(r))) => {
+9 -2
View File
@@ -18,8 +18,8 @@
// dispatch's own unreachable_patterns internally. // dispatch's own unreachable_patterns internally.
use smarm::gen_statem; use smarm::gen_statem;
use smarm::run;
use smarm::gen_statem::Reply; use smarm::gen_statem::Reply;
use smarm::run;
// === user types (identical to gen_statem_expanded.rs) ========================= // === user types (identical to gen_statem_expanded.rs) =========================
@@ -135,7 +135,14 @@ gen_statem! {
fn main() { fn main() {
run(|| { run(|| {
let door = DoorSm::start(Door::Closed, Data { enters: 0, pushes: 0, knocks: 0 }); let door = DoorSm::start(
Door::Closed,
Data {
enters: 0,
pushes: 0,
knocks: 0,
},
);
door.send(Ev::Cast(Cast::Lock)).unwrap(); // Closed -> Locked door.send(Ev::Cast(Cast::Lock)).unwrap(); // Closed -> Locked
door.send(Ev::Cast(Cast::Knock)).unwrap(); // Locked: postponed (not yet counted) door.send(Ev::Cast(Cast::Knock)).unwrap(); // Locked: postponed (not yet counted)
+139
View File
@@ -0,0 +1,139 @@
//! Graceful shutdown, end to end: a supervised app tree, a server that
//! drains before it exits, and the two ways the whole thing winds down.
//!
//! Stopping an actor comes in two strengths, as in OTP:
//! - `request_stop(pid)` = `exit(Pid, kill)`: cooperative hard stop,
//! unwinds at the next observation point.
//! - `request_shutdown(pid)` = `exit(Pid, shutdown)`: a trapping target gets
//! an `ExitSignal { reason: Shutdown }` and winds
//! down on its own terms; a non-trapping one is
//! stopped outright.
//!
//! A supervisor traps exits. `request_shutdown(sup)` runs its ordered
//! shutdown — children in reverse start order, each per its `ChildSpec`
//! `Shutdown` policy (`Timeout(d)` default 5s, `Infinity`, `BrutalKill`) —
//! and the supervisor then returns normally.
//!
//! Two triggers are shown:
//! 1. **Root exit.** The run's root actor returning means "the program is
//! done": the runtime delivers `request_shutdown` to every top-level actor
//! (here: the supervisor). Trapping actors may keep running to drain and
//! end the run when they stop themselves; non-trapping ones are stopped.
//! 2. **An outside thread** (e.g. a signal handler) driving it via
//! `RuntimeHandle::request_shutdown` on the supervisor — the root then
//! just waits for the tree to come down.
use smarm::gen_server::{
GenServer, GenServerBuilder, GenServerCtx, GenServerName, ShutdownAction, StopHandle,
TimerHandle,
};
use smarm::supervisor::{ChildSpec, OneForOne, Restart, Shutdown};
use smarm::{sleep, spawn};
use std::thread;
use std::time::Duration;
/// A server with in-flight work: on shutdown it stops accepting, finishes what
/// it has (simulated with a ticking timer), then ends itself.
struct Drainer {
pending: u32,
stop: Option<StopHandle<Drainer>>,
timer: Option<TimerHandle<Drainer>>,
}
impl GenServer for Drainer {
type Call = ();
type Reply = ();
type Cast = ();
type Info = ();
type Timer = ();
fn init(&mut self, ctx: &GenServerCtx<Self>) {
ctx.trap_exit(); // opt in: shutdown arrives as handle_shutdown
self.stop = Some(ctx.stop_handle());
self.timer = Some(ctx.timer());
}
fn handle_call(&mut self, _: ()) {}
fn handle_cast(&mut self, _: ()) {}
fn handle_shutdown(&mut self) -> ShutdownAction {
println!(
"drainer: shutdown requested, {} items pending",
self.pending
);
self.timer
.as_ref()
.unwrap()
.tick_every(Duration::from_millis(20), ());
ShutdownAction::Continue // keep serving until drained
}
fn handle_timer(&mut self, _: ()) {
self.pending -= 1;
if self.pending == 0 {
println!("drainer: drained, stopping");
self.stop.as_ref().unwrap().stop(); // normal exit
}
}
fn terminate(&mut self) {
// Graceful path: this runs on the normal path and may block.
println!("drainer: terminate");
}
}
/// The server's name: how the rest of the app reaches it (and the only handle
/// that survives a restart).
const DRAINER: GenServerName<Drainer> = GenServerName::new("drainer");
fn app_tree() -> OneForOne {
OneForOne::new()
.child(
ChildSpec::new(Restart::Permanent, || {
// A plain worker that does not trap: stopped outright on shutdown.
loop {
sleep(Duration::from_millis(10));
}
})
.shutdown(Shutdown::Timeout(Duration::from_millis(100))),
)
// A gen_server is a direct child: `named(N).run()` runs the loop as
// the child actor itself, so the supervisor's shutdown arrives as
// `handle_shutdown` and a restart re-binds the name.
.child(
ChildSpec::new(Restart::Permanent, || {
GenServerBuilder::new(Drainer {
pending: 3,
stop: None,
timer: None,
})
.named(DRAINER)
.run()
.expect("drainer name is free");
})
.shutdown(Shutdown::Infinity),
)
}
fn main() {
println!("--- 1. root exit drives the shutdown ---");
smarm::run(|| {
spawn(|| app_tree().run());
sleep(Duration::from_millis(50)); // the app "runs" for a while
// Returning here asks the supervisor to shut down; the run ends when
// the tree — drainer included — is gone.
});
println!("--- 2. an outside thread drives the shutdown ---");
let rt = smarm::init(smarm::Config::default());
let handle = rt.handle(); // Send + Sync; grab it before run
rt.run(move || {
let sup = spawn(|| app_tree().run());
let sup_pid = sup.pid();
// Stand-in for a SIGTERM handler thread.
thread::spawn(move || {
thread::sleep(Duration::from_millis(50));
println!("signal thread: requesting shutdown");
handle.request_shutdown(sup_pid);
});
sup.join()
.expect("supervisor returns normally after ordered shutdown");
println!("supervisor down; root returns");
});
}
+9 -1
View File
@@ -6,7 +6,9 @@
//! every use — so the address keeps working across a supervised restart, with //! every use — so the address keeps working across a supervised restart, with
//! no stale [`GenServerRef`] to refresh. //! no stale [`GenServerRef`] to refresh.
use smarm::{call, cast, run, whereis_server, GenServer, GenServerBuilder, GenServerName, GenServerRef}; use smarm::{
call, cast, run, whereis_server, GenServer, GenServerBuilder, GenServerName, GenServerRef,
};
/// A counter server: synchronous `Get`, asynchronous `Inc` / `Add`. /// A counter server: synchronous `Get`, asynchronous `Inc` / `Add`.
struct Counter { struct Counter {
@@ -64,6 +66,12 @@ fn main() {
let svc: Option<GenServerRef<Counter>> = whereis_server(COUNTER); let svc: Option<GenServerRef<Counter>> = whereis_server(COUNTER);
if let Some(svc) = svc { if let Some(svc) = svc {
let _ = svc.call(Query::Get); let _ = svc.call(Query::Get);
// A named server is pinned alive by the registry, so dropping refs
// does not end it. Stop it explicitly: `shutdown()` asks politely
// (a trapping server drains first; this one is stopped outright)
// and waits until it is gone. Left running, the root's return
// would shut it down the same way — see examples/graceful_shutdown.rs.
svc.shutdown();
} }
}); });
} }
+13 -3
View File
@@ -15,7 +15,9 @@
//! `call`, nothing more. //! `call`, nothing more.
use smarm::observer::{self, ObserverReply, ObserverRequest}; use smarm::observer::{self, ObserverReply, ObserverRequest};
use smarm::{channel, register, run, spawn, ActorState, Name, RuntimeSnapshot, RuntimeTree, TreeNode}; use smarm::{
channel, register, run, spawn, ActorState, Name, RuntimeSnapshot, RuntimeTree, TreeNode,
};
const ECHO: Name<u64> = Name::new("echo"); const ECHO: Name<u64> = Name::new("echo");
@@ -31,7 +33,11 @@ fn state_glyph(s: ActorState) -> &'static str {
/// A `ps`-style table over the flat snapshot. /// A `ps`-style table over the flat snapshot.
fn print_snapshot(snap: &RuntimeSnapshot) { fn print_snapshot(snap: &RuntimeSnapshot) {
println!("snapshot (format v{}, {} actors)", snap.format_version, snap.actors.len()); println!(
"snapshot (format v{}, {} actors)",
snap.format_version,
snap.actors.len()
);
println!( println!(
" {:<10} {:<9} {:<10} {:>4} {:>4} {:>4} {:>4} {:>5} {}", " {:<10} {:<9} {:<10} {:>4} {:>4} {:>4} {:>4} {:>5} {}",
"pid", "state", "parent", "mon", "lnk", "joi", "mbox", "msgs", "names" "pid", "state", "parent", "mon", "lnk", "joi", "mbox", "msgs", "names"
@@ -52,7 +58,11 @@ fn print_snapshot(snap: &RuntimeSnapshot) {
a.joiners, a.joiners,
a.mailbox_depth, a.mailbox_depth,
a.messages_received, a.messages_received,
if a.names.is_empty() { "-".to_string() } else { a.names.join(",") }, if a.names.is_empty() {
"-".to_string()
} else {
a.names.join(",")
},
); );
} }
} }
+20 -10
View File
@@ -342,8 +342,7 @@ mod inner {
// Count the loss in would-be delta terms so the audit's columns // Count the loss in would-be delta terms so the audit's columns
// compare directly against `injected_cycles`. // compare directly against `injected_cycles`.
DISCARD_OVERMAX_N.fetch_add(1, Ordering::Relaxed); DISCARD_OVERMAX_N.fetch_add(1, Ordering::Relaxed);
DISCARD_OVERMAX_CYCLES DISCARD_OVERMAX_CYCLES.fetch_add(interval.saturating_mul(pct) / 100, Ordering::Relaxed);
.fetch_add(interval.saturating_mul(pct) / 100, Ordering::Relaxed);
return; return;
} }
let delta = interval.saturating_mul(pct) / 100; let delta = interval.saturating_mul(pct) / 100;
@@ -419,8 +418,7 @@ mod inner {
let gap = preempt::rdtsc() let gap = preempt::rdtsc()
.saturating_sub(desched_tsc) .saturating_sub(desched_tsc)
.min(MAX_SAMPLE_CYCLES); .min(MAX_SAMPLE_CYCLES);
OFFCPU_IN_SITE_CYCLES OFFCPU_IN_SITE_CYCLES.fetch_add(gap.saturating_mul(pct) / 100, Ordering::Relaxed);
.fetch_add(gap.saturating_mul(pct) / 100, Ordering::Relaxed);
OFFCPU_IN_SITE_N.fetch_add(1, Ordering::Relaxed); OFFCPU_IN_SITE_N.fetch_add(1, Ordering::Relaxed);
} }
} }
@@ -533,7 +531,9 @@ mod inner {
park_forgiven_cycles: self park_forgiven_cycles: self
.park_forgiven_cycles .park_forgiven_cycles
.saturating_sub(before.park_forgiven_cycles), .saturating_sub(before.park_forgiven_cycles),
drop_park_cycles: self.drop_park_cycles.saturating_sub(before.drop_park_cycles), drop_park_cycles: self
.drop_park_cycles
.saturating_sub(before.drop_park_cycles),
drop_park_n: self.drop_park_n.saturating_sub(before.drop_park_n), drop_park_n: self.drop_park_n.saturating_sub(before.drop_park_n),
drop_yield_cycles: self drop_yield_cycles: self
.drop_yield_cycles .drop_yield_cycles
@@ -542,12 +542,18 @@ mod inner {
discard_overmax_cycles: self discard_overmax_cycles: self
.discard_overmax_cycles .discard_overmax_cycles
.saturating_sub(before.discard_overmax_cycles), .saturating_sub(before.discard_overmax_cycles),
discard_overmax_n: self.discard_overmax_n.saturating_sub(before.discard_overmax_n), discard_overmax_n: self
discard_unarmed_n: self.discard_unarmed_n.saturating_sub(before.discard_unarmed_n), .discard_overmax_n
.saturating_sub(before.discard_overmax_n),
discard_unarmed_n: self
.discard_unarmed_n
.saturating_sub(before.discard_unarmed_n),
offcpu_in_site_cycles: self offcpu_in_site_cycles: self
.offcpu_in_site_cycles .offcpu_in_site_cycles
.saturating_sub(before.offcpu_in_site_cycles), .saturating_sub(before.offcpu_in_site_cycles),
offcpu_in_site_n: self.offcpu_in_site_n.saturating_sub(before.offcpu_in_site_n), offcpu_in_site_n: self
.offcpu_in_site_n
.saturating_sub(before.offcpu_in_site_n),
} }
} }
} }
@@ -795,7 +801,9 @@ mod inner {
let cell = results let cell = results
.iter() .iter()
.find(|r| r.site == site && r.speedup_pct == speedup_pct)?; .find(|r| r.site == site && r.speedup_pct == speedup_pct)?;
let base = results.iter().find(|r| r.site == site && r.speedup_pct == 0)?; let base = results
.iter()
.find(|r| r.site == site && r.speedup_pct == 0)?;
let rate = normalized_rate(cell, point)?; let rate = normalized_rate(cell, point)?;
let b = normalized_rate(base, point)?; let b = normalized_rate(base, point)?;
if b <= 0.0 { if b <= 0.0 {
@@ -946,7 +954,9 @@ macro_rules! progress {
macro_rules! causal_site { macro_rules! causal_site {
($name:literal) => {{ ($name:literal) => {{
static __SMARM_SITE: ::std::sync::OnceLock<u32> = ::std::sync::OnceLock::new(); static __SMARM_SITE: ::std::sync::OnceLock<u32> = ::std::sync::OnceLock::new();
$crate::causal::SiteGuard::enter(*__SMARM_SITE.get_or_init(|| $crate::causal::site_id($name))) $crate::causal::SiteGuard::enter(
*__SMARM_SITE.get_or_init(|| $crate::causal::site_id($name)),
)
}}; }};
} }
+77 -40
View File
@@ -90,8 +90,9 @@
use crate::pid::Pid; use crate::pid::Pid;
use crate::raw_mutex::RawMutex; use crate::raw_mutex::RawMutex;
use crate::runtime::RuntimeInner;
use std::collections::VecDeque; use std::collections::VecDeque;
use std::sync::Arc; use std::sync::{Arc, Weak};
/// Create a new channel and return its `(Sender, Receiver)` halves. /// Create a new channel and return its `(Sender, Receiver)` halves.
/// ///
@@ -104,17 +105,25 @@ pub fn channel<T>() -> (Sender<T>, Receiver<T>) {
senders: 1, senders: 1,
receiver_alive: true, receiver_alive: true,
})); }));
(Sender { inner: inner.clone() }, Receiver { inner }) (
Sender {
inner: inner.clone(),
},
Receiver { inner },
)
} }
struct Inner<T> { struct Inner<T> {
queue: VecDeque<T>, queue: VecDeque<T>,
/// The parked receiver's `(pid, park-epoch)`, if one is currently /// The parked receiver's `(pid, park-epoch, runtime)`, if one is currently
/// waiting. The epoch identifies exactly which wait this is, so a waker /// waiting. The epoch identifies exactly which wait this is, so a waker
/// left over from a wait that already ended (a losing `select` arm, a /// left over from a wait that already ended (a losing `select` arm, a
/// `recv_timeout` whose timer fired after it was already satisfied) is /// `recv_timeout` whose timer fired after it was already satisfied) is
/// inert and does nothing when it fires. /// inert and does nothing when it fires. The `Weak<RuntimeInner>` is the
parked_receiver: Option<(Pid, u32)>, /// receiver's runtime, captured while it parked (so provably alive then);
/// it lets a sender on a foreign OS thread wake the receiver without the
/// `RUNTIME` thread-local, which is unset off a scheduler thread.
parked_receiver: Option<(Pid, u32, Weak<RuntimeInner>)>,
senders: usize, senders: usize,
receiver_alive: bool, receiver_alive: bool,
} }
@@ -178,7 +187,9 @@ impl std::error::Error for RecvTimeoutError {}
impl<T> Clone for Sender<T> { impl<T> Clone for Sender<T> {
fn clone(&self) -> Self { fn clone(&self) -> Self {
self.inner.lock().senders += 1; self.inner.lock().senders += 1;
Sender { inner: self.inner.clone() } Sender {
inner: self.inner.clone(),
}
} }
} }
@@ -199,8 +210,8 @@ impl<T> Drop for Sender<T> {
None None
} }
}; };
if let Some((pid, epoch)) = unpark { if let Some((pid, epoch, rt)) = unpark {
crate::scheduler::unpark_at(pid, epoch); crate::scheduler::unpark_at_via(pid, epoch, &rt);
} }
} }
} }
@@ -247,11 +258,19 @@ impl<T> Sender<T> {
g.queue.push_back(value); g.queue.push_back(value);
g.parked_receiver.take() g.parked_receiver.take()
}; };
if let Some((pid, epoch)) = unpark { if let Some((pid, epoch, rt)) = unpark {
crate::te!(crate::trace::Event::Send { sender: crate::actor::current_pid().unwrap_or(crate::pid::Pid::new(u32::MAX, u32::MAX)), receiver: Some(pid) }); crate::te!(crate::trace::Event::Send {
crate::scheduler::unpark_at(pid, epoch); sender: crate::actor::current_pid()
.unwrap_or(crate::pid::Pid::new(u32::MAX, u32::MAX)),
receiver: Some(pid)
});
crate::scheduler::unpark_at_via(pid, epoch, &rt);
} else { } else {
crate::te!(crate::trace::Event::Send { sender: crate::actor::current_pid().unwrap_or(crate::pid::Pid::new(u32::MAX, u32::MAX)), receiver: None }); crate::te!(crate::trace::Event::Send {
sender: crate::actor::current_pid()
.unwrap_or(crate::pid::Pid::new(u32::MAX, u32::MAX)),
receiver: None
});
} }
Ok(()) Ok(())
} }
@@ -278,22 +297,28 @@ impl<T> Receiver<T> {
None => panic!("smarm: recv() called outside an actor"), None => panic!("smarm: recv() called outside an actor"),
}; };
debug_assert!( debug_assert!(
g.parked_receiver.is_none_or(|(p, _)| p == me), g.parked_receiver.as_ref().is_none_or(|(p, _, _)| *p == me),
"channel has more than one receiver" "channel has more than one receiver"
); );
// begin_wait is lock-free, so it's legal under the Channel lock; // begin_wait is lock-free, so it's legal under the Channel lock;
// registering in the same critical section makes the epoch // registering in the same critical section makes the epoch
// atomic with the senders' view of the registration. // atomic with the senders' view of the registration.
g.parked_receiver = Some((me, crate::scheduler::begin_wait())); g.parked_receiver = Some((
me,
crate::scheduler::begin_wait(),
crate::scheduler::runtime_weak(),
));
crate::te!(crate::trace::Event::RecvPark(me)); crate::te!(crate::trace::Event::RecvPark(me));
} }
// Release the lock before parking: the unparker will need it. // Release the lock before parking: the unparker will need it.
crate::scheduler::park_current(); crate::scheduler::park_current();
// Woken up. Record it before looping to check the queue. // Woken up. Record it before looping to check the queue.
crate::te!(crate::trace::Event::RecvWake(match crate::actor::current_pid() { crate::te!(crate::trace::Event::RecvWake(
match crate::actor::current_pid() {
Some(p) => p, Some(p) => p,
None => panic!("smarm: RecvWake outside an actor (core corrupt)"), None => panic!("smarm: RecvWake outside an actor (core corrupt)"),
})); }
));
} }
} }
@@ -330,11 +355,11 @@ impl<T> Receiver<T> {
return Err(RecvTimeoutError::Disconnected); return Err(RecvTimeoutError::Disconnected);
} }
debug_assert!( debug_assert!(
g.parked_receiver.is_none_or(|(p, _)| p == me), g.parked_receiver.as_ref().is_none_or(|(p, _, _)| *p == me),
"channel has more than one receiver" "channel has more than one receiver"
); );
epoch = crate::scheduler::begin_wait(); epoch = crate::scheduler::begin_wait();
g.parked_receiver = Some((me, epoch)); g.parked_receiver = Some((me, epoch, crate::scheduler::runtime_weak()));
crate::te!(crate::trace::Event::RecvPark(me)); crate::te!(crate::trace::Event::RecvPark(me));
} }
@@ -347,10 +372,12 @@ impl<T> Receiver<T> {
crate::scheduler::insert_wait_timer(deadline, me, target, epoch); crate::scheduler::insert_wait_timer(deadline, me, target, epoch);
crate::scheduler::park_current(); crate::scheduler::park_current();
crate::te!(crate::trace::Event::RecvWake(match crate::actor::current_pid() { crate::te!(crate::trace::Event::RecvWake(
match crate::actor::current_pid() {
Some(p) => p, Some(p) => p,
None => panic!("smarm: RecvWake outside an actor (core corrupt)"), None => panic!("smarm: RecvWake outside an actor (core corrupt)"),
})); }
));
let mut g = self.inner.lock(); let mut g = self.inner.lock();
if let Some(v) = g.queue.pop_front() { if let Some(v) = g.queue.pop_front() {
crate::preempt::note_message_received(); crate::preempt::note_message_received();
@@ -404,18 +431,24 @@ impl<T> Receiver<T> {
None => panic!("smarm: recv_match() called outside an actor"), None => panic!("smarm: recv_match() called outside an actor"),
}; };
debug_assert!( debug_assert!(
g.parked_receiver.is_none_or(|(p, _)| p == me), g.parked_receiver.as_ref().is_none_or(|(p, _, _)| *p == me),
"channel has more than one receiver" "channel has more than one receiver"
); );
g.parked_receiver = Some((me, crate::scheduler::begin_wait())); g.parked_receiver = Some((
me,
crate::scheduler::begin_wait(),
crate::scheduler::runtime_weak(),
));
crate::te!(crate::trace::Event::RecvPark(me)); crate::te!(crate::trace::Event::RecvPark(me));
} }
// Release the lock before parking: the unparker will need it. // Release the lock before parking: the unparker will need it.
crate::scheduler::park_current(); crate::scheduler::park_current();
crate::te!(crate::trace::Event::RecvWake(match crate::actor::current_pid() { crate::te!(crate::trace::Event::RecvWake(
match crate::actor::current_pid() {
Some(p) => p, Some(p) => p,
None => panic!("smarm: RecvWake outside an actor (core corrupt)"), None => panic!("smarm: RecvWake outside an actor (core corrupt)"),
})); }
));
} }
} }
@@ -476,11 +509,12 @@ impl<T: Send + 'static> crate::timer::TimerTarget for RawMutex<Inner<T>> {
// keeps the registration bookkeeping exact.) // keeps the registration bookkeeping exact.)
let unpark = { let unpark = {
let mut g = self.lock(); let mut g = self.lock();
if g.parked_receiver == Some((pid, epoch)) { match g.parked_receiver {
Some((p, e, _)) if p == pid && e == epoch => {
g.parked_receiver = None; g.parked_receiver = None;
true true
} else { }
false _ => false,
} }
}; };
// Unpark outside the channel lock: it may take the run-queue lock; // Unpark outside the channel lock: it may take the run-queue lock;
@@ -541,10 +575,10 @@ impl<T> Selectable for Receiver<T> {
return Ok(false); return Ok(false);
} }
debug_assert!( debug_assert!(
g.parked_receiver.is_none_or(|(p, _)| p == pid), g.parked_receiver.as_ref().is_none_or(|(p, _, _)| *p == pid),
"channel has more than one receiver" "channel has more than one receiver"
); );
g.parked_receiver = Some((pid, epoch)); g.parked_receiver = Some((pid, epoch, crate::scheduler::runtime_weak()));
Ok(true) Ok(true)
} }
@@ -616,7 +650,12 @@ pub fn try_select(arms: &[&dyn Selectable]) -> std::io::Result<usize> {
// Channel-only selects skip all of it: `eager` is false, the guard // Channel-only selects skip all of it: `eager` is false, the guard
// is disarmed, and the loser-arm self-cleaning story is unchanged. // is disarmed, and the loser-arm self-cleaning story is unchanged.
let eager = arms.iter().any(|a| a.sel_eager_cleanup()); let eager = arms.iter().any(|a| a.sel_eager_cleanup());
let mut guard = UnregisterGuard { arms, me, epoch, armed: eager }; let mut guard = UnregisterGuard {
arms,
me,
epoch,
armed: eager,
};
crate::scheduler::park_current(); crate::scheduler::park_current();
@@ -687,11 +726,7 @@ impl Drop for UnregisterGuard<'_> {
// unregistered eagerly so none are left dangling. `Err` = an arm failed to // unregistered eagerly so none are left dangling. `Err` = an arm failed to
// register; same unwind (earlier fd arms unregistered, wait retired). // register; same unwind (earlier fd arms unregistered, wait retired).
// `Ok(None)` = every arm registered successfully; the caller parks. // `Ok(None)` = every arm registered successfully; the caller parks.
fn register_arms( fn register_arms(me: Pid, epoch: u32, arms: &[&dyn Selectable]) -> std::io::Result<Option<usize>> {
me: Pid,
epoch: u32,
arms: &[&dyn Selectable],
) -> std::io::Result<Option<usize>> {
for (i, arm) in arms.iter().enumerate() { for (i, arm) in arms.iter().enumerate() {
let registered = match arm.sel_register(me, epoch) { let registered = match arm.sel_register(me, epoch) {
Ok(r) => r, Ok(r) => r,
@@ -736,10 +771,7 @@ impl crate::timer::TimerTarget for SelectTimeout {
/// Panics if `arms` is empty, if called outside an actor, or if an fd arm /// Panics if `arms` is empty, if called outside an actor, or if an fd arm
/// fails to register (see [`try_select_timeout`] for the fallible form; a /// fails to register (see [`try_select_timeout`] for the fallible form; a
/// channel-only select can never fail). /// channel-only select can never fail).
pub fn select_timeout( pub fn select_timeout(arms: &[&dyn Selectable], timeout: std::time::Duration) -> Option<usize> {
arms: &[&dyn Selectable],
timeout: std::time::Duration,
) -> Option<usize> {
match try_select_timeout(arms, timeout) { match try_select_timeout(arms, timeout) {
Ok(r) => r, Ok(r) => r,
Err(e) => panic!( Err(e) => panic!(
@@ -776,7 +808,12 @@ pub fn try_select_timeout(
// would leave those fds unusable until a kernel event happened to // would leave those fds unusable until a kernel event happened to
// clear them. // clear them.
let eager = arms.iter().any(|a| a.sel_eager_cleanup()); let eager = arms.iter().any(|a| a.sel_eager_cleanup());
let mut guard = UnregisterGuard { arms, me, epoch, armed: eager }; let mut guard = UnregisterGuard {
arms,
me,
epoch,
armed: eager,
};
crate::scheduler::park_current(); crate::scheduler::park_current();
+26 -11
View File
@@ -16,10 +16,18 @@ thread_local! {
static ACTOR_SP: Cell<usize> = const { Cell::new(0) }; static ACTOR_SP: Cell<usize> = const { Cell::new(0) };
} }
fn get_scheduler_sp() -> usize { SCHEDULER_SP.with(|c| c.get()) } fn get_scheduler_sp() -> usize {
fn set_scheduler_sp(v: usize) { SCHEDULER_SP.with(|c| c.set(v)) } SCHEDULER_SP.with(|c| c.get())
pub fn get_actor_sp() -> usize { ACTOR_SP.with(|c| c.get()) } }
pub fn set_actor_sp(v: usize) { ACTOR_SP.with(|c| c.set(v)) } fn set_scheduler_sp(v: usize) {
SCHEDULER_SP.with(|c| c.set(v))
}
pub fn get_actor_sp() -> usize {
ACTOR_SP.with(|c| c.get())
}
pub fn set_actor_sp(v: usize) {
ACTOR_SP.with(|c| c.set(v))
}
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// Initial stack layout // Initial stack layout
@@ -49,13 +57,20 @@ pub fn set_actor_sp(v: usize) { ACTOR_SP.with(|c| c.set(v)) }
pub fn init_actor_stack(top: *mut u8, entry: extern "C-unwind" fn()) -> usize { pub fn init_actor_stack(top: *mut u8, entry: extern "C-unwind" fn()) -> usize {
unsafe { unsafe {
let mut sp = (top as usize & !15) - 8; let mut sp = (top as usize & !15) - 8;
sp -= 8; (sp as *mut usize).write(entry as usize); // ret target sp -= 8;
sp -= 8; (sp as *mut usize).write(0); // rbx (sp as *mut usize).write(entry as usize); // ret target
sp -= 8; (sp as *mut usize).write(0); // rbp sp -= 8;
sp -= 8; (sp as *mut usize).write(0); // r12 (sp as *mut usize).write(0); // rbx
sp -= 8; (sp as *mut usize).write(0); // r13 sp -= 8;
sp -= 8; (sp as *mut usize).write(0); // r14 (sp as *mut usize).write(0); // rbp
sp -= 8; (sp as *mut usize).write(0); // r15 sp -= 8;
(sp as *mut usize).write(0); // r12
sp -= 8;
(sp as *mut usize).write(0); // r13
sp -= 8;
(sp as *mut usize).write(0); // r14
sp -= 8;
(sp as *mut usize).write(0); // r15
sp sp
} }
} }
+313 -63
View File
@@ -127,16 +127,51 @@
//! - [`GenServer::init`] runs once before the first message. Use it to start //! - [`GenServer::init`] runs once before the first message. Use it to start
//! timers or set up monitors; see the [`GenServerCtx`] it receives. //! timers or set up monitors; see the [`GenServerCtx`] it receives.
//! - [`GenServer::terminate`] runs when the server is about to exit. It fires //! - [`GenServer::terminate`] runs when the server is about to exit. It fires
//! on every exit path (all `GenServerRef`s dropped, a handler panic, or an //! on every exit path (a self-stop, a graceful shutdown, a cooperative hard
//! explicit [`GenServerRef::shutdown`]), not only on clean shutdown. Keep it //! stop, or a handler panic), not only on clean shutdown.
//! short and non-blocking: if `terminate` panics while the server is already //! Keep it non-panicking: on the panic and hard-stop paths it runs
//! unwinding from a handler panic, the process aborts. //! mid-unwind, where a second panic aborts the process and where it must
//! not block (any park re-observes the stop). Only on the graceful path
//! (see below) may it do real work.
//! //!
//! ## When the server stops //! ## When the server stops
//! //!
//! The server runs as long as at least one [`GenServerRef`] exists. When the last //! A server lives until it stops, is shut down, or is killed — as an OTP
//! one is dropped, the inbox closes and the loop exits gracefully. To stop a //! process does. A [`GenServerRef`] is an *address*: cloning and dropping it
//! server explicitly and wait for it to finish, call [`GenServerRef::shutdown`]. //! never changes the server's lifetime, and a ref that nobody holds is not a
//! leak — a forgotten server idles until the run ends, when the root-exit
//! shutdown (see [`Runtime::run`](crate::Runtime::run)) takes it down with
//! every other unsupervised actor. Anything meant to live long should be
//! supervised (see *Supervised servers* below); [`start`] / [`start_under`]
//! are for scripts, tests and short-lived helpers, and the explicit close is
//! [`GenServerRef::shutdown`].
//!
//! A server can end itself: clone a [`StopHandle`] from
//! [`GenServerCtx::stop_handle`] in `init` and call [`StopHandle::stop`] from
//! any handler — the loop breaks after the current message and exits
//! *normally* (OTP's `{stop, normal}`). This is distinct from
//! `request_stop(self_pid())`, which is an abnormal `Stopped` and gets a
//! `Transient` child restarted.
//!
//! ## Graceful shutdown
//!
//! From outside, [`GenServerRef::shutdown`] (or a plain
//! [`request_shutdown`](crate::request_shutdown), which is what a supervisor
//! sends) asks the server to stop. What happens next is the server's choice:
//!
//! - By default a server does not trap exits, and the request stops it
//! outright at its next observation point — `terminate` runs mid-unwind.
//! - A server that calls [`GenServerCtx::trap_exit`] in `init` receives the
//! request as [`GenServer::handle_shutdown`]. Return
//! [`ShutdownAction::Exit`] (the default) to have the loop break and
//! `terminate` run on the normal path, where it may block; return
//! [`ShutdownAction::Continue`] to keep serving — e.g. to drain in-flight
//! work — and end the server later with a [`StopHandle`]. The supervisor's
//! [`Shutdown`](crate::supervisor::Shutdown) policy bounds how long that
//! may take before it falls back to a hard stop.
//!
//! A trapping server also receives the deaths of its linked peers as
//! [`GenServer::handle_exit`] messages instead of dying with them.
//! //!
//! If the server panics inside a handler, the panic unwinds the server thread. //! If the server panics inside a handler, the panic unwinds the server thread.
//! Any caller currently waiting in `call` sees `Err(ServerDown)`: the reply //! Any caller currently waiting in `call` sees `Err(ServerDown)`: the reply
@@ -167,9 +202,23 @@
//! Names and registration: a server can be given a static name so other //! Names and registration: a server can be given a static name so other
//! actors can reach it without holding a `GenServerRef`. Use //! actors can reach it without holding a `GenServerRef`. Use
//! [`GenServerBuilder::named`] to register on start, and the free functions //! [`GenServerBuilder::named`] to register on start, and the free functions
//! [`call`], [`cast`], and [`whereis_server`] to address it by name. Registered //! [`call`], [`cast`], and [`whereis_server`] to address it by name.
//! servers are a natural fit for supervision; see `supervisor` for how to //!
//! build a tree that restarts servers on failure. //! ## Supervised servers
//!
//! The supervised shape is [`NamedGenServerBuilder::run`]: it runs the loop
//! **inline, as the current actor**, so the closure of a
//! [`ChildSpec`](crate::supervisor::ChildSpec) *is* the server — the
//! supervisor's shutdown arrives as [`GenServer::handle_shutdown`], a restart
//! runs the factory again and re-binds the name, and the rest of the program
//! addresses it by name (a held ref would go stale on restart anyway).
//!
//! ```ignore
//! const COUNTER: GenServerName<Counter> = GenServerName::new("counter");
//! OneForOne::new().child(ChildSpec::new(Restart::Permanent, || {
//! GenServerBuilder::new(Counter::default()).named(COUNTER).run().unwrap();
//! }));
//! ```
//! //!
//! ## Limitations //! ## Limitations
//! //!
@@ -178,11 +227,15 @@
//! from any handler via [`Watcher::watch`]) because monitors are inherently //! from any handler via [`Watcher::watch`]) because monitors are inherently
//! created at runtime. The idle window is set once, in `init`. //! created at runtime. The idle window is set once, in `init`.
use crate::channel::{channel, select, select_timeout, Receiver, RecvTimeoutError, Selectable, Sender}; use crate::channel::{
channel, select, select_timeout, Receiver, RecvTimeoutError, Selectable, Sender,
};
use crate::link::ExitSignal;
use crate::monitor::DownReason;
use crate::monitor::{demonitor, monitor, Down, Monitor}; use crate::monitor::{demonitor, monitor, Down, Monitor};
use crate::pid::Pid; use crate::pid::Pid;
use crate::registry::{register_with, resolve_named_sender, RegisterError}; use crate::registry::{register_with, resolve_named_sender, RegisterError};
use crate::scheduler::{cancel_timer, request_stop, send_after_to, spawn, spawn_under}; use crate::scheduler::{cancel_timer, request_shutdown, send_after_to};
use crate::timer::TimerId; use crate::timer::TimerId;
use std::cell::Cell; use std::cell::Cell;
use std::collections::HashMap; use std::collections::HashMap;
@@ -252,10 +305,34 @@ pub trait GenServer: Send + 'static {
/// Default: no-op. /// Default: no-op.
fn handle_idle(&mut self) {} fn handle_idle(&mut self) {}
/// A graceful shutdown request (a [`request_shutdown`](crate::request_shutdown)
/// reaching this server), delivered only if `init` called
/// [`GenServerCtx::trap_exit`]. Return [`ShutdownAction::Exit`] to stop
/// now (the default), or [`ShutdownAction::Continue`] to keep serving and
/// end the server later with a [`StopHandle`].
fn handle_shutdown(&mut self) -> ShutdownAction {
ShutdownAction::Exit
}
/// A linked peer's abnormal death (an [`ExitSignal`] that is not a
/// shutdown request), delivered only if `init` called
/// [`GenServerCtx::trap_exit`]. Default: drop it.
fn handle_exit(&mut self, _sig: ExitSignal) {}
/// Runs as the server actor exits, on any exit path (see module docs). /// Runs as the server actor exits, on any exit path (see module docs).
fn terminate(&mut self) {} fn terminate(&mut self) {}
} }
/// What a server does with a shutdown request; see [`GenServer::handle_shutdown`].
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum ShutdownAction {
/// Break the loop now. `terminate` runs on the normal path and may block.
Exit,
/// Keep dispatching. The server is expected to end itself with a
/// [`StopHandle`] once it is done winding down.
Continue,
}
/// What travels the server's single inbox channel: a synchronous call (with a /// What travels the server's single inbox channel: a synchronous call (with a
/// reply sender) or an asynchronous cast. Private — callers use [`GenServerRef`]. /// reply sender) or an asynchronous cast. Private — callers use [`GenServerRef`].
enum Envelope<G: GenServer> { enum Envelope<G: GenServer> {
@@ -263,9 +340,9 @@ enum Envelope<G: GenServer> {
Cast(G::Cast), Cast(G::Cast),
} }
/// A clonable handle to a running server. Cloning yields another sender to the /// A clonable handle to a running server: an *address*, not an owner. Cloning
/// same inbox; the server lives until the last `GenServerRef` is dropped, at which /// yields another sender to the same inbox; dropping refs never ends the
/// point its inbox closes and the loop exits normally. /// server (see the module docs, *When the server stops*).
pub struct GenServerRef<G: GenServer> { pub struct GenServerRef<G: GenServer> {
tx: Sender<Envelope<G>>, tx: Sender<Envelope<G>>,
pid: Pid, pid: Pid,
@@ -273,7 +350,10 @@ pub struct GenServerRef<G: GenServer> {
impl<G: GenServer> Clone for GenServerRef<G> { impl<G: GenServer> Clone for GenServerRef<G> {
fn clone(&self) -> Self { fn clone(&self) -> Self {
GenServerRef { tx: self.tx.clone(), pid: self.pid } GenServerRef {
tx: self.tx.clone(),
pid: self.pid,
}
} }
} }
@@ -357,18 +437,22 @@ impl<G: GenServer> GenServerRef<G> {
/// Stop the server and block until it has fully exited. /// Stop the server and block until it has fully exited.
/// ///
/// Sends a cooperative stop signal to the server actor and waits for it to /// Asks the server to shut down (a [`request_shutdown`](crate::request_shutdown))
/// exit, so [`GenServer::terminate`] has run by the time this returns. /// and waits for it to exit, so [`GenServer::terminate`] has run by the
/// Returns immediately if the server is already gone. /// time this returns. A trapping server gets to wind down via
/// [`GenServer::handle_shutdown`]; any other is stopped outright. Returns
/// immediately if the server is already gone. Waits as long as the server
/// takes — the caller, not the server, decides whether that is acceptable;
/// a supervisor uses its child's [`Shutdown`](crate::supervisor::Shutdown)
/// policy to bound it.
/// ///
/// This is the right teardown for a server kept alive by a registered /// This is the explicit close: dropping refs never ends a server. Like
/// [`GenServerName`], where dropping every external `GenServerRef` is not enough /// all cooperative cancellation, it is best-effort:
/// to close the inbox. Like all cooperative cancellation, it is best-effort:
/// a server wedged in a tight loop with no observation point cannot be /// a server wedged in a tight loop with no observation point cannot be
/// stopped this way. Panics if called outside `Runtime::run()`. /// stopped this way. Panics if called outside `Runtime::run()`.
pub fn shutdown(&self) { pub fn shutdown(&self) {
let mon = monitor(self.pid); let mon = monitor(self.pid);
request_stop(self.pid); request_shutdown(self.pid);
// The Down lands when the server finalizes; an already-dead target makes // The Down lands when the server finalizes; an already-dead target makes
// `monitor` deliver NoProc immediately, so this never blocks forever. // `monitor` deliver NoProc immediately, so this never blocks forever.
let _ = mon.rx.recv(); let _ = mon.rx.recv();
@@ -391,6 +475,8 @@ enum Sys<G: GenServer> {
/// the payload factory, dispatches it to [`GenServer::handle_timer`], and /// the payload factory, dispatches it to [`GenServer::handle_timer`], and
/// re-arms the next tick before returning. /// re-arms the next tick before returning.
Tick(crate::timer::TimerId), Tick(crate::timer::TimerId),
/// The state asked to end the server (via [`StopHandle::stop`]).
Stop,
} }
/// The server loop's runtime hook, passed to [`GenServer::init`]. Hands out the /// The server loop's runtime hook, passed to [`GenServer::init`]. Hands out the
@@ -406,13 +492,35 @@ pub struct GenServerCtx<G: GenServer> {
/// because `init` holds only `&ctx`; not `Send`, but `GenServerCtx` is only ever /// because `init` holds only `&ctx`; not `Send`, but `GenServerCtx` is only ever
/// borrowed on the actor's own stack during `init`, never sent. /// borrowed on the actor's own stack during `init`, never sent.
idle: Cell<Option<Duration>>, idle: Cell<Option<Duration>>,
/// Whether the loop should trap exits (set via [`trap_exit`](Self::trap_exit)
/// during `init`, read by the loop after).
trap: Cell<bool>,
} }
impl<G: GenServer> GenServerCtx<G> { impl<G: GenServer> GenServerCtx<G> {
/// Trap exits for the server's lifetime: shutdown requests then arrive as
/// [`GenServer::handle_shutdown`] and linked-peer deaths as
/// [`GenServer::handle_exit`], instead of stopping the server outright.
/// Call this once during `init`.
pub fn trap_exit(&self) {
self.trap.set(true);
}
/// A clonable handle that lets the state end the server from any handler
/// (a normal exit; see the module docs). Store it on the state during
/// `init`.
pub fn stop_handle(&self) -> StopHandle<G> {
StopHandle {
sys_tx: self.sys_tx.clone(),
}
}
/// A clonable handle to the loop's monitor intake. Store it in the state /// A clonable handle to the loop's monitor intake. Store it in the state
/// during `init` to watch monitors from later handlers. /// during `init` to watch monitors from later handlers.
pub fn watcher(&self) -> Watcher<G> { pub fn watcher(&self) -> Watcher<G> {
Watcher { tx: self.sys_tx.clone() } Watcher {
tx: self.sys_tx.clone(),
}
} }
/// Shorthand for `ctx.watcher().watch(m)` when watching during `init`. /// Shorthand for `ctx.watcher().watch(m)` when watching during `init`.
@@ -426,7 +534,10 @@ impl<G: GenServer> GenServerCtx<G> {
/// [`tick_every`](TimerHandle::tick_every) / /// [`tick_every`](TimerHandle::tick_every) /
/// [`cancel`](TimerHandle::cancel) from any later handler. /// [`cancel`](TimerHandle::cancel) from any later handler.
pub fn timer(&self) -> TimerHandle<G> { pub fn timer(&self) -> TimerHandle<G> {
TimerHandle { sys_tx: self.sys_tx.clone(), reg: self.reg.clone() } TimerHandle {
sys_tx: self.sys_tx.clone(),
reg: self.reg.clone(),
}
} }
/// Set a quiet-period window: if the loop goes `after` without dispatching /// Set a quiet-period window: if the loop goes `after` without dispatching
@@ -442,6 +553,30 @@ impl<G: GenServer> GenServerCtx<G> {
} }
} }
/// Lets a server's state end the server, cloned from
/// [`GenServerCtx::stop_handle`] during `init`. [`stop`](Self::stop) makes the
/// loop break after the current message and exit normally; `terminate` runs on
/// the normal path.
pub struct StopHandle<G: GenServer> {
sys_tx: Sender<Sys<G>>,
}
impl<G: GenServer> Clone for StopHandle<G> {
fn clone(&self) -> Self {
StopHandle {
sys_tx: self.sys_tx.clone(),
}
}
}
impl<G: GenServer> StopHandle<G> {
/// End the server after the current message. Idempotent; a no-op once the
/// server is gone.
pub fn stop(&self) {
let _ = self.sys_tx.send(Sys::Stop);
}
}
/// Per-server timer bookkeeping, shared between the loop and every /// Per-server timer bookkeeping, shared between the loop and every
/// [`TimerHandle`] clone. A gen_server actor is single-threaded — handlers /// [`TimerHandle`] clone. A gen_server actor is single-threaded — handlers
/// and the loop never run concurrently — so this `Mutex` is always /// and the loop never run concurrently — so this `Mutex` is always
@@ -518,7 +653,10 @@ pub struct TimerHandle<G: GenServer> {
// Manual Clone for the same reason as `Watcher`: no `G: Clone` needed. // Manual Clone for the same reason as `Watcher`: no `G: Clone` needed.
impl<G: GenServer> Clone for TimerHandle<G> { impl<G: GenServer> Clone for TimerHandle<G> {
fn clone(&self) -> Self { fn clone(&self) -> Self {
TimerHandle { sys_tx: self.sys_tx.clone(), reg: self.reg.clone() } TimerHandle {
sys_tx: self.sys_tx.clone(),
reg: self.reg.clone(),
}
} }
} }
@@ -569,8 +707,18 @@ impl<G: GenServer> TimerHandle<G> {
// First instance fires after `every`; the payload is produced loop-side // First instance fires after `every`; the payload is produced loop-side
// from `make` on fire, so the tick carries only the stable id. // from `make` on fire, so the tick carries only the stable id.
let sub = send_after_to(every, self.sys_tx.clone(), Sys::Tick(local)); let sub = send_after_to(every, self.sys_tx.clone(), Sys::Tick(local));
reg.periodics.insert(local, Periodic { every, live: sub, make }); reg.periodics.insert(
debug_assert!(reg.rearm_tx.is_some(), "rearm_tx must be Some while periodics is non-empty"); local,
Periodic {
every,
live: sub,
make,
},
);
debug_assert!(
reg.rearm_tx.is_some(),
"rearm_tx must be Some while periodics is non-empty"
);
local local
} }
@@ -617,7 +765,9 @@ pub struct Watcher<G: GenServer> {
// regardless of the server type (it clones only the inner sender). // regardless of the server type (it clones only the inner sender).
impl<G: GenServer> Clone for Watcher<G> { impl<G: GenServer> Clone for Watcher<G> {
fn clone(&self) -> Self { fn clone(&self) -> Self {
Watcher { tx: self.tx.clone() } Watcher {
tx: self.tx.clone(),
}
} }
} }
@@ -643,11 +793,17 @@ pub struct GenServerBuilder<G: GenServer> {
state: G, state: G,
infos: Vec<Receiver<G::Info>>, infos: Vec<Receiver<G::Info>>,
supervisor: Option<Pid>, supervisor: Option<Pid>,
stack_opts: crate::scheduler::SpawnOpts,
} }
impl<G: GenServer> GenServerBuilder<G> { impl<G: GenServer> GenServerBuilder<G> {
pub fn new(state: G) -> Self { pub fn new(state: G) -> Self {
GenServerBuilder { state, infos: Vec::new(), supervisor: None } GenServerBuilder {
state,
infos: Vec::new(),
supervisor: None,
stack_opts: crate::scheduler::SpawnOpts::default(),
}
} }
/// Add an out-of-band channel; messages arriving on it are dispatched to /// Add an out-of-band channel; messages arriving on it are dispatched to
@@ -665,9 +821,18 @@ impl<G: GenServer> GenServerBuilder<G> {
self self
} }
/// Spawn the server actor and hand back its [`GenServerRef`]. The server's /// Stack shape for the server actor (RFC 019) — see
/// lifetime is governed by its refs, not by joining, so the backing join /// [`SpawnOpts`](crate::SpawnOpts). Useful for servers that recurse
/// handle is dropped. /// deeply or call into FFI with large C frames.
pub fn stack_opts(mut self, opts: crate::scheduler::SpawnOpts) -> Self {
self.stack_opts = opts;
self
}
/// Spawn the server actor and hand back its [`GenServerRef`] (an address;
/// the server's lifetime is its own, see the module docs). The backing
/// join handle is dropped. For a supervised server use
/// [`named`](Self::named) + [`NamedGenServerBuilder::run`] instead.
pub fn start(self) -> GenServerRef<G> { pub fn start(self) -> GenServerRef<G> {
self.spawn_server() self.spawn_server()
} }
@@ -677,7 +842,10 @@ impl<G: GenServer> GenServerBuilder<G> {
/// live server). Consumes the builder, carrying its `with_info` / `under` /// live server). Consumes the builder, carrying its `with_info` / `under`
/// configuration through. /// configuration through.
pub fn named(self, name: GenServerName<G>) -> NamedGenServerBuilder<G> { pub fn named(self, name: GenServerName<G>) -> NamedGenServerBuilder<G> {
NamedGenServerBuilder { builder: self, name: name.as_str() } NamedGenServerBuilder {
builder: self,
name: name.as_str(),
}
} }
/// Private shared body behind [`start`](Self::start) and /// Private shared body behind [`start`](Self::start) and
@@ -686,12 +854,25 @@ impl<G: GenServer> GenServerBuilder<G> {
/// under the name before returning. /// under the name before returning.
fn spawn_server(self) -> GenServerRef<G> { fn spawn_server(self) -> GenServerRef<G> {
let (tx, rx) = channel::<Envelope<G>>(); let (tx, rx) = channel::<Envelope<G>>();
let GenServerBuilder { state, infos, supervisor } = self; let GenServerBuilder {
state,
infos,
supervisor,
stack_opts,
} = self;
let keep = tx.clone();
let handle = match supervisor { let handle = match supervisor {
Some(sup) => spawn_under(sup, move || server_loop::<G>(rx, state, infos)), Some(sup) => crate::scheduler::spawn_under_with(sup, stack_opts, move || {
None => spawn(move || server_loop::<G>(rx, state, infos)), server_loop::<G>(keep, rx, state, infos)
}),
None => crate::scheduler::spawn_with(stack_opts, move || {
server_loop::<G>(keep, rx, state, infos)
}),
}; };
GenServerRef { tx, pid: handle.pid() } GenServerRef {
tx,
pid: handle.pid(),
}
} }
} }
@@ -719,7 +900,10 @@ impl<G> GenServerName<G> {
/// associated constants at call sites. /// associated constants at call sites.
#[inline] #[inline]
pub const fn new(name: &'static str) -> Self { pub const fn new(name: &'static str) -> Self {
Self { name, _marker: PhantomData } Self {
name,
_marker: PhantomData,
}
} }
/// The underlying registry key. /// The underlying registry key.
@@ -758,6 +942,12 @@ impl<G: GenServer> NamedGenServerBuilder<G> {
self self
} }
/// Stack shape for the server actor (see [`GenServerBuilder::stack_opts`]).
pub fn stack_opts(mut self, opts: crate::scheduler::SpawnOpts) -> Self {
self.builder = self.builder.stack_opts(opts);
self
}
/// Spawn the server and bind its name in one step. Fallible: returns /// Spawn the server and bind its name in one step. Fallible: returns
/// [`RegisterError::NameTaken`] if the name is already held by a different /// [`RegisterError::NameTaken`] if the name is already held by a different
/// live server. /// live server.
@@ -765,19 +955,38 @@ impl<G: GenServer> NamedGenServerBuilder<G> {
/// The inbox sender is published under the name **from the parent side, /// The inbox sender is published under the name **from the parent side,
/// before this returns**, so a by-name `call` / `cast` resolves the instant /// before this returns**, so a by-name `call` / `cast` resolves the instant
/// `start()` returns — no race with the server body. On a name clash the /// `start()` returns — no race with the server body. On a name clash the
/// just-spawned server is wound down (its only ref is dropped, closing the /// just-spawned server is stopped, so a failed bind leaks no actor.
/// inbox), so a failed bind leaks no actor.
pub fn start(self) -> Result<GenServerRef<G>, RegisterError> { pub fn start(self) -> Result<GenServerRef<G>, RegisterError> {
let NamedGenServerBuilder { builder, name } = self; let NamedGenServerBuilder { builder, name } = self;
let server = builder.spawn_server(); let server = builder.spawn_server();
match register_with::<Envelope<G>>(server.pid, name, server.tx.clone()) { match register_with::<Envelope<G>>(server.pid, name, server.tx.clone()) {
Ok(()) => Ok(server), Ok(()) => Ok(server),
Err(e) => { Err(e) => {
drop(server); // inbox closes → loop exits gracefully crate::scheduler::request_stop(server.pid); // never ran init
Err(e) Err(e)
} }
} }
} }
/// Run the server **inline, as the current actor**, bound to its name.
/// This is the supervised shape: the closure of a
/// [`ChildSpec`](crate::supervisor::ChildSpec) *is* the server, so the
/// supervisor's shutdown reaches it as [`GenServer::handle_shutdown`], a
/// restart runs the factory again and re-binds the name, and clients
/// address it by name ([`call`], [`cast`], [`whereis_server`]). Returns
/// when the server exits; [`RegisterError::NameTaken`] (before `init`) if
/// the name is held by a different live actor.
///
/// `under` / `stack_opts` are spawn options and do not apply here — the
/// actor already exists.
pub fn run(self) -> Result<(), RegisterError> {
let NamedGenServerBuilder { builder, name } = self;
let GenServerBuilder { state, infos, .. } = builder;
let (tx, rx) = channel::<Envelope<G>>();
register_with::<Envelope<G>>(crate::scheduler::self_pid(), name, tx.clone())?;
server_loop::<G>(tx, rx, state, infos);
Ok(())
}
} }
/// Resolve a [`GenServerName`] to a [`GenServerRef`] when you want a handle to hold or /// Resolve a [`GenServerName`] to a [`GenServerRef`] when you want a handle to hold or
@@ -836,10 +1045,17 @@ pub fn start_under<G: GenServer>(supervisor: Pid, state: G) -> GenServerRef<G> {
} }
fn server_loop<G: GenServer>( fn server_loop<G: GenServer>(
keep: Sender<Envelope<G>>,
rx: Receiver<Envelope<G>>, rx: Receiver<Envelope<G>>,
state: G, state: G,
mut infos: Vec<Receiver<G::Info>>, mut infos: Vec<Receiver<G::Info>>,
) { ) {
// The loop holds one inbox sender for its whole life: the inbox never
// closes, so refs are addresses and the server's lifetime is the actor's
// (stop handle, shutdown, stop, panic). The `Disconnected` arms below are
// defensive only.
let _keep = keep;
// Drop guard — owns the server state and the timer registry. // Drop guard — owns the server state and the timer registry.
// //
// Why a guard rather than code after the loop: // Why a guard rather than code after the loop:
@@ -907,9 +1123,18 @@ fn server_loop<G: GenServer>(
// Bind the ctx so the idle window set during init can be read back, then // Bind the ctx so the idle window set during init can be read back, then
// drop it — that drops the loop's own Sys sender, so a state that cloned no // drop it — that drops the loop's own Sys sender, so a state that cloned no
// Watcher/TimerHandle lets the arm auto-close (the unused-ctx behaviour). // Watcher/TimerHandle lets the arm auto-close (the unused-ctx behaviour).
let ctx = GenServerCtx { sys_tx, reg: reg.clone(), idle: Cell::new(None) }; let ctx = GenServerCtx {
sys_tx,
reg: reg.clone(),
idle: Cell::new(None),
trap: Cell::new(false),
};
guard.0.init(&ctx); guard.0.init(&ctx);
let idle = ctx.idle.get(); let idle = ctx.idle.get();
// Trapping is opted into during init and fixed for the loop's life. The
// inbox is armed only when set: an untrapped server keeps the fast path,
// and a shutdown request simply stops it as `request_stop` would.
let exits: Option<Receiver<ExitSignal>> = ctx.trap.get().then(crate::link::trap_exit);
drop(ctx); drop(ctx);
let mut monitors: Vec<Monitor> = Vec::new(); let mut monitors: Vec<Monitor> = Vec::new();
@@ -925,7 +1150,7 @@ fn server_loop<G: GenServer>(
}; };
loop { loop {
if monitors.is_empty() && !sys_open && infos.is_empty() { if exits.is_none() && monitors.is_empty() && !sys_open && infos.is_empty() {
// Fast path: no extra arms, no select overhead — park directly on // Fast path: no extra arms, no select overhead — park directly on
// the inbox. Mirrors the inbox arm of the select path below; any // the inbox. Mirrors the inbox arm of the select path below; any
// change there must be applied here too. // change there must be applied here too.
@@ -942,7 +1167,7 @@ fn server_loop<G: GenServer>(
guard.0.handle_idle(); guard.0.handle_idle();
reset_idle(&mut idle_deadline); reset_idle(&mut idle_deadline);
} }
// All ServerRefs dropped → inbox closed → shutdown. // Defensive: the loop holds a sender, so unreachable.
Err(RecvTimeoutError::Disconnected) => break, Err(RecvTimeoutError::Disconnected) => break,
} }
} }
@@ -953,17 +1178,21 @@ fn server_loop<G: GenServer>(
} }
} else { } else {
// Slow path: one or more extra arms live — build the arm slice and // Slow path: one or more extra arms live — build the arm slice and
// select. Arm order encodes priority: downs → system → infos → // select. Arm order encodes priority: exits → downs → system →
// inbox. The slice is rebuilt each iteration because the monitor // infos → inbox (a shutdown request is noticed under any load).
// and info sets shrink/grow. Mirrors the fast-path inbox park // The slice is rebuilt each iteration because the monitor and info
// above; keep them in sync. // sets shrink/grow. Mirrors the fast-path inbox park above; keep
let nd = monitors.len(); // monitor band: [0, nd) // them in sync.
let ne = exits.is_some() as usize; // exit arm: [0, ne)
let nd = ne + monitors.len(); // monitor band: [ne, nd)
let nw = sys_open as usize; // system arm: [nd, nd+nw) let nw = sys_open as usize; // system arm: [nd, nd+nw)
// info band: [nd+nw, nd+nw+ni) // info band: [nd+nw, nd+nw+ni)
// inbox arm: [nd+nw+ni] // inbox arm: [nd+nw+ni]
let sel = { let sel = {
let mut arms: Vec<&dyn Selectable> = let mut arms: Vec<&dyn Selectable> = Vec::with_capacity(nd + nw + infos.len() + 1);
Vec::with_capacity(nd + nw + infos.len() + 1); if let Some(e) = &exits {
arms.push(e);
}
for m in &monitors { for m in &monitors {
arms.push(&m.rx); arms.push(&m.rx);
} }
@@ -975,9 +1204,7 @@ fn server_loop<G: GenServer>(
} }
arms.push(&rx); arms.push(&rx);
match idle_deadline { match idle_deadline {
Some(dl) => { Some(dl) => select_timeout(&arms, dl.saturating_duration_since(Instant::now())),
select_timeout(&arms, dl.saturating_duration_since(Instant::now()))
}
None => Some(select(&arms)), None => Some(select(&arms)),
} }
}; };
@@ -991,10 +1218,25 @@ fn server_loop<G: GenServer>(
continue; continue;
} }
}; };
if i < nd { if i < ne {
// Exit arm: a shutdown request or a linked peer's death.
// The inbox lives for the loop's life, so it never closes.
let sig = exits.as_ref().and_then(|e| e.try_recv().ok().flatten());
if let Some(sig) = sig {
if sig.reason == DownReason::Shutdown {
match guard.0.handle_shutdown() {
ShutdownAction::Exit => break,
ShutdownAction::Continue => {}
}
} else {
guard.0.handle_exit(sig);
}
reset_idle(&mut idle_deadline);
}
} else if i < nd {
// Monitor band: a Down retires its arm either way (one-shot) // Monitor band: a Down retires its arm either way (one-shot)
// or closes without delivering (defensive; shouldn't happen). // or closes without delivering (defensive; shouldn't happen).
let m = monitors.remove(i); let m = monitors.remove(i - ne);
if let Ok(Some(down)) = m.rx.try_recv() { if let Ok(Some(down)) = m.rx.try_recv() {
guard.0.handle_down(down); guard.0.handle_down(down);
reset_idle(&mut idle_deadline); reset_idle(&mut idle_deadline);
@@ -1003,13 +1245,19 @@ fn server_loop<G: GenServer>(
match sys_rx.try_recv() { match sys_rx.try_recv() {
// Control intake, not a dispatched message: no idle reset. // Control intake, not a dispatched message: no idle reset.
Ok(Some(Sys::Watch(m))) => monitors.push(m), Ok(Some(Sys::Watch(m))) => monitors.push(m),
// The state ended the server: a normal exit.
Ok(Some(Sys::Stop)) => break,
Ok(Some(Sys::Timer(id, msg))) => { Ok(Some(Sys::Timer(id, msg))) => {
// The one-shot fired: retire its registry entry so the // The one-shot fired: retire its registry entry so the
// live set tracks only still-pending timers, then // live set tracks only still-pending timers, then
// dispatch. // dispatch.
match reg.lock() { match reg.lock() {
Ok(mut g) => { g.oneshots.remove(&id); } Ok(mut g) => {
Err(e) => panic!("smarm: gen_server reg lock poisoned (core corrupt): {e}"), g.oneshots.remove(&id);
}
Err(e) => {
panic!("smarm: gen_server reg lock poisoned (core corrupt): {e}")
}
} }
guard.0.handle_timer(msg); guard.0.handle_timer(msg);
reset_idle(&mut idle_deadline); reset_idle(&mut idle_deadline);
@@ -1023,7 +1271,9 @@ fn server_loop<G: GenServer>(
let msg = { let msg = {
let mut g = match reg.lock() { let mut g = match reg.lock() {
Ok(g) => g, Ok(g) => g,
Err(e) => panic!("smarm: gen_server reg lock poisoned (core corrupt): {e}"), Err(e) => panic!(
"smarm: gen_server reg lock poisoned (core corrupt): {e}"
),
}; };
let r = &mut *g; let r = &mut *g;
if let Some(p) = r.periodics.get_mut(&id) { if let Some(p) = r.periodics.get_mut(&id) {
@@ -1031,9 +1281,9 @@ fn server_loop<G: GenServer>(
let msg = (p.make)(); let msg = (p.make)();
let tx = match r.rearm_tx.as_ref() { let tx = match r.rearm_tx.as_ref() {
Some(tx) => tx.clone(), Some(tx) => tx.clone(),
None => panic!( None => {
"smarm: live periodic without rearm_tx (logic bug)" panic!("smarm: live periodic without rearm_tx (logic bug)")
), }
}; };
p.live = send_after_to(every, tx, Sys::Tick(id)); p.live = send_after_to(every, tx, Sys::Tick(id));
Some(msg) Some(msg)
+503 -72
View File
@@ -68,10 +68,38 @@
//! inbox or timer event; a replayed event may postpone again (it re-queues for //! inbox or timer event; a replayed event may postpone again (it re-queues for
//! the next transition). See the macro docs for the row surface and [`Step`] //! the next transition). See the macro docs for the row surface and [`Step`]
//! for how a postpone surfaces to the loop. //! for how a postpone surfaces to the loop.
//!
//! ## Stopping, shutdown, and exits
//!
//! A machine lives until it stops, is shut down, or is killed; a
//! [`GenStatemRef`] is an address, and dropping refs never ends it (the
//! gen_server rule — see its *When the server stops*). The supervised shape
//! is [`run_named`], which runs the machine inline as the current actor so it
//! is a direct `ChildSpec` child addressed by [`GenStatemName`]. A machine
//! can end itself: any row
//! body may call [`cx.stop()`](Cx::stop) (or use the `stop` tail keyword,
//! sugar for `{ cx.stop(); prev }`) — the loop breaks after that event and the
//! actor exits *normally* (OTP's `{stop, normal}`; a `Transient` child is not
//! restarted). [`terminate`](Machine::terminate) — the optional `terminate
//! { … }` macro block — runs on every exit path.
//!
//! From outside, [`GenStatemRef::shutdown`] (or a plain
//! [`request_shutdown`](crate::request_shutdown), which is what a supervisor
//! sends) asks the machine to stop. By default a machine does not trap exits
//! and the request stops it outright. A machine that calls
//! [`cx.trap_exit()`](Cx::trap_exit) in its initial `enter` instead receives
//! it as the **`shutdown` event**, routed by state like any other — a
//! `Connected` state may transition into `Draining` and stop later from a
//! timeout row, while a state with no `shutdown` row takes the macro's
//! default, `stop`. Linked-peer deaths reach a trapping machine as
//! `exit <pat>` events; an unmatched one is dropped like an unmatched info.
use crate::channel::{channel, select, Receiver, Sender}; use crate::channel::{channel, select, Receiver, Selectable, Sender};
use crate::link::ExitSignal;
use crate::monitor::{monitor, DownReason};
use crate::pid::Pid; use crate::pid::Pid;
use crate::scheduler::{cancel_timer, send_after_to, spawn as spawn_actor}; use crate::registry::{register_with, resolve_named_sender, RegisterError};
use crate::scheduler::{cancel_timer, request_shutdown, send_after_to};
use crate::timer::TimerId; use crate::timer::TimerId;
use std::collections::{HashMap, VecDeque}; use std::collections::{HashMap, VecDeque};
use std::marker::PhantomData; use std::marker::PhantomData;
@@ -119,6 +147,33 @@ pub trait Machine: Send + 'static {
/// state's `enter`, and returns [`Step::Transitioned`]; a stay or unmatched /// state's `enter`, and returns [`Step::Transitioned`]; a stay or unmatched
/// event returns [`Step::Stayed`]. /// event returns [`Step::Stayed`].
fn handle(&mut self, ev: Self::Ev, cx: &mut Cx<Self::Ev>) -> Step<Self::Ev>; fn handle(&mut self, ev: Self::Ev, cx: &mut Cx<Self::Ev>) -> Step<Self::Ev>;
/// Wrap a graceful shutdown request into this machine's event, so a
/// trapping machine (see [`Cx::trap_exit`]) can route it **by state**. The
/// macro generates it as `Ev::Shutdown` and matches it in `shutdown` rows;
/// its default for a state that writes no such row is `stop`. A
/// hand-written machine that returns `None` (the default) is simply
/// stopped — the loop breaks and `terminate` runs on the normal path.
fn shutdown_ev() -> Option<Self::Ev> {
None
}
/// Wrap a linked peer's death (an [`ExitSignal`] that is not a shutdown
/// request, delivered only when trapping) into this machine's event. The
/// macro generates it as `Ev::Exit(sig)` and matches it in `exit <pat>`
/// rows; an unmatched exit is silently dropped, like an unmatched info.
/// A hand-written machine that returns `None` (the default) drops it.
fn exit_ev(_sig: ExitSignal) -> Option<Self::Ev> {
None
}
/// Runs as the machine actor exits, on any exit path (a `stop`, a
/// graceful shutdown, a handler panic, a hard stop). Like
/// `gen_server::terminate`: on the panic and hard-stop paths it runs
/// mid-unwind — do not panic or park there; only on the normal path
/// (`stop`, shutdown rows) may it do real work. The macro's optional
/// `terminate { … }` block generates it.
fn terminate(&mut self) {}
} }
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -219,7 +274,11 @@ struct Timers {
impl Timers { impl Timers {
fn new() -> Self { fn new() -> Self {
Timers { next_local: 0, state: None, named: HashMap::new() } Timers {
next_local: 0,
state: None,
named: HashMap::new(),
}
} }
fn mint(&mut self) -> u64 { fn mint(&mut self) -> u64 {
@@ -243,12 +302,41 @@ impl Timers {
pub struct Cx<Ev> { pub struct Cx<Ev> {
sys_tx: Sender<Sys>, sys_tx: Sender<Sys>,
reg: Arc<Mutex<Timers>>, reg: Arc<Mutex<Timers>>,
/// Set by [`trap_exit`](Self::trap_exit) during `on_start`; read once by
/// the loop right after, fixed for the machine's life.
trap: bool,
/// Set by [`stop`](Self::stop); the loop breaks after the current event.
stop: bool,
_ev: PhantomData<fn() -> Ev>, _ev: PhantomData<fn() -> Ev>,
} }
impl<Ev> Cx<Ev> { impl<Ev> Cx<Ev> {
fn new(sys_tx: Sender<Sys>, reg: Arc<Mutex<Timers>>) -> Self { fn new(sys_tx: Sender<Sys>, reg: Arc<Mutex<Timers>>) -> Self {
Cx { sys_tx, reg, _ev: PhantomData } Cx {
sys_tx,
reg,
trap: false,
stop: false,
_ev: PhantomData,
}
}
/// Trap exits for the machine's lifetime: a shutdown request then arrives
/// as the `shutdown` event (routed by state) and linked-peer deaths as
/// `exit` events, instead of stopping the machine outright. Call it in the
/// initial state's `enter` (i.e. during `on_start`); later calls have no
/// effect.
pub fn trap_exit(&mut self) {
self.trap = true;
}
/// End the machine after the current event: the loop breaks and the actor
/// exits *normally* (OTP's `{stop, normal}`); `terminate` runs on the
/// normal path. Anything still queued or postponed is dropped. The `stop`
/// tail keyword in a macro row is sugar for `{ cx.stop(); prev }`.
/// Mirrors gen_server's [`StopHandle`](crate::gen_server::StopHandle).
pub fn stop(&mut self) {
self.stop = true;
} }
/// Arm the **state timeout**: fire a `state_timeout` event after `after` in /// Arm the **state timeout**: fire a `state_timeout` event after `after` in
@@ -378,8 +466,7 @@ pub enum SendError {
} }
/// A clonable handle to a running machine. Cloning yields another sender to the /// A clonable handle to a running machine. Cloning yields another sender to the
/// same inbox; the machine lives until the last `GenStatemRef` is dropped, at which /// same inbox. An address, not an owner: dropping refs never ends the machine.
/// point its inbox closes and the loop exits.
pub struct GenStatemRef<M: Machine> { pub struct GenStatemRef<M: Machine> {
tx: Sender<M::Ev>, tx: Sender<M::Ev>,
pid: Pid, pid: Pid,
@@ -387,7 +474,10 @@ pub struct GenStatemRef<M: Machine> {
impl<M: Machine> Clone for GenStatemRef<M> { impl<M: Machine> Clone for GenStatemRef<M> {
fn clone(&self) -> Self { fn clone(&self) -> Self {
GenStatemRef { tx: self.tx.clone(), pid: self.pid } GenStatemRef {
tx: self.tx.clone(),
pid: self.pid,
}
} }
} }
@@ -422,6 +512,19 @@ impl<M: Machine> GenStatemRef<M> {
self.send(ev).map_err(|_| CallError::Down)?; self.send(ev).map_err(|_| CallError::Down)?;
rx.recv().map_err(|_| CallError::Down) rx.recv().map_err(|_| CallError::Down)
} }
/// Ask the machine to shut down and block until it has fully exited, so
/// [`Machine::terminate`] has run by the time this returns. A trapping
/// machine winds down through its `shutdown` rows; any other is stopped
/// outright. Returns immediately if the machine is already gone. Waits as
/// long as the machine takes — a supervisor bounds that with its child's
/// [`Shutdown`](crate::supervisor::Shutdown) policy. Mirrors
/// [`GenServerRef::shutdown`](crate::gen_server::GenServerRef::shutdown).
pub fn shutdown(&self) {
let mon = monitor(self.pid);
request_shutdown(self.pid);
let _ = mon.rx.recv();
}
} }
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -430,42 +533,204 @@ impl<M: Machine> GenStatemRef<M> {
/// Spawn `machine` as an actor and hand back its [`GenStatemRef`]. Shape mirrors /// Spawn `machine` as an actor and hand back its [`GenStatemRef`]. Shape mirrors
/// `gen_server::start`: make the inbox, spawn the loop, return the ref; the /// `gen_server::start`: make the inbox, spawn the loop, return the ref; the
/// backing join handle is dropped (lifetime is governed by refs, not joining). /// backing join handle is dropped (the machine's lifetime is its own). For a
/// supervised machine use [`run_named`].
/// ///
/// Panics if called outside `Runtime::run()`. /// Panics if called outside `Runtime::run()`.
pub fn spawn<M: Machine>(machine: M) -> GenStatemRef<M> { pub fn spawn<M: Machine>(machine: M) -> GenStatemRef<M> {
spawn_with(crate::scheduler::SpawnOpts::default(), machine)
}
/// [`spawn`] with per-actor stack shape overrides (RFC 019) for the machine's
/// actor — see [`SpawnOpts`](crate::SpawnOpts). gen_statem has no builder
/// (its one-shot `spawn(machine)` shape predates RFC 019), so the opts ride
/// a `_with` variant like the scheduler's own spawns.
///
/// Panics if called outside `Runtime::run()`.
pub fn spawn_with<M: Machine>(opts: crate::scheduler::SpawnOpts, machine: M) -> GenStatemRef<M> {
let (tx, rx) = channel::<M::Ev>(); let (tx, rx) = channel::<M::Ev>();
let handle = spawn_actor(move || statem_loop(rx, machine)); let keep = tx.clone();
GenStatemRef { tx, pid: handle.pid() } let handle = crate::scheduler::spawn_with(opts, move || statem_loop(keep, rx, machine));
GenStatemRef {
tx,
pid: handle.pid(),
}
}
/// A typed, static name for a gen_statem, used to address a machine through
/// the registry without holding a [`GenStatemRef`]. Mirrors
/// [`GenServerName`](crate::gen_server::GenServerName): declare it as a
/// constant and bind it with [`run_named`].
pub struct GenStatemName<M> {
name: &'static str,
_marker: PhantomData<fn() -> M>,
}
impl<M> GenStatemName<M> {
/// Bind a static string as a machine name.
#[inline]
pub const fn new(name: &'static str) -> Self {
Self {
name,
_marker: PhantomData,
}
}
/// The underlying registry key.
#[inline]
pub const fn as_str(self) -> &'static str {
self.name
}
}
impl<M> Copy for GenStatemName<M> {}
impl<M> Clone for GenStatemName<M> {
fn clone(&self) -> Self {
*self
}
}
/// Run `machine` **inline, as the current actor**, bound to `name`. The
/// supervised shape: the closure of a
/// [`ChildSpec`](crate::supervisor::ChildSpec) *is* the machine, so the
/// supervisor's shutdown reaches it as a `shutdown` row, a restart runs the
/// factory again and re-binds the name, and clients address it by name
/// ([`send`], [`call`], [`whereis_machine`]). Returns when the machine exits;
/// [`RegisterError::NameTaken`] (before `on_start`) if the name is held by a
/// different live actor. Mirrors
/// [`NamedGenServerBuilder::run`](crate::gen_server::NamedGenServerBuilder::run).
pub fn run_named<M: Machine>(name: GenStatemName<M>, machine: M) -> Result<(), RegisterError> {
let (tx, rx) = channel::<M::Ev>();
register_with::<M::Ev>(crate::scheduler::self_pid(), name.as_str(), tx.clone())?;
statem_loop(tx, rx, machine);
Ok(())
}
/// Resolve a [`GenStatemName`] to a [`GenStatemRef`]; `None` if no live
/// machine holds the name. Panics if called outside `Runtime::run()`.
pub fn whereis_machine<M: Machine>(name: GenStatemName<M>) -> Option<GenStatemRef<M>> {
resolve_named_sender::<M::Ev>(name.as_str()).map(|(pid, tx)| GenStatemRef { tx, pid })
}
/// Push an event to the machine registered under `name`, resolving per send.
/// [`SendError::Down`] if no live machine holds the name.
pub fn send<M: Machine>(name: GenStatemName<M>, ev: M::Ev) -> Result<(), SendError> {
match whereis_machine(name) {
Some(m) => m.send(ev),
None => Err(SendError::Down),
}
}
/// Synchronous request-reply to the machine registered under `name`,
/// resolving per call (a machine restarted under the same name is reached
/// transparently). [`CallError::Down`] if no live machine holds the name.
pub fn call<M, T, F>(name: GenStatemName<M>, make: F) -> Result<T, CallError>
where
M: Machine,
T: Send + 'static,
F: FnOnce(Reply<T>) -> M::Ev,
{
match whereis_machine(name) {
Some(m) => m.call(make),
None => Err(CallError::Down),
}
}
/// Shut down the machine registered under `name` and wait for it (see
/// [`GenStatemRef::shutdown`]). A no-op if no live machine holds the name.
pub fn shutdown<M: Machine>(name: GenStatemName<M>) {
if let Some(m) = whereis_machine(name) {
m.shutdown();
}
} }
/// The machine actor body: `on_start`, then one `handle` per event until the /// The machine actor body: `on_start`, then one `handle` per event until the
/// inbox closes (all refs dropped → graceful shutdown). /// row resolves to `stop`, a shutdown row stops it, or the actor is stopped
/// from outside.
/// ///
/// Two intake sources are selected each iteration with the **timer arm above /// Intake arms are selected each iteration in priority order — **exits**
/// the inbox**, so a timeout fire is never starved by inbox traffic: `sys_rx` /// (only when trapping) above **timers** above the **inbox** — so a shutdown
/// carries timer fires armed through `cx`, `rx` is the user inbox. A fire is /// request or a timeout fire is never starved by inbox traffic. `sys_rx`
/// turned into the matching internal event (`state_timeout` / `timeout(name)`) /// carries timer fires armed through `cx`; a fire is turned into the matching
/// and run through the same `handle` dispatch as an inbox event — the /// internal event (`state_timeout` / `timeout(name)`) and run through the same
/// gen_statem model, where timeouts surface as ordinary events. /// `handle` dispatch as an inbox event — the gen_statem model, where timeouts
/// (and, when trapping, shutdown and exits) surface as ordinary events.
/// ///
/// The loop owns the **postpone queue**: a `handle` that defers its event hands /// The loop owns the **postpone queue**: a `handle` that defers its event hands
/// it back ([`Step::Postponed`]) for the queue; a `handle` that transitions /// it back ([`Step::Postponed`]) for the queue; a `handle` that transitions
/// ([`Step::Transitioned`]) triggers a [`replay`] of the queue in the new /// ([`Step::Transitioned`]) triggers a [`replay`] of the queue in the new
/// state, ahead of the next intake. /// state, ahead of the next intake.
fn statem_loop<M: Machine>(rx: Receiver<M::Ev>, mut machine: M) { fn statem_loop<M: Machine>(keep: Sender<M::Ev>, rx: Receiver<M::Ev>, machine: M) {
// One inbox sender lives with the loop: the inbox never closes, refs are
// addresses, the machine's lifetime is the actor's (stop row, shutdown,
// stop, panic). The `Disconnected` inbox arm below is defensive only.
let _keep = keep;
// Drop guard — owns the machine and the timer registry, so `terminate`
// fires on every exit path (clean close, `stop`, a handler panic, a hard
// stop) and the timer drain is sequenced before it. Same shape as
// gen_server's guard; see the rationale there.
struct Terminate<M: Machine>(M, Arc<Mutex<Timers>>);
impl<M: Machine> Drop for Terminate<M> {
fn drop(&mut self) {
{
let mut reg = match self.1.lock() {
Ok(g) => g,
Err(e) => panic!("smarm: gen_statem reg lock poisoned (core corrupt): {e}"),
};
if let Some((_, sub)) = reg.state.take() {
cancel_timer(sub);
}
for (_, (_, sub)) in reg.named.drain() {
cancel_timer(sub);
}
}
self.0.terminate();
}
}
let (sys_tx, sys_rx) = channel::<Sys>(); let (sys_tx, sys_rx) = channel::<Sys>();
let reg = Arc::new(Mutex::new(Timers::new())); let reg = Arc::new(Mutex::new(Timers::new()));
let mut guard = Terminate(machine, reg.clone());
// The loop owns `cx` (and through it a `sys_tx` clone) for its whole life, // The loop owns `cx` (and through it a `sys_tx` clone) for its whole life,
// so the sys arm never closes from under us — no auto-close dance needed. // so the sys arm never closes from under us — no auto-close dance needed.
let mut cx = Cx::new(sys_tx, reg.clone()); let mut cx = Cx::new(sys_tx, reg.clone());
// Events deferred by `postpone` rows, replayed FIFO on the next transition. // Events deferred by `postpone` rows, replayed FIFO on the next transition.
let mut postpone: VecDeque<M::Ev> = VecDeque::new(); let mut postpone: VecDeque<M::Ev> = VecDeque::new();
machine.on_start(&mut cx); guard.0.on_start(&mut cx);
// Trapping is opted into during on_start and fixed for the loop's life.
// The inbox is armed only when set: an untrapped machine keeps the
// two-arm select, and a shutdown request simply stops it as
// `request_stop` would.
let exits: Option<Receiver<ExitSignal>> = cx.trap.then(crate::link::trap_exit);
loop { loop {
// Timer arm first: a ready fire is taken in preference to the inbox. // Arm order encodes priority: exits → timers → inbox.
let i = select(&[&sys_rx, &rx]); let ne = exits.is_some() as usize;
if i == 0 { let i = {
let mut arms: Vec<&dyn Selectable> = Vec::with_capacity(3);
if let Some(e) = &exits {
arms.push(e);
}
arms.push(&sys_rx);
arms.push(&rx);
select(&arms)
};
if i < ne {
// Exit arm: a shutdown request or a linked peer's death. The trap
// inbox lives for the loop's life, so it never closes.
let sig = exits.as_ref().and_then(|e| e.try_recv().ok().flatten());
match sig {
Some(sig) if sig.reason == DownReason::Shutdown => match M::shutdown_ev() {
Some(ev) => dispatch(&mut guard.0, &mut cx, &mut postpone, ev),
None => cx.stop(),
},
Some(sig) => {
if let Some(ev) = M::exit_ev(sig) {
dispatch(&mut guard.0, &mut cx, &mut postpone, ev);
}
}
None => {}
}
} else if i == ne {
match sys_rx.try_recv() { match sys_rx.try_recv() {
Ok(Some(fire)) => { Ok(Some(fire)) => {
// Confirm the fire is still the live one before dispatching: // Confirm the fire is still the live one before dispatching:
@@ -475,7 +740,9 @@ fn statem_loop<M: Machine>(rx: Receiver<M::Ev>, mut machine: M) {
Sys::StateTimeout(local) => { Sys::StateTimeout(local) => {
let mut t = match reg.lock() { let mut t = match reg.lock() {
Ok(g) => g, Ok(g) => g,
Err(e) => panic!("smarm: gen_statem reg lock poisoned (core corrupt): {e}"), Err(e) => panic!(
"smarm: gen_statem reg lock poisoned (core corrupt): {e}"
),
}; };
match t.state { match t.state {
Some((live, _)) if live == local => { Some((live, _)) if live == local => {
@@ -488,7 +755,9 @@ fn statem_loop<M: Machine>(rx: Receiver<M::Ev>, mut machine: M) {
Sys::Timeout(name, local) => { Sys::Timeout(name, local) => {
let mut t = match reg.lock() { let mut t = match reg.lock() {
Ok(g) => g, Ok(g) => g,
Err(e) => panic!("smarm: gen_statem reg lock poisoned (core corrupt): {e}"), Err(e) => panic!(
"smarm: gen_statem reg lock poisoned (core corrupt): {e}"
),
}; };
match t.named.get(name) { match t.named.get(name) {
Some(&(live, _)) if live == local => { Some(&(live, _)) if live == local => {
@@ -500,7 +769,7 @@ fn statem_loop<M: Machine>(rx: Receiver<M::Ev>, mut machine: M) {
} }
}; };
if let Some(ev) = ev { if let Some(ev) = ev {
dispatch(&mut machine, &mut cx, &mut postpone, ev); dispatch(&mut guard.0, &mut cx, &mut postpone, ev);
} }
} }
// Single-receiver: nothing can drain the arm between select's // Single-receiver: nothing can drain the arm between select's
@@ -512,12 +781,16 @@ fn statem_loop<M: Machine>(rx: Receiver<M::Ev>, mut machine: M) {
} }
} else { } else {
match rx.try_recv() { match rx.try_recv() {
Ok(Some(ev)) => dispatch(&mut machine, &mut cx, &mut postpone, ev), Ok(Some(ev)) => dispatch(&mut guard.0, &mut cx, &mut postpone, ev),
Ok(None) => debug_assert!(false, "ready inbox was empty"), Ok(None) => debug_assert!(false, "ready inbox was empty"),
// All GenStatemRefs dropped → inbox closed → shutdown. // Defensive: the loop holds a sender, so unreachable.
Err(_) => break, Err(_) => break,
} }
} }
// A handler (or a replay) asked to stop: a normal exit.
if cx.stop {
break;
}
// Observation point so a machine fed a hot inbox stays preemptible and // Observation point so a machine fed a hot inbox stays preemptible and
// cancellable. // cancellable.
crate::check!(); crate::check!();
@@ -526,7 +799,8 @@ fn statem_loop<M: Machine>(rx: Receiver<M::Ev>, mut machine: M) {
/// Run one event through `handle` and act on its [`Step`]: stash a deferred /// Run one event through `handle` and act on its [`Step`]: stash a deferred
/// event on the postpone queue, or — on a real transition — [`replay`] the /// event on the postpone queue, or — on a real transition — [`replay`] the
/// queue in the new state. A stay/unmatched event needs nothing further. /// queue in the new state. A stay/unmatched event needs nothing further. A
/// [`Cx::stop`] raised by the handler skips the replay; the loop breaks next.
fn dispatch<M: Machine>( fn dispatch<M: Machine>(
machine: &mut M, machine: &mut M,
cx: &mut Cx<M::Ev>, cx: &mut Cx<M::Ev>,
@@ -536,14 +810,19 @@ fn dispatch<M: Machine>(
match machine.handle(ev, cx) { match machine.handle(ev, cx) {
Step::Postponed(ev) => postpone.push_back(ev), Step::Postponed(ev) => postpone.push_back(ev),
Step::Stayed => {} Step::Stayed => {}
Step::Transitioned => replay(machine, cx, postpone), Step::Transitioned => {
if !cx.stop {
replay(machine, cx, postpone)
}
}
} }
} }
/// Replay deferred events after a real transition: each goes back through /// Replay deferred events after a real transition: each goes back through
/// `handle` in FIFO order, in the now-current state. An event that postpones /// `handle` in FIFO order, in the now-current state. An event that postpones
/// again re-queues (to wait for the *next* transition); one that transitions /// again re-queues (to wait for the *next* transition); one that transitions
/// re-arms the replay, so a later state can in turn drain what is still pending. /// re-arms the replay, so a later state can in turn drain what is still pending;
/// one that raises [`Cx::stop`] ends the replay (and the machine) at once.
/// Subsequent events in a batch already see the post-transition state, since /// Subsequent events in a batch already see the post-transition state, since
/// `handle` reads the live state cell — the outer loop only re-runs to give /// `handle` reads the live state cell — the outer loop only re-runs to give
/// re-queued events another pass once a transition has occurred within a batch. /// re-queued events another pass once a transition has occurred within a batch.
@@ -562,6 +841,9 @@ fn replay<M: Machine>(machine: &mut M, cx: &mut Cx<M::Ev>, postpone: &mut VecDeq
Step::Stayed => {} Step::Stayed => {}
Step::Transitioned => transitioned = true, Step::Transitioned => transitioned = true,
} }
if cx.stop {
return;
}
} }
if !transitioned { if !transitioned {
return; return;
@@ -663,8 +945,9 @@ fn replay<M: Machine>(machine: &mut M, cx: &mut Cx<M::Ev>, postpone: &mut VecDeq
/// // The transition table. Group rows by current state with `on <pat>`. /// // The transition table. Group rows by current state with `on <pat>`.
/// // A row is: <kind> <event-pattern> [if <guard>] => <tail> , /// // A row is: <kind> <event-pattern> [if <guard>] => <tail> ,
/// // where <kind> is one of `cast`, `call`, `info`, `state_timeout` /// // where <kind> is one of `cast`, `call`, `info`, `state_timeout`
/// // (no pattern — it is a unit event), or `timeout <name-pattern>`, /// // (no pattern — it is a unit event), `timeout <name-pattern>`,
/// // and the tail is one of: /// // `shutdown` (unit; trapping machines only) or `exit <sig-pattern>`
/// // (trapping only), and the tail is one of:
/// // * a state tag `Door::Closed` (transition, or "stay" /// // * a state tag `Door::Closed` (transition, or "stay"
/// // if it equals current) /// // if it equals current)
/// // * a block ending in one `{ data.enters += 1; Door::Closed }` /// // * a block ending in one `{ data.enters += 1; Door::Closed }`
@@ -673,6 +956,9 @@ fn replay<M: Machine>(machine: &mut M, cx: &mut Cx<M::Ev>, postpone: &mut VecDeq
/// // * the keyword `postpone` (defer until next /// // * the keyword `postpone` (defer until next
/// // transition; cast/call/ /// // transition; cast/call/
/// // info only) /// // info only)
/// // * the keyword `stop` (end the machine
/// // normally; sugar for
/// // `{ cx.stop(); prev }`)
/// on Door::Open => { /// on Door::Open => {
/// cast DoorCast::Push => Door::Closed, /// cast DoorCast::Push => Door::Closed,
/// // An armed state-timeout surfaces as an ordinary event: /// // An armed state-timeout surfaces as an ordinary event:
@@ -696,6 +982,9 @@ fn replay<M: Machine>(machine: &mut M, cx: &mut Cx<M::Ev>, postpone: &mut VecDeq
/// on _ => { /// on _ => {
/// call DoorCall::GetState(r) => { r.reply(prev); prev }, /// call DoorCall::GetState(r) => { r.reply(prev); prev },
/// } /// }
///
/// // Optional: runs as the machine exits, on every exit path.
/// terminate { data.enters = 0; }
/// } /// }
/// ``` /// ```
/// ///
@@ -720,6 +1009,12 @@ fn replay<M: Machine>(machine: &mut M, cx: &mut Cx<M::Ev>, postpone: &mut VecDeq
/// ignores info writes no `info` rows at all. State-timeouts and named /// ignores info writes no `info` rows at all. State-timeouts and named
/// timeouts have **no** such default: a state that can see one must handle it /// timeouts have **no** such default: a state that can see one must handle it
/// (or `unhandled` it) or the match is non-exhaustive. /// (or `unhandled` it) or the match is non-exhaustive.
/// * **`shutdown` defaults to `stop`, `exit` to a silent drop.** Both reach a
/// machine only if its initial `enter` called `cx.trap_exit()`; a
/// non-trapping machine is stopped outright by a shutdown request. Write
/// `shutdown => …` rows only in the states that want to wind down first
/// (transition into a draining state, `stop` later); write `exit sig => …`
/// rows to react to linked-peer deaths. Neither is postponable.
/// * **Stay** = return the current tag. The `prev` you named in `context` is /// * **Stay** = return the current tag. The `prev` you named in `context` is
/// bound to the pre-handler state for exactly this — handy in any-state /// bound to the pre-handler state for exactly this — handy in any-state
/// (`on _`) rows where there is no single literal tag to write. /// (`on _`) rows where there is no single literal tag to write.
@@ -745,12 +1040,12 @@ fn replay<M: Machine>(machine: &mut M, cx: &mut Cx<M::Ev>, postpone: &mut VecDeq
/// # What it emits /// # What it emits
/// ///
/// The unified `enum $Ev` (the `Cast`/`Call`/`Info` wrappers plus the internal /// The unified `enum $Ev` (the `Cast`/`Call`/`Info` wrappers plus the internal
/// `StateTimeout` / `Timeout` events), `struct $Sm { state, data }`, /// `StateTimeout` / `Timeout` / `Shutdown` / `Exit` events), `struct $Sm {
/// `$Sm::start(init, data) -> GenStatemRef<$Sm>`, and the `Machine` impl: /// state, data }`, `$Sm::start(init, data) -> GenStatemRef<$Sm>`, and the
/// `on_start` runs the initial `enter`; `handle` is the dispatch match plus the /// `Machine` impl: `on_start` runs the initial `enter`; `handle` is the
/// stay/transition/unhandled apply-tail (the cell's sole writer, which also /// dispatch match plus the stay/transition/unhandled apply-tail (the cell's
/// auto-resets the state-timeout on every real transition); and the `enter` /// sole writer, which also auto-resets the state-timeout on every real
/// dispatch. /// transition); the `enter` dispatch; and `terminate` when the block is given.
/// ///
/// # Limitation /// # Limitation
/// ///
@@ -766,9 +1061,10 @@ macro_rules! gen_statem {
context ( $data:ident , $cur:ident , $cx:ident ) ; context ( $data:ident , $cur:ident , $cx:ident ) ;
enter { $( $est:pat => $ebody:expr ),+ $(,)? } enter { $( $est:pat => $ebody:expr ),+ $(,)? }
$( on $st:pat => { $($rows:tt)* } )+ $( on $st:pat => { $($rows:tt)* } )+
$( terminate { $($tbody:tt)* } )?
) => { ) => {
/// Unified inbox payload: the user's `cast`/`call`/`info` enums folded /// Unified inbox payload: the user's `cast`/`call`/`info` enums folded
/// together with the runtime's internal timeout events. /// together with the runtime's internal events.
enum $Ev { enum $Ev {
Cast($Cast), Cast($Cast),
Call($Call), Call($Call),
@@ -779,6 +1075,12 @@ macro_rules! gen_statem {
StateTimeout, StateTimeout,
/// A named timeout fired (matched in `timeout <pat>` rows). /// A named timeout fired (matched in `timeout <pat>` rows).
Timeout(&'static str), Timeout(&'static str),
/// A graceful shutdown request reached this (trapping) machine
/// (matched in `shutdown` rows). A state with no such row `stop`s.
Shutdown,
/// A linked peer died (trapping only; matched in `exit <pat>`
/// rows). An unmatched exit is silently dropped.
Exit($crate::ExitSignal),
} }
struct $sm { struct $sm {
@@ -787,8 +1089,14 @@ macro_rules! gen_statem {
} }
impl $sm { impl $sm {
/// The machine value, for [`gen_statem::run_named`]
/// (`$crate::gen_statem::run_named`) or `spawn`.
fn new(init: $State, data: $Data) -> $sm {
$sm { state: init, data }
}
fn start(init: $State, data: $Data) -> $crate::gen_statem::GenStatemRef<$sm> { fn start(init: $State, data: $Data) -> $crate::gen_statem::GenStatemRef<$sm> {
$crate::gen_statem::spawn($sm { state: init, data }) $crate::gen_statem::spawn($sm::new(init, data))
} }
#[allow(unused_variables)] #[allow(unused_variables)]
@@ -812,11 +1120,27 @@ macro_rules! gen_statem {
$Ev::Timeout(name) $Ev::Timeout(name)
} }
fn shutdown_ev() -> Option<$Ev> {
Some($Ev::Shutdown)
}
fn exit_ev(sig: $crate::ExitSignal) -> Option<$Ev> {
Some($Ev::Exit(sig))
}
fn on_start(&mut self, $cx: &mut $crate::gen_statem::Cx<$Ev>) { fn on_start(&mut self, $cx: &mut $crate::gen_statem::Cx<$Ev>) {
let s = self.state; let s = self.state;
self.enter(s, $cx); self.enter(s, $cx);
} }
$(
#[allow(unused_variables)]
fn terminate(&mut self) {
let $data = &mut self.data;
$($tbody)*
}
)?
#[allow(unused_variables)] #[allow(unused_variables)]
#[deny(unreachable_patterns)] // conflicting rows must fail even though #[deny(unreachable_patterns)] // conflicting rows must fail even though
// this match is external-macro-expanded // this match is external-macro-expanded
@@ -832,7 +1156,7 @@ macro_rules! gen_statem {
// untouched), then the consuming `match (state, event)` whose // untouched), then the consuming `match (state, event)` whose
// value is this `Resolution`. // value is this `Resolution`.
let next: $crate::gen_statem::Resolution<$State> = let next: $crate::gen_statem::Resolution<$State> =
$crate::gen_statem!(@arms ($Ev) ($cur, ev) [ ] [ ] $crate::gen_statem!(@arms ($Ev) ($cur, ev, $cx) [ ] [ ]
$( on $st => { $($rows)* } )+); $( on $st => { $($rows)* } )+);
match next { match next {
$crate::gen_statem::Resolution::To(s) if s == $cur => { $crate::gen_statem::Resolution::To(s) if s == $cur => {
@@ -868,7 +1192,7 @@ macro_rules! gen_statem {
// is the `Info` silent-drop — last and broadest, so per-state `info` rows // is the `Info` silent-drop — last and broadest, so per-state `info` rows
// stay reachable; cast/call/timeouts get no fallback, so a forgotten pair is // stay reachable; cast/call/timeouts get no fallback, so a forgotten pair is
// still E0004). // still E0004).
(@arms ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ]) => { (@arms ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ]) => {
{ {
// Phase 1 — postpone routing (borrow-only). A guard on a postpone // Phase 1 — postpone routing (borrow-only). A guard on a postpone
// row runs here, against by-ref bindings, so it must not depend on // row runs here, against by-ref bindings, so it must not depend on
@@ -885,140 +1209,247 @@ macro_rules! gen_statem {
// returned for it. // returned for it.
match ($ss, $se) { match ($ss, $se) {
$($arms)* $($arms)*
// Macro-injected defaults, last and broadest so per-state rows
// stay reachable; a user's own catch-all row may shadow them
// entirely, hence the allow.
#[allow(unreachable_patterns)]
(_, $Ev::Info(_)) => $crate::gen_statem::Resolution::Unhandled, (_, $Ev::Info(_)) => $crate::gen_statem::Resolution::Unhandled,
#[allow(unreachable_patterns)]
(_, $Ev::Exit(_)) => $crate::gen_statem::Resolution::Unhandled,
#[allow(unreachable_patterns)]
(_, $Ev::Shutdown) => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) },
} }
} }
}; };
// Open an on-block: remember its state pat, drain its rows, then continue. // Open an on-block: remember its state pat, drain its rows, then continue.
(@arms ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] (@arms ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ]
on $st:pat => { $($rows:tt)* } $($more:tt)* on $st:pat => { $($rows:tt)* } $($more:tt)*
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) [ $($arms)* ] [ $($post)* ] ($st) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx) [ $($arms)* ] [ $($post)* ] ($st)
{ $($rows)* } { $($more)* }) { $($rows)* } { $($more)* })
}; };
// ===== @rows: drain one on-block's rows, threading both accs ============= // ===== @rows: drain one on-block's rows, threading both accs =============
// cast, explicit refusal // cast, explicit refusal
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ cast $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* } { cast $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Cast($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ] [ $($arms)* ($st, $Ev::Cast($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ]
[ $($post)* ] [ $($post)* ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// cast, stop (end the machine normally after this event)
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ cast $ev:pat $(if $g:expr)? => stop , $($rows:tt)* } { $($more:tt)* }
) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Cast($ev)) $(if $g)? => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) }, ]
[ $($post)* ]
($st) { $($rows)* } { $($more)* })
};
// cast, postpone (defer the event; the replay in a later state handles it) // cast, postpone (defer the event; the replay in a later state handles it)
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ cast $ev:pat $(if $g:expr)? => postpone , $($rows:tt)* } { $($more:tt)* } { cast $ev:pat $(if $g:expr)? => postpone , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Cast($ev)) $(if $g)? => unreachable!("postponed event is replayed, not dispatched here"), ] [ $($arms)* ($st, $Ev::Cast($ev)) $(if $g)? => unreachable!("postponed event is replayed, not dispatched here"), ]
[ $($post)* ($st, $Ev::Cast($ev)) $(if $g)? => true, ] [ $($post)* ($st, $Ev::Cast($ev)) $(if $g)? => true, ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// cast, transition / stay / branch // cast, transition / stay / branch
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ cast $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* } { cast $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Cast($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ] [ $($arms)* ($st, $Ev::Cast($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ]
[ $($post)* ] [ $($post)* ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// call, explicit refusal // call, explicit refusal
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ call $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* } { call $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Call($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ] [ $($arms)* ($st, $Ev::Call($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ]
[ $($post)* ] [ $($post)* ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// call, stop (end the machine normally after this event)
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ call $ev:pat $(if $g:expr)? => stop , $($rows:tt)* } { $($more:tt)* }
) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Call($ev)) $(if $g)? => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) }, ]
[ $($post)* ]
($st) { $($rows)* } { $($more)* })
};
// call, postpone (the Reply rides inside the event onto the queue) // call, postpone (the Reply rides inside the event onto the queue)
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ call $ev:pat $(if $g:expr)? => postpone , $($rows:tt)* } { $($more:tt)* } { call $ev:pat $(if $g:expr)? => postpone , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Call($ev)) $(if $g)? => unreachable!("postponed event is replayed, not dispatched here"), ] [ $($arms)* ($st, $Ev::Call($ev)) $(if $g)? => unreachable!("postponed event is replayed, not dispatched here"), ]
[ $($post)* ($st, $Ev::Call($ev)) $(if $g)? => true, ] [ $($post)* ($st, $Ev::Call($ev)) $(if $g)? => true, ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// call, transition / stay / branch // call, transition / stay / branch
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ call $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* } { call $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Call($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ] [ $($arms)* ($st, $Ev::Call($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ]
[ $($post)* ] [ $($post)* ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// info, explicit refusal // info, explicit refusal
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ info $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* } { info $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Info($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ] [ $($arms)* ($st, $Ev::Info($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ]
[ $($post)* ] [ $($post)* ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// info, stop (end the machine normally after this event)
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ info $ev:pat $(if $g:expr)? => stop , $($rows:tt)* } { $($more:tt)* }
) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Info($ev)) $(if $g)? => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) }, ]
[ $($post)* ]
($st) { $($rows)* } { $($more)* })
};
// info, postpone // info, postpone
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ info $ev:pat $(if $g:expr)? => postpone , $($rows:tt)* } { $($more:tt)* } { info $ev:pat $(if $g:expr)? => postpone , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Info($ev)) $(if $g)? => unreachable!("postponed event is replayed, not dispatched here"), ] [ $($arms)* ($st, $Ev::Info($ev)) $(if $g)? => unreachable!("postponed event is replayed, not dispatched here"), ]
[ $($post)* ($st, $Ev::Info($ev)) $(if $g)? => true, ] [ $($post)* ($st, $Ev::Info($ev)) $(if $g)? => true, ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// info, transition / stay / branch // info, transition / stay / branch
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ info $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* } { info $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Info($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ] [ $($arms)* ($st, $Ev::Info($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ]
[ $($post)* ] [ $($post)* ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// state_timeout, explicit refusal (unit event — no pattern; not postponable) // state_timeout, explicit refusal (unit event — no pattern; not postponable)
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ state_timeout $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* } { state_timeout $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::StateTimeout) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ] [ $($arms)* ($st, $Ev::StateTimeout) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ]
[ $($post)* ] [ $($post)* ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// state_timeout, stop (end the machine normally after this event)
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ state_timeout $(if $g:expr)? => stop , $($rows:tt)* } { $($more:tt)* }
) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::StateTimeout) $(if $g)? => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) }, ]
[ $($post)* ]
($st) { $($rows)* } { $($more)* })
};
// state_timeout, transition / stay / branch // state_timeout, transition / stay / branch
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ state_timeout $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* } { state_timeout $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::StateTimeout) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ] [ $($arms)* ($st, $Ev::StateTimeout) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ]
[ $($post)* ] [ $($post)* ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// timeout, explicit refusal (pattern matches the name; not postponable) // timeout, explicit refusal (pattern matches the name; not postponable)
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ timeout $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* } { timeout $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Timeout($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ] [ $($arms)* ($st, $Ev::Timeout($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ]
[ $($post)* ] [ $($post)* ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// timeout, stop (end the machine normally after this event)
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ timeout $ev:pat $(if $g:expr)? => stop , $($rows:tt)* } { $($more:tt)* }
) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Timeout($ev)) $(if $g)? => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) }, ]
[ $($post)* ]
($st) { $($rows)* } { $($more)* })
};
// timeout, transition / stay / branch // timeout, transition / stay / branch
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ timeout $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* } { timeout $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Timeout($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ] [ $($arms)* ($st, $Ev::Timeout($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ]
[ $($post)* ] [ $($post)* ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// shutdown, explicit refusal (unit event — no pattern; not postponable; a state with no shutdown row defaults to `stop`)
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ shutdown $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* }
) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Shutdown) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ]
[ $($post)* ]
($st) { $($rows)* } { $($more)* })
};
// shutdown, stop (end the machine normally after this event)
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ shutdown $(if $g:expr)? => stop , $($rows:tt)* } { $($more:tt)* }
) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Shutdown) $(if $g)? => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) }, ]
[ $($post)* ]
($st) { $($rows)* } { $($more)* })
};
// shutdown, transition / stay / branch
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ shutdown $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* }
) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Shutdown) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ]
[ $($post)* ]
($st) { $($rows)* } { $($more)* })
};
// exit, explicit refusal (pattern matches the ExitSignal; not postponable)
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ exit $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* }
) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Exit($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ]
[ $($post)* ]
($st) { $($rows)* } { $($more)* })
};
// exit, stop (end the machine normally after this event)
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ exit $ev:pat $(if $g:expr)? => stop , $($rows:tt)* } { $($more:tt)* }
) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Exit($ev)) $(if $g)? => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) }, ]
[ $($post)* ]
($st) { $($rows)* } { $($more)* })
};
// exit, transition / stay / branch
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ exit $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* }
) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Exit($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ]
[ $($post)* ]
($st) { $($rows)* } { $($more)* })
};
// this block is drained: hand the remaining on-blocks back to @arms // this block is drained: hand the remaining on-blocks back to @arms
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ } { $($more:tt)* } { } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@arms ($Ev) ($ss, $se) [ $($arms)* ] [ $($post)* ] $($more)*) $crate::gen_statem!(@arms ($Ev) ($ss, $se, $cx) [ $($arms)* ] [ $($post)* ] $($more)*)
}; };
} }
+70 -4
View File
@@ -179,6 +179,34 @@ pub struct ActorInfo {
/// `budget-accounting` feature is enabled, since measuring it costs a /// `budget-accounting` feature is enabled, since measuring it costs a
/// timestamp read on every resume. /// timestamp read on every resume.
pub budget_cycles: u64, pub budget_cycles: u64,
/// RFC 019 §8 — this actor's stack, as the runtime sees it. All fields
/// are lock-free atomic reads, coherent for this incarnation via the
/// same generation check as the counters above. Exact RSS is
/// deliberately absent: `mincore` is debug tooling, never a runtime
/// path.
pub stack: StackInfo,
}
/// RFC 019 §8 — per-actor stack introspection. Sizes are page-rounded, as
/// [`Stack::new`](crate::stack::Stack::new) rounds them.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub struct StackInfo {
/// Usable stack size ([`SpawnOpts::stack_reserve`]
/// (crate::SpawnOpts::stack_reserve) or the Config/default).
pub reserve: usize,
/// PROT_NONE guard below the usable region.
pub guard: usize,
/// Sampled high-water depth in bytes: `top − lowest saved sp`. Sampled,
/// not exact — the context save at yields/parks/preemptions is the
/// sampler (RFC 019 §2), so a spike the actor never yielded inside is
/// invisible. 0 depth means "never descheduled at any depth", not
/// "never ran".
pub depth_high_water: usize,
/// Parks on this incarnation since its last shrink (or since install if
/// it has never shrunk) — the §3 cooldown counter, live.
pub parks_since_shrink: u32,
/// §3 shrinks performed on this incarnation.
pub shrinks: u32,
} }
/// A snapshot of every actor in the runtime at (approximately) one moment. /// A snapshot of every actor in the runtime at (approximately) one moment.
@@ -211,7 +239,10 @@ pub fn snapshot() -> RuntimeSnapshot {
actors.push(info); actors.push(info);
} }
} }
RuntimeSnapshot { format_version: SNAPSHOT_FORMAT_VERSION, actors } RuntimeSnapshot {
format_version: SNAPSHOT_FORMAT_VERSION,
actors,
}
}) })
} }
@@ -220,6 +251,22 @@ pub fn snapshot() -> RuntimeSnapshot {
/// slot was reused by another), out of range, or was never a real pid at /// slot was reused by another), out of range, or was never a real pid at
/// all. Unlike [`snapshot`], every field of the result describes the same /// all. Unlike [`snapshot`], every field of the result describes the same
/// instant, since there is only one actor to read. /// instant, since there is only one actor to read.
/// The stack shape `(reserve, guard)` of a live actor, page-rounded — the
/// RFC 019 introspection surface's first field (depth sampling and shrink
/// counters land with the shrink machinery). `None` if `pid` no longer names
/// a live actor. Takes the actor's cold lock briefly; debugging/assertion
/// use, not a hot-path call.
pub fn stack_shape(pid: Pid) -> Option<(usize, usize)> {
with_runtime(|inner| {
let slot = inner.slot_at(pid)?;
let cold = slot.cold.lock();
if slot.generation() != pid.generation() {
return None;
}
cold.actor.as_ref().map(|a| a.stack.shape())
})
}
pub fn actor_info(pid: Pid) -> Option<ActorInfo> { pub fn actor_info(pid: Pid) -> Option<ActorInfo> {
with_runtime(|inner| { with_runtime(|inner| {
let slot = inner.slot_at(pid)?; let slot = inner.slot_at(pid)?;
@@ -265,6 +312,14 @@ fn read_slot(slot: &Slot, idx: u32, mail: Option<&MailboxInfo>) -> Option<ActorI
drop(cold); drop(cold);
// Counters are plain atomics, read lock-free. // Counters are plain atomics, read lock-free.
let (reserve, guard, top, hwm, parks_since_shrink, shrinks) = slot.stack_introspect();
let stack = StackInfo {
reserve,
guard,
depth_high_water: top.saturating_sub(hwm),
parks_since_shrink,
shrinks,
};
let overruns = slot.overruns(); let overruns = slot.overruns();
let messages_received = slot.messages_received(); let messages_received = slot.messages_received();
let budget_cycles = slot.budget_cycles(); let budget_cycles = slot.budget_cycles();
@@ -290,6 +345,7 @@ fn read_slot(slot: &Slot, idx: u32, mail: Option<&MailboxInfo>) -> Option<ActorI
overruns, overruns,
messages_received, messages_received,
budget_cycles, budget_cycles,
stack,
}) })
} }
@@ -334,7 +390,10 @@ pub fn tree() -> RuntimeTree {
/// want to inspect again) and want the tree view of it without re-reading /// want to inspect again) and want the tree view of it without re-reading
/// the runtime. /// the runtime.
pub fn tree_from(snap: RuntimeSnapshot) -> RuntimeTree { pub fn tree_from(snap: RuntimeSnapshot) -> RuntimeTree {
let RuntimeSnapshot { format_version, actors } = snap; let RuntimeSnapshot {
format_version,
actors,
} = snap;
let mut index_of: HashMap<Pid, usize> = HashMap::with_capacity(actors.len()); let mut index_of: HashMap<Pid, usize> = HashMap::with_capacity(actors.len());
for (i, a) in actors.iter().enumerate() { for (i, a) in actors.iter().enumerate() {
@@ -365,7 +424,10 @@ pub fn tree_from(snap: RuntimeSnapshot) -> RuntimeTree {
.into_iter() .into_iter()
.filter_map(|i| build_node(i, &children_of, &orphaned, &mut slots)) .filter_map(|i| build_node(i, &children_of, &orphaned, &mut slots))
.collect(); .collect();
RuntimeTree { format_version, roots: root_nodes } RuntimeTree {
format_version,
roots: root_nodes,
}
} }
fn build_node( fn build_node(
@@ -383,5 +445,9 @@ fn build_node(
.collect() .collect()
}) })
.unwrap_or_default(); .unwrap_or_default();
Some(TreeNode { info, orphaned: orphaned[i], children }) Some(TreeNode {
info,
orphaned: orphaned[i],
children,
})
} }
+155 -240
View File
@@ -13,44 +13,68 @@
//! leaves the actor, no copying through an intermediary thread. Built on //! leaves the actor, no copying through an intermediary thread. Built on
//! these are the conveniences `read(fd, &mut buf)` and `write(fd, &buf)`. //! these are the conveniences `read(fd, &mut buf)` and `write(fd, &buf)`.
//! //!
//! Architecture //! Architecture (RFC 018: driver-enqueues)
//! ============ //! =======================================
//! Per `run()`, two OS threads: //! Per `run()`, two OS threads, each a *producer* behind the runtime's
//! - **epoll thread**: owns the epollfd. Loops in `epoll_wait`. On a //! two-call contract — make the actor runnable (`unpark_at`, whose enqueue
//! ready fd, pushes `Completion::FdReady { pid, fd, events }` to the //! tail wakes a parked scheduler), nothing else:
//! shared completion queue and writes the scheduler-wake pipe. On the
//! shutdown pipe (also registered in epollfd), exits.
//! - **pool thread**: blocks on the request mpsc. Runs the closure
//! inside `catch_unwind`, pushes `Completion::Blocking { pid, result }`,
//! writes the scheduler-wake pipe.
//! //!
//! Both threads share a single `completions: Arc<Mutex<VecDeque<Completion>>>` //! - **epoll thread**: owns `epoll_wait` on the epollfd. On a ready fd it
//! and the same scheduler-wake pipe. //! removes the parked waiter from the shared `waiters` map and DELs the
//! fd (both under the waiters lock — see below), then unparks the
//! actor directly. On the shutdown pipe (also registered in the
//! epollfd), exits.
//! - **pool thread**: blocks on the request mpsc. Runs the closure inside
//! `catch_unwind`, stashes the result in the actor's slot
//! (`pending_io_result`, under the cold lock, generation-checked),
//! decrements the runtime's `io_outstanding`, and unparks the actor.
//! //!
//! `epoll_ctl` (register/unregister fd interest) is called by the //! There is no shared completion queue and no wake pipe: each producer
//! scheduler thread *directly* on the epollfd. That's well-defined per //! routes its own completion, so the whole byte-vs-completion visibility
//! `epoll_ctl(2)`: a thread may be calling `epoll_wait` on the epollfd //! discipline of the drain era — and the stranded-completion hazards it
//! while another thread calls `epoll_ctl`. Avoids needing a second mpsc //! defended against — is unrepresentable. Producers reach the runtime
//! and a second wake mechanism. //! through a `Weak<RuntimeInner>`: upgraded per completion (the path is
//! syscall-bound; the refcount op is noise) and avoiding an Arc cycle
//! through `RuntimeInner::io`.
//!
//! `epoll_ctl` (register fd interest) is called by the scheduler thread
//! directly on the epollfd. That's well-defined per `epoll_ctl(2)`: a
//! thread may be calling `epoll_wait` on the epollfd while another thread
//! calls `epoll_ctl`.
//! //!
//! Epoll mode //! Epoll mode
//! ========== //! ==========
//! Level-triggered with EPOLLONESHOT. After a wakeup the kernel //! Level-triggered with EPOLLONESHOT. After a wakeup the kernel
//! auto-disarms the fd, so we never get two wakeups for one //! auto-disarms the fd, so we never get two wakeups for one
//! `wait_readable` call. The scheduler explicitly `EPOLL_CTL_DEL`s the fd //! `wait_readable` call. The epoll thread explicitly `EPOLL_CTL_DEL`s the
//! on completion to free the slot for re-registration. Net effect: each //! fd on readiness to free the slot for re-registration. Net effect: each
//! `wait_readable(fd)` is one ADD, one wakeup, one DEL — symmetric and //! `wait_readable(fd)` is one ADD, one wakeup, one DEL — symmetric and
//! stateless between calls. //! stateless between calls.
//! //!
//! ## The waiters lock is the ADD/DEL serialization
//!
//! Registration (scheduler thread: check-vacant, defensive DEL, ADD,
//! insert) and readiness consumption (epoll thread: remove, DEL) each run
//! entirely under the `waiters` mutex. This is what makes the
//! oneshot-rearm race unrepresentable: a woken actor re-registering the
//! same fd cannot interleave with the epoll thread's DEL for the *previous*
//! registration — whichever takes the lock second sees a consistent
//! kernel-side state. Lock order: `io` (the runtime's outer mutex, held by
//! scheduler-side callers) → `waiters` → slot/queue leaves via `unpark_at`.
//! The epoll thread takes `waiters` without `io` — it must never take
//! `io`, both for lock-order hygiene and because teardown holds `io` while
//! joining it.
//!
//! Fd hygiene //! Fd hygiene
//! ========== //! ==========
//! An actor stopped while waiting on an fd unwinds out of `wait_fd`'s park; //! An actor stopped while waiting on an fd unwinds out of `wait_fd`'s park;
//! a drop guard there (armed after a successful register, forgotten on a //! a drop guard there (armed after a successful register, forgotten on a
//! normal wake) removes the `waiters` entry iff it is still that wait's //! normal wake) calls [`IoThread::cancel_waiter`], which removes the
//! `(pid, epoch)` and only then `EPOLL_CTL_DEL`s the fd — an entry already //! `waiters` entry iff it is still that wait's `(pid, epoch)` and only then
//! consumed by a racing `FdReady` means the fd may carry someone else's //! `EPOLL_CTL_DEL`s the fd — an entry already consumed by the epoll thread
//! fresh registration, which must be left alone. `epoll_register` keeps a //! means the fd may carry someone else's fresh registration, which must be
//! defensive bare DEL before ADD as belt-and-braces. //! left alone. `epoll_register` keeps a defensive bare DEL before ADD as
//! belt-and-braces.
//! //!
//! Buffers used with `read`/`write` should be on fds opened with //! Buffers used with `read`/`write` should be on fds opened with
//! `O_NONBLOCK`. If they aren't, the syscall may block the scheduler //! `O_NONBLOCK`. If they aren't, the syscall may block the scheduler
@@ -68,13 +92,14 @@
//! they have no equivalent panic-propagation path. //! they have no equivalent panic-propagation path.
use crate::pid::Pid; use crate::pid::Pid;
use crate::runtime::RuntimeInner;
use std::any::Any; use std::any::Any;
use std::collections::{HashMap, VecDeque}; use std::collections::HashMap;
use std::io; use std::io;
use std::os::fd::RawFd; use std::os::fd::RawFd;
use std::panic; use std::panic;
use std::sync::mpsc; use std::sync::atomic::Ordering;
use std::sync::{Arc, Mutex}; use std::sync::{mpsc, Arc, Mutex, Weak};
use std::thread::JoinHandle as OsJoinHandle; use std::thread::JoinHandle as OsJoinHandle;
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -86,45 +111,31 @@ use std::thread::JoinHandle as OsJoinHandle;
pub type IoResult = Result<Box<dyn Any + Send>, Box<dyn Any + Send>>; pub type IoResult = Result<Box<dyn Any + Send>, Box<dyn Any + Send>>;
struct Request { struct Request {
/// The submitter's park-epoch — carried through to the `Blocking` /// The submitter's park-epoch — the eventual wake is epoch-matched.
/// completion so the wake is epoch-matched.
epoch: u32, epoch: u32,
pid: Pid, pid: Pid,
/// The work to perform. Returns the wire-form result directly. /// The work to perform. Returns the wire-form result directly.
work: Box<dyn FnOnce() -> IoResult + Send>, work: Box<dyn FnOnce() -> IoResult + Send>,
} }
/// Completion message from either IO thread back to the scheduler. /// The parked-waiter map, shared between scheduler-side registration and
pub enum Completion { /// the epoll thread's readiness consumption. See the module docs on why
/// A `block_on_io` closure has finished (Ok = return value, Err = panic /// this single lock is the ADD/DEL serialization.
/// payload). type Waiters = Arc<Mutex<HashMap<RawFd, (Pid, u32)>>>;
Blocking { pid: Pid, epoch: u32, result: IoResult },
/// An fd registered via `wait_readable`/`wait_writable` is ready. The
/// scheduler looks up the parked pid in `waiters`, unparks it, and
/// removes the entry. `pid` isn't in this variant because the epoll
/// thread doesn't have access to the `waiters` map; the scheduler
/// thread owns that.
FdReady { fd: RawFd, events: u32 },
}
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// IoThread — created per `run()`, owned by `SchedulerState`. // IoThread — created per `run()`, owned by `RuntimeInner::io`.
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
pub struct IoThread { pub struct IoThread {
// ----- Channels & queues -----
/// Submission queue into the blocking-work pool. /// Submission queue into the blocking-work pool.
tx: mpsc::Sender<Request>, tx: mpsc::Sender<Request>,
/// Shared completion queue, fed by both the pool and the epoll thread. /// One parked actor per registered fd. Populated by `epoll_register`,
completions: Arc<Mutex<VecDeque<Completion>>>, /// consumed by the epoll thread on readiness or `cancel_waiter` on an
/// Pipe the scheduler polls in its idle path. Both IO threads write to /// unwound wait.
/// `wake_write` after pushing a completion. waiters: Waiters,
wake_read: RawFd,
wake_write: RawFd,
// ----- Epoll machinery ----- // ----- Epoll machinery -----
/// The epollfd, owned by `IoThread`. Callable cross-thread via /// The epollfd, owned by `IoThread`. Callable cross-thread via
/// `epoll_ctl` per the man page. /// `epoll_ctl` per the man page.
epollfd: RawFd, epollfd: RawFd,
@@ -133,39 +144,24 @@ pub struct IoThread {
/// shutdown. /// shutdown.
shutdown_read: RawFd, shutdown_read: RawFd,
shutdown_write: RawFd, shutdown_write: RawFd,
/// One parked actor per registered fd. Populated by `wait_readable` /
/// `wait_writable` and drained by the scheduler when a `FdReady`
/// completion is processed.
pub waiters: HashMap<RawFd, (Pid, u32)>,
// ----- Threads ----- // ----- Threads -----
pool_thread: Option<OsJoinHandle<()>>, pool_thread: Option<OsJoinHandle<()>>,
epoll_thread: Option<OsJoinHandle<()>>, epoll_thread: Option<OsJoinHandle<()>>,
/// Number of `block_on_io` requests in-flight. Used by the scheduler's
/// idle path to decide whether to wait on the pipe or exit. Fd waits
/// are not counted here; they're counted by `waiters.len()`.
pub outstanding: u32,
} }
impl IoThread { impl IoThread {
pub fn start() -> io::Result<Self> { /// Start the pool and epoll threads. `rt` is the producers' route back
// Scheduler-facing wake pipe. /// into the runtime (slot table + unpark protocol); a `Weak` so the
let (wake_read, wake_write) = make_pipe()?; /// `RuntimeInner → IoThread → RuntimeInner` cycle never forms.
// Pool submission channel + shared completion queue. pub(crate) fn start(rt: Weak<RuntimeInner>) -> io::Result<Self> {
// Pool submission channel.
let (tx, rx) = mpsc::channel::<Request>(); let (tx, rx) = mpsc::channel::<Request>();
let completions: Arc<Mutex<VecDeque<Completion>>> = let waiters: Waiters = Arc::new(Mutex::new(HashMap::new()));
Arc::new(Mutex::new(VecDeque::new()));
// Epoll machinery. // Epoll machinery.
let epollfd = unsafe { libc::epoll_create1(libc::EPOLL_CLOEXEC) }; let epollfd = unsafe { libc::epoll_create1(libc::EPOLL_CLOEXEC) };
if epollfd < 0 { if epollfd < 0 {
// Best-effort fd cleanup before bailing.
unsafe {
libc::close(wake_read);
libc::close(wake_write);
}
return Err(io::Error::last_os_error()); return Err(io::Error::last_os_error());
} }
@@ -174,8 +170,6 @@ impl IoThread {
Err(e) => { Err(e) => {
unsafe { unsafe {
libc::close(epollfd); libc::close(epollfd);
libc::close(wake_read);
libc::close(wake_write);
} }
return Err(e); return Err(e);
} }
@@ -202,42 +196,37 @@ impl IoThread {
libc::close(epollfd); libc::close(epollfd);
libc::close(shutdown_read); libc::close(shutdown_read);
libc::close(shutdown_write); libc::close(shutdown_write);
libc::close(wake_read);
libc::close(wake_write);
} }
return Err(e); return Err(e);
} }
// Spawn pool thread. // Spawn pool thread.
let pool_comps = completions.clone(); let pool_rt = rt.clone();
let pool_thread = std::thread::Builder::new() let pool_thread = std::thread::Builder::new()
.name("smarm-io-pool".into()) .name("smarm-io-pool".into())
.spawn(move || pool_loop(rx, pool_comps, wake_write))?; .spawn(move || pool_loop(rx, pool_rt))?;
// Spawn epoll thread. // Spawn epoll thread.
let epoll_comps = completions.clone(); let epoll_waiters = waiters.clone();
let epoll_thread = std::thread::Builder::new() let epoll_thread = std::thread::Builder::new()
.name("smarm-io-epoll".into()) .name("smarm-io-epoll".into())
.spawn(move || epoll_loop(epollfd, epoll_comps, wake_write))?; .spawn(move || epoll_loop(epollfd, epoll_waiters, rt))?;
Ok(Self { Ok(Self {
tx, tx,
completions, waiters,
wake_read,
wake_write,
epollfd, epollfd,
shutdown_read, shutdown_read,
shutdown_write, shutdown_write,
waiters: HashMap::new(),
pool_thread: Some(pool_thread), pool_thread: Some(pool_thread),
epoll_thread: Some(epoll_thread), epoll_thread: Some(epoll_thread),
outstanding: 0,
}) })
} }
/// Hand a request to the pool. Increments `outstanding`. /// Hand a request to the pool. The caller (scheduler.rs) increments
/// `io_outstanding` BEFORE calling — the pool decrements on completion,
/// and an increment that trailed the completion would underflow.
pub fn submit(&mut self, pid: Pid, epoch: u32, work: Box<dyn FnOnce() -> IoResult + Send>) { pub fn submit(&mut self, pid: Pid, epoch: u32, work: Box<dyn FnOnce() -> IoResult + Send>) {
self.outstanding += 1;
// Send can only fail if the pool has hung up, which only happens // Send can only fail if the pool has hung up, which only happens
// on shutdown. submit during shutdown is a bug. // on shutdown. submit during shutdown is a bug.
if self.tx.send(Request { pid, epoch, work }).is_err() { if self.tx.send(Request { pid, epoch, work }).is_err() {
@@ -245,39 +234,13 @@ impl IoThread {
} }
} }
/// Drain every available completion. Caller (the scheduler) routes the
/// results and updates `outstanding` / `waiters` accordingly.
pub fn drain_completions(&mut self) -> Vec<Completion> {
let mut q = match self.completions.lock() {
Ok(g) => g,
Err(e) => panic!("smarm: io completions lock poisoned (core corrupt): {e}"),
};
let mut out = Vec::with_capacity(q.len());
while let Some(c) = q.pop_front() {
out.push(c);
}
out
}
pub fn wake_fd(&self) -> RawFd {
self.wake_read
}
/// Write the wake pipe directly: rouse every scheduler thread blocked in
/// its idle `poll_wake`. Used by the terminal (AllDone) path — an idle
/// sibling may be blocked on a snapshot that nothing will ever refresh
/// (an orphaned timer deadline, or `io_outstanding` from a waiter that
/// was stop-cancelled and so never produces a completion).
pub fn wake(&self) {
wake_scheduler(self.wake_write);
}
/// Register interest in `fd` becoming readable/writable; record `pid` /// Register interest in `fd` becoming readable/writable; record `pid`
/// as the parked waiter. The epoll thread will push a `FdReady` /// as the parked waiter. The epoll thread unparks it on readiness.
/// completion when the kernel signals. /// The caller increments `io_fd_waiters` BEFORE calling (mirror of
/// `submit`'s contract) and decrements it again if this errors.
/// ///
/// EPOLLONESHOT: one wakeup per registration. The scheduler must /// EPOLLONESHOT: one wakeup per registration; the epoll thread DELs on
/// `epoll_del` on completion to free the slot for re-registration. /// readiness, `cancel_waiter` DELs on an unwound wait.
pub fn epoll_register( pub fn epoll_register(
&mut self, &mut self,
fd: RawFd, fd: RawFd,
@@ -286,20 +249,24 @@ impl IoThread {
readable: bool, readable: bool,
writable: bool, writable: bool,
) -> io::Result<()> { ) -> io::Result<()> {
let mut waiters = match self.waiters.lock() {
Ok(g) => g,
Err(e) => panic!("smarm: io waiters lock poisoned (core corrupt): {e}"),
};
// Two actors waiting on the same fd would be a misuse: the kernel // Two actors waiting on the same fd would be a misuse: the kernel
// delivers exactly one EPOLLONESHOT wakeup, so the second waiter // delivers exactly one EPOLLONESHOT wakeup, so the second waiter
// would hang. Reject up front. // would hang. Reject up front.
if self.waiters.contains_key(&fd) { if waiters.contains_key(&fd) {
return Err(io::Error::new( return Err(io::Error::new(
io::ErrorKind::AlreadyExists, io::ErrorKind::AlreadyExists,
"fd already has a parked waiter", "fd already has a parked waiter",
)); ));
} }
// Belt-and-braces: the unwind guard in `wait_fd` is responsible for // Belt-and-braces: `cancel_waiter` is responsible for cleaning up a
// cleaning up a stopped waiter's registration, but a bare DEL is // stopped waiter's registration, but a bare DEL is harmless if the
// harmless if the fd isn't registered (ENOENT) and removes any leak // fd isn't registered (ENOENT) and removes any leak a path we
// a path we haven't thought of might leave behind. // haven't thought of might leave behind.
unsafe { unsafe {
libc::epoll_ctl(self.epollfd, libc::EPOLL_CTL_DEL, fd, std::ptr::null_mut()); libc::epoll_ctl(self.epollfd, libc::EPOLL_CTL_DEL, fd, std::ptr::null_mut());
} }
@@ -315,26 +282,35 @@ impl IoThread {
events, events,
u64: fd as u64, u64: fd as u64,
}; };
let r = unsafe { let r =
libc::epoll_ctl(self.epollfd, libc::EPOLL_CTL_ADD, fd, &mut ev as *mut _) unsafe { libc::epoll_ctl(self.epollfd, libc::EPOLL_CTL_ADD, fd, &mut ev as *mut _) };
};
if r < 0 { if r < 0 {
return Err(io::Error::last_os_error()); return Err(io::Error::last_os_error());
} }
self.waiters.insert(fd, (pid, epoch)); waiters.insert(fd, (pid, epoch));
Ok(()) Ok(())
} }
/// Remove `fd` from the epollfd. Called by the scheduler after a /// Remove `fd`'s waiter iff it is still `(pid, epoch)`, DELing the fd
/// `FdReady` completion, so the next `wait_readable(fd)` can ADD again. /// from the epollfd in the same critical section. Returns whether the
/// /// entry was removed (the caller then decrements `io_fd_waiters`).
/// Does NOT touch `waiters` — that's the scheduler's bookkeeping; this /// `false` means the epoll thread consumed the registration first —
/// is purely the kernel-side cleanup. /// the fd may already carry someone else's fresh ADD; hands off.
pub fn epoll_deregister(&mut self, fd: RawFd) { pub fn cancel_waiter(&mut self, fd: RawFd, pid: Pid, epoch: u32) -> bool {
let mut waiters = match self.waiters.lock() {
Ok(g) => g,
Err(e) => panic!("smarm: io waiters lock poisoned (core corrupt): {e}"),
};
if waiters.get(&fd) == Some(&(pid, epoch)) {
waiters.remove(&fd);
// EPOLL_CTL_DEL of an already-removed fd returns ENOENT; ignore. // EPOLL_CTL_DEL of an already-removed fd returns ENOENT; ignore.
unsafe { unsafe {
libc::epoll_ctl(self.epollfd, libc::EPOLL_CTL_DEL, fd, std::ptr::null_mut()); libc::epoll_ctl(self.epollfd, libc::EPOLL_CTL_DEL, fd, std::ptr::null_mut());
} }
true
} else {
false
}
} }
} }
@@ -354,7 +330,10 @@ impl Drop for IoThread {
let real_tx = std::mem::replace(&mut self.tx, dead_tx); let real_tx = std::mem::replace(&mut self.tx, dead_tx);
drop(real_tx); drop(real_tx);
// 3. Join both threads. // 3. Join both threads. Safe even while the caller holds the
// runtime's `io` mutex: neither thread ever takes it (they reach
// the runtime through a Weak they upgrade per completion, and
// the epoll thread's only lock is `waiters`).
if let Some(h) = self.epoll_thread.take() { if let Some(h) = self.epoll_thread.take() {
let _ = h.join(); let _ = h.join();
} }
@@ -367,8 +346,6 @@ impl Drop for IoThread {
libc::close(self.epollfd); libc::close(self.epollfd);
libc::close(self.shutdown_read); libc::close(self.shutdown_read);
libc::close(self.shutdown_write); libc::close(self.shutdown_write);
libc::close(self.wake_read);
libc::close(self.wake_write);
} }
} }
} }
@@ -379,36 +356,38 @@ impl Drop for IoThread {
const SHUTDOWN_EPOLL_TOKEN: u64 = u64::MAX; const SHUTDOWN_EPOLL_TOKEN: u64 = u64::MAX;
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// Pool loop // Pool loop (producer: Blocking completions)
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
fn pool_loop( fn pool_loop(rx: mpsc::Receiver<Request>, rt: Weak<RuntimeInner>) {
rx: mpsc::Receiver<Request>,
completions: Arc<Mutex<VecDeque<Completion>>>,
wake_write: RawFd,
) {
while let Ok(Request { pid, epoch, work }) = rx.recv() { while let Ok(Request { pid, epoch, work }) = rx.recv() {
let result: IoResult = match panic::catch_unwind(panic::AssertUnwindSafe(work)) { let result: IoResult = match panic::catch_unwind(panic::AssertUnwindSafe(work)) {
Ok(r) => r, Ok(r) => r,
Err(payload) => Err(payload), Err(payload) => Err(payload),
}; };
match completions.lock() { let Some(inner) = rt.upgrade() else { return };
Ok(mut g) => g.push_back(Completion::Blocking { pid, epoch, result }), // Stash the result under the cold lock (generation-checked: an
Err(e) => panic!("smarm: io completions lock poisoned (core corrupt): {e}"), // actor stopped with the op in flight discards it), decrement the
// in-flight count, then wake through the epoch-matched unpark. The
// unpark's enqueue tail wakes a parked scheduler; the actor stays
// `live` until it resumes and finalizes, so the decrement's
// ordering against the termination verdict is not load-bearing.
if let Some(slot) = inner.slot_at(pid) {
let mut cold = slot.cold.lock();
if slot.generation() == pid.generation() {
cold.pending_io_result = Some(result);
} }
wake_scheduler(wake_write); }
inner.io_outstanding.fetch_sub(1, Ordering::AcqRel);
inner.unpark_at(pid, epoch);
} }
} }
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// Epoll loop // Epoll loop (producer: FdReady completions)
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
fn epoll_loop( fn epoll_loop(epollfd: RawFd, waiters: Waiters, rt: Weak<RuntimeInner>) {
epollfd: RawFd,
completions: Arc<Mutex<VecDeque<Completion>>>,
wake_write: RawFd,
) {
// Buffer for epoll_wait. 64 is plenty for our scale; if a real load // Buffer for epoll_wait. 64 is plenty for our scale; if a real load
// appears that needs more, this is a one-line change. // appears that needs more, this is a one-line change.
const MAX_EVENTS: usize = 64; const MAX_EVENTS: usize = 64;
@@ -416,12 +395,7 @@ fn epoll_loop(
loop { loop {
let n = unsafe { let n = unsafe {
libc::epoll_wait( libc::epoll_wait(epollfd, events.as_mut_ptr(), MAX_EVENTS as libc::c_int, -1)
epollfd,
events.as_mut_ptr(),
MAX_EVENTS as libc::c_int,
-1,
)
}; };
if n < 0 { if n < 0 {
@@ -436,29 +410,36 @@ fn epoll_loop(
} }
let mut shutdown_requested = false; let mut shutdown_requested = false;
let mut pushed_any = false;
{
let mut q = match completions.lock() {
Ok(g) => g,
Err(e) => panic!("smarm: io completions lock poisoned (core corrupt): {e}"),
};
for ev in events.iter().take(n as usize) { for ev in events.iter().take(n as usize) {
if ev.u64 == SHUTDOWN_EPOLL_TOKEN { if ev.u64 == SHUTDOWN_EPOLL_TOKEN {
shutdown_requested = true; shutdown_requested = true;
continue; continue;
} }
let fd = ev.u64 as RawFd; let fd = ev.u64 as RawFd;
let evs = ev.events; // Consume the registration: remove + DEL under the waiters
q.push_back(Completion::FdReady { // lock (the ADD/DEL serialization — see module docs). A
fd, // vanished entry means `cancel_waiter` beat us: the wake is
events: evs, // already moot.
}); let entry = {
pushed_any = true; let mut w = match waiters.lock() {
Ok(g) => g,
Err(e) => {
panic!("smarm: io waiters lock poisoned (core corrupt): {e}")
}
};
let entry = w.remove(&fd);
if entry.is_some() {
unsafe {
libc::epoll_ctl(epollfd, libc::EPOLL_CTL_DEL, fd, std::ptr::null_mut());
} }
} }
entry
if pushed_any { };
wake_scheduler(wake_write); if let Some((pid, epoch)) = entry {
let Some(inner) = rt.upgrade() else { return };
inner.io_fd_waiters.fetch_sub(1, Ordering::AcqRel);
inner.unpark_at(pid, epoch);
}
} }
if shutdown_requested { if shutdown_requested {
return; return;
@@ -466,27 +447,8 @@ fn epoll_loop(
} }
} }
/// Write one byte to the scheduler's wake pipe. Retries on EINTR; ignores
/// EAGAIN (pipe full means there's already an outstanding wake we haven't
/// consumed yet, which is sufficient).
fn wake_scheduler(wake_write: RawFd) {
let buf: [u8; 1] = [0];
unsafe {
loop {
let n = libc::write(wake_write, buf.as_ptr() as *const _, 1);
if n < 0 {
let e = *libc::__errno_location();
if e == libc::EINTR {
continue;
}
}
break;
}
}
}
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// Pipe helpers (unchanged from v0.2) // Pipe helper
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
fn make_pipe() -> io::Result<(RawFd, RawFd)> { fn make_pipe() -> io::Result<(RawFd, RawFd)> {
@@ -497,50 +459,3 @@ fn make_pipe() -> io::Result<(RawFd, RawFd)> {
} }
Ok((fds[0], fds[1])) Ok((fds[0], fds[1]))
} }
/// Drain pending bytes from the wake pipe. Nonblocking (pipe is O_NONBLOCK).
///
/// DISCIPLINE: called only by the phase-1 drain-lock winner, immediately
/// before `drain_completions`. Bytes are the notification channel for
/// completions; consuming one anywhere else can strand the completion it
/// announces (see the lost-wakeup note at the call site in `schedule_loop`).
pub fn drain_wake_pipe(fd: RawFd) {
let mut buf = [0u8; 64];
loop {
let n = unsafe { libc::read(fd, buf.as_mut_ptr() as *mut _, buf.len()) };
if n <= 0 {
break;
}
}
}
/// Block on `fd` for up to `timeout`, returning when either there's data
/// to read or the timeout elapses. `None` for `timeout` means wait forever.
pub fn poll_wake(fd: RawFd, timeout: Option<std::time::Duration>) {
let timeout_ms: libc::c_int = match timeout {
None => -1,
Some(d) => {
let ms = d.as_millis();
if ms > i32::MAX as u128 {
i32::MAX
} else {
ms as i32
}
}
};
let mut pfd = libc::pollfd {
fd,
events: libc::POLLIN,
revents: 0,
};
loop {
let r = unsafe { libc::poll(&mut pfd as *mut _, 1, timeout_ms) };
if r < 0 {
let e = unsafe { *libc::__errno_location() };
if e == libc::EINTR {
continue;
}
}
break;
}
}
+41 -33
View File
@@ -11,34 +11,36 @@
//! //!
//! See `LOOM.md` for the design intent and the deferred-for-later list. //! See `LOOM.md` for the design intent and the deferred-for-later list.
pub mod stack;
pub mod context;
pub mod preempt;
pub mod pid;
pub mod actor; pub mod actor;
pub mod causal;
pub mod channel; pub mod channel;
pub mod scheduler; pub mod context;
pub mod supervisor;
pub mod timer;
pub mod io;
pub mod mutex;
pub mod monitor;
pub mod registry;
pub mod pg;
pub mod link;
pub mod gen_server; pub mod gen_server;
pub mod gen_statem; pub mod gen_statem;
pub mod introspect; pub mod introspect;
pub mod io;
pub mod link;
pub mod monitor;
pub mod mutex;
#[cfg(feature = "observer")] #[cfg(feature = "observer")]
pub mod observer; pub mod observer;
pub mod runtime; pub(crate) mod park;
pub mod pg;
pub mod pid;
pub mod preempt;
pub(crate) mod raw_mutex; pub(crate) mod raw_mutex;
pub(crate) mod slot_state; pub mod registry;
pub(crate) mod sync_shim;
#[doc(hidden)] // pub only so benches/rq_micro.rs can drive the raw structures #[doc(hidden)] // pub only so benches/rq_micro.rs can drive the raw structures
pub mod run_queue; pub mod run_queue;
pub mod runtime;
pub mod scheduler;
pub(crate) mod signal;
pub(crate) mod slot_state;
pub mod stack;
pub mod supervisor;
pub(crate) mod sync_shim;
pub mod timer;
pub mod trace; pub mod trace;
pub mod causal;
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// Global allocator // Global allocator
@@ -57,35 +59,41 @@ pub use channel::{
}; };
pub use gen_server::{ pub use gen_server::{
call, cast, shutdown, whereis_server, CallError, CallTimeoutError, CastError, GenServer, call, cast, shutdown, whereis_server, CallError, CallTimeoutError, CastError, GenServer,
NamedGenServerBuilder, GenServerBuilder, GenServerCtx, GenServerName, GenServerRef, TimerHandle, Watcher, GenServerBuilder, GenServerCtx, GenServerName, GenServerRef, NamedGenServerBuilder,
ShutdownAction, StopHandle, TimerHandle, Watcher,
}; };
pub use gen_statem::{ pub use gen_statem::{
CallError as GenStatemCallError, Cx, Machine, Reply, Resolution, SendError as GenStatemSendError, CallError as GenStatemCallError, Cx, GenStatemName, GenStatemRef, Machine, Reply, Resolution,
GenStatemRef, SendError as GenStatemSendError,
}; };
pub use introspect::{ pub use introspect::{
actor_info, snapshot, tree, tree_from, ActorInfo, ActorState, RuntimeSnapshot, RuntimeTree, actor_info, snapshot, tree, tree_from, ActorInfo, ActorState, RuntimeSnapshot, RuntimeTree,
TreeNode, SNAPSHOT_FORMAT_VERSION, StackInfo, TreeNode, SNAPSHOT_FORMAT_VERSION,
}; };
pub use link::{link, trap_exit, unlink, ExitSignal};
pub use monitor::{
demonitor, mark_watchable, monitor, terminal_reason, Down, DownReason, Monitor, MonitorId,
};
pub use mutex::{LockTimeout, Mutex, MutexGuard};
#[cfg(feature = "observer")] #[cfg(feature = "observer")]
pub use observer::{ObserverReply, ObserverRequest}; pub use observer::{ObserverReply, ObserverRequest};
pub use link::{link, trap_exit, unlink, ExitSignal}; pub use pg::{
pub use monitor::{demonitor, monitor, Down, DownReason, Monitor, MonitorId}; dispatch, join, leave, members, members_as, pick, pick_as, Incarnation, Member, NodeId,
pub use mutex::{LockTimeout, Mutex, MutexGuard}; };
pub use pid::{Addressable, Erased, Name, Pid, RawPid}; pub use pid::{Addressable, Erased, Name, Pid, RawPid};
pub use pg::{dispatch, join, leave, members, members_as, pick, pick_as, Incarnation, Member, NodeId};
pub use registry::{ pub use registry::{
install, lookup_as, register, send, send_dyn, send_to, unregister, whereis, RegisterError, install, lookup_as, register, resolve_name, send, send_dyn, send_to, unregister, whereis,
SendError, NameResolution, RegisterError, SendError,
}; };
pub use runtime::{init, Config, Runtime}; pub use runtime::{init, Config, Runtime, RuntimeHandle};
pub use scheduler::{ pub use scheduler::{
block_on_io, cancel_timer, request_stop, run, self_pid, send_after, send_after_named, block_on_io, cancel_timer, request_shutdown, request_stop, run, self_pid, send_after,
send_after_named_wall, send_after_wall, sleep, sleep_wall, send_after_named, send_after_named_wall, send_after_wall, sleep, sleep_wall, spawn, spawn_addr,
spawn, spawn_addr, spawn_under, wait_readable, wait_readable_timeout, wait_writable, spawn_addr_with, spawn_under, spawn_under_with, spawn_with, try_spawn, try_spawn_under_with,
wait_writable_timeout, yield_now, FdArm, JoinError, JoinHandle, wait_readable, wait_readable_timeout, wait_writable, wait_writable_timeout, yield_now, FdArm,
JoinError, JoinHandle, SpawnError, SpawnOpts,
}; };
pub use supervisor::{ChildSpec, OneForOne, Restart, Signal, Strategy}; pub use supervisor::{ChildSpec, OneForOne, Restart, Shutdown, Signal, Strategy};
pub use timer::TimerId; pub use timer::TimerId;
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
+4 -1
View File
@@ -157,7 +157,10 @@ pub fn link<A>(target: Pid<A>) {
}); });
match my_trap { match my_trap {
Some(tx) => { Some(tx) => {
let _ = tx.send(ExitSignal { from: target, reason: DownReason::NoProc }); let _ = tx.send(ExitSignal {
from: target,
reason: DownReason::NoProc,
});
} }
None => request_stop(me), None => request_stop(me),
} }
+67 -1
View File
@@ -98,6 +98,11 @@ pub enum DownReason {
Panic, Panic,
/// The target was cooperatively cancelled via `request_stop`. /// The target was cooperatively cancelled via `request_stop`.
Stopped, Stopped,
/// A graceful shutdown was requested via `request_shutdown`. Only ever
/// appears in an [`ExitSignal`](crate::link::ExitSignal) delivered to a
/// trapping actor — never in a [`Down`]: a target that honours the request
/// exits *normally*, one that does not trap is `Stopped`.
Shutdown,
/// The target was already gone (finished and reclaimed, or never alive) /// The target was already gone (finished and reclaimed, or never alive)
/// at the moment `monitor()` was called. /// at the moment `monitor()` was called.
NoProc, NoProc,
@@ -171,12 +176,73 @@ pub fn monitor<A>(target: Pid<A>) -> Monitor {
}); });
if !registered { if !registered {
let _ = tx.send(Down { pid: target, reason: DownReason::NoProc }); let _ = tx.send(Down {
pid: target,
reason: DownReason::NoProc,
});
} }
Monitor { id, target, rx } Monitor { id, target, rx }
} }
/// Flag `target`'s tenancy as watchable: its death will stamp the slot's
/// terminal record (see [`terminal_reason`]), exactly as registering a name
/// does. The bridge calls this wherever a smarm pid is *encoded across the
/// boundary* — a contract reply, an introspection listing — because BEAM can
/// only watch pids it holds, and can only hold pids that crossed. Keeping the
/// bit rare is what keeps the record alive: anonymous never-exported churn
/// (holder threads, egress tasks) stays ineligible and cannot evict a
/// watchable tenancy's record from a LIFO-recycled slot.
///
/// Generation-checked and live-screened: marking a pid whose tenancy already
/// ended is a no-op — its record either exists (it was flagged before dying)
/// or is honestly unknowable. Same `Runtime::run()` context contract as
/// [`monitor`].
pub fn mark_watchable<A>(target: Pid<A>) {
let target = target.erase();
with_runtime(|inner| {
if let Some(slot) = inner.slot_at(target) {
// Cold lock FIRST: finalize publishes Done and checks the
// watchable bit under this same lock, so the mark either lands
// before finalize reads it (the death stamps) or observes the
// tenancy already dead (no-op). No lost-stamp window between an
// unlocked liveness read and the flag set.
let mut cold = slot.cold.lock();
if slot.is_live_for(target) {
cold.watchable = true;
}
}
});
}
/// The terminal [`DownReason`] of the tenancy `target` names, if that tenancy
/// ever registered a name and is the *most recent named* death of its slot:
/// finalize stamps the slot with `(generation, reason)` for once-registered
/// tenancies (anonymous green-thread churn does not stamp — nor evict), and
/// the record survives reclaim and the next tenant's install, until the next
/// *named* tenant of the slot itself dies. `None` means the pid never lived,
/// is still alive, never held a name, or its record was overwritten by a
/// later named tenancy's death — callers fall back to `NoProc` semantics.
///
/// This exists for watch-installers that raced their target's death (bridge
/// soak signature 4): a `NoProc` observed at install time can be upgraded to
/// the real reason while the record still matches, which is exactly what an
/// install that had won the race would have delivered. It does NOT change
/// [`monitor`]'s own semantics — monitoring a stale pid still queues `NoProc`,
/// the same shape Erlang gives — the upgrade is the caller's deliberate act.
/// Same context contract as [`monitor`]: must run inside `Runtime::run()`.
pub fn terminal_reason<A>(target: Pid<A>) -> Option<DownReason> {
let target = target.erase();
with_runtime(|inner| {
let slot = inner.slot_at(target)?;
let cold = slot.cold.lock();
match cold.terminal {
Some((generation, reason)) if generation == target.generation() => Some(reason),
_ => None,
}
})
}
/// Cancel the monitor `m`. Returns `Some(id)` if a live registration was found /// Cancel the monitor `m`. Returns `Some(id)` if a live registration was found
/// and removed, so no `Down` will arrive on `m.rx` from here on. Returns /// and removed, so no `Down` will arrive on `m.rx` from here on. Returns
/// `None` if there was nothing left to remove: the target had already gone /// `None` if there was nothing left to remove: the target had already gone
+29 -10
View File
@@ -158,7 +158,11 @@ impl TimerTarget for MutexCore {
if st.holder == Some(pid) { if st.holder == Some(pid) {
return; return;
} }
match st.waiters.iter().position(|w| w.pid == pid && w.epoch == epoch) { match st
.waiters
.iter()
.position(|w| w.pid == pid && w.epoch == epoch)
{
Some(pos) => { Some(pos) => {
st.waiters.remove(pos); st.waiters.remove(pos);
true true
@@ -246,7 +250,10 @@ impl<T> Mutex<T> {
Some(v) => v, Some(v) => v,
None => panic!("smarm: Mutex value missing on free fast path (core corrupt)"), None => panic!("smarm: Mutex value missing on free fast path (core corrupt)"),
}; };
return Ok(MutexGuard { mutex: self, value: Some(value) }); return Ok(MutexGuard {
mutex: self,
value: Some(value),
});
} }
} }
@@ -287,7 +294,10 @@ impl<T> Mutex<T> {
Some(v) => v, Some(v) => v,
None => panic!("smarm: Mutex value missing after grant (core corrupt)"), None => panic!("smarm: Mutex value missing after grant (core corrupt)"),
}; };
Ok(MutexGuard { mutex: self, value: Some(value) }) Ok(MutexGuard {
mutex: self,
value: Some(value),
})
} else { } else {
Err(LockTimeout) Err(LockTimeout)
} }
@@ -315,7 +325,10 @@ impl<T> Mutex<T> {
Some(v) => v, Some(v) => v,
None => panic!("smarm: Mutex value missing on try_lock free path (core corrupt)"), None => panic!("smarm: Mutex value missing on try_lock free path (core corrupt)"),
}; };
Some(MutexGuard { mutex: self, value: Some(value) }) Some(MutexGuard {
mutex: self,
value: Some(value),
})
} }
/// Blocking fallback used when called outside the smarm runtime. /// Blocking fallback used when called outside the smarm runtime.
@@ -329,10 +342,15 @@ impl<T> Mutex<T> {
Ok(mut g) => g.take(), Ok(mut g) => g.take(),
Err(e) => panic!("smarm: mutex value lock poisoned (core corrupt): {e}"), Err(e) => panic!("smarm: mutex value lock poisoned (core corrupt): {e}"),
}; };
if let Some(v) = v { break v; } if let Some(v) = v {
break v;
}
std::thread::yield_now(); std::thread::yield_now();
}; };
Ok(MutexGuard { mutex: self, value: Some(value) }) Ok(MutexGuard {
mutex: self,
value: Some(value),
})
} }
} }
@@ -342,7 +360,10 @@ impl<T> Clone for Mutex<T> {
/// lock and one protected value; locking through any clone excludes /// lock and one protected value; locking through any clone excludes
/// every other clone. /// every other clone.
fn clone(&self) -> Self { fn clone(&self) -> Self {
Self { core: self.core.clone(), value: self.value.clone() } Self {
core: self.core.clone(),
value: self.value.clone(),
}
} }
} }
@@ -388,9 +409,7 @@ impl<T: std::fmt::Debug> std::fmt::Debug for MutexGuard<'_, T> {
Some(v) => v, Some(v) => v,
None => panic!("smarm: MutexGuard value missing (core corrupt)"), None => panic!("smarm: MutexGuard value missing (core corrupt)"),
}; };
f.debug_tuple("MutexGuard") f.debug_tuple("MutexGuard").field(value).finish()
.field(value)
.finish()
} }
} }
+1017
View File
File diff suppressed because it is too large Load Diff
+80 -17
View File
@@ -221,7 +221,9 @@ pub(crate) struct ProcessGroups {
impl ProcessGroups { impl ProcessGroups {
pub(crate) fn new() -> Self { pub(crate) fn new() -> Self {
Self { groups: HashMap::new() } Self {
groups: HashMap::new(),
}
} }
/// Insert `ms` into `group`. Idempotent on the *member*: if the member is /// Insert `ms` into `group`. Idempotent on the *member*: if the member is
@@ -323,20 +325,33 @@ impl ProcessGroups {
fn members_where(&self, group: &str, mut is_live: impl FnMut(Pid) -> bool) -> Vec<Pid> { fn members_where(&self, group: &str, mut is_live: impl FnMut(Pid) -> bool) -> Vec<Pid> {
self.groups self.groups
.get(group) .get(group)
.map(|v| v.iter().map(|e| e.member.pid).filter(|&p| is_live(p)).collect()) .map(|v| {
v.iter()
.map(|e| e.member.pid)
.filter(|&p| is_live(p))
.collect()
})
.unwrap_or_default() .unwrap_or_default()
} }
/// The first live member of `group` in insertion order — stateless /// The first live member of `group` in insertion order — stateless
/// first-live `pick`, with the same read-path backstop as `members_where`. /// first-live `pick`, with the same read-path backstop as `members_where`.
fn first_member_where(&self, group: &str, mut is_live: impl FnMut(Pid) -> bool) -> Option<Pid> { fn first_member_where(&self, group: &str, mut is_live: impl FnMut(Pid) -> bool) -> Option<Pid> {
self.groups.get(group)?.iter().map(|e| e.member.pid).find(|&p| is_live(p)) self.groups
.get(group)?
.iter()
.map(|e| e.member.pid)
.find(|&p| is_live(p))
} }
} }
/// Build the full member identity for `pid` from runtime identity. /// Build the full member identity for `pid` from runtime identity.
fn member_for(inner: &crate::runtime::RuntimeInner, pid: Pid) -> Member { fn member_for(inner: &crate::runtime::RuntimeInner, pid: Pid) -> Member {
Member { node: inner.node_id, incarnation: inner.incarnation, pid } Member {
node: inner.node_id,
incarnation: inner.incarnation,
pid,
}
} }
/// Is `pid` a live actor right now? Generation-checked atomic slot-word read, /// Is `pid` a live actor right now? Generation-checked atomic slot-word read,
@@ -367,7 +382,10 @@ pub fn join<A>(group: impl Into<String>, pid: Pid<A>) -> bool {
let mon = monitor(pid); let mon = monitor(pid);
let (rejected, reaped) = with_runtime(|inner| { let (rejected, reaped) = with_runtime(|inner| {
let ms = Membership { member: member_for(inner, pid), monitor: mon }; let ms = Membership {
member: member_for(inner, pid),
monitor: mon,
};
let mut pg = inner.process_groups.lock(); let mut pg = inner.process_groups.lock();
let reaped = pg.reap_group(&group); let reaped = pg.reap_group(&group);
let rejected = pg.join(&group, ms); let rejected = pg.join(&group, ms);
@@ -507,7 +525,11 @@ mod tests {
let (tx, rx) = channel::<Down>(); let (tx, rx) = channel::<Down>();
let ms = Membership { let ms = Membership {
member: member(index, generation), member: member(index, generation),
monitor: Monitor { id: MonitorId(0), target: pid, rx }, monitor: Monitor {
id: MonitorId(0),
target: pid,
rx,
},
}; };
(ms, tx) (ms, tx)
} }
@@ -518,7 +540,10 @@ mod tests {
let (a, _ta) = synth(1, 0); let (a, _ta) = synth(1, 0);
let (b, _tb) = synth(1, 0); let (b, _tb) = synth(1, 0);
assert!(pg.join("workers", a).is_none(), "first join inserts"); assert!(pg.join("workers", a).is_none(), "first join inserts");
assert!(pg.join("workers", b).is_some(), "second identical join is handed back"); assert!(
pg.join("workers", b).is_some(),
"second identical join is handed back"
);
assert_eq!(pg.members_of("workers"), vec![member(1, 0)]); assert_eq!(pg.members_of("workers"), vec![member(1, 0)]);
} }
@@ -542,7 +567,10 @@ mod tests {
let (a, _ta) = synth(1, 0); let (a, _ta) = synth(1, 0);
let (b, _tb) = synth(1, 1); let (b, _tb) = synth(1, 1);
assert!(pg.join("g", a).is_none()); assert!(pg.join("g", a).is_none());
assert!(pg.join("g", b).is_none(), "different generation is a distinct member"); assert!(
pg.join("g", b).is_none(),
"different generation is a distinct member"
);
assert_eq!(pg.members_of("g"), vec![member(1, 0), member(1, 1)]); assert_eq!(pg.members_of("g"), vec![member(1, 0), member(1, 1)]);
} }
@@ -555,16 +583,27 @@ mod tests {
pg.join("g", b); pg.join("g", b);
assert!(pg.leave("g", member(1, 0)).is_some()); assert!(pg.leave("g", member(1, 0)).is_some());
assert_eq!(pg.members_of("g"), vec![member(2, 0)]); assert_eq!(pg.members_of("g"), vec![member(2, 0)]);
assert!(pg.leave("g", member(1, 0)).is_none(), "second leave finds nothing"); assert!(
pg.leave("g", member(1, 0)).is_none(),
"second leave finds nothing"
);
assert!(pg.leave("g", member(2, 0)).is_some()); assert!(pg.leave("g", member(2, 0)).is_some());
assert!(pg.members_of("g").is_empty(), "group is now empty"); assert!(pg.members_of("g").is_empty(), "group is now empty");
assert!(pg.leave("never", member(9, 0)).is_none(), "leaving an unknown group is a no-op"); assert!(
pg.leave("never", member(9, 0)).is_none(),
"leaving an unknown group is a no-op"
);
} }
#[test] #[test]
fn remove_where_sweeps_every_group() { fn remove_where_sweeps_every_group() {
let mut pg = ProcessGroups::new(); let mut pg = ProcessGroups::new();
for (g, (m, _t)) in [("a", synth(1, 0)), ("a", synth(2, 0)), ("b", synth(1, 0)), ("c", synth(3, 0))] { for (g, (m, _t)) in [
("a", synth(1, 0)),
("a", synth(2, 0)),
("b", synth(1, 0)),
("c", synth(3, 0)),
] {
pg.join(g, m); pg.join(g, m);
} }
// Death of pid index 1 (any generation) evicts it everywhere. // Death of pid index 1 (any generation) evicts it everywhere.
@@ -582,8 +621,16 @@ mod tests {
let pid = Pid::new(1, 0); let pid = Pid::new(1, 0);
let (tx, rx) = channel::<Down>(); let (tx, rx) = channel::<Down>();
let dead = Membership { let dead = Membership {
member: Member { node: DEFAULT_NODE_ID, incarnation: Incarnation::new(7), pid }, member: Member {
monitor: Monitor { id: MonitorId(0), target: pid, rx }, node: DEFAULT_NODE_ID,
incarnation: Incarnation::new(7),
pid,
},
monitor: Monitor {
id: MonitorId(0),
target: pid,
rx,
},
}; };
let _keep = tx; let _keep = tx;
let (live, _tl) = synth(2, 0); let (live, _tl) = synth(2, 0);
@@ -614,9 +661,17 @@ mod tests {
pg.join("b", b1); pg.join("b", b1);
// pid 1 dies: its group-a monitor receives a Down. Its group-b monitor // pid 1 dies: its group-a monitor receives a Down. Its group-b monitor
// has not — reap must still sweep pid 1 out of b by the pid predicate. // has not — reap must still sweep pid 1 out of b by the pid predicate.
ta1.send(Down { pid: Pid::new(1, 0), reason: DownReason::Exit }).unwrap(); ta1.send(Down {
pid: Pid::new(1, 0),
reason: DownReason::Exit,
})
.unwrap();
let evicted = pg.reap_group("a"); let evicted = pg.reap_group("a");
assert_eq!(evicted.len(), 2, "pid 1's memberships in both a and b are evicted"); assert_eq!(
evicted.len(),
2,
"pid 1's memberships in both a and b are evicted"
);
assert_eq!(pg.members_of("a"), vec![member(2, 0)]); assert_eq!(pg.members_of("a"), vec![member(2, 0)]);
assert!(pg.members_of("b").is_empty(), "swept from b too; pruned"); assert!(pg.members_of("b").is_empty(), "swept from b too; pruned");
} }
@@ -646,8 +701,16 @@ mod tests {
let dead = Pid::new(1, 0); let dead = Pid::new(1, 0);
let oracle = |pid: Pid| pid != dead; let oracle = |pid: Pid| pid != dead;
assert_eq!(pg.members_where("g", oracle), vec![Pid::new(2, 0)], "dead pid filtered from read"); assert_eq!(
assert_eq!(pg.first_member_where("g", oracle), Some(Pid::new(2, 0)), "pick skips the dead first member"); pg.members_where("g", oracle),
vec![Pid::new(2, 0)],
"dead pid filtered from read"
);
assert_eq!(
pg.first_member_where("g", oracle),
Some(Pid::new(2, 0)),
"pick skips the dead first member"
);
// Backstop does not evict — that stays the monitor's job; raw storage // Backstop does not evict — that stays the monitor's job; raw storage
// still holds both until reap runs. // still holds both until reap runs.
+12 -3
View File
@@ -79,7 +79,10 @@ impl Pid<Erased> {
/// here; typing happens at typed-actor boundaries via [`Pid::from_raw`]. /// here; typing happens at typed-actor boundaries via [`Pid::from_raw`].
#[inline] #[inline]
pub const fn new(index: u32, generation: u32) -> Self { pub const fn new(index: u32, generation: u32) -> Self {
Self { raw: RawPid::new(index, generation), _marker: PhantomData } Self {
raw: RawPid::new(index, generation),
_marker: PhantomData,
}
} }
} }
@@ -90,7 +93,10 @@ impl<A> Pid<A> {
/// resolution paths. /// resolution paths.
#[inline] #[inline]
pub(crate) const fn from_raw(raw: RawPid) -> Self { pub(crate) const fn from_raw(raw: RawPid) -> Self {
Self { raw, _marker: PhantomData } Self {
raw,
_marker: PhantomData,
}
} }
/// The raw identity, dropping the actor type — the key for identity-only /// The raw identity, dropping the actor type — the key for identity-only
@@ -192,7 +198,10 @@ impl<M> Name<M> {
/// associated constants at call sites. /// associated constants at call sites.
#[inline] #[inline]
pub const fn new(name: &'static str) -> Self { pub const fn new(name: &'static str) -> Self {
Self { name, _marker: PhantomData } Self {
name,
_marker: PhantomData,
}
} }
/// The underlying registry key. /// The underlying registry key.
+7 -4
View File
@@ -98,10 +98,13 @@ pub(crate) fn clear_current_slot() {
CURRENT_SLOT.with(|c| c.set(std::ptr::null())); CURRENT_SLOT.with(|c| c.set(std::ptr::null()));
} }
/// RFC 007 (`smarm-causal`) — raw pointer to the on-CPU actor's slot, null on /// Raw pointer to the on-CPU actor's slot, null on the scheduler's own
/// the scheduler's own stack. Same lifetime argument as `note_overrun`: the /// stack. Same lifetime argument as `note_overrun`: the slot is never
/// slot is never reclaimed while its actor is on-CPU. /// reclaimed while its actor is on-CPU. Consumers: the `smarm-causal`
#[cfg(feature = "smarm-causal")] /// profiler (RFC 007) and — unconditionally — the SIGSEGV classifier
/// (RFC 019 §7), which additionally relies on this being a plain load of a
/// const-initialized TLS Cell (no lazy init, no allocation, no dtor): safe
/// from a signal handler.
#[inline] #[inline]
pub(crate) fn current_slot_ptr() -> *const crate::runtime::Slot { pub(crate) fn current_slot_ptr() -> *const crate::runtime::Slot {
CURRENT_SLOT.with(|c| c.get()) CURRENT_SLOT.with(|c| c.get())
+4 -1
View File
@@ -166,7 +166,10 @@ impl<T> RawMutex<T> {
{ {
self.lock_slow(); self.lock_slow();
} }
RawMutexGuard { m: self, prev_preempt } RawMutexGuard {
m: self,
prev_preempt,
}
} }
#[cold] #[cold]
+107 -15
View File
@@ -1,4 +1,3 @@
//! Give an actor a name so other actors can find it and message it. //! Give an actor a name so other actors can find it and message it.
//! //!
//! Without the registry, the only way to reach an actor is to already be //! Without the registry, the only way to reach an actor is to already be
@@ -196,7 +195,9 @@ impl<M> std::fmt::Display for SendError<M> {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
match self { match self {
SendError::Unresolved(_) => write!(f, "no live actor registered under that name"), SendError::Unresolved(_) => write!(f, "no live actor registered under that name"),
SendError::Dead(_) => write!(f, "the addressed actor is no longer the live incarnation"), SendError::Dead(_) => {
write!(f, "the addressed actor is no longer the live incarnation")
}
SendError::NoChannel(_) => write!(f, "actor has no channel for this message type"), SendError::NoChannel(_) => write!(f, "actor has no channel for this message type"),
SendError::Closed(_) => write!(f, "the actor's channel for this type is closed"), SendError::Closed(_) => write!(f, "the actor's channel for this type is closed"),
SendError::NoMember(_) => write!(f, "no live member in the process group"), SendError::NoMember(_) => write!(f, "no live member in the process group"),
@@ -243,7 +244,10 @@ struct Mailbox {
impl Mailbox { impl Mailbox {
fn new(pid: Pid) -> Self { fn new(pid: Pid) -> Self {
Self { pid, channels: HashMap::new() } Self {
pid,
channels: HashMap::new(),
}
} }
/// Clone the `Sender<M>` for this actor, if it has one. Called **under the /// Clone the `Sender<M>` for this actor, if it has one. Called **under the
@@ -294,7 +298,10 @@ pub(crate) struct Registry {
impl Registry { impl Registry {
pub(crate) fn new() -> Self { pub(crate) fn new() -> Self {
Self { by_index: HashMap::new(), by_name: HashMap::new() } Self {
by_index: HashMap::new(),
by_name: HashMap::new(),
}
} }
/// Drop a dead holder's artifacts: every name bound to it, and its /// Drop a dead holder's artifacts: every name bound to it, and its
@@ -303,7 +310,11 @@ impl Registry {
/// wholesale on pid mismatch) and is left untouched. /// wholesale on pid mismatch) and is left untouched.
fn prune_holder(&mut self, holder: Pid) { fn prune_holder(&mut self, holder: Pid) {
self.by_name.retain(|_, p| *p != holder); self.by_name.retain(|_, p| *p != holder);
if self.by_index.get(&holder.index()).is_some_and(|mb| mb.pid == holder) { if self
.by_index
.get(&holder.index())
.is_some_and(|mb| mb.pid == holder)
{
self.by_index.remove(&holder.index()); self.by_index.remove(&holder.index());
} }
} }
@@ -351,7 +362,11 @@ impl Registry {
.iter() .iter()
.filter_map(|(&n, &p)| (p == mb.pid).then_some(n)) .filter_map(|(&n, &p)| (p == mb.pid).then_some(n))
.collect(); .collect();
Some(MailboxInfo { pid: mb.pid, names, depth: depth.min(u32::MAX as usize) as u32 }) Some(MailboxInfo {
pid: mb.pid,
names,
depth: depth.min(u32::MAX as usize) as u32,
})
} }
} }
@@ -388,6 +403,16 @@ pub(crate) fn register_with<M: Send + 'static>(
tx: Sender<M>, tx: Sender<M>,
) -> Result<(), RegisterError> { ) -> Result<(), RegisterError> {
with_runtime(|inner| { with_runtime(|inner| {
// Stamp-eligibility for the terminal record (soak sig 4): flag the
// tenancy BEFORE the binding lands and outside the registry lock (no
// nesting), so no successfully-registered actor can die unflagged.
// A register that then fails leaves a harmless overshoot; a stale
// `me` is screened by the same live() the binding requires below.
if live(inner, me) {
if let Some(slot) = inner.slot_at(me) {
slot.cold.lock().watchable = true;
}
}
let mut reg = inner.registry.lock(); let mut reg = inner.registry.lock();
if !live(inner, me) { if !live(inner, me) {
return Err(RegisterError::NoProc); return Err(RegisterError::NoProc);
@@ -418,13 +443,19 @@ pub(crate) fn register_with<M: Send + 'static>(
/// index from a dead prior incarnation (pid mismatch) is replaced wholesale. /// index from a dead prior incarnation (pid mismatch) is replaced wholesale.
/// Caller holds the registry lock and has established that `me` is live. /// Caller holds the registry lock and has established that `me` is live.
fn publish_channel<M: Send + 'static>(reg: &mut Registry, me: Pid, tx: Sender<M>) { fn publish_channel<M: Send + 'static>(reg: &mut Registry, me: Pid, tx: Sender<M>) {
let mb = reg.by_index.entry(me.index()).or_insert_with(|| Mailbox::new(me)); let mb = reg
.by_index
.entry(me.index())
.or_insert_with(|| Mailbox::new(me));
if mb.pid != me { if mb.pid != me {
*mb = Mailbox::new(me); *mb = Mailbox::new(me);
} }
mb.channels.insert( mb.channels.insert(
TypeId::of::<M>(), TypeId::of::<M>(),
Channel { sender: Box::new(tx), msg_type: type_name::<M>() }, Channel {
sender: Box::new(tx),
msg_type: type_name::<M>(),
},
); );
} }
@@ -463,7 +494,10 @@ pub fn install<A: Addressable>(tx: Sender<A::Msg>) -> Pid<A> {
pub(crate) fn install_for<M: Send + 'static>(pid: Pid, tx: Sender<M>) { pub(crate) fn install_for<M: Send + 'static>(pid: Pid, tx: Sender<M>) {
with_runtime(|inner| { with_runtime(|inner| {
let mut reg = inner.registry.lock(); let mut reg = inner.registry.lock();
debug_assert!(live(inner, pid), "install_for: pid must be a freshly spawned, live actor"); debug_assert!(
live(inner, pid),
"install_for: pid must be a freshly spawned, live actor"
);
publish_channel::<M>(&mut reg, pid, tx); publish_channel::<M>(&mut reg, pid, tx);
}); });
} }
@@ -486,6 +520,47 @@ pub fn whereis(name: &str) -> Option<Pid> {
}) })
} }
/// What a name is bound to, three-valued (bridge soak signature 4).
///
/// [`Live`](NameResolution::Live) is [`whereis`]'s `Some`.
/// [`Corpse`](NameResolution::Corpse) carries the *stored* holder pid of a
/// dead-but-unpruned binding — a state Erlang cannot represent (its name
/// death unregisters atomically; smarm's prune is lazy), captured here before
/// the prune that `whereis` performs discards it, so the caller can consult
/// [`terminal_reason`](crate::monitor::terminal_reason) for the tenancy's
/// real down reason. [`Unbound`](NameResolution::Unbound) matches Erlang's
/// unregistered name.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum NameResolution {
/// The stored holder is live (generation-checked); the binding stands.
Live(Pid),
/// The stored holder is dead. The binding was pruned on the way out —
/// the name heals exactly as `whereis` heals it; only the evidence is
/// returned instead of discarded. A second resolve is `Unbound`.
Corpse(Pid),
/// No binding stored (never registered, or already pruned by any reader).
Unbound,
}
/// Resolve `name` like [`whereis`], but keep the corpse: the dead-holder arm
/// returns the stored pid it pruned instead of a bare `None`. Same lock
/// discipline and pruning behavior as `whereis`; same `Runtime::run()`
/// context contract.
pub fn resolve_name(name: &str) -> NameResolution {
with_runtime(|inner| {
let mut reg = inner.registry.lock();
let Some(&pid) = reg.by_name.get(name) else {
return NameResolution::Unbound;
};
if live(inner, pid) {
NameResolution::Live(pid)
} else {
reg.prune_holder(pid);
NameResolution::Corpse(pid)
}
})
}
/// Like [`whereis`], but returns a *typed* [`Pid<A>`] instead of a bare /// Like [`whereis`], but returns a *typed* [`Pid<A>`] instead of a bare
/// [`Pid`], so a follow-up [`send_to`] is compile-checked instead of needing /// [`Pid`], so a follow-up [`send_to`] is compile-checked instead of needing
/// the untyped [`send_dyn`] escape hatch. `None` if the name is unbound or its /// the untyped [`send_dyn`] escape hatch. `None` if the name is unbound or its
@@ -522,7 +597,10 @@ pub(crate) fn resolve_named_sender<M: Send + 'static>(name: &str) -> Option<(Pid
} }
// A live holder's mailbox is its own (publish replaces wholesale on // A live holder's mailbox is its own (publish replaces wholesale on
// pid mismatch, and one live actor per slot), so index lookup is safe. // pid mismatch, and one live actor per slot), so index lookup is safe.
let tx = reg.by_index.get(&pid.index()).and_then(Mailbox::clone_sender::<M>)?; let tx = reg
.by_index
.get(&pid.index())
.and_then(Mailbox::clone_sender::<M>)?;
Some((pid, tx)) Some((pid, tx))
}) })
} }
@@ -535,7 +613,11 @@ pub fn unregister(name: &str) -> Option<Pid> {
with_runtime(|inner| { with_runtime(|inner| {
let mut reg = inner.registry.lock(); let mut reg = inner.registry.lock();
let pid = reg.by_name.remove(name)?; let pid = reg.by_name.remove(name)?;
if live(inner, pid) { Some(pid) } else { None } if live(inner, pid) {
Some(pid)
} else {
None
}
}) })
} }
@@ -567,12 +649,17 @@ pub fn send<M: Send + 'static>(name: Name<M>, msg: M) -> Result<(), SendError<M>
reg.prune_holder(pid); reg.prune_holder(pid);
return Err(SendError::Unresolved(msg)); return Err(SendError::Unresolved(msg));
} }
match reg.by_index.get(&pid.index()).and_then(Mailbox::clone_sender::<M>) { match reg
.by_index
.get(&pid.index())
.and_then(Mailbox::clone_sender::<M>)
{
Some(tx) => tx, Some(tx) => tx,
None => return Err(SendError::NoChannel(msg)), None => return Err(SendError::NoChannel(msg)),
} }
}; };
tx.send(msg).map_err(|crate::channel::SendError(m)| SendError::Closed(m)) tx.send(msg)
.map_err(|crate::channel::SendError(m)| SendError::Closed(m))
}) })
} }
@@ -596,7 +683,11 @@ fn send_to_pid<M: Send + 'static>(
match reg.by_index.get(&pid.index()).map(|m| m.pid) { match reg.by_index.get(&pid.index()).map(|m| m.pid) {
// Exact incarnation, still alive: its `M` channel, or NoChannel. // Exact incarnation, still alive: its `M` channel, or NoChannel.
Some(stored) if stored == pid && live(inner, pid) => { Some(stored) if stored == pid && live(inner, pid) => {
match reg.by_index.get(&pid.index()).and_then(Mailbox::clone_sender::<M>) { match reg
.by_index
.get(&pid.index())
.and_then(Mailbox::clone_sender::<M>)
{
Some(tx) => tx, Some(tx) => tx,
None => return Err(SendError::NoChannel(msg)), None => return Err(SendError::NoChannel(msg)),
} }
@@ -611,7 +702,8 @@ fn send_to_pid<M: Send + 'static>(
_ => return Err(SendError::Dead(msg)), _ => return Err(SendError::Dead(msg)),
} }
}; };
tx.send(msg).map_err(|crate::channel::SendError(m)| SendError::Closed(m)) tx.send(msg)
.map_err(|crate::channel::SendError(m)| SendError::Closed(m))
} }
/// Deliver `msg` directly to the exact actor identified by `pid`. Unlike /// Deliver `msg` directly to the exact actor identified by `pid`. Unlike
+28 -5
View File
@@ -222,7 +222,10 @@ impl MpmcRing {
if diff == 0 { if diff == 0 {
// Our turn: claim the position. // Our turn: claim the position.
match self.enqueue_pos.0.compare_exchange_weak( match self.enqueue_pos.0.compare_exchange_weak(
pos, pos + 1, Ordering::Relaxed, Ordering::Relaxed, pos,
pos + 1,
Ordering::Relaxed,
Ordering::Relaxed,
) { ) {
Ok(_) => { Ok(_) => {
// SAFETY: the claim gives us exclusive write access // SAFETY: the claim gives us exclusive write access
@@ -250,7 +253,10 @@ impl MpmcRing {
let diff = seq as isize - (pos + 1) as isize; let diff = seq as isize - (pos + 1) as isize;
if diff == 0 { if diff == 0 {
match self.dequeue_pos.0.compare_exchange_weak( match self.dequeue_pos.0.compare_exchange_weak(
pos, pos + 1, Ordering::Relaxed, Ordering::Relaxed, pos,
pos + 1,
Ordering::Relaxed,
Ordering::Relaxed,
) { ) {
Ok(_) => { Ok(_) => {
// SAFETY: the claim gives us exclusive read access; // SAFETY: the claim gives us exclusive read access;
@@ -464,19 +470,36 @@ mod tests {
let popped = popped.lock().unwrap(); let popped = popped.lock().unwrap();
assert_eq!(popped.len(), total, "count mismatch"); assert_eq!(popped.len(), total, "count mismatch");
let set: HashSet<u64> = popped.iter().map(|p| ((p.index() as u64) << 32) | p.generation() as u64).collect(); let set: HashSet<u64> = popped
.iter()
.map(|p| ((p.index() as u64) << 32) | p.generation() as u64)
.collect();
assert_eq!(set.len(), total, "duplicate or lost element"); assert_eq!(set.len(), total, "duplicate or lost element");
assert_eq!(pop(&q), None); assert_eq!(pop(&q), None);
} }
#[test] #[test]
fn mpmc_exactly_once_contended() { fn mpmc_exactly_once_contended() {
exactly_once(MpmcRing::new(8, 4096), |q, p| q.push(p), |q| q.pop(), 4, 4, 1000); exactly_once(
MpmcRing::new(8, 4096),
|q, p| q.push(p),
|q| q.pop(),
4,
4,
1000,
);
} }
#[test] #[test]
fn striped_exactly_once_contended() { fn striped_exactly_once_contended() {
exactly_once(StripedRing::new(8, 4096), |q, p| q.push(p), |q| q.pop(), 4, 4, 1000); exactly_once(
StripedRing::new(8, 4096),
|q, p| q.push(p),
|q| q.pop(),
4,
4,
1000,
);
} }
#[test] #[test]
+724 -270
View File
File diff suppressed because it is too large Load Diff
+336 -72
View File
@@ -68,11 +68,10 @@
use crate::actor::current_pid; use crate::actor::current_pid;
use crate::channel::Sender; use crate::channel::Sender;
use crate::pid::{Name, Pid}; use crate::pid::{Name, Pid};
use crate::runtime::{ use crate::runtime::{self, RuntimeInner, YieldIntent, RUNTIME};
self, RuntimeInner, YieldIntent, RUNTIME,
};
use crate::supervisor::Signal; use crate::supervisor::Signal;
use std::sync::Arc; use std::sync::atomic::Ordering;
use std::sync::{Arc, Weak};
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// with_runtime / try_with_runtime // with_runtime / try_with_runtime
@@ -151,7 +150,9 @@ pub struct JoinHandle {
impl JoinHandle { impl JoinHandle {
/// The identity of the actor this handle refers to. /// The identity of the actor this handle refers to.
pub fn pid(&self) -> Pid { self.pid } pub fn pid(&self) -> Pid {
self.pid
}
/// Block the calling actor until the spawned actor finishes, then /// Block the calling actor until the spawned actor finishes, then
/// report how it finished: `Ok(())` if it returned normally or stopped /// report how it finished: `Ok(())` if it returned normally or stopped
@@ -181,12 +182,10 @@ impl JoinHandle {
crate::slot_state::Status::Stale => { crate::slot_state::Status::Stale => {
panic!("join: target slot has been reused") panic!("join: target slot has been reused")
} }
crate::slot_state::Status::Done => { crate::slot_state::Status::Done => Some(match cold.outcome.take() {
Some(match cold.outcome.take() {
Some(outcome) => outcome, Some(outcome) => outcome,
None => panic!("Done slot must have outcome"), None => panic!("Done slot must have outcome"),
}) }),
}
crate::slot_state::Status::Live => { crate::slot_state::Status::Live => {
// begin_wait is lock-free, legal under the cold lock; // begin_wait is lock-free, legal under the cold lock;
// registering under it makes the epoch atomic with // registering under it makes the epoch atomic with
@@ -226,8 +225,7 @@ impl JoinHandle {
match slot.status_for(self.pid) { match slot.status_for(self.pid) {
crate::slot_state::Status::Stale => false, crate::slot_state::Status::Stale => false,
status => { status => {
cold.outstanding_handles = cold.outstanding_handles = cold.outstanding_handles.saturating_sub(1);
cold.outstanding_handles.saturating_sub(1);
cold.outstanding_handles == 0 cold.outstanding_handles == 0
&& status == crate::slot_state::Status::Done && status == crate::slot_state::Status::Done
} }
@@ -258,6 +256,59 @@ impl Drop for JoinHandle {
// spawn / spawn_under / self_pid // spawn / spawn_under / self_pid
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
/// Per-spawn stack shape overrides (RFC 019). `None` fields resolve to the
/// runtime's [`Config`](crate::runtime::Config) defaults at spawn time, so
/// struct-update syntax works anywhere without a runtime handle:
///
/// ```
/// use smarm::SpawnOpts;
/// let opts = SpawnOpts { stack_reserve: Some(8 * 1024 * 1024), ..SpawnOpts::default() };
/// ```
///
/// Both sizes are page-rounded. The reserve is *virtual* (demand-paged):
/// an 8 MiB reserve costs address space, not memory — RSS follows touched
/// pages. The guard is PROT_NONE below the stack; raise it for FFI code
/// with unusually large C frames. Custom-shaped stacks bypass the recycle
/// pool: they are mmapped fresh at spawn and munmapped at death.
#[derive(Clone, Copy, Debug, Default, PartialEq, Eq)]
pub struct SpawnOpts {
/// Usable stack reservation. `None` ⇒ [`Config::stack_reserve`](crate::runtime::Config::stack_reserve).
pub stack_reserve: Option<usize>,
/// PROT_NONE guard below the stack. `None` ⇒ [`Config::stack_guard`](crate::runtime::Config::stack_guard).
pub guard_size: Option<usize>,
}
/// Why [`try_spawn`] could not start an actor.
///
/// Marked `non_exhaustive`: today the only refusal is a full slab, but a
/// future variant (say, a shutdown-in-progress refusal) must not be a
/// breaking change for shed-path `match`es.
#[non_exhaustive]
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub enum SpawnError {
/// The fixed actor slab ([`Config::max_actors`]
/// (crate::runtime::Config::max_actors)) is full: every slot is claimed
/// by a live actor. This is a routine overload condition, not an
/// invariant violation — shed the unit of work (close the socket,
/// return a 503) and try again once actors have died.
AtCapacity,
}
impl core::fmt::Display for SpawnError {
fn fmt(&self, f: &mut core::fmt::Formatter<'_>) -> core::fmt::Result {
match self {
SpawnError::AtCapacity => {
write!(
f,
"actor slab at capacity (`Config::max_actors` live actors)"
)
}
}
}
}
impl std::error::Error for SpawnError {}
/// Start a new actor running `f`, and return a [`JoinHandle`] for it. /// Start a new actor running `f`, and return a [`JoinHandle`] for it.
/// ///
/// The new actor runs concurrently with its caller and with every other /// The new actor runs concurrently with its caller and with every other
@@ -280,22 +331,32 @@ pub fn spawn(f: impl FnOnce() + Send + 'static) -> JoinHandle {
spawn_under(parent, f) spawn_under(parent, f)
} }
/// [`spawn`] with per-actor stack shape overrides (RFC 019).
pub fn spawn_with(opts: SpawnOpts, f: impl FnOnce() + Send + 'static) -> JoinHandle {
let parent = current_pid().unwrap_or_else(|| with_runtime(|_| crate::runtime::ROOT_PID));
spawn_under_with(parent, opts, f)
}
/// Like [`spawn`], but explicitly attaches the new actor to `supervisor` /// Like [`spawn`], but explicitly attaches the new actor to `supervisor`
/// instead of the calling actor. Ordinary code should reach for [`spawn`]; /// instead of the calling actor. Ordinary code should reach for [`spawn`];
/// this exists for supervision trees (see [`supervisor`](crate::supervisor)) /// this exists for supervision trees (see [`supervisor`](crate::supervisor))
/// and other cases that need to place a child under a specific ancestor /// and other cases that need to place a child under a specific ancestor
/// rather than its true caller. /// rather than its true caller.
pub fn spawn_under<A>(supervisor: Pid<A>, f: impl FnOnce() + Send + 'static) -> JoinHandle { pub fn spawn_under<A>(supervisor: Pid<A>, f: impl FnOnce() + Send + 'static) -> JoinHandle {
let supervisor = supervisor.erase(); spawn_under_with(supervisor, SpawnOpts::default(), f)
// Stack + closure boxing happen before ANY runtime lock is taken: no
// syscall and no allocation ever stalls another scheduler thread.
let stack = with_runtime(|inner| inner.stack_pool.lock().pop())
.unwrap_or_else(|| {
match crate::stack::Stack::new(crate::runtime::ACTOR_STACK_SIZE) {
Ok(stack) => stack,
Err(e) => panic!("stack allocation failed: {e}"),
} }
});
/// [`spawn_under`] with per-actor stack shape overrides (RFC 019).
pub fn spawn_under_with<A>(
supervisor: Pid<A>,
opts: SpawnOpts,
f: impl FnOnce() + Send + 'static,
) -> JoinHandle {
let supervisor = supervisor.erase();
// Stack + closure boxing happen before the slot locks are taken; the
// pool lock inside acquire_stack is dropped before any mmap, so no
// syscall ever stalls another scheduler thread.
let stack = with_runtime(|inner| crate::runtime::acquire_stack(inner, opts));
let sp = init_actor_stack(stack.top(), crate::actor::trampoline); let sp = init_actor_stack(stack.top(), crate::actor::trampoline);
let closure: crate::runtime::Closure = Box::new(f); let closure: crate::runtime::Closure = Box::new(f);
@@ -304,7 +365,77 @@ pub fn spawn_under<A>(supervisor: Pid<A>, f: impl FnOnce() + Send + 'static) ->
crate::runtime::install_actor(inner, idx, sp, stack, supervisor, closure) crate::runtime::install_actor(inner, idx, sp, stack, supervisor, closure)
}); });
JoinHandle { pid, consumed: false } JoinHandle {
pid,
consumed: false,
}
}
/// [`spawn`] that reports a full actor slab instead of panicking.
///
/// Behaviour parity with [`spawn`] in every case except one: when the fixed
/// slab ([`Config::max_actors`](crate::runtime::Config::max_actors)) is
/// full, this returns [`Err(SpawnError::AtCapacity)`](SpawnError::AtCapacity)
/// where `spawn` panics the calling actor. Use it at load-shedding call
/// sites — an accept loop spawning one actor per connection, a request
/// admission point — where "at capacity" is a routine overload condition to
/// handle (reject the unit of work), not an invariant violation. Internal
/// and bounded spawn sites should keep [`spawn`]: there, the panic is a
/// correct loud invariant check.
///
/// The claim is atomic (claim-or-report): there is no
/// check-then-spawn race against other spawners for the last slot, so no
/// headroom margin is needed.
pub fn try_spawn(f: impl FnOnce() + Send + 'static) -> Result<JoinHandle, SpawnError> {
let parent = current_pid().unwrap_or_else(|| with_runtime(|_| crate::runtime::ROOT_PID));
try_spawn_under_with(parent, SpawnOpts::default(), f)
}
/// [`try_spawn`] with an explicit supervisor and per-actor stack shape
/// overrides — the full-control core the other `try_` surface is built on
/// (mirrors [`spawn_under_with`]).
pub fn try_spawn_under_with<A>(
supervisor: Pid<A>,
opts: SpawnOpts,
f: impl FnOnce() + Send + 'static,
) -> Result<JoinHandle, SpawnError> {
let supervisor = supervisor.erase();
// Slot FIRST — deliberately the reverse of `spawn`'s stack-first order:
// under overload the Err arm is the HOT path, and a rejection must cost
// one mutex pop, not an mmap/pool-pop + init + recycle per shed unit of
// work. The claim is a single atomic pop (no TOCTOU; see
// `try_allocate_slot`).
let idx = match with_runtime(|inner| inner.try_allocate_slot()) {
Some(idx) => idx,
None => return Err(SpawnError::AtCapacity),
};
// Between claim and install the slot is owned by this frame alone; if
// stack allocation panics in that window the slot must go back or it
// leaks for the life of the runtime (and would trip the run()-teardown
// slot-leak debug_assert).
struct ReturnOnUnwind(Option<u32>);
impl Drop for ReturnOnUnwind {
fn drop(&mut self) {
if let Some(idx) = self.0 {
with_runtime(|inner| inner.return_vacant_slot(idx));
}
}
}
let mut claimed = ReturnOnUnwind(Some(idx));
let stack = with_runtime(|inner| crate::runtime::acquire_stack(inner, opts));
let sp = init_actor_stack(stack.top(), crate::actor::trampoline);
let closure: crate::runtime::Closure = Box::new(f);
claimed.0 = None; // install_actor takes ownership of the slot from here
let pid = with_runtime(|inner| {
crate::runtime::install_actor(inner, idx, sp, stack, supervisor, closure)
});
Ok(JoinHandle {
pid,
consumed: false,
})
} }
/// Spawn an actor that other actors can message directly by its [`Pid<A>`], /// Spawn an actor that other actors can message directly by its [`Pid<A>`],
@@ -335,6 +466,18 @@ pub fn spawn_addr<A: crate::pid::Addressable>(
crate::pid::assert_type::<A>(pid) crate::pid::assert_type::<A>(pid)
} }
/// [`spawn_addr`] with per-actor stack shape overrides (RFC 019).
pub fn spawn_addr_with<A: crate::pid::Addressable>(
opts: SpawnOpts,
body: impl FnOnce(crate::channel::Receiver<A::Msg>) + Send + 'static,
) -> Pid<A> {
let (tx, rx) = crate::channel::channel::<A::Msg>();
let handle = spawn_with(opts, move || body(rx));
let pid = handle.pid();
crate::registry::install_for::<A::Msg>(pid, tx);
crate::pid::assert_type::<A>(pid)
}
use crate::context::init_actor_stack; use crate::context::init_actor_stack;
/// The identity of the actor currently running. Use it to hand your own /// The identity of the actor currently running. Use it to hand your own
@@ -401,6 +544,29 @@ pub(crate) fn unpark_at(pid: Pid, epoch: u32) {
let _ = try_with_runtime(|inner| inner.unpark_at(pid, epoch)); let _ = try_with_runtime(|inner| inner.unpark_at(pid, epoch));
} }
// The current actor's runtime as a `Weak`, for a waker that must reach the
// runtime from a foreign thread later. A channel captures this when its
// receiver parks, so a cross-thread `send` can wake without the `RUNTIME`
// thread-local (unset off a scheduler thread). Panics outside `Runtime::run()`,
// the same contract as `begin_wait`.
pub(crate) fn runtime_weak() -> Weak<RuntimeInner> {
with_runtime(Arc::downgrade)
}
// Epoch-matched wake of `pid` from a waker that may or may not be on a
// scheduler thread. On a scheduler thread we take the thread-local path
// (preemption-gated, slot-eligible); off one that path is a silent no-op, so
// we reach the runtime through `rt` — the `Weak` the waker captured while it
// was in-runtime. Mirrors the IO backend's cross-context wake (io.rs, RFC 018).
pub(crate) fn unpark_at_via(pid: Pid, epoch: u32, rt: &Weak<RuntimeInner>) {
if try_with_runtime(|inner| inner.unpark_at(pid, epoch)).is_some() {
return;
}
if let Some(inner) = rt.upgrade() {
inner.unpark_at(pid, epoch);
}
}
// Open a new wait for the current actor and return its wait identity // Open a new wait for the current actor and return its wait identity
// ("epoch"). Call once per wait, before registering with any waker. Lock-free, // ("epoch"). Call once per wait, before registering with any waker. Lock-free,
// so it's legal to call while already holding another internal lock. // so it's legal to call while already holding another internal lock.
@@ -434,10 +600,12 @@ pub(crate) fn retire_wait() {
/// [`JoinHandle::join`] reports it as a normal, non-error exit: cooperative /// [`JoinHandle::join`] reports it as a normal, non-error exit: cooperative
/// stop is a controlled shutdown, not a failure. /// stop is a controlled shutdown, not a failure.
/// ///
/// This is exactly the mechanism `gen_server` shutdown, supervisor restarts, /// This is the *hard* stop — OTP's `exit(Pid, kill)`. It is what a supervisor
/// and structured teardown are built from: reach for [`GenServerRef::shutdown`](crate::GenServerRef::shutdown) /// falls back to when a child overstays its [`Shutdown`](crate::supervisor::Shutdown)
/// or a [`supervisor`](crate::supervisor) instead of calling this directly /// grace period. For a stop the target gets to prepare for, use
/// where those apply. /// [`request_shutdown`]; for structured teardown, reach for
/// [`GenServerRef::shutdown`](crate::GenServerRef::shutdown) or a
/// [`supervisor`](crate::supervisor) instead of calling this directly.
/// ///
/// Because it's cooperative, an actor stuck in a tight loop with no /// Because it's cooperative, an actor stuck in a tight loop with no
/// blocking call, no [`check!`](crate::check), and no allocation cannot be /// blocking call, no [`check!`](crate::check), and no allocation cannot be
@@ -449,7 +617,8 @@ pub fn request_stop<A>(pid: Pid<A>) {
} }
// The core of `request_stop`, taking the runtime directly so it can also be // The core of `request_stop`, taking the runtime directly so it can also be
// driven from inside the runtime itself (the root-exit sweep) without // driven from inside the runtime itself (the RuntimeHandle path, supervisor
// sweeps) without
// re-borrowing the thread-local. Sets the stop flag under the target's lock // re-borrowing the thread-local. Sets the stop flag under the target's lock
// (a generation mismatch, or no live actor there, makes it a no-op) and // (a generation mismatch, or no live actor there, makes it a no-op) and
// wakes the target. // wakes the target.
@@ -469,6 +638,68 @@ pub(crate) fn request_stop_inner(inner: &RuntimeInner, pid: Pid) {
} }
} }
/// Ask an actor to shut down gracefully — OTP's `exit(Pid, shutdown)`, where
/// [`request_stop`] is `exit(Pid, kill)`.
///
/// If the target has called [`trap_exit`](crate::trap_exit), it receives an
/// [`ExitSignal`](crate::ExitSignal) with reason
/// [`DownReason::Shutdown`](crate::DownReason::Shutdown) on its trap inbox and
/// keeps running: the request is advisory, and the target is expected to wind
/// down and exit normally in its own time (a supervisor bounds that time with
/// its child's [`Shutdown`](crate::supervisor::Shutdown) policy and falls back
/// to `request_stop`). A target that is not trapping is stopped exactly as by
/// `request_stop`. A dead pid is a no-op.
///
/// The signal's `from` is the calling actor, or `ROOT_PID` when driven from
/// outside the runtime (see [`RuntimeHandle::request_shutdown`](crate::RuntimeHandle::request_shutdown)).
pub fn request_shutdown<A>(pid: Pid<A>) {
let pid = pid.erase();
let from = current_pid().unwrap_or(crate::runtime::ROOT_PID);
let _ = try_with_runtime(|inner| request_shutdown_inner(inner, pid, from));
}
// The core of `request_shutdown`. Reads the target's trap sender under its
// cold lock (generation-verified), then acts outside the lock: a trap send
// may unpark the receiver, and `request_stop_inner` re-takes the lock.
pub(crate) fn request_shutdown_inner(inner: &RuntimeInner, pid: Pid, from: Pid) {
request_shutdown_inner_probe(inner, pid, from);
}
/// [`request_shutdown_inner`], reporting what it found: `Some(true)` if the
/// target was trapping (got the signal), `Some(false)` if it was stopped
/// outright, `None` if there was nothing live at `pid`.
pub(crate) fn request_shutdown_inner_probe(
inner: &RuntimeInner,
pid: Pid,
from: Pid,
) -> Option<bool> {
let trap = match inner.slot_at(pid) {
Some(slot) => {
let cold = slot.cold.lock();
if slot.generation() == pid.generation() {
cold.actor.as_ref().map(|a| a.trap.clone())
} else {
None // stale pid: nothing there to shut down
}
}
None => None,
};
match trap {
Some(Some(tx)) => {
let _ = tx.send(crate::link::ExitSignal {
from,
reason: crate::monitor::DownReason::Shutdown,
});
Some(true)
}
Some(None) => {
request_stop_inner(inner, pid);
Some(false)
}
None => None,
}
}
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// NoPreempt // NoPreempt
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -519,11 +750,9 @@ pub fn sleep(duration: std::time::Duration) {
let _np = NoPreempt::enter(); let _np = NoPreempt::enter();
let epoch = begin_wait(); let epoch = begin_wait();
let deadline = crate::timer::deadline_from_now(duration); let deadline = crate::timer::deadline_from_now(duration);
with_runtime(|inner| { with_runtime(|inner| match inner.timers.lock() {
match inner.timers.lock() {
Ok(mut timers) => timers.insert_sleep(deadline, me, epoch), Ok(mut timers) => timers.insert_sleep(deadline, me, epoch),
Err(e) => panic!("smarm: timers lock poisoned (core corrupt): {e}"), Err(e) => panic!("smarm: timers lock poisoned (core corrupt): {e}"),
}
}); });
park_current(); park_current();
} }
@@ -541,11 +770,9 @@ pub fn sleep_wall(duration: std::time::Duration) {
let _np = NoPreempt::enter(); let _np = NoPreempt::enter();
let epoch = begin_wait(); let epoch = begin_wait();
let deadline = crate::timer::deadline_from_now(duration); let deadline = crate::timer::deadline_from_now(duration);
with_runtime(|inner| { with_runtime(|inner| match inner.timers.lock() {
match inner.timers.lock() {
Ok(mut timers) => timers.insert_sleep_wall(deadline, me, epoch), Ok(mut timers) => timers.insert_sleep_wall(deadline, me, epoch),
Err(e) => panic!("smarm: timers lock poisoned (core corrupt): {e}"), Err(e) => panic!("smarm: timers lock poisoned (core corrupt): {e}"),
}
}); });
park_current(); park_current();
} }
@@ -561,15 +788,13 @@ pub fn insert_wait_timer(
target: std::sync::Arc<dyn crate::timer::TimerTarget>, target: std::sync::Arc<dyn crate::timer::TimerTarget>,
epoch: u32, epoch: u32,
) { ) {
with_runtime(|inner| { with_runtime(|inner| match inner.timers.lock() {
match inner.timers.lock() {
Ok(mut timers) => timers.insert( Ok(mut timers) => timers.insert(
deadline, deadline,
pid, pid,
crate::timer::Reason::WaitTimeout { target, epoch }, crate::timer::Reason::WaitTimeout { target, epoch },
), ),
Err(e) => panic!("smarm: timers lock poisoned (core corrupt): {e}"), Err(e) => panic!("smarm: timers lock poisoned (core corrupt): {e}"),
}
}); });
} }
@@ -599,11 +824,9 @@ pub fn send_after<A: crate::pid::Addressable>(
let fire = Box::new(move || { let fire = Box::new(move || {
let _ = crate::registry::send_to(dest, msg); let _ = crate::registry::send_to(dest, msg);
}); });
with_runtime(|inner| { with_runtime(|inner| match inner.timers.lock() {
match inner.timers.lock() {
Ok(mut timers) => timers.insert_send(deadline, dest.erase(), fire), Ok(mut timers) => timers.insert_send(deadline, dest.erase(), fire),
Err(e) => panic!("smarm: timers lock poisoned (core corrupt): {e}"), Err(e) => panic!("smarm: timers lock poisoned (core corrupt): {e}"),
}
}) })
} }
@@ -623,11 +846,9 @@ pub fn send_after_named<M: Send + 'static>(
let fire = Box::new(move || { let fire = Box::new(move || {
let _ = crate::registry::send(dest, msg); let _ = crate::registry::send(dest, msg);
}); });
with_runtime(|inner| { with_runtime(|inner| match inner.timers.lock() {
match inner.timers.lock() {
Ok(mut timers) => timers.insert_send(deadline, armed_by, fire), Ok(mut timers) => timers.insert_send(deadline, armed_by, fire),
Err(e) => panic!("smarm: timers lock poisoned (core corrupt): {e}"), Err(e) => panic!("smarm: timers lock poisoned (core corrupt): {e}"),
}
}) })
} }
@@ -646,11 +867,9 @@ pub fn send_after_wall<A: crate::pid::Addressable>(
let fire = Box::new(move || { let fire = Box::new(move || {
let _ = crate::registry::send_to(dest, msg); let _ = crate::registry::send_to(dest, msg);
}); });
with_runtime(|inner| { with_runtime(|inner| match inner.timers.lock() {
match inner.timers.lock() {
Ok(mut timers) => timers.insert_send_wall(deadline, dest.erase(), fire), Ok(mut timers) => timers.insert_send_wall(deadline, dest.erase(), fire),
Err(e) => panic!("smarm: timers lock poisoned (core corrupt): {e}"), Err(e) => panic!("smarm: timers lock poisoned (core corrupt): {e}"),
}
}) })
} }
@@ -667,11 +886,9 @@ pub fn send_after_named_wall<M: Send + 'static>(
let fire = Box::new(move || { let fire = Box::new(move || {
let _ = crate::registry::send(dest, msg); let _ = crate::registry::send(dest, msg);
}); });
with_runtime(|inner| { with_runtime(|inner| match inner.timers.lock() {
match inner.timers.lock() {
Ok(mut timers) => timers.insert_send_wall(deadline, armed_by, fire), Ok(mut timers) => timers.insert_send_wall(deadline, armed_by, fire),
Err(e) => panic!("smarm: timers lock poisoned (core corrupt): {e}"), Err(e) => panic!("smarm: timers lock poisoned (core corrupt): {e}"),
}
}) })
} }
@@ -691,11 +908,9 @@ pub(crate) fn send_after_to<T: Send + 'static>(
let fire = Box::new(move || { let fire = Box::new(move || {
let _ = tx.send(msg); let _ = tx.send(msg);
}); });
with_runtime(|inner| { with_runtime(|inner| match inner.timers.lock() {
match inner.timers.lock() {
Ok(mut timers) => timers.insert_send(deadline, armed_by, fire), Ok(mut timers) => timers.insert_send(deadline, armed_by, fire),
Err(e) => panic!("smarm: timers lock poisoned (core corrupt): {e}"), Err(e) => panic!("smarm: timers lock poisoned (core corrupt): {e}"),
}
}) })
} }
@@ -703,11 +918,9 @@ pub(crate) fn send_after_to<T: Send + 'static>(
/// it fires. Returns `true` if the timer was still pending and delivery is /// it fires. Returns `true` if the timer was still pending and delivery is
/// now prevented, `false` if it had already fired or was already cancelled. /// now prevented, `false` if it had already fired or was already cancelled.
pub fn cancel_timer(id: crate::timer::TimerId) -> bool { pub fn cancel_timer(id: crate::timer::TimerId) -> bool {
with_runtime(|inner| { with_runtime(|inner| match inner.timers.lock() {
match inner.timers.lock() {
Ok(mut timers) => timers.cancel(id), Ok(mut timers) => timers.cancel(id),
Err(e) => panic!("smarm: timers lock poisoned (core corrupt): {e}"), Err(e) => panic!("smarm: timers lock poisoned (core corrupt): {e}"),
}
}) })
} }
@@ -753,7 +966,14 @@ where
Err(e) => panic!("smarm: io lock poisoned (core corrupt): {e}"), Err(e) => panic!("smarm: io lock poisoned (core corrupt): {e}"),
}; };
match io.as_mut() { match io.as_mut() {
Some(io) => io.submit(me, epoch, work), Some(io) => {
// RFC 018: count the op in flight BEFORE submit — the
// pool decrements on completion, and an increment that
// trailed the completion would underflow. Under the io
// lock, so ordered against the same-lock submit.
inner.io_outstanding.fetch_add(1, Ordering::AcqRel);
io.submit(me, epoch, work);
}
None => panic!("io thread not started"), None => panic!("io thread not started"),
} }
}); });
@@ -766,7 +986,8 @@ where
}; };
let mut cold = slot.cold.lock(); let mut cold = slot.cold.lock();
debug_assert_eq!( debug_assert_eq!(
slot.generation(), me.generation(), slot.generation(),
me.generation(),
"block_on_io: own slot reused mid-park" "block_on_io: own slot reused mid-park"
); );
match cold.pending_io_result.take() { match cold.pending_io_result.take() {
@@ -813,7 +1034,17 @@ fn wait_fd(fd: std::os::fd::RawFd, readable: bool, writable: bool) -> std::io::R
Err(e) => panic!("smarm: io lock poisoned (core corrupt): {e}"), Err(e) => panic!("smarm: io lock poisoned (core corrupt): {e}"),
}; };
match io.as_mut() { match io.as_mut() {
Some(io) => io.epoll_register(fd, me, epoch, readable, writable), Some(io) => {
// RFC 018: count the waiter BEFORE the ADD (mirror of
// submit); roll back if the registration fails so a
// rejected wait leaves the verdict counters clean.
inner.io_fd_waiters.fetch_add(1, Ordering::AcqRel);
let r = io.epoll_register(fd, me, epoch, readable, writable);
if r.is_err() {
inner.io_fd_waiters.fetch_sub(1, Ordering::AcqRel);
}
r
}
None => panic!("io thread not started"), None => panic!("io thread not started"),
} }
})?; })?;
@@ -838,9 +1069,12 @@ fn wait_fd(fd: std::os::fd::RawFd, readable: bool, writable: bool) -> std::io::R
Err(e) => panic!("smarm: io lock poisoned (core corrupt): {e}"), Err(e) => panic!("smarm: io lock poisoned (core corrupt): {e}"),
}; };
if let Some(io) = io.as_mut() { if let Some(io) = io.as_mut() {
if io.waiters.get(&self.fd) == Some(&(self.me, self.epoch)) { // `cancel_waiter` removes + DELs iff still ours, all
io.waiters.remove(&self.fd); // under the waiters lock (the ADD/DEL serialization);
io.epoll_deregister(self.fd); // decrement only when we actually removed it — a
// FdReady that consumed it already did the decrement.
if io.cancel_waiter(self.fd, self.me, self.epoch) {
inner.io_fd_waiters.fetch_sub(1, Ordering::AcqRel);
} }
} }
}); });
@@ -879,12 +1113,20 @@ pub struct FdArm {
impl FdArm { impl FdArm {
/// An arm that becomes ready when `fd` is readable. /// An arm that becomes ready when `fd` is readable.
pub fn readable(fd: std::os::fd::RawFd) -> Self { pub fn readable(fd: std::os::fd::RawFd) -> Self {
FdArm { fd, readable: true, writable: false } FdArm {
fd,
readable: true,
writable: false,
}
} }
/// An arm that becomes ready when `fd` is writable. /// An arm that becomes ready when `fd` is writable.
pub fn writable(fd: std::os::fd::RawFd) -> Self { pub fn writable(fd: std::os::fd::RawFd) -> Self {
FdArm { fd, readable: false, writable: true } FdArm {
fd,
readable: false,
writable: true,
}
} }
} }
@@ -908,7 +1150,12 @@ impl crate::channel::Selectable for FdArm {
}; };
match io.as_mut() { match io.as_mut() {
Some(io) => { Some(io) => {
io.epoll_register(self.fd, pid, epoch, self.readable, self.writable) inner.io_fd_waiters.fetch_add(1, Ordering::AcqRel);
let r = io.epoll_register(self.fd, pid, epoch, self.readable, self.writable);
if r.is_err() {
inner.io_fd_waiters.fetch_sub(1, Ordering::AcqRel);
}
r
} }
None => panic!("io thread not started"), None => panic!("io thread not started"),
} }
@@ -936,9 +1183,8 @@ impl crate::channel::Selectable for FdArm {
Err(e) => panic!("smarm: io lock poisoned (core corrupt): {e}"), Err(e) => panic!("smarm: io lock poisoned (core corrupt): {e}"),
}; };
if let Some(io) = io.as_mut() { if let Some(io) = io.as_mut() {
if io.waiters.get(&self.fd) == Some(&(pid, epoch)) { if io.cancel_waiter(self.fd, pid, epoch) {
io.waiters.remove(&self.fd); inner.io_fd_waiters.fetch_sub(1, Ordering::AcqRel);
io.epoll_deregister(self.fd);
} }
} }
}); });
@@ -960,7 +1206,11 @@ fn poll_events(fd: std::os::fd::RawFd, readable: bool, writable: bool) -> std::i
if writable { if writable {
events |= libc::POLLOUT; events |= libc::POLLOUT;
} }
let mut pfd = libc::pollfd { fd, events, revents: 0 }; let mut pfd = libc::pollfd {
fd,
events,
revents: 0,
};
loop { loop {
let r = unsafe { libc::poll(&mut pfd, 1, 0) }; let r = unsafe { libc::poll(&mut pfd, 1, 0) };
if r < 0 { if r < 0 {
@@ -1008,7 +1258,11 @@ pub fn wait_writable_timeout(
pub fn read(fd: std::os::fd::RawFd, buf: &mut [u8]) -> std::io::Result<usize> { pub fn read(fd: std::os::fd::RawFd, buf: &mut [u8]) -> std::io::Result<usize> {
wait_readable(fd)?; wait_readable(fd)?;
let n = unsafe { libc::read(fd, buf.as_mut_ptr() as *mut _, buf.len()) }; let n = unsafe { libc::read(fd, buf.as_mut_ptr() as *mut _, buf.len()) };
if n < 0 { Err(std::io::Error::last_os_error()) } else { Ok(n as usize) } if n < 0 {
Err(std::io::Error::last_os_error())
} else {
Ok(n as usize)
}
} }
/// Convenience wrapper: park until `fd` is writable, then perform the /// Convenience wrapper: park until `fd` is writable, then perform the
@@ -1017,7 +1271,11 @@ pub fn read(fd: std::os::fd::RawFd, buf: &mut [u8]) -> std::io::Result<usize> {
pub fn write(fd: std::os::fd::RawFd, buf: &[u8]) -> std::io::Result<usize> { pub fn write(fd: std::os::fd::RawFd, buf: &[u8]) -> std::io::Result<usize> {
wait_writable(fd)?; wait_writable(fd)?;
let n = unsafe { libc::write(fd, buf.as_ptr() as *const _, buf.len()) }; let n = unsafe { libc::write(fd, buf.as_ptr() as *const _, buf.len()) };
if n < 0 { Err(std::io::Error::last_os_error()) } else { Ok(n as usize) } if n < 0 {
Err(std::io::Error::last_os_error())
} else {
Ok(n as usize)
}
} }
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -1026,12 +1284,15 @@ pub fn write(fd: std::os::fd::RawFd, buf: &[u8]) -> std::io::Result<usize> {
pub fn register_supervisor_channel(pid: Pid, sender: Sender<Signal>) { pub fn register_supervisor_channel(pid: Pid, sender: Sender<Signal>) {
with_runtime(|inner| { with_runtime(|inner| {
let slot = inner.slot_at(pid) let slot = inner
.slot_at(pid)
.unwrap_or_else(|| panic!("register_supervisor_channel: pid {:?} not found", pid)); .unwrap_or_else(|| panic!("register_supervisor_channel: pid {:?} not found", pid));
let mut cold = slot.cold.lock(); let mut cold = slot.cold.lock();
assert_eq!( assert_eq!(
slot.generation(), pid.generation(), slot.generation(),
"register_supervisor_channel: pid {:?} not found", pid pid.generation(),
"register_supervisor_channel: pid {:?} not found",
pid
); );
cold.supervisor_channel = Some(sender); cold.supervisor_channel = Some(sender);
}); });
@@ -1104,6 +1365,9 @@ mod send_after_to_tests {
crate::sleep(Duration::from_millis(30)); crate::sleep(Duration::from_millis(30));
r2.store(true, Ordering::SeqCst); r2.store(true, Ordering::SeqCst);
}); });
assert!(reached.load(Ordering::SeqCst), "runtime survived the dead-channel fire"); assert!(
reached.load(Ordering::SeqCst),
"runtime survived the dead-channel fire"
);
} }
} }
+330
View File
@@ -0,0 +1,330 @@
//! RFC 019 §7 — overflow diagnostics.
//!
//! One process-global SIGSEGV handler, installed once at [`crate::runtime::init`]
//! (before any scheduler thread exists, so the PRIOR save is unracing), plus a
//! per-scheduler-thread `sigaltstack` registered at `schedule_loop` entry — a
//! guard hit means the faulting stack has no room to run anything, so the
//! altstack is not optional.
//!
//! The handler classifies `si_addr` against the *current* actor only, reached
//! through `preempt::CURRENT_SLOT` — a const-initialized `Cell<*const Slot>`
//! whose access is a plain TLS load (no lazy init, no allocation, no dtor
//! registration), and which every scheduler thread has materialized before an
//! actor can run on it. The slot's diag atomics (`diag_stack_top` & co) are
//! written in `install_actor` before the Release publish and are only consulted
//! here while the actor is on-CPU, so they cannot be stale.
//!
//! Two classification tiers:
//! - **In-guard**: definitive. Rust frames probe pages in order
//! (`__rust_probestack`), so Rust overflow always lands here; so does any C
//! built with `-fstack-clash-protection` (distro-packaged libraries), and —
//! with the 1 MiB default guard — nearly every unprobed frame too.
//! - **Overshoot**: within [`OVERSHOOT_SLOP`] *below* the guard. An unprobed
//! frame (cargo-built C via `cc` almost never enables clash protection)
//! large enough to step over the guard in one `sub rsp`. Attribution is
//! "probable": the address is in unmapped VA that nothing else owns, an
//! actor was on-CPU, and the distance fits a frame — the diagnostic says so.
//!
//! Classified faults print one line (async-signal-safe: stack buffer +
//! `write(2)`, no fmt, no alloc, no locks) and re-raise with default
//! disposition — no unwind, no resume, no fail-soft (jarred; UB-adjacent from
//! a handler). Unclassified faults reinstate the PRIOR handler and refault, so
//! std's own "thread ... has overflowed its stack" diagnostics for OS-thread
//! stacks survive our presence. Reinstating deregisters us for good, which is
//! fine: the process is dying either way.
use std::cell::Cell;
use std::mem::MaybeUninit;
use std::sync::atomic::Ordering;
use std::sync::Once;
/// Tier-2 window below the guard. Matches the guard default (and the kernel's
/// `stack_guard_gap`): a frame that out-jumps both the guard and this window
/// in one displacement is past what a diagnostic can honestly attribute.
pub(crate) const OVERSHOOT_SLOP: usize = 1024 * 1024;
/// Per-scheduler-thread signal stack. MINSIGSTKSZ is ~11 KiB on AVX-512
/// hardware; 64 KiB leaves the formatter room without mattering to anyone.
/// One per OS thread, never freed: scheduler threads live for the process in
/// practice, and repeated `run()`s on reused threads re-use the registration
/// (the TLS flag), so the leak is bounded by the OS thread count.
const ALTSTACK_SIZE: usize = 64 * 1024;
static INSTALL: Once = Once::new();
/// The handler that was installed before ours (std's, typically). Written
/// exactly once inside INSTALL — which completes in `runtime::init` before
/// any scheduler thread (and thus any classifiable fault) can exist — and
/// only read from the handler afterwards.
static mut PRIOR: MaybeUninit<libc::sigaction> = MaybeUninit::uninit();
thread_local! {
/// Whether this OS thread has registered its altstack.
static ALTSTACK_SET: Cell<bool> = const { Cell::new(false) };
}
/// Where a fault landed relative to the current actor's stack.
#[derive(Debug, PartialEq, Eq)]
pub(crate) enum FaultClass {
/// Inside `[top − reserve − guard, top − reserve)`: the guard region.
Guard,
/// Within `OVERSHOOT_SLOP` below the guard: stepped over it. Payload is
/// the distance below `guard_lo`.
Overshoot(usize),
/// Not ours to explain.
Foreign,
}
/// Pure classifier — all edges unit-tested below. `top` is the stack's usable
/// top, `reserve`/`guard` its shape; both page-rounded by `Stack::new`.
pub(crate) fn classify(addr: usize, top: usize, reserve: usize, guard: usize) -> FaultClass {
let guard_hi = top.wrapping_sub(reserve);
let guard_lo = guard_hi.wrapping_sub(guard);
if addr >= guard_lo && addr < guard_hi {
FaultClass::Guard
} else if addr < guard_lo && addr >= guard_lo.saturating_sub(OVERSHOOT_SLOP) {
FaultClass::Overshoot(guard_lo - addr)
} else {
FaultClass::Foreign
}
}
/// Install the process-global handler. Idempotent; called from
/// `runtime::init`.
pub(crate) fn install_once() {
INSTALL.call_once(|| unsafe {
let mut sa: libc::sigaction = std::mem::zeroed();
sa.sa_sigaction = handler as *const () as usize;
sa.sa_flags = libc::SA_SIGINFO | libc::SA_ONSTACK;
libc::sigemptyset(&mut sa.sa_mask);
let prior = &mut *std::ptr::addr_of_mut!(PRIOR);
libc::sigaction(libc::SIGSEGV, &sa, prior.as_mut_ptr());
});
}
/// Register this OS thread's altstack (idempotent per thread). Called at
/// `schedule_loop` entry, so every thread that can run an actor has one.
pub(crate) fn register_altstack() {
ALTSTACK_SET.with(|set| {
if set.get() {
return;
}
unsafe {
let sp = libc::mmap(
std::ptr::null_mut(),
ALTSTACK_SIZE,
libc::PROT_READ | libc::PROT_WRITE,
libc::MAP_PRIVATE | libc::MAP_ANONYMOUS,
-1,
0,
);
if sp == libc::MAP_FAILED {
// Degrade: no altstack means a guard hit dies without the
// message (handler can't run) — the pre-RFC behavior, never
// incorrectness.
return;
}
let ss = libc::stack_t {
ss_sp: sp,
ss_flags: 0,
ss_size: ALTSTACK_SIZE,
};
libc::sigaltstack(&ss, std::ptr::null_mut());
}
set.set(true);
});
}
// ---------------------------------------------------------------------------
// The handler
// ---------------------------------------------------------------------------
unsafe extern "C" fn handler(
_sig: libc::c_int,
info: *mut libc::siginfo_t,
_ctx: *mut libc::c_void,
) {
let slot_ptr = crate::preempt::current_slot_ptr();
if !slot_ptr.is_null() {
let slot = &*slot_ptr;
let top = slot.diag_stack_top.load(Ordering::Relaxed);
if top != 0 {
let reserve = slot.diag_stack_reserve.load(Ordering::Relaxed);
let guard = slot.diag_stack_guard.load(Ordering::Relaxed);
let pid = slot.diag_pid.load(Ordering::Relaxed);
let addr = (*info).si_addr() as usize;
match classify(addr, top, reserve, guard) {
FaultClass::Guard => {
let mut b = Buf::new();
b.s("smarm: actor ");
b.pid(pid);
b.s(" overflowed its stack: fault in the guard region, depth-at-fault=");
b.u(top - addr);
b.s(" bytes (reserve=");
b.u(reserve);
b.s(", guard=");
b.u(guard);
b.s("). Raise stack_reserve (SpawnOpts or Config).\n");
b.emit();
die_by_default();
return;
}
FaultClass::Overshoot(below) => {
let mut b = Buf::new();
b.s("smarm: actor ");
b.pid(pid);
b.s(" probably overflowed its stack: fault ");
b.u(below);
b.s(" bytes below the guard - an unprobed (FFI?) frame stepped over it (reserve=");
b.u(reserve);
b.s(", guard=");
b.u(guard);
b.s("). Raise stack_guard or stack_reserve.\n");
b.emit();
die_by_default();
return;
}
FaultClass::Foreign => {}
}
}
}
// Not ours: put back whoever was there before us and refault into them.
let prior = &*std::ptr::addr_of!(PRIOR);
libc::sigaction(libc::SIGSEGV, prior.as_ptr(), std::ptr::null_mut());
}
/// Reset SIGSEGV to default disposition; returning from the handler then
/// refaults at the same instruction and the process dies the normal death
/// (core-dumpable, correct wait status), exactly as if we were never here —
/// but with the message already on stderr.
unsafe fn die_by_default() {
let mut dfl: libc::sigaction = std::mem::zeroed();
dfl.sa_sigaction = libc::SIG_DFL;
libc::sigemptyset(&mut dfl.sa_mask);
libc::sigaction(libc::SIGSEGV, &dfl, std::ptr::null_mut());
}
// ---------------------------------------------------------------------------
// Async-signal-safe formatting: fixed buffer, decimal itoa, one write(2).
// ---------------------------------------------------------------------------
struct Buf {
b: [u8; 320],
len: usize,
}
impl Buf {
fn new() -> Self {
Buf {
b: [0; 320],
len: 0,
}
}
fn s(&mut self, s: &str) {
for &c in s.as_bytes() {
if self.len < self.b.len() {
self.b[self.len] = c;
self.len += 1;
}
}
}
fn u(&mut self, mut n: usize) {
let mut tmp = [0u8; 20];
let mut i = tmp.len();
loop {
i -= 1;
tmp[i] = b'0' + (n % 10) as u8;
n /= 10;
if n == 0 {
break;
}
}
for &c in &tmp[i..] {
if self.len < self.b.len() {
self.b[self.len] = c;
self.len += 1;
}
}
}
/// `idx.gen`, unpacked from the install-time packing.
fn pid(&mut self, packed: u64) {
self.u((packed >> 32) as usize);
self.s(".");
self.u((packed & 0xffff_ffff) as usize);
}
fn emit(&self) {
unsafe {
libc::write(2, self.b.as_ptr() as *const libc::c_void, self.len);
}
}
}
// ---------------------------------------------------------------------------
// Classifier units — the arithmetic edges, before anything integrates.
// ---------------------------------------------------------------------------
#[cfg(test)]
mod tests {
use super::{classify, FaultClass, OVERSHOOT_SLOP};
const PG: usize = 4096;
// A synthetic stack far from address-space edges: top at 1 GiB.
const TOP: usize = 1 << 30;
const RESERVE: usize = 16 * PG;
const GUARD: usize = 4 * PG;
const GUARD_HI: usize = TOP - RESERVE;
const GUARD_LO: usize = GUARD_HI - GUARD;
#[test]
fn inside_guard_both_edges() {
assert_eq!(classify(GUARD_LO, TOP, RESERVE, GUARD), FaultClass::Guard);
assert_eq!(
classify(GUARD_HI - 1, TOP, RESERVE, GUARD),
FaultClass::Guard
);
assert_eq!(
classify(GUARD_LO + GUARD / 2, TOP, RESERVE, GUARD),
FaultClass::Guard
);
}
#[test]
fn usable_region_is_foreign() {
// A fault inside the RW stack itself isn't a guard hit and must not
// be explained as one.
assert_eq!(classify(GUARD_HI, TOP, RESERVE, GUARD), FaultClass::Foreign);
assert_eq!(classify(TOP - 1, TOP, RESERVE, GUARD), FaultClass::Foreign);
}
#[test]
fn above_top_is_foreign() {
assert_eq!(classify(TOP, TOP, RESERVE, GUARD), FaultClass::Foreign);
assert_eq!(classify(TOP + PG, TOP, RESERVE, GUARD), FaultClass::Foreign);
}
#[test]
fn overshoot_window_edges() {
assert_eq!(
classify(GUARD_LO - 1, TOP, RESERVE, GUARD),
FaultClass::Overshoot(1)
);
assert_eq!(
classify(GUARD_LO - OVERSHOOT_SLOP, TOP, RESERVE, GUARD),
FaultClass::Overshoot(OVERSHOOT_SLOP)
);
assert_eq!(
classify(GUARD_LO - OVERSHOOT_SLOP - 1, TOP, RESERVE, GUARD),
FaultClass::Foreign
);
}
#[test]
fn low_address_stack_saturates_not_wraps() {
// A stack mapped so low that the slop window would underflow: the
// window clips to 0 instead of wrapping around the address space.
let top = RESERVE + GUARD + PG; // guard_lo == PG
assert_eq!(classify(0, top, RESERVE, GUARD), FaultClass::Overshoot(PG));
// Null-page fault still classified only because it IS within slop
// here; with a normal-height stack it is Foreign (covered above by
// the window-edge test at realistic addresses).
}
}
+9 -9
View File
@@ -188,8 +188,7 @@ impl StateWord {
loop { loop {
let w = self.load(); let w = self.load();
debug_assert!( debug_assert!(
matches!(word_state(w), ST_RUNNING | ST_RUNNING_NOTIFIED) matches!(word_state(w), ST_RUNNING | ST_RUNNING_NOTIFIED) && word_gen(w) == gen,
&& word_gen(w) == gen,
"yield return from invalid word {w:#x}" "yield return from invalid word {w:#x}"
); );
if self if self
@@ -247,8 +246,7 @@ impl StateWord {
loop { loop {
let w = self.load(); let w = self.load();
debug_assert!( debug_assert!(
matches!(word_state(w), ST_RUNNING | ST_RUNNING_NOTIFIED) matches!(word_state(w), ST_RUNNING | ST_RUNNING_NOTIFIED) && word_gen(w) == gen,
&& word_gen(w) == gen,
"begin_wait from invalid word {w:#x}" "begin_wait from invalid word {w:#x}"
); );
let next = word_epoch(w).wrapping_add(1) & EPOCH_MASK; let next = word_epoch(w).wrapping_add(1) & EPOCH_MASK;
@@ -342,8 +340,7 @@ impl StateWord {
loop { loop {
let w = self.load(); let w = self.load();
debug_assert!( debug_assert!(
matches!(word_state(w), ST_RUNNING | ST_RUNNING_NOTIFIED) matches!(word_state(w), ST_RUNNING | ST_RUNNING_NOTIFIED) && word_gen(w) == gen,
&& word_gen(w) == gen,
"clear_notify from invalid word {w:#x}" "clear_notify from invalid word {w:#x}"
); );
if word_state(w) != ST_RUNNING_NOTIFIED { if word_state(w) != ST_RUNNING_NOTIFIED {
@@ -372,8 +369,7 @@ impl StateWord {
pub(crate) fn set_done(&self, gen: u32) { pub(crate) fn set_done(&self, gen: u32) {
let prev = self.0.swap(pack(gen, 0, ST_DONE), Ordering::AcqRel); let prev = self.0.swap(pack(gen, 0, ST_DONE), Ordering::AcqRel);
debug_assert!( debug_assert!(
matches!(word_state(prev), ST_RUNNING | ST_RUNNING_NOTIFIED) matches!(word_state(prev), ST_RUNNING | ST_RUNNING_NOTIFIED) && word_gen(prev) == gen,
&& word_gen(prev) == gen,
"finalize from invalid word {prev:#x}" "finalize from invalid word {prev:#x}"
); );
} }
@@ -538,7 +534,11 @@ mod loom_tests {
// not a pending notification. // not a pending notification.
assert!(word.try_claim(0)); assert!(word.try_claim(0));
assert_eq!(word.unpark(0, Some(epoch)), Unpark::Noop); assert_eq!(word.unpark(0, Some(epoch)), Unpark::Noop);
assert_eq!(word_state(word.load()), ST_RUNNING, "stale epoch notified a live run"); assert_eq!(
word_state(word.load()),
ST_RUNNING,
"stale epoch notified a live run"
);
}); });
} }
+236 -16
View File
@@ -1,32 +1,45 @@
//! mmap-based growable stack with a guard page below. //! mmap-based actor stack with a PROT_NONE guard region below (RFC 019).
//! //!
//! Layout (low → high address): //! Layout (low → high address):
//! [ guard page (PROT_NONE) | stack region ] //! [ guard region (PROT_NONE) | stack region ]
//! ^ top() — initial stack pointer //! ^ top() — initial stack pointer
//! //!
//! Stacks grow downward. Overflow lands in the guard page → SIGSEGV. //! Stacks grow downward. Overflow lands in the guard region → SIGSEGV.
//!
//! Both the usable reserve and the guard are caller-chosen (page-rounded).
//! The reserve is a *virtual* reservation: anonymous mmap is demand-paged,
//! so RSS is touched-pages, not reserve × actors. The guard costs address
//! space only. A wide guard (the runtime defaults to 64 KiB) exists for
//! unprobed FFI frames: Rust frames touch pages in order (probestack), so
//! one page catches Rust overflow, but a C frame with a large local can
//! step over a single page in one `sub rsp`.
use std::io; use std::io;
pub struct Stack { pub struct Stack {
/// Bottom of the entire mmap'd region (start of guard page). /// Bottom of the entire mmap'd region (start of the guard).
base: *mut u8, base: *mut u8,
/// Total mmap'd size: guard_size + stack_size. /// Total mmap'd size: guard_size + stack_size.
total_size: usize, total_size: usize,
/// Usable stack size (excluding guard page). /// Usable stack size (excluding the guard).
stack_size: usize, stack_size: usize,
/// PROT_NONE region below the usable stack.
guard_size: usize,
} }
// Stack owns its memory; safe to send across threads. // Stack owns its memory; safe to send across threads.
unsafe impl Send for Stack {} unsafe impl Send for Stack {}
impl Stack { impl Stack {
/// Allocate a new stack. `stack_size` is the usable region; one page is /// Allocate a new stack. `stack_size` is the usable region; `guard_size`
/// added below as a guard page. Both are rounded up to the page size. /// is mapped PROT_NONE below it. Both are rounded up to the page size
pub fn new(stack_size: usize) -> io::Result<Self> { /// and must be non-zero.
pub fn new(stack_size: usize, guard_size: usize) -> io::Result<Self> {
assert!(stack_size > 0, "stack_size must be non-zero");
assert!(guard_size > 0, "guard_size must be non-zero");
let page = page_size(); let page = page_size();
let stack_size = round_up(stack_size, page); let stack_size = round_up(stack_size, page);
let guard_size = page; let guard_size = round_up(guard_size, page);
let total_size = guard_size + stack_size; let total_size = guard_size + stack_size;
let base = unsafe { let base = unsafe {
@@ -44,16 +57,19 @@ impl Stack {
} }
let base = base as *mut u8; let base = base as *mut u8;
let ret = unsafe { let ret = unsafe { libc::mprotect(base as *mut libc::c_void, guard_size, libc::PROT_NONE) };
libc::mprotect(base as *mut libc::c_void, guard_size, libc::PROT_NONE)
};
if ret != 0 { if ret != 0 {
let err = io::Error::last_os_error(); let err = io::Error::last_os_error();
unsafe { libc::munmap(base as *mut libc::c_void, total_size) }; unsafe { libc::munmap(base as *mut libc::c_void, total_size) };
return Err(err); return Err(err);
} }
Ok(Self { base, total_size, stack_size }) Ok(Self {
base,
total_size,
stack_size,
guard_size,
})
} }
/// 16-byte-aligned top of the usable region. /// 16-byte-aligned top of the usable region.
@@ -62,14 +78,54 @@ impl Stack {
(raw_top & !15) as *mut u8 (raw_top & !15) as *mut u8
} }
/// Pointer to the bottom of the usable region (just above the guard page). /// Pointer to the bottom of the usable region (just above the guard).
pub fn usable_base(&self) -> *mut u8 { pub fn usable_base(&self) -> *mut u8 {
unsafe { self.base.add(page_size()) } unsafe { self.base.add(self.guard_size) }
} }
pub fn stack_size(&self) -> usize { pub fn stack_size(&self) -> usize {
self.stack_size self.stack_size
} }
pub fn guard_size(&self) -> usize {
self.guard_size
}
/// `(stack_size, guard_size)` after page rounding. The pool rule
/// (RFC 019 §1) compares this against the runtime defaults: only
/// default-shaped stacks are pooled.
pub fn shape(&self) -> (usize, usize) {
(self.stack_size, self.guard_size)
}
/// Pool-recycle zap (RFC 019 §6): `MADV_DONTNEED` everything below the
/// retained entry end `[top − retain, top)` — the span the next actor's
/// shallow frames land in stays resident, the dead spike below it is
/// released. The stack is unowned at the call site (its actor is dead),
/// so a synchronous eager zap races nothing and the RSS drop is
/// immediate — a museum of worst-case spikes is exactly what a pool must
/// not be; DONTNEED's ~8× per-page cost vs FREE is irrelevant off the
/// hot path. Advisory like the park-path shrink: a failure degrades to
/// "the pool keeps RSS", never to incorrectness. No-op (no syscall) when
/// `retain` covers the whole usable region — i.e. always, at the 64 KiB
/// default reserve.
pub(crate) fn recycle_zap(&self, retain: usize) {
if let Some((off, len)) = retain_range(self.stack_size, retain, page_size()) {
unsafe {
libc::madvise(
self.usable_base().add(off) as *mut libc::c_void,
len,
libc::MADV_DONTNEED,
);
}
}
}
}
/// Round `n` up to whole pages — the same rounding `Stack::new` applies, so
/// runtime defaults stored pre-rounded compare exactly against [`Stack::shape`].
pub(crate) fn round_to_pages(n: usize) -> usize {
round_up(n, page_size())
} }
impl Drop for Stack { impl Drop for Stack {
@@ -80,10 +136,174 @@ impl Drop for Stack {
} }
} }
fn page_size() -> usize { pub(crate) fn page_size() -> usize {
unsafe { libc::sysconf(libc::_SC_PAGESIZE) as usize } unsafe { libc::sysconf(libc::_SC_PAGESIZE) as usize }
} }
fn round_up(n: usize, align: usize) -> usize { fn round_up(n: usize, align: usize) -> usize {
(n + align - 1) & !(align - 1) (n + align - 1) & !(align - 1)
} }
/// The whole-page span the park-path shrink may `MADV_FREE` (RFC 019 §3):
/// `[page_up(hwm), page_down(sp − redzone))`, or `None` if no full page fits.
///
/// `hwm` is the sampled high-water (deepest observed `sp`); everything in
/// `[hwm, sp)` is below the live frame and dead by definition. One page of
/// redzone stays resident under live `sp` — it covers the SysV 128-byte red
/// zone plus spill margin with room to spare. Rounding is inward on both
/// ends so the result can never touch the redzone, cross `sp`, or dip below
/// `hwm`; all arithmetic is checked so adversarial inputs (`sp < redzone`,
/// `hwm ≥ sp`, values near the address-space edges) collapse to `None`
/// rather than a wild or negative-length range.
pub(crate) fn shrink_range(hwm: usize, sp: usize, page: usize) -> Option<(usize, usize)> {
debug_assert!(page.is_power_of_two());
if hwm >= sp {
return None;
}
let redzone = page;
let end = sp.checked_sub(redzone)? & !(page - 1); // page_down(sp − redzone)
let start = hwm.checked_add(page - 1)? & !(page - 1); // page_up(hwm)
if end > start {
Some((start, end - start))
} else {
None
}
}
/// The `(offset_from_usable_base, len)` span the pool recycle DONTNEEDs
/// (RFC 019 §6): everything below the retained entry end. "Bottom RETAIN of
/// the stack" is read stack-wise (entry frames = highest addresses of a
/// downward stack): the retained span is `[top − page_up(retain), top)`, the
/// zapped span is the rest — retaining the low-address deep end instead
/// would keep the coldest pages and release the ones the next actor faults
/// first. `retain` rounds *up* to whole pages (retain more, zap less), so
/// with `stack_size` page-rounded by `Stack::new` the result is always
/// page-aligned. Checked math: `retain ≥ stack_size` (notably the default
/// 64 KiB reserve with the 64 KiB RETAIN) and overflow collapse to `None`.
pub(crate) fn retain_range(
stack_size: usize,
retain: usize,
page: usize,
) -> Option<(usize, usize)> {
debug_assert!(page.is_power_of_two());
let retain = retain.checked_add(page - 1)? & !(page - 1); // page_up(retain)
let len = stack_size.checked_sub(retain)?;
if len == 0 {
return None;
}
Some((0, len))
}
#[cfg(test)]
mod tests {
use super::{retain_range, shrink_range};
const PG: usize = 4096;
#[test]
fn retain_covers_whole_stack_is_a_noop() {
// The default config: reserve == RETAIN == 64 KiB. No zap, no syscall.
assert_eq!(retain_range(16 * PG, 16 * PG, PG), None);
assert_eq!(retain_range(PG, PG, PG), None);
}
#[test]
fn retain_larger_than_stack_is_a_noop() {
assert_eq!(retain_range(16 * PG, 17 * PG, PG), None);
assert_eq!(retain_range(PG, usize::MAX, PG), None); // page_up overflows
}
#[test]
fn retain_zero_zaps_everything() {
assert_eq!(retain_range(16 * PG, 0, PG), Some((0, 16 * PG)));
}
#[test]
fn retain_rounds_up_zapping_less() {
// 1 byte of retain keeps a whole page.
assert_eq!(retain_range(16 * PG, 1, PG), Some((0, 15 * PG)));
assert_eq!(retain_range(16 * PG, PG + 1, PG), Some((0, 14 * PG)));
}
#[test]
fn retain_one_page_short_of_stack() {
assert_eq!(retain_range(2 * PG, PG, PG), Some((0, PG)));
}
#[test]
fn retain_range_is_page_aligned() {
for size_pg in [1usize, 2, 3, 16, 1024] {
for retain in [0usize, 1, PG - 1, PG, PG + 1, 4 * PG, size_pg * PG] {
if let Some((off, len)) = retain_range(size_pg * PG, retain, PG) {
assert_eq!(off, 0);
assert_eq!(len % PG, 0);
assert!(len <= size_pg * PG);
assert!(len > 0);
}
}
}
}
#[test]
fn empty_and_inverted_spans_are_none() {
assert_eq!(shrink_range(0x8000_0000, 0x8000_0000, PG), None); // hwm == sp
assert_eq!(shrink_range(0x8000_1000, 0x8000_0000, PG), None); // hwm > sp
}
#[test]
fn span_smaller_than_redzone_plus_page_is_none() {
let sp = 0x8000_0000;
// Everything within redzone+1 page of sp: no full page clears both
// the redzone and the page_up(hwm) rounding.
assert_eq!(shrink_range(sp - PG, sp, PG), None);
assert_eq!(shrink_range(sp - 2 * PG + 1, sp, PG), None);
}
#[test]
fn exact_two_pages_frees_one() {
let sp = 0x8000_0000;
let hwm = sp - 2 * PG;
// [hwm, hwm+PG) frees; [sp−PG, sp) is redzone.
assert_eq!(shrink_range(hwm, sp, PG), Some((hwm, PG)));
}
#[test]
fn unaligned_ends_round_inward() {
let sp = 0x8000_0123; // live sp mid-page
let hwm = 0x7f00_0abc; // high-water mid-page
let (start, len) = shrink_range(hwm, sp, PG).unwrap();
assert_eq!(start % PG, 0);
assert_eq!(len % PG, 0);
assert!(start >= hwm); // never below the sampled high-water
assert!(start + len <= (sp - PG) & !(PG - 1)); // never into the redzone
}
#[test]
fn result_never_crosses_sp() {
// Sweep hwm across every offset of the page straddling the boundary.
let sp = 0x8000_0000 + 137;
for hwm in (sp - 4 * PG)..(sp) {
if let Some((start, len)) = shrink_range(hwm, sp, PG) {
assert!(start >= hwm);
assert!(start + len + PG <= sp + PG); // end ≤ page_down(sp − PG) < sp
assert!(len > 0);
}
}
}
#[test]
fn underflow_near_zero_is_none() {
assert_eq!(shrink_range(0, PG - 1, PG), None); // sp < redzone
assert_eq!(shrink_range(0, 0, PG), None);
}
#[test]
fn big_span_frees_interior() {
let sp = 0x8000_0000;
let spike = 4 * 1024 * 1024;
let hwm = sp - spike;
let (start, len) = shrink_range(hwm, sp, PG).unwrap();
assert_eq!(start, hwm); // aligned input: starts exactly at hwm
assert_eq!(len, spike - PG); // everything but the redzone page
}
}
+191 -68
View File
@@ -140,7 +140,8 @@ impl Signal {
} }
} }
use crate::channel::channel; use crate::channel::{channel, RecvTimeoutError};
use crate::monitor::DownReason;
use std::collections::{HashMap, VecDeque}; use std::collections::{HashMap, VecDeque};
use std::sync::Arc; use std::sync::Arc;
use std::time::{Duration, Instant}; use std::time::{Duration, Instant};
@@ -165,11 +166,52 @@ pub enum Restart {
pub struct ChildSpec { pub struct ChildSpec {
start: Arc<dyn Fn() + Send + Sync + 'static>, start: Arc<dyn Fn() + Send + Sync + 'static>,
restart: Restart, restart: Restart,
shutdown: Shutdown,
} }
impl ChildSpec { impl ChildSpec {
/// A child with the given restart policy and the default
/// [`Shutdown::Timeout`] of 5 seconds.
pub fn new(restart: Restart, start: impl Fn() + Send + Sync + 'static) -> Self { pub fn new(restart: Restart, start: impl Fn() + Send + Sync + 'static) -> Self {
Self { start: Arc::new(start), restart } Self {
start: Arc::new(start),
restart,
shutdown: Shutdown::default(),
}
}
/// Set how the supervisor stops this child (see [`Shutdown`]). A child
/// that is itself a supervisor should use [`Shutdown::Infinity`] so its
/// own subtree gets its full grace periods.
pub fn shutdown(mut self, shutdown: Shutdown) -> Self {
self.shutdown = shutdown;
self
}
}
/// How a supervisor stops a child it is taking down — the OTP child-spec
/// `shutdown` value. Applies to every supervisor-initiated stop: the ordered
/// shutdown of the whole set and the sibling cycling of
/// [`Strategy::OneForAll`] / [`Strategy::RestForOne`].
///
/// A graceful stop is a [`request_shutdown`](crate::request_shutdown): a child
/// that traps exits receives the request as a message and winds down in its
/// own time; one that does not is stopped outright.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum Shutdown {
/// `request_stop` immediately; no request, no grace period.
BrutalKill,
/// `request_shutdown`, wait up to the duration for the child to exit, then
/// `request_stop` it. The default, at 5 seconds.
Timeout(Duration),
/// `request_shutdown` and wait however long the child takes. Use for a
/// child supervisor, whose subtree has its own timeouts.
Infinity,
}
impl Default for Shutdown {
fn default() -> Self {
Shutdown::Timeout(Duration::from_secs(5))
} }
} }
@@ -241,23 +283,35 @@ impl OneForOne {
} }
/// Run the supervision loop on the current actor. Returns when every child /// Run the supervision loop on the current actor. Returns when every child
/// has reached a terminal, non-restartable state, or when the restart /// has reached a terminal, non-restartable state, when the restart
/// intensity cap is tripped. /// intensity cap is tripped, or when the supervisor is asked to shut down
/// (a [`request_shutdown`](crate::request_shutdown) — from its own
/// supervisor, or from the app). On every one of those exits the survivors
/// are stopped in reverse start order, each per its
/// [`Shutdown`] policy, before this returns.
///
/// The supervisor traps exits for the length of the loop (that is how the
/// shutdown request reaches it as a message). Should the supervisor itself
/// be hard-stopped with [`request_stop`](crate::request_stop), it unwinds
/// without waiting for anything — but a drop guard hard-stops its live
/// children on the way out, so the subtree is not orphaned (a child
/// supervisor unwinds the same way, recursively).
pub fn run(self) { pub fn run(self) {
let me = crate::scheduler::self_pid(); let me = crate::scheduler::self_pid();
let (tx, rx) = channel::<Signal>(); let (tx, rx) = channel::<Signal>();
crate::scheduler::register_supervisor_channel(me, tx); crate::scheduler::register_supervisor_channel(me, tx);
let exits = crate::link::trap_exit();
// pid -> index into `self.children`, for the children currently alive. // pid -> index into `self.children`, for the children currently alive.
let mut by_pid: HashMap<Pid, usize> = HashMap::new(); let mut live = Live::default();
let mut active: usize = 0; let mut active: usize = 0;
// Sliding window of recent restart instants, for the intensity cap. // Sliding window of recent restart instants, for the intensity cap.
let mut restarts: Vec<Instant> = Vec::new(); let mut restarts: Vec<Instant> = Vec::new();
let start_child = |idx: usize, by_pid: &mut HashMap<Pid, usize>| { let start_child = |idx: usize, live: &mut Live| {
let start = self.children[idx].start.clone(); let start = self.children[idx].start.clone();
let h = crate::scheduler::spawn_under(me, move || (start)()); let h = crate::scheduler::spawn_under(me, move || (start)());
by_pid.insert(h.pid(), idx); live.insert(h.pid(), idx);
// We supervise via the signal funnel, not by joining; drop the // We supervise via the signal funnel, not by joining; drop the
// handle so the child's slot is reclaimed promptly on death (the // handle so the child's slot is reclaimed promptly on death (the
// termination Signal is delivered before reclamation regardless). // termination Signal is delivered before reclamation regardless).
@@ -265,28 +319,105 @@ impl OneForOne {
}; };
for idx in 0..self.children.len() { for idx in 0..self.children.len() {
start_child(idx, &mut by_pid); start_child(idx, &mut live);
active += 1; active += 1;
} }
// A signal that arrives while we are awaiting stop-confirmations (for a // A signal that arrives while we are awaiting stop-confirmations (for a
// child we are *not* currently stopping) is stashed here and processed // child we are *not* currently stopping) is stashed here and processed
// by the main loop before it blocks on `recv` again. // by the main loop before it blocks again.
let mut pending: VecDeque<Signal> = VecDeque::new(); let mut pending: VecDeque<Signal> = VecDeque::new();
let next_signal = |pending: &mut VecDeque<Signal>| -> Option<Signal> {
// Stop one child per its policy and wait for its termination signal.
// Signals for other pids that arrive meanwhile are stashed. Bounded by
// construction: `request_stop` (used directly, or as the fallback once
// the grace period lapses) always produces a signal.
let stop_child = |pid: Pid, idx: usize, pending: &mut VecDeque<Signal>| {
let await_one = |deadline: Option<Instant>, pending: &mut VecDeque<Signal>| -> bool {
loop {
let sig = match pending.iter().position(|s| s.pid() == pid) {
Some(i) => pending.remove(i),
None => match deadline {
None => rx.recv().ok(),
Some(dl) => {
match rx.recv_timeout(dl.saturating_duration_since(Instant::now()))
{
Ok(s) => Some(s),
Err(RecvTimeoutError::Timeout) => return false,
Err(RecvTimeoutError::Disconnected) => None,
}
}
},
};
match sig {
Some(s) if s.pid() == pid => return true,
Some(s) => pending.push_back(s),
None => return true, // funnel closed: nothing more can arrive
}
}
};
match self.children[idx].shutdown {
Shutdown::BrutalKill => {
crate::scheduler::request_stop(pid);
await_one(None, pending);
}
Shutdown::Timeout(grace) => {
crate::scheduler::request_shutdown(pid);
if !await_one(Some(Instant::now() + grace), pending) {
crate::scheduler::request_stop(pid);
await_one(None, pending);
}
}
Shutdown::Infinity => {
crate::scheduler::request_shutdown(pid);
await_one(None, pending);
}
}
};
// Stop a set of children in reverse start order, one at a time.
let stop_set =
|set: &mut Vec<(Pid, usize)>, live: &mut Live, pending: &mut VecDeque<Signal>| {
set.sort_unstable_by_key(|x| std::cmp::Reverse(x.1));
for (pid, idx) in set.iter() {
live.remove(pid);
stop_child(*pid, *idx, pending);
}
};
// Wait for the next event: a stashed signal, a child signal, or a
// shutdown request. `Ok(sig)`, or `Err(())` when we must wind down.
let next_event = |pending: &mut VecDeque<Signal>| -> Result<Signal, ()> {
loop {
if let Some(s) = pending.pop_front() { if let Some(s) = pending.pop_front() {
Some(s) return Ok(s);
} else { }
rx.recv().ok() // The trap inbox is arm 0: a shutdown request is noticed even
// under a flood of child signals.
match crate::channel::select(&[&exits, &rx]) {
0 => match exits.try_recv() {
Ok(Some(sig)) if sig.reason == DownReason::Shutdown => return Err(()),
// Any other exit signal (a linked peer's death — a
// supervisor links nothing itself, but may be linked
// to) is not ours to act on; a closed trap inbox is
// impossible while `exits` is held here.
_ => {}
},
_ => match rx.try_recv() {
Ok(Some(s)) => return Ok(s),
Ok(None) => {}
Err(_) => return Err(()), // funnel closed: nothing left to supervise
},
}
} }
}; };
while active > 0 { while active > 0 {
let sig = match next_signal(&mut pending) { let sig = match next_event(&mut pending) {
Some(s) => s, Ok(s) => s,
None => break, // mailbox closed: nothing left to supervise Err(()) => break,
}; };
let idx = match by_pid.remove(&sig.pid()) { let idx = match live.remove(&sig.pid()) {
Some(i) => i, Some(i) => i,
None => continue, // stray/duplicate signal None => continue, // stray/duplicate signal
}; };
@@ -318,78 +449,70 @@ impl OneForOne {
restarts.push(now); restarts.push(now);
// Which *live* siblings get cycled along with the failed child. // Which *live* siblings get cycled along with the failed child.
// (The failed child is already gone — removed from `by_pid` above.) // (The failed child is already gone — removed from `live` above.)
let mut to_stop: Vec<(Pid, usize)> = match self.strategy { let mut to_stop: Vec<(Pid, usize)> = match self.strategy {
Strategy::OneForOne => Vec::new(), Strategy::OneForOne => Vec::new(),
Strategy::OneForAll => by_pid.iter().map(|(p, i)| (*p, *i)).collect(), Strategy::OneForAll => live.iter().map(|(p, i)| (*p, *i)).collect(),
Strategy::RestForOne => by_pid Strategy::RestForOne => live
.iter() .iter()
.filter(|(_, i)| **i > idx) .filter(|(_, i)| **i > idx)
.map(|(p, i)| (*p, *i)) .map(|(p, i)| (*p, *i))
.collect(), .collect(),
}; };
// Stop survivors in reverse start order (highest child index first).
to_stop.sort_unstable_by_key(|x| std::cmp::Reverse(x.1));
// The set we will restart: the failed child plus every sibling we // The set we will restart: the failed child plus every sibling we
// are about to stop, restarted in start (ascending index) order. // are about to stop, restarted in start (ascending index) order.
let mut restart_set: Vec<usize> = Vec::with_capacity(to_stop.len() + 1); let mut restart_set: Vec<usize> = Vec::with_capacity(to_stop.len() + 1);
restart_set.push(idx); restart_set.push(idx);
restart_set.extend(to_stop.iter().map(|(_, i)| *i));
// Request stops, then await each survivor's termination signal // Stop the survivors (each per its policy, reverse start order),
// before restarting. `request_stop` on an already-dead pid is a // then restart the whole set in start order. Net effect on
// no-op; in that case its (already-sent) Exit signal serves as the // `active`: one child died (idx), `to_stop.len()` were stopped,
// confirmation. Any signal for a pid we are *not* awaiting is // and `restart_set.len() == 1 + to_stop.len()` are started — so
// stashed for the main loop.
let mut awaiting: Vec<Pid> = Vec::with_capacity(to_stop.len());
for (pid, cidx) in &to_stop {
by_pid.remove(pid);
restart_set.push(*cidx);
crate::scheduler::request_stop(*pid);
awaiting.push(*pid);
}
while !awaiting.is_empty() {
let s = match next_signal(&mut pending) {
Some(s) => s,
None => break, // mailbox closed mid-await; stop waiting
};
if let Some(pos) = awaiting.iter().position(|p| *p == s.pid()) {
awaiting.swap_remove(pos);
} else {
pending.push_back(s);
}
}
// Restart the whole set in start order. Net effect on `active`:
// one child died (idx), `to_stop.len()` were stopped, and
// `restart_set.len() == 1 + to_stop.len()` are started — so
// `active` is unchanged and needs no adjustment here. // `active` is unchanged and needs no adjustment here.
stop_set(&mut to_stop, &mut live, &mut pending);
restart_set.sort_unstable(); restart_set.sort_unstable();
for cidx in restart_set { for cidx in restart_set {
start_child(cidx, &mut by_pid); start_child(cidx, &mut live);
} }
} }
// Ordered shutdown: stop any survivors in reverse start order and await // Ordered shutdown: stop any survivors in reverse start order, each per
// their termination. On the normal `active == 0` exit `by_pid` is empty // its policy. On the normal `active == 0` exit `live` is empty and this
// and this is a no-op; on a cap-trip or mailbox-closed break it tears // is a no-op; on a shutdown request, a cap-trip, or a closed funnel it
// the remaining children down deterministically instead of leaking them. // tears the remaining children down deterministically.
let mut survivors: Vec<(Pid, usize)> = by_pid.iter().map(|(p, i)| (*p, *i)).collect(); let mut survivors: Vec<(Pid, usize)> = live.iter().map(|(p, i)| (*p, *i)).collect();
survivors.sort_unstable_by_key(|x| std::cmp::Reverse(x.1)); stop_set(&mut survivors, &mut live, &mut pending);
let mut awaiting: Vec<Pid> = Vec::with_capacity(survivors.len()); }
for (pid, _) in &survivors { }
/// The live children of a supervisor, with a drop guard: if the supervisor is
/// unwound (a hard `request_stop`, or a panic in the loop) its children are
/// hard-stopped rather than orphaned. Fire-and-forget by necessity — a guard
/// running mid-unwind cannot park to await anything.
#[derive(Default)]
struct Live(HashMap<Pid, usize>);
impl std::ops::Deref for Live {
type Target = HashMap<Pid, usize>;
fn deref(&self) -> &Self::Target {
&self.0
}
}
impl std::ops::DerefMut for Live {
fn deref_mut(&mut self) -> &mut Self::Target {
&mut self.0
}
}
impl Drop for Live {
fn drop(&mut self) {
if std::thread::panicking() {
for pid in self.0.keys() {
crate::scheduler::request_stop(*pid); crate::scheduler::request_stop(*pid);
awaiting.push(*pid);
}
while !awaiting.is_empty() {
let s = match next_signal(&mut pending) {
Some(s) => s,
None => break,
};
if let Some(pos) = awaiting.iter().position(|p| *p == s.pid()) {
awaiting.swap_remove(pos);
} }
} }
} }
} }
+11 -2
View File
@@ -6,10 +6,19 @@
//! Build the loom models with: `RUSTFLAGS="--cfg loom" cargo test --lib --release` //! Build the loom models with: `RUSTFLAGS="--cfg loom" cargo test --lib --release`
#[cfg(loom)] #[cfg(loom)]
pub(crate) use loom::sync::atomic::{AtomicU64, AtomicUsize, Ordering}; pub(crate) use loom::sync::atomic::{fence, AtomicU64, AtomicUsize, Ordering};
#[cfg(not(loom))] #[cfg(not(loom))]
pub(crate) use std::sync::atomic::{AtomicU64, AtomicUsize, Ordering}; pub(crate) use std::sync::atomic::{fence, AtomicU64, AtomicUsize, Ordering};
// park.rs condvar-parker (loom + non-Linux builds only; the Linux non-loom
// build parks on a futex and never touches these — gating them identically
// keeps the default build free of unused imports).
#[cfg(loom)]
pub(crate) use loom::sync::{Condvar, Mutex};
#[cfg(all(not(loom), not(target_os = "linux")))]
pub(crate) use std::sync::{Condvar, Mutex};
/// `UnsafeCell` with loom's `with`/`with_mut` access API; pass-through cost /// `UnsafeCell` with loom's `with`/`with_mut` access API; pass-through cost
/// is zero in normal builds (`#[inline]`, newtype over std's cell). /// is zero in normal builds (`#[inline]`, newtype over std's cell).
+40 -2
View File
@@ -129,7 +129,9 @@ impl Ord for Entry {
// Earlier deadline first; ties broken by insertion order so the // Earlier deadline first; ties broken by insertion order so the
// ordering is total. `Reason` and `Pid` deliberately don't // ordering is total. `Reason` and `Pid` deliberately don't
// participate. // participate.
self.deadline.cmp(&other.deadline).then_with(|| self.seq.cmp(&other.seq)) self.deadline
.cmp(&other.deadline)
.then_with(|| self.seq.cmp(&other.seq))
} }
} }
@@ -141,6 +143,14 @@ impl PartialOrd for Entry {
#[derive(Default)] #[derive(Default)]
pub struct Timers { pub struct Timers {
/// RFC 018: the scheduler coordination layer. Attached once at
/// `RuntimeInner::new`; every insert notes its deadline (min-maintained
/// snapshot for the busy-path due-check + the timekeeper re-arm wake)
/// and every pop/clear re-anchors the snapshot to the heap minimum.
/// All calls happen under the timers mutex — the serialization the
/// coordinator's timer protocol mandates. `None` only in unit tests
/// that construct a bare `Timers`.
coord: Option<std::sync::Arc<crate::park::Coordinator>>,
/// Reverse-wrapped so the smallest deadline is at the top. /// Reverse-wrapped so the smallest deadline is at the top.
heap: BinaryHeap<Reverse<Entry>>, heap: BinaryHeap<Reverse<Entry>>,
/// Monotonic counter for the tiebreaker `seq` field (and the `TimerId` of a /// Monotonic counter for the tiebreaker `seq` field (and the `TimerId` of a
@@ -157,7 +167,18 @@ pub struct Timers {
impl Timers { impl Timers {
pub fn new() -> Self { pub fn new() -> Self {
Self { heap: BinaryHeap::new(), next_seq: 0, armed: std::collections::HashSet::new() } Self {
coord: None,
heap: BinaryHeap::new(),
next_seq: 0,
armed: std::collections::HashSet::new(),
}
}
/// Attach the scheduler coordination layer (RFC 018). Called once, at
/// runtime construction, before any scheduler thread exists.
pub(crate) fn attach_coordinator(&mut self, c: std::sync::Arc<crate::park::Coordinator>) {
self.coord = Some(c);
} }
/// Insert a `Sleep` timer. Convenience for the common case. /// Insert a `Sleep` timer. Convenience for the common case.
@@ -242,6 +263,13 @@ impl Timers {
#[cfg(feature = "smarm-causal")] #[cfg(feature = "smarm-causal")]
wall, wall,
})); }));
// RFC 018: publish the (possibly new-minimum) deadline to the
// busy-path snapshot and wake the timekeeper if it is parked
// toward a later one. We hold the timers mutex — the mandated
// serialization for both.
if let Some(c) = &self.coord {
c.note_deadline(deadline);
}
seq seq
} }
@@ -255,6 +283,9 @@ impl Timers {
pub fn clear(&mut self) { pub fn clear(&mut self) {
self.heap.clear(); self.heap.clear();
self.armed.clear(); self.armed.clear();
if let Some(c) = &self.coord {
c.refresh_deadline(None);
}
} }
/// Soonest pending deadline, or `None` if the heap is empty. /// Soonest pending deadline, or `None` if the heap is empty.
@@ -324,6 +355,13 @@ impl Timers {
} }
out.push(entry); out.push(entry);
} }
// RFC 018: re-anchor the busy-path snapshot to the new heap minimum
// (still under the timers mutex). A causal-shift re-queue above went
// through `heap.push` directly, so this peek is the one place the
// snapshot is guaranteed to catch up.
if let Some(c) = &self.coord {
c.refresh_deadline(self.peek_deadline());
}
out out
} }
} }
+52 -14
View File
@@ -16,13 +16,17 @@
#[cfg(feature = "smarm-trace")] #[cfg(feature = "smarm-trace")]
#[macro_export] #[macro_export]
macro_rules! te { macro_rules! te {
($kind:expr) => { $crate::trace::record($kind) }; ($kind:expr) => {
$crate::trace::record($kind)
};
} }
#[cfg(not(feature = "smarm-trace"))] #[cfg(not(feature = "smarm-trace"))]
#[macro_export] #[macro_export]
macro_rules! te { macro_rules! te {
($kind:expr) => { () }; ($kind:expr) => {
()
};
} }
#[cfg(feature = "smarm-trace")] #[cfg(feature = "smarm-trace")]
@@ -42,17 +46,32 @@ mod inner {
#[derive(Clone, Debug)] #[derive(Clone, Debug)]
pub enum Event { pub enum Event {
// Actor lifecycle // Actor lifecycle
Spawn { parent: Pid, child: Pid }, Spawn {
parent: Pid,
child: Pid,
},
Resume(Pid), Resume(Pid),
Yield(Pid), Yield(Pid),
Park(Pid), Park(Pid),
Done(Pid), Done(Pid),
/// Root exit found a live forest root (an actor nobody supervises)
/// and delivered `request_shutdown` to it. `trapping` says whether it
/// got the chance to drain (true) or was stopped outright (false).
/// Every such line is an actor whose lifetime was nobody's business
/// but the runtime's — the way to *see* unsupervised leftovers.
RootSweep {
target: Pid,
trapping: bool,
},
// Wakeup paths // Wakeup paths
UnparkDirect(Pid), // unpark() saw Parked -> re-queued immediately UnparkDirect(Pid), // unpark() saw Parked -> re-queued immediately
UnparkDeferred(Pid), // unpark() saw Runnable -> set pending_unpark flag UnparkDeferred(Pid), // unpark() saw Runnable -> set pending_unpark flag
UnparkFlagConsumed(Pid), // scheduler saw flag on Park -> re-queued instead UnparkFlagConsumed(Pid), // scheduler saw flag on Park -> re-queued instead
// Channel // Channel
Send { sender: Pid, receiver: Option<Pid> }, Send {
sender: Pid,
receiver: Option<Pid>,
},
RecvPark(Pid), RecvPark(Pid),
RecvWake(Pid), RecvWake(Pid),
// Queue // Queue
@@ -109,8 +128,8 @@ mod inner {
// ----------------------------------------------------------------------- // -----------------------------------------------------------------------
pub fn open() { pub fn open() {
let path = std::env::var("SMARM_TRACE_FILE") let path =
.unwrap_or_else(|_| "smarm_trace.json".to_owned()); std::env::var("SMARM_TRACE_FILE").unwrap_or_else(|_| "smarm_trace.json".to_owned());
let (tx, rx) = mpsc::channel::<Msg>(); let (tx, rx) = mpsc::channel::<Msg>();
let start = Instant::now(); let start = Instant::now();
@@ -164,8 +183,11 @@ mod inner {
// which would try to re-acquire inner.shared (already held at many // which would try to re-acquire inner.shared (already held at many
// te!() call sites) -> deadlock. Guard at the very top, before any // te!() call sites) -> deadlock. Guard at the very top, before any
// allocation-capable call. // allocation-capable call.
let was_enabled = crate::preempt::PREEMPTION_ENABLED let was_enabled = crate::preempt::PREEMPTION_ENABLED.with(|e| {
.with(|e| { let v = e.get(); e.set(false); v }); let v = e.get();
e.set(false);
v
});
LOCAL_STATE.with(|cell| { LOCAL_STATE.with(|cell| {
let mut opt = cell.borrow_mut(); let mut opt = cell.borrow_mut();
@@ -197,7 +219,10 @@ mod inner {
fn drain_thread(rx: mpsc::Receiver<Msg>, path: &str) { fn drain_thread(rx: mpsc::Receiver<Msg>, path: &str) {
let f = match std::fs::File::create(path) { let f = match std::fs::File::create(path) {
Ok(f) => f, Ok(f) => f,
Err(e) => { eprintln!("[smarm-trace] create failed: {}", e); return; } Err(e) => {
eprintln!("[smarm-trace] create failed: {}", e);
return;
}
}; };
let mut w = std::io::BufWriter::new(f); let mut w = std::io::BufWriter::new(f);
let _ = writeln!(w, "{{\"traceEvents\":["); let _ = writeln!(w, "{{\"traceEvents\":[");
@@ -210,7 +235,9 @@ mod inner {
Ok(Msg::Event(r)) => { Ok(Msg::Event(r)) => {
let (name, actor_idx) = chrome_fields(&r.event); let (name, actor_idx) = chrome_fields(&r.event);
let ts_us = r.nanos as f64 / 1000.0; let ts_us = r.nanos as f64 / 1000.0;
if !first { let _ = w.write_all(b",\n"); } if !first {
let _ = w.write_all(b",\n");
}
first = false; first = false;
let _ = write!(w, let _ = write!(w,
"{{\"ph\":\"i\",\"ts\":{:.3},\"pid\":{},\"tid\":{},\"name\":{:?},\"s\":\"g\"}}", "{{\"ph\":\"i\",\"ts\":{:.3},\"pid\":{},\"tid\":{},\"name\":{:?},\"s\":\"g\"}}",
@@ -234,19 +261,30 @@ mod inner {
fn chrome_fields(ev: &Event) -> (String, u32) { fn chrome_fields(ev: &Event) -> (String, u32) {
match ev { match ev {
Event::Spawn { parent, child } => Event::Spawn { parent, child } => {
(format!("spawn c={}", child.index()), parent.index()), (format!("spawn c={}", child.index()), parent.index())
}
Event::Resume(p) => ("resume".into(), p.index()), Event::Resume(p) => ("resume".into(), p.index()),
Event::Yield(p) => ("yield".into(), p.index()), Event::Yield(p) => ("yield".into(), p.index()),
Event::Park(p) => ("park".into(), p.index()), Event::Park(p) => ("park".into(), p.index()),
Event::Done(p) => ("done".into(), p.index()), Event::Done(p) => ("done".into(), p.index()),
Event::RootSweep { target, trapping } => (
format!(
"root_sweep {}",
if *trapping { "shutdown" } else { "stopped" }
),
target.index(),
),
Event::UnparkDirect(p) => ("unpark_direct".into(), p.index()), Event::UnparkDirect(p) => ("unpark_direct".into(), p.index()),
Event::UnparkDeferred(p) => ("unpark_deferred".into(), p.index()), Event::UnparkDeferred(p) => ("unpark_deferred".into(), p.index()),
Event::UnparkFlagConsumed(p) => ("unpark_flag_consumed".into(), p.index()), Event::UnparkFlagConsumed(p) => ("unpark_flag_consumed".into(), p.index()),
Event::Send { sender, receiver } => ( Event::Send { sender, receiver } => (
format!("send rx={}", receiver format!(
"send rx={}",
receiver
.map(|p| p.index().to_string()) .map(|p| p.index().to_string())
.unwrap_or_else(|| "none".into())), .unwrap_or_else(|| "none".into())
),
sender.index(), sender.index(),
), ),
Event::RecvPark(p) => ("recv_park".into(), p.index()), Event::RecvPark(p) => ("recv_park".into(), p.index()),
+24 -6
View File
@@ -49,8 +49,14 @@ fn looping_actor_on_check_is_stopped() {
} }
let _ = h.join(); let _ = h.join();
}); });
assert!(saw_stopped.load(Ordering::SeqCst), "expected DownReason::Stopped"); assert!(
assert!(dropped.load(Ordering::SeqCst), "Drop guard must run during the cancellation unwind"); saw_stopped.load(Ordering::SeqCst),
"expected DownReason::Stopped"
);
assert!(
dropped.load(Ordering::SeqCst),
"Drop guard must run during the cancellation unwind"
);
} }
#[test] #[test]
@@ -79,8 +85,14 @@ fn parked_on_recv_actor_is_stopped() {
} }
let _ = h.join(); let _ = h.join();
}); });
assert!(saw_stopped.load(Ordering::SeqCst), "expected DownReason::Stopped"); assert!(
assert!(dropped.load(Ordering::SeqCst), "Drop guard must run on cancellation of a parked actor"); saw_stopped.load(Ordering::SeqCst),
"expected DownReason::Stopped"
);
assert!(
dropped.load(Ordering::SeqCst),
"Drop guard must run on cancellation of a parked actor"
);
} }
#[test] #[test]
@@ -185,6 +197,12 @@ fn stop_flagged_while_queued_lands_at_first_park() {
.recv_timeout(Duration::from_secs(10)) .recv_timeout(Duration::from_secs(10))
.expect("runtime deadlocked: stop against a QUEUED actor was lost at its first park"); .expect("runtime deadlocked: stop against a QUEUED actor was lost at its first park");
assert!(saw_stopped.load(Ordering::SeqCst), "expected DownReason::Stopped"); assert!(
assert!(dropped.load(Ordering::SeqCst), "Drop guard must run during the cancellation unwind"); saw_stopped.load(Ordering::SeqCst),
"expected DownReason::Stopped"
);
assert!(
dropped.load(Ordering::SeqCst),
"Drop guard must run during the cancellation unwind"
);
} }
+25 -26
View File
@@ -24,7 +24,11 @@ fn progress_point_counts() {
h.join().unwrap(); h.join().unwrap();
let after = smarm::causal::progress_snapshot(); let after = smarm::causal::progress_snapshot();
let delta = |name: &str| { let delta = |name: &str| {
after.iter().find(|(n, _)| n == name).map(|(_, c)| *c).unwrap() after
.iter()
.find(|(n, _)| n == name)
.map(|(_, c)| *c)
.unwrap()
- before - before
.iter() .iter()
.find(|(n, _)| n == name) .find(|(n, _)| n == name)
@@ -45,21 +49,12 @@ fn site_guard_nesting_restores() {
assert_eq!(smarm::causal::current_site_name(), None); assert_eq!(smarm::causal::current_site_name(), None);
{ {
let _outer = smarm::causal_site!("outer"); let _outer = smarm::causal_site!("outer");
assert_eq!( assert_eq!(smarm::causal::current_site_name().as_deref(), Some("outer"));
smarm::causal::current_site_name().as_deref(),
Some("outer")
);
{ {
let _inner = smarm::causal_site!("inner"); let _inner = smarm::causal_site!("inner");
assert_eq!( assert_eq!(smarm::causal::current_site_name().as_deref(), Some("inner"));
smarm::causal::current_site_name().as_deref(),
Some("inner")
);
} }
assert_eq!( assert_eq!(smarm::causal::current_site_name().as_deref(), Some("outer"));
smarm::causal::current_site_name().as_deref(),
Some("outer")
);
} }
assert_eq!(smarm::causal::current_site_name(), None); assert_eq!(smarm::causal::current_site_name(), None);
}); });
@@ -88,10 +83,7 @@ fn virtual_speedup_ledger() {
let bystander = smarm::spawn(move || { let bystander = smarm::spawn(move || {
while !stop2.load(Ordering::Relaxed) { while !stop2.load(Ordering::Relaxed) {
smarm::check!(); smarm::check!();
out2.store( out2.store(smarm::causal::my_absorbed_delay_cycles(), Ordering::Relaxed);
smarm::causal::my_absorbed_delay_cycles(),
Ordering::Relaxed,
);
} }
}); });
@@ -166,10 +158,7 @@ fn runnable_bystander_pays_delay() {
while !stop_b.load(Ordering::Relaxed) { while !stop_b.load(Ordering::Relaxed) {
iters2.fetch_add(1, Ordering::Relaxed); iters2.fetch_add(1, Ordering::Relaxed);
smarm::check!(); smarm::check!();
absorbed2.store( absorbed2.store(smarm::causal::my_absorbed_delay_cycles(), Ordering::Relaxed);
smarm::causal::my_absorbed_delay_cycles(),
Ordering::Relaxed,
);
} }
}); });
@@ -177,8 +166,7 @@ fn runnable_bystander_pays_delay() {
let i0 = iters.load(Ordering::Relaxed); let i0 = iters.load(Ordering::Relaxed);
let t = std::time::Instant::now(); let t = std::time::Instant::now();
smarm::sleep(Duration::from_millis(150)); smarm::sleep(Duration::from_millis(150));
let rate = let rate = (iters.load(Ordering::Relaxed) - i0) as f64 / t.elapsed().as_secs_f64();
(iters.load(Ordering::Relaxed) - i0) as f64 / t.elapsed().as_secs_f64();
out.store(rate as u64, Ordering::Relaxed); out.store(rate as u64, Ordering::Relaxed);
}; };
@@ -448,7 +436,10 @@ fn timer_deadline_shifts_with_injected_delay() {
// Raw deadline passed, effective deadline not: nothing fires, entry kept. // Raw deadline passed, effective deadline not: nothing fires, entry kept.
assert!(t.pop_due(now + Duration::from_millis(60)).is_empty()); assert!(t.pop_due(now + Duration::from_millis(60)).is_empty());
assert!(!t.is_empty(), "shifted entry must be re-queued, not dropped"); assert!(
!t.is_empty(),
"shifted entry must be re-queued, not dropped"
);
// Past raw + injected (with margin) it must fire. Chase in case a // Past raw + injected (with margin) it must fire. Chase in case a
// parallel test injected more debt meanwhile. // parallel test injected more debt meanwhile.
@@ -503,7 +494,11 @@ fn wall_timer_ignores_injected_delay() {
// Just past the raw deadline: the wall entry fires, the virtual one is // Just past the raw deadline: the wall entry fires, the virtual one is
// re-queued at its shifted deadline. // re-queued at its shifted deadline.
let due = t.pop_due(now + Duration::from_millis(60)); let due = t.pop_due(now + Duration::from_millis(60));
assert_eq!(due.len(), 1, "exactly the wall entry must fire at raw deadline"); assert_eq!(
due.len(),
1,
"exactly the wall entry must fire at raw deadline"
);
assert_eq!(due[0].pid, Pid::new(0, 0)); assert_eq!(due[0].pid, Pid::new(0, 0));
assert!(!t.is_empty(), "virtual sibling must remain queued, shifted"); assert!(!t.is_empty(), "virtual sibling must remain queued, shifted");
} }
@@ -623,7 +618,11 @@ fn wall_send_after_ignores_injected_delay() {
// Just past the raw deadline: only the wall send pops; run its thunk. // Just past the raw deadline: only the wall send pops; run its thunk.
let due = t.pop_due(now + Duration::from_millis(60)); let due = t.pop_due(now + Duration::from_millis(60));
assert_eq!(due.len(), 1, "exactly the wall send must fire at raw deadline"); assert_eq!(
due.len(),
1,
"exactly the wall send must fire at raw deadline"
);
for e in due { for e in due {
if let smarm::timer::Reason::Send { fire } = e.reason { if let smarm::timer::Reason::Send { fire } = e.reason {
fire(); fire();
+11 -3
View File
@@ -154,7 +154,10 @@ fn channel_ops_interleaved_with_monitor_churn_multi_thread() {
} }
consumer.join().unwrap(); consumer.join().unwrap();
}); });
assert_eq!(total.load(std::sync::atomic::Ordering::Relaxed), (0..32).sum::<i64>()); assert_eq!(
total.load(std::sync::atomic::Ordering::Relaxed),
(0..32).sum::<i64>()
);
} }
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -220,7 +223,10 @@ fn recv_timeout_reports_disconnected_on_close() {
fn recv_timeout_zero_duration_is_a_bounded_poll() { fn recv_timeout_zero_duration_is_a_bounded_poll() {
run(|| { run(|| {
let (_tx, rx) = channel::<i64>(); let (_tx, rx) = channel::<i64>();
assert_eq!(rx.recv_timeout(Duration::ZERO), Err(RecvTimeoutError::Timeout)); assert_eq!(
rx.recv_timeout(Duration::ZERO),
Err(RecvTimeoutError::Timeout)
);
}); });
} }
@@ -262,7 +268,8 @@ fn recv_timeout_many_waiters_multi_thread() {
let (tx, rx) = channel::<i64>(); let (tx, rx) = channel::<i64>();
let got = got2.clone(); let got = got2.clone();
let timed_out = timed_out2.clone(); let timed_out = timed_out2.clone();
handles.push(spawn(move || match rx.recv_timeout(Duration::from_millis(100)) { handles.push(spawn(move || {
match rx.recv_timeout(Duration::from_millis(100)) {
Ok(v) => { Ok(v) => {
assert_eq!(v, i); assert_eq!(v, i);
got.fetch_add(1, Ordering::Relaxed); got.fetch_add(1, Ordering::Relaxed);
@@ -271,6 +278,7 @@ fn recv_timeout_many_waiters_multi_thread() {
timed_out.fetch_add(1, Ordering::Relaxed); timed_out.fetch_add(1, Ordering::Relaxed);
} }
Err(e) => panic!("unexpected: {e}"), Err(e) => panic!("unexpected: {e}"),
}
})); }));
if i % 2 == 0 { if i % 2 == 0 {
handles.push(spawn(move || { handles.push(spawn(move || {
+34 -15
View File
@@ -11,9 +11,15 @@ thread_local! {
static LOG: Cell<u64> = const { Cell::new(0) }; static LOG: Cell<u64> = const { Cell::new(0) };
} }
fn log(v: u64) { LOG.with(|c| c.set(c.get() | v)); } fn log(v: u64) {
fn get_log() -> u64 { LOG.with(|c| c.get()) } LOG.with(|c| c.set(c.get() | v));
fn reset_log() { LOG.with(|c| c.set(0)); } }
fn get_log() -> u64 {
LOG.with(|c| c.get())
}
fn reset_log() {
LOG.with(|c| c.set(0));
}
extern "C-unwind" fn actor_simple() { extern "C-unwind" fn actor_simple() {
log(0x1); log(0x1);
@@ -23,7 +29,7 @@ extern "C-unwind" fn actor_simple() {
#[test] #[test]
fn actor_runs_and_returns_to_scheduler() { fn actor_runs_and_returns_to_scheduler() {
reset_log(); reset_log();
let stack = Stack::new(64 * 1024).unwrap(); let stack = Stack::new(64 * 1024, 4096).unwrap();
let sp = init_actor_stack(stack.top(), actor_simple); let sp = init_actor_stack(stack.top(), actor_simple);
set_actor_sp(sp); set_actor_sp(sp);
unsafe { switch_to_actor() }; unsafe { switch_to_actor() };
@@ -40,7 +46,7 @@ extern "C-unwind" fn actor_two_steps() {
#[test] #[test]
fn actor_yields_and_resumes() { fn actor_yields_and_resumes() {
reset_log(); reset_log();
let stack = Stack::new(64 * 1024).unwrap(); let stack = Stack::new(64 * 1024, 4096).unwrap();
let sp = init_actor_stack(stack.top(), actor_two_steps); let sp = init_actor_stack(stack.top(), actor_two_steps);
set_actor_sp(sp); set_actor_sp(sp);
@@ -73,7 +79,10 @@ extern "C-unwind" fn actor_reg_check() {
REG_BEFORE.set([s0, s1, s2, s3]).ok(); REG_BEFORE.set([s0, s1, s2, s3]).ok();
switch_to_scheduler(); switch_to_scheduler();
let a0: u64; let a1: u64; let a2: u64; let a3: u64; let a0: u64;
let a1: u64;
let a2: u64;
let a3: u64;
core::arch::asm!( core::arch::asm!(
"mov {a0}, r12", "mov {a1}, r13", "mov {a2}, r14", "mov {a3}, r15", "mov {a0}, r12", "mov {a1}, r13", "mov {a2}, r14", "mov {a3}, r15",
a0 = out(reg) a0, a1 = out(reg) a1, a2 = out(reg) a2, a3 = out(reg) a3, a0 = out(reg) a0, a1 = out(reg) a1, a2 = out(reg) a2, a3 = out(reg) a3,
@@ -85,11 +94,17 @@ extern "C-unwind" fn actor_reg_check() {
#[test] #[test]
fn callee_saved_registers_survive_yield() { fn callee_saved_registers_survive_yield() {
let stack = Stack::new(64 * 1024).unwrap(); let stack = Stack::new(64 * 1024, 4096).unwrap();
let sp = init_actor_stack(stack.top(), actor_reg_check); let sp = init_actor_stack(stack.top(), actor_reg_check);
set_actor_sp(sp); set_actor_sp(sp);
unsafe { switch_to_actor(); switch_to_actor(); } unsafe {
assert_eq!(REG_BEFORE.get().copied().unwrap(), REG_AFTER.get().copied().unwrap()); switch_to_actor();
switch_to_actor();
}
assert_eq!(
REG_BEFORE.get().copied().unwrap(),
REG_AFTER.get().copied().unwrap()
);
} }
// Two actors, independent stacks. // Two actors, independent stacks.
@@ -117,20 +132,24 @@ extern "C-unwind" fn actor_b() {
#[test] #[test]
fn two_actors_dont_corrupt_each_other() { fn two_actors_dont_corrupt_each_other() {
let stack_a = Stack::new(64 * 1024).unwrap(); let stack_a = Stack::new(64 * 1024, 4096).unwrap();
let stack_b = Stack::new(64 * 1024).unwrap(); let stack_b = Stack::new(64 * 1024, 4096).unwrap();
let sp_a = init_actor_stack(stack_a.top(), actor_a); let sp_a = init_actor_stack(stack_a.top(), actor_a);
let sp_b = init_actor_stack(stack_b.top(), actor_b); let sp_b = init_actor_stack(stack_b.top(), actor_b);
set_actor_sp(sp_a); unsafe { switch_to_actor() }; set_actor_sp(sp_a);
unsafe { switch_to_actor() };
let sp_a = get_actor_sp(); let sp_a = get_actor_sp();
set_actor_sp(sp_b); unsafe { switch_to_actor() }; set_actor_sp(sp_b);
unsafe { switch_to_actor() };
let sp_b = get_actor_sp(); let sp_b = get_actor_sp();
set_actor_sp(sp_a); unsafe { switch_to_actor() }; set_actor_sp(sp_a);
set_actor_sp(sp_b); unsafe { switch_to_actor() }; unsafe { switch_to_actor() };
set_actor_sp(sp_b);
unsafe { switch_to_actor() };
assert_eq!(A_VAL.with(|c| c.get()), 0xA00D); assert_eq!(A_VAL.with(|c| c.get()), 0xA00D);
assert_eq!(B_VAL.with(|c| c.get()), 0xB00D); assert_eq!(B_VAL.with(|c| c.get()), 0xB00D);
+125
View File
@@ -0,0 +1,125 @@
//! Cross-thread wake: a thread that is *not* a smarm scheduler thread must be
//! able to wake (and stop) a parked actor.
//!
//! The gap this pins down: every off-runtime wake primitive (`unpark`,
//! `unpark_at`, `request_stop`) reaches the runtime through the `RUNTIME`
//! thread-local, which is `None` on any non-scheduler thread — so a wake
//! issued from a foreign OS thread is a silent no-op and the parked actor
//! sleeps forever. Both failure modes below manifest as `Runtime::run` never
//! returning, so each test is wrapped in a watchdog: a timeout is the failure.
//!
//! The fix mirrors RFC 018's IO backend — the waker reaches the runtime
//! through a `Weak<RuntimeInner>` it already holds (the receiver captures one
//! when it parks; `Runtime::handle()` hands one to an app thread).
use std::sync::mpsc;
use std::thread;
use std::time::Duration;
const WATCHDOG: Duration = Duration::from_secs(10);
/// Give the target actor time to actually park before the foreign thread pokes
/// it, so we exercise the *wake* of a parked actor rather than the entry-side
/// stop check.
const SETTLE: Duration = Duration::from_millis(200);
fn assert_send_sync<T: Send + Sync>() {}
/// A cross-thread `send` from a plain OS thread must wake a receiver parked in
/// `recv`. Under the thread-local-only wake path the send enqueues the message
/// but never wakes the receiver, so `recv` — and therefore `run` — hangs.
#[test]
fn foreign_thread_send_wakes_parked_receiver() {
let (done_tx, done_rx) = mpsc::channel();
thread::spawn(move || {
let rt = smarm::init(smarm::Config::exact(2));
rt.run(|| {
let (tx, rx) = smarm::channel::<u32>();
// Receiver actor: parks on recv until the foreign thread sends.
let h = smarm::spawn(move || {
assert_eq!(rx.recv().expect("recv"), 42);
});
// Foreign (non-scheduler) OS thread owns the Sender and sends
// after the receiver has parked.
let sender = thread::spawn(move || {
thread::sleep(SETTLE);
tx.send(42).expect("send");
});
let _ = h.join();
sender.join().expect("sender thread");
});
let _ = done_tx.send(());
});
done_rx
.recv_timeout(WATCHDOG)
.expect("run did not return: a foreign-thread send never woke the parked receiver");
}
/// A cross-thread `request_stop` through a `RuntimeHandle` must wake and stop a
/// parked actor. The actor parks on a long sleep (only a stop can end it); the
/// handle is grabbed before `run` and driven from a foreign thread.
#[test]
fn foreign_thread_request_stop_wakes_parked_actor() {
assert_send_sync::<smarm::RuntimeHandle>();
let rt = smarm::init(smarm::Config::exact(2));
let handle = rt.handle();
// Foreign thread: learn the target pid from inside the run, let it park,
// then stop it through the handle.
let (pid_tx, pid_rx) = mpsc::channel::<smarm::Pid>();
let stopper = thread::spawn(move || {
let pid = pid_rx.recv().expect("pid");
thread::sleep(SETTLE);
handle.request_stop(pid);
});
let (done_tx, done_rx) = mpsc::channel();
thread::spawn(move || {
rt.run(move || {
let h = smarm::spawn(|| {
// Parks indefinitely; only a cooperative stop unwinds it.
smarm::sleep(Duration::from_secs(3600));
});
pid_tx.send(h.pid()).expect("send pid");
let _ = h.join();
});
let _ = done_tx.send(());
});
done_rx
.recv_timeout(WATCHDOG)
.expect("run did not return: a foreign-thread request_stop never woke the parked actor");
stopper.join().expect("stopper thread");
}
/// A `RuntimeHandle` held across (and beyond) a run must not keep the runtime
/// alive or block all-done: `run` still returns, and once the `Runtime` is
/// dropped the handle degrades to a harmless no-op (Weak lifecycle) rather than
/// panicking or touching freed memory.
#[test]
fn lingering_handle_does_not_block_all_done() {
let rt = smarm::init(smarm::Config::exact(1));
let handle = rt.handle(); // outlives the run below
let (pid_tx, pid_rx) = mpsc::channel::<smarm::Pid>();
let (done_tx, done_rx) = mpsc::channel();
let runner = thread::spawn(move || {
rt.run(move || {
let h = smarm::spawn(|| {});
pid_tx.send(h.pid()).expect("send pid");
let _ = h.join();
});
// `rt` is dropped here, at the end of this thread.
let _ = done_tx.send(());
});
done_rx
.recv_timeout(WATCHDOG)
.expect("run did not return while a RuntimeHandle was held live");
runner.join().expect("runner thread");
// Runtime is now dropped. A stop through the lingering handle must be a
// silent no-op, not a panic or use-after-free.
let dead_pid = pid_rx.recv().expect("pid");
handle.request_stop(dead_pid);
}
+22 -7
View File
@@ -11,8 +11,8 @@
//! OUTSIDE `run` — an in-actor assertion alone passes vacuously. //! OUTSIDE `run` — an in-actor assertion alone passes vacuously.
use smarm::{ use smarm::{
channel, run, select, select_timeout, spawn, try_select, wait_readable, channel, run, select, select_timeout, spawn, try_select, wait_readable, wait_readable_timeout,
wait_readable_timeout, wait_writable_timeout, yield_now, FdArm, wait_writable_timeout, yield_now, FdArm,
}; };
use std::os::fd::RawFd; use std::os::fd::RawFd;
use std::sync::atomic::{AtomicBool, AtomicU32, Ordering}; use std::sync::atomic::{AtomicBool, AtomicU32, Ordering};
@@ -33,7 +33,10 @@ impl Pipe {
let mut fds: [libc::c_int; 2] = [0; 2]; let mut fds: [libc::c_int; 2] = [0; 2];
let r = unsafe { libc::pipe2(fds.as_mut_ptr(), libc::O_CLOEXEC | libc::O_NONBLOCK) }; let r = unsafe { libc::pipe2(fds.as_mut_ptr(), libc::O_CLOEXEC | libc::O_NONBLOCK) };
assert_eq!(r, 0, "pipe2 failed"); assert_eq!(r, 0, "pipe2 failed");
Pipe { read: fds[0], write: fds[1] } Pipe {
read: fds[0],
write: fds[1],
}
} }
} }
@@ -253,12 +256,18 @@ fn wait_readable_timeout_times_out_then_succeeds_with_data() {
let (rfd, wfd) = (p.read, p.write); let (rfd, wfd) = (p.read, p.write);
let start = Instant::now(); let start = Instant::now();
assert_eq!(wait_readable_timeout(rfd, Duration::from_millis(30)).unwrap(), false); assert_eq!(
wait_readable_timeout(rfd, Duration::from_millis(30)).unwrap(),
false
);
assert!(start.elapsed() >= Duration::from_millis(30)); assert!(start.elapsed() >= Duration::from_millis(30));
// Timed-out wait must leave the fd clean; ready path returns true. // Timed-out wait must leave the fd clean; ready path returns true.
assert_eq!(raw_write(wfd, b"d"), 1); assert_eq!(raw_write(wfd, b"d"), 1);
assert_eq!(wait_readable_timeout(rfd, Duration::from_secs(5)).unwrap(), true); assert_eq!(
wait_readable_timeout(rfd, Duration::from_secs(5)).unwrap(),
true
);
let mut buf = [0u8; 1]; let mut buf = [0u8; 1];
assert_eq!(raw_read(rfd, &mut buf), 1); assert_eq!(raw_read(rfd, &mut buf), 1);
ok2.store(true, Ordering::SeqCst); ok2.store(true, Ordering::SeqCst);
@@ -274,7 +283,10 @@ fn wait_readable_timeout_wakes_on_late_data() {
let p = Pipe::new(); let p = Pipe::new();
let (rfd, wfd) = (p.read, p.write); let (rfd, wfd) = (p.read, p.write);
let h = spawn(move || { let h = spawn(move || {
assert_eq!(wait_readable_timeout(rfd, Duration::from_secs(5)).unwrap(), true); assert_eq!(
wait_readable_timeout(rfd, Duration::from_secs(5)).unwrap(),
true
);
let mut buf = [0u8; 1]; let mut buf = [0u8; 1];
assert_eq!(raw_read(rfd, &mut buf), 1); assert_eq!(raw_read(rfd, &mut buf), 1);
got2.store(buf[0] as u32, Ordering::SeqCst); got2.store(buf[0] as u32, Ordering::SeqCst);
@@ -292,7 +304,10 @@ fn wait_writable_timeout_ready_now_on_empty_pipe() {
run(move || { run(move || {
let p = Pipe::new(); let p = Pipe::new();
// An empty pipe's write end is writable: ready-now path, no park. // An empty pipe's write end is writable: ready-now path, no park.
assert_eq!(wait_writable_timeout(p.write, Duration::from_secs(5)).unwrap(), true); assert_eq!(
wait_writable_timeout(p.write, Duration::from_secs(5)).unwrap(),
true
);
ok2.store(true, Ordering::SeqCst); ok2.store(true, Ordering::SeqCst);
}); });
assert!(ok.load(Ordering::SeqCst)); assert!(ok.load(Ordering::SeqCst));
+63 -19
View File
@@ -88,7 +88,7 @@ impl GenServer for Lifecycle {
} }
} }
// init -> handle_call -> (drop last ref closes inbox) -> terminate. // init -> handle_call -> shutdown -> terminate.
#[test] #[test]
fn init_and_terminate_run() { fn init_and_terminate_run() {
let log = Arc::new(Mutex::new(Vec::new())); let log = Arc::new(Mutex::new(Vec::new()));
@@ -96,9 +96,9 @@ fn init_and_terminate_run() {
run(move || { run(move || {
let server = start(Lifecycle { log: log2 }); let server = start(Lifecycle { log: log2 });
server.call(()).unwrap(); server.call(()).unwrap();
// Dropping the only ref closes the inbox; the server breaks out of its // Refs are addresses: dropping one does not end the server. The
// recv loop and runs terminate. run() will not return until it has. // explicit close does, and waits for terminate.
drop(server); server.shutdown();
}); });
assert_eq!(*log.lock().unwrap(), vec!["init", "call", "terminate"]); assert_eq!(*log.lock().unwrap(), vec!["init", "call", "terminate"]);
} }
@@ -403,7 +403,10 @@ fn worker_pool_down_reaches_handle_down() {
let got = Arc::new(Mutex::new(Vec::new())); let got = Arc::new(Mutex::new(Vec::new()));
let got2 = got.clone(); let got2 = got.clone();
run(move || { run(move || {
let server = start(Pool { watcher: None, log: Vec::new() }); let server = start(Pool {
watcher: None,
log: Vec::new(),
});
server.cast(PoolCast::SpawnDoomedWorker).unwrap(); server.cast(PoolCast::SpawnDoomedWorker).unwrap();
let _ = server.call(()).unwrap(); // sync point: cast handled, worker live let _ = server.call(()).unwrap(); // sync point: cast handled, worker live
*got2.lock().unwrap() = server.call(()).unwrap(); *got2.lock().unwrap() = server.call(()).unwrap();
@@ -421,7 +424,10 @@ fn watch_dead_pid_is_noproc_down() {
let h = spawn(|| {}); let h = spawn(|| {});
let dead = h.pid(); let dead = h.pid();
h.join().unwrap(); h.join().unwrap();
let server = start(Pool { watcher: None, log: Vec::new() }); let server = start(Pool {
watcher: None,
log: Vec::new(),
});
server.cast(PoolCast::Watch(dead)).unwrap(); server.cast(PoolCast::Watch(dead)).unwrap();
*got2.lock().unwrap() = server.call(()).unwrap(); *got2.lock().unwrap() = server.call(()).unwrap();
}); });
@@ -497,7 +503,12 @@ impl GenServer for Timed {
} }
fn timed(fired: Arc<Mutex<Vec<u32>>>, cancel_won: Arc<Mutex<Option<bool>>>) -> Timed { fn timed(fired: Arc<Mutex<Vec<u32>>>, cancel_won: Arc<Mutex<Option<bool>>>) -> Timed {
Timed { timer: None, fired, cancel_won, last: None } Timed {
timer: None,
fired,
cancel_won,
last: None,
}
} }
// A one-shot armed from a handler fires into handle_timer with its payload. // A one-shot armed from a handler fires into handle_timer with its payload.
@@ -534,7 +545,11 @@ fn cancel_before_fire_suppresses_it() {
let count = server.call(()).unwrap(); let count = server.call(()).unwrap();
assert_eq!(count, 0, "cancelled timer must not fire"); assert_eq!(count, 0, "cancelled timer must not fire");
}); });
assert_eq!(*cancel_won.lock().unwrap(), Some(true), "cancel beat the fire"); assert_eq!(
*cancel_won.lock().unwrap(),
Some(true),
"cancel beat the fire"
);
assert!(fired.lock().unwrap().is_empty()); assert!(fired.lock().unwrap().is_empty());
} }
@@ -549,11 +564,16 @@ fn tick_every_rearms_repeatedly() {
run(move || { run(move || {
let cw = Arc::new(Mutex::new(None)); let cw = Arc::new(Mutex::new(None));
let server = start(timed(f2, cw)); let server = start(timed(f2, cw));
server.cast(TkCast::Tick(Duration::from_millis(20))).unwrap(); server
.cast(TkCast::Tick(Duration::from_millis(20)))
.unwrap();
let _ = server.call(()).unwrap(); // sync: periodic armed let _ = server.call(()).unwrap(); // sync: periodic armed
smarm::sleep(Duration::from_millis(130)); // ~6 periods smarm::sleep(Duration::from_millis(130)); // ~6 periods
let count = server.call(()).unwrap(); let count = server.call(()).unwrap();
assert!(count >= 3, "periodic should have re-armed several times, got {count}"); assert!(
count >= 3,
"periodic should have re-armed several times, got {count}"
);
}); });
// Every tick delivered the same payload. // Every tick delivered the same payload.
assert!(fired.lock().unwrap().iter().all(|&v| v == 9)); assert!(fired.lock().unwrap().iter().all(|&v| v == 9));
@@ -568,7 +588,9 @@ fn cancel_stops_a_periodic() {
let c2 = cancel_won.clone(); let c2 = cancel_won.clone();
run(move || { run(move || {
let server = start(timed(f2, c2)); let server = start(timed(f2, c2));
server.cast(TkCast::Tick(Duration::from_millis(20))).unwrap(); server
.cast(TkCast::Tick(Duration::from_millis(20)))
.unwrap();
let _ = server.call(()).unwrap(); let _ = server.call(()).unwrap();
smarm::sleep(Duration::from_millis(70)); // a few ticks smarm::sleep(Duration::from_millis(70)); // a few ticks
server.cast(TkCast::CancelLast).unwrap(); server.cast(TkCast::CancelLast).unwrap();
@@ -616,11 +638,17 @@ fn idle_fires_repeatedly_on_quiet() {
let idles = Arc::new(Mutex::new(0)); let idles = Arc::new(Mutex::new(0));
let i2 = idles.clone(); let i2 = idles.clone();
run(move || { run(move || {
let server = start(Idler { window: Duration::from_millis(25), idles: i2 }); let server = start(Idler {
window: Duration::from_millis(25),
idles: i2,
});
smarm::sleep(Duration::from_millis(130)); // quiet ⇒ ~5 windows smarm::sleep(Duration::from_millis(130)); // quiet ⇒ ~5 windows
drop(server); // keep the server alive across the quiet span drop(server); // keep the server alive across the quiet span
}); });
assert!(*idles.lock().unwrap() >= 2, "idle should re-arm and fire several times"); assert!(
*idles.lock().unwrap() >= 2,
"idle should re-arm and fire several times"
);
} }
// Traffic within the window keeps idle from firing; only once the inbox goes // Traffic within the window keeps idle from firing; only once the inbox goes
@@ -632,7 +660,10 @@ fn traffic_resets_the_idle_window() {
let before_quiet = Arc::new(Mutex::new(u32::MAX)); let before_quiet = Arc::new(Mutex::new(u32::MAX));
let bq = before_quiet.clone(); let bq = before_quiet.clone();
run(move || { run(move || {
let server = start(Idler { window: Duration::from_millis(60), idles: i2 }); let server = start(Idler {
window: Duration::from_millis(60),
idles: i2,
});
// Poke every 25ms (< 60ms window) for ~100ms: each cast resets the // Poke every 25ms (< 60ms window) for ~100ms: each cast resets the
// window before it can elapse. // window before it can elapse.
for _ in 0..4 { for _ in 0..4 {
@@ -643,8 +674,15 @@ fn traffic_resets_the_idle_window() {
smarm::sleep(Duration::from_millis(140)); // now genuinely quiet smarm::sleep(Duration::from_millis(140)); // now genuinely quiet
drop(server); drop(server);
}); });
assert_eq!(*before_quiet.lock().unwrap(), 0, "steady traffic must suppress idle"); assert_eq!(
assert!(*idles.lock().unwrap() >= 1, "idle fires once the inbox falls quiet"); *before_quiet.lock().unwrap(),
0,
"steady traffic must suppress idle"
);
assert!(
*idles.lock().unwrap() >= 1,
"idle fires once the inbox falls quiet"
);
} }
// RFC 015 §4.7 — no armed timer survives loop exit. A server with a live // RFC 015 §4.7 — no armed timer survives loop exit. A server with a live
@@ -658,15 +696,21 @@ fn no_timer_survives_exit() {
let f_read = fired.clone(); let f_read = fired.clone();
run(move || { run(move || {
let server = start(timed(f_server, Arc::new(Mutex::new(None)))); let server = start(timed(f_server, Arc::new(Mutex::new(None))));
server.cast(TkCast::Tick(Duration::from_millis(15))).unwrap(); server
.cast(TkCast::Tick(Duration::from_millis(15)))
.unwrap();
let _ = server.call(()).unwrap(); // sync: periodic armed let _ = server.call(()).unwrap(); // sync: periodic armed
smarm::sleep(Duration::from_millis(45)); // a couple of ticks smarm::sleep(Duration::from_millis(45)); // a couple of ticks
let mon = smarm::monitor(server.pid()); let mon = smarm::monitor(server.pid());
drop(server); // inbox closes → loop exits → guard drains timers server.shutdown(); // loop exits → guard drains timers
// Clean Down ⇒ the loop returned without the no-leak assert aborting. // Clean Down ⇒ the loop returned without the no-leak assert aborting.
assert!(mon.rx.recv().is_ok()); assert!(mon.rx.recv().is_ok());
let at_exit = f_read.lock().unwrap().len(); let at_exit = f_read.lock().unwrap().len();
smarm::sleep(Duration::from_millis(90)); // would be several more ticks smarm::sleep(Duration::from_millis(90)); // would be several more ticks
assert_eq!(f_read.lock().unwrap().len(), at_exit, "no tick may fire after exit"); assert_eq!(
f_read.lock().unwrap().len(),
at_exit,
"no tick may fire after exit"
);
}); });
} }
+243
View File
@@ -0,0 +1,243 @@
//! gen_server lifetime is the actor's, not its refs' (OTP: a pid is an
//! address, a process lives until it stops, is shut down, or is killed).
//!
//! - Dropping the last `GenServerRef` does NOT end the server. It ends via
//! `StopHandle::stop`, `request_shutdown` / `GenServerRef::shutdown`,
//! `request_stop`, or a handler panic.
//! - `GenServerBuilder::named(N).run()` runs the loop inline as the *current*
//! actor, so a server is a direct `ChildSpec` child: the supervisor's
//! shutdown reaches it as `handle_shutdown`, a restart re-binds the name,
//! and by-name `call`/`cast` reach whichever incarnation is live.
use smarm::gen_server::{
self, GenServer, GenServerBuilder, GenServerCtx, GenServerName, ShutdownAction, StopHandle,
};
use smarm::registry::RegisterError;
use smarm::supervisor::{ChildSpec, OneForOne, Restart, Shutdown};
use smarm::{monitor, request_shutdown, run, sleep, spawn, DownReason};
use std::sync::atomic::{AtomicUsize, Ordering};
use std::sync::{Arc, Mutex};
use std::time::Duration;
#[derive(Default, Clone)]
struct Log(Arc<Mutex<Vec<String>>>);
impl Log {
fn push(&self, s: impl Into<String>) {
self.0.lock().unwrap().push(s.into());
}
fn get(&self) -> Vec<String> {
self.0.lock().unwrap().clone()
}
}
struct Counter {
log: Log,
n: u64,
trap: bool,
stop: Option<StopHandle<Counter>>,
}
enum Call {
Get,
}
enum Cast {
Inc,
Stop,
}
impl GenServer for Counter {
type Call = Call;
type Reply = u64;
type Cast = Cast;
type Info = ();
type Timer = ();
fn init(&mut self, ctx: &GenServerCtx<Self>) {
if self.trap {
ctx.trap_exit();
}
self.stop = Some(ctx.stop_handle());
self.log.push("init");
}
fn handle_call(&mut self, Call::Get: Call) -> u64 {
self.n
}
fn handle_cast(&mut self, c: Cast) {
match c {
Cast::Inc => self.n += 1,
Cast::Stop => self.stop.as_ref().unwrap().stop(),
}
}
fn handle_shutdown(&mut self) -> ShutdownAction {
self.log.push("handle_shutdown");
ShutdownAction::Exit
}
fn terminate(&mut self) {
self.log.push("terminate");
}
}
fn counter(log: &Log, trap: bool) -> Counter {
Counter {
log: log.clone(),
n: 0,
trap,
stop: None,
}
}
// ---------------------------------------------------------------------------
// Refs are addresses: dropping the last one does not end the server.
// ---------------------------------------------------------------------------
#[test]
fn dropping_last_ref_does_not_end_server() {
let log = Log::default();
let l = log.clone();
run(move || {
let srv = gen_server::start(counter(&l, true));
let pid = srv.pid();
srv.cast(Cast::Inc).unwrap();
assert_eq!(srv.call(Call::Get).unwrap(), 1);
let mon = monitor(pid);
drop(srv);
sleep(Duration::from_millis(30));
assert!(
mon.rx.try_recv().unwrap().is_none(),
"server must outlive its last ref"
);
assert_eq!(l.get(), vec!["init"], "terminate must not have run");
// Explicit teardown still works, and is what ends it.
request_shutdown(pid);
let down = mon.rx.recv().unwrap();
assert_eq!(down.reason, DownReason::Exit);
});
assert_eq!(log.get(), vec!["init", "handle_shutdown", "terminate"]);
}
#[test]
fn ref_shutdown_is_the_explicit_close() {
let log = Log::default();
let l = log.clone();
run(move || {
let srv = gen_server::start(counter(&l, true));
srv.call(Call::Get).unwrap(); // sync: init (and trap_exit) has run
srv.shutdown(); // graceful, waits
assert_eq!(l.get(), vec!["init", "handle_shutdown", "terminate"]);
});
}
#[test]
fn forgotten_server_is_shut_down_at_root_exit() {
// A ref-less server is not a hung run: root exit shuts it down.
let log = Log::default();
let l = log.clone();
run(move || {
let srv = gen_server::start(counter(&l, false));
drop(srv);
});
assert_eq!(log.get(), vec!["init", "terminate"]);
}
// ---------------------------------------------------------------------------
// Inline run: a gen_server as a direct ChildSpec child.
// ---------------------------------------------------------------------------
const COUNTER: GenServerName<Counter> = GenServerName::new("lifetime-counter");
#[test]
fn named_run_is_a_direct_supervised_child_and_gets_shutdown() {
let log = Log::default();
let l = log.clone();
run(move || {
let l2 = l.clone();
let sup = spawn(move || {
let l3 = l2.clone();
OneForOne::new()
.child(
ChildSpec::new(Restart::Permanent, move || {
GenServerBuilder::new(counter(&l3, true))
.named(COUNTER)
.run()
.expect("name free");
})
.shutdown(Shutdown::Infinity),
)
.run();
});
sleep(Duration::from_millis(10));
gen_server::cast(COUNTER, Cast::Inc).unwrap();
assert_eq!(gen_server::call(COUNTER, Call::Get).unwrap(), 1);
request_shutdown(sup.pid());
sup.join()
.expect("ordered shutdown, supervisor returns normally");
assert_eq!(l.get(), vec!["init", "handle_shutdown", "terminate"]);
assert!(gen_server::whereis_server(COUNTER).is_none());
});
}
#[test]
fn named_run_child_restarts_and_rebinds_name() {
let log = Log::default();
let inits = Arc::new(AtomicUsize::new(0));
let l = log.clone();
let i = inits.clone();
run(move || {
let l2 = l.clone();
let i2 = i.clone();
let sup = spawn(move || {
let l3 = l2.clone();
let i3 = i2.clone();
OneForOne::new()
.child(ChildSpec::new(Restart::Permanent, move || {
i3.fetch_add(1, Ordering::SeqCst);
GenServerBuilder::new(counter(&l3, false))
.named(COUNTER)
.run()
.expect("name free on (re)start");
}))
.run();
});
sleep(Duration::from_millis(10));
gen_server::cast(COUNTER, Cast::Inc).unwrap();
assert_eq!(gen_server::call(COUNTER, Call::Get).unwrap(), 1);
// Normal self-exit → Permanent restarts it, fresh state, same name.
gen_server::cast(COUNTER, Cast::Stop).unwrap();
sleep(Duration::from_millis(30));
assert_eq!(i.load(Ordering::SeqCst), 2, "restarted once");
assert_eq!(gen_server::call(COUNTER, Call::Get).unwrap(), 0);
request_shutdown(sup.pid());
sup.join().unwrap();
});
assert_eq!(log.get(), vec!["init", "terminate", "init", "terminate"]);
}
#[test]
fn named_run_name_clash_fails_before_init() {
let log = Log::default();
let l = log.clone();
run(move || {
let first = GenServerBuilder::new(counter(&l, false))
.named(COUNTER)
.start()
.unwrap();
let l2 = l.clone();
let res = Arc::new(Mutex::new(None));
let r2 = res.clone();
let first_pid = first.pid();
spawn(move || {
let r = GenServerBuilder::new(counter(&l2, false))
.named(COUNTER)
.run();
*r2.lock().unwrap() = Some(r);
})
.join()
.unwrap();
assert_eq!(
*res.lock().unwrap(),
Some(Err(RegisterError::NameTaken { holder: first_pid }))
);
assert_eq!(l.get(), vec!["init"], "clashing server never ran init");
first.shutdown();
});
}
+222
View File
@@ -0,0 +1,222 @@
//! gen_server graceful shutdown.
//!
//! - A server that does not opt in (`ctx.trap_exit()` in `init`) is stopped
//! outright by `request_shutdown`, exactly as by `request_stop`.
//! - A trapping server receives the request as `handle_shutdown`. The default
//! returns `ShutdownAction::Exit`: the loop breaks and `terminate` runs on
//! the normal (non-unwind) path, so it may block. `Continue` keeps the loop
//! dispatching; the state later ends itself with a `StopHandle` — the only
//! way for a gen_server to exit *normally* on its own (`request_stop` on
//! self is an abnormal `Stopped`, which `Transient` restarts).
//! - Other exit signals (linked peers dying) reach a trapping server via
//! `handle_exit`.
use smarm::gen_server::{
start, GenServer, GenServerBuilder, GenServerCtx, GenServerRef, ShutdownAction, StopHandle,
};
use smarm::supervisor::{ChildSpec, OneForOne, Restart};
use smarm::{link, monitor, request_shutdown, run, self_pid, sleep, spawn, DownReason, ExitSignal};
use std::sync::atomic::{AtomicBool, AtomicUsize, Ordering};
use std::sync::{Arc, Mutex};
use std::time::Duration;
#[derive(Default, Clone)]
struct Log {
events: Arc<Mutex<Vec<&'static str>>>,
}
impl Log {
fn push(&self, e: &'static str) {
self.events.lock().unwrap().push(e);
}
fn get(&self) -> Vec<&'static str> {
self.events.lock().unwrap().clone()
}
}
/// A server with configurable shutdown behaviour.
struct Srv {
log: Log,
trap: bool,
action: ShutdownAction,
stop: Option<StopHandle<Srv>>,
exits: Arc<Mutex<Vec<ExitSignal>>>,
}
impl Srv {
fn new(log: &Log, trap: bool, action: ShutdownAction) -> Self {
Srv {
log: log.clone(),
trap,
action,
stop: None,
exits: Default::default(),
}
}
}
enum Cast {
Note(&'static str),
StopNow,
}
impl GenServer for Srv {
type Call = ();
type Reply = ();
type Cast = Cast;
type Info = ();
type Timer = ();
fn init(&mut self, ctx: &GenServerCtx<Self>) {
if self.trap {
ctx.trap_exit();
}
self.stop = Some(ctx.stop_handle());
}
fn handle_call(&mut self, _: ()) {}
fn handle_cast(&mut self, c: Cast) {
match c {
Cast::Note(s) => self.log.push(s),
Cast::StopNow => self.stop.as_ref().unwrap().stop(),
}
}
fn handle_shutdown(&mut self) -> ShutdownAction {
self.log.push("handle_shutdown");
self.action
}
fn handle_exit(&mut self, sig: ExitSignal) {
self.log.push("handle_exit");
self.exits.lock().unwrap().push(sig);
}
fn terminate(&mut self) {
// Allowed to block on the graceful path.
if self.trap {
sleep(Duration::from_millis(10));
}
self.log.push("terminate");
}
}
fn spawn_settled<G: GenServer>(state: G) -> GenServerRef<G> {
let r = start(state);
sleep(Duration::from_millis(20)); // let init (trap_exit) run
r
}
#[test]
fn non_trapping_server_is_stopped_outright() {
let log = Log::default();
let l = log.clone();
run(move || {
let r = spawn_settled(Srv::new(&l, false, ShutdownAction::Exit));
let mon = monitor(r.pid());
request_shutdown(r.pid());
let d = mon.rx.recv().unwrap();
assert_eq!(d.reason, DownReason::Stopped);
});
assert_eq!(log.get(), vec!["terminate"]);
}
#[test]
fn trapping_server_exits_normally_via_handle_shutdown() {
let log = Log::default();
let l = log.clone();
run(move || {
let r = spawn_settled(Srv::new(&l, true, ShutdownAction::Exit));
let mon = monitor(r.pid());
request_shutdown(r.pid());
let d = mon.rx.recv().unwrap();
assert_eq!(d.reason, DownReason::Exit);
});
assert_eq!(log.get(), vec!["handle_shutdown", "terminate"]);
}
#[test]
fn continue_keeps_dispatching_until_stop_handle() {
let log = Log::default();
let l = log.clone();
run(move || {
let r = spawn_settled(Srv::new(&l, true, ShutdownAction::Continue));
let mon = monitor(r.pid());
request_shutdown(r.pid());
sleep(Duration::from_millis(20));
r.cast(Cast::Note("after-shutdown-request")).unwrap();
r.cast(Cast::StopNow).unwrap();
let d = mon.rx.recv().unwrap();
assert_eq!(d.reason, DownReason::Exit);
});
assert_eq!(
log.get(),
vec!["handle_shutdown", "after-shutdown-request", "terminate"]
);
}
#[test]
fn stop_handle_is_a_normal_exit_that_transient_does_not_restart() {
let starts = Arc::new(AtomicUsize::new(0));
let s = starts.clone();
run(move || {
let s2 = s.clone();
let sup = spawn(move || {
let s3 = s2.clone();
OneForOne::new()
.child(ChildSpec::new(Restart::Transient, move || {
s3.fetch_add(1, Ordering::SeqCst);
let log = Log::default();
let r = GenServerBuilder::new(Srv::new(&log, false, ShutdownAction::Exit))
.under(self_pid())
.start();
r.cast(Cast::StopNow).unwrap();
// Block until the server is gone; a bare spawn parent
// returning would not itself end the server.
let mon = monitor(r.pid());
let _ = mon.rx.recv();
}))
.run();
});
sup.join().unwrap(); // returns only if the child was not restarted forever
});
assert_eq!(starts.load(Ordering::SeqCst), 1);
}
#[test]
fn linked_peer_death_reaches_handle_exit() {
let log = Log::default();
let l = log.clone();
let alive = Arc::new(AtomicBool::new(false));
let a = alive.clone();
run(move || {
let r = spawn_settled(Srv::new(&l, true, ShutdownAction::Exit));
let pid = r.pid();
let peer = spawn(move || {
link(pid);
panic!("peer dies");
});
let _ = peer.join();
sleep(Duration::from_millis(20));
r.cast(Cast::Note("still-serving")).unwrap();
sleep(Duration::from_millis(20));
a.store(true, Ordering::SeqCst);
r.shutdown(); // graceful; waits for terminate
});
assert!(alive.load(Ordering::SeqCst));
assert_eq!(
log.get(),
vec![
"handle_exit",
"still-serving",
"handle_shutdown",
"terminate"
]
);
}
#[test]
fn gen_server_ref_shutdown_is_graceful_for_a_trapping_server() {
let log = Log::default();
let l = log.clone();
run(move || {
let r = spawn_settled(Srv::new(&l, true, ShutdownAction::Exit));
r.shutdown();
});
assert_eq!(log.get(), vec!["handle_shutdown", "terminate"]);
}
+88 -12
View File
@@ -100,7 +100,15 @@ fn state_timeout_fires() {
let got = Arc::new(Mutex::new(0u32)); let got = Arc::new(Mutex::new(0u32));
let got2 = got.clone(); let got2 = got.clone();
run(move || { run(move || {
let m = TimerSm::start(T::Idle, TData { enters: 0, st_fires: 0, named_fires: 0, st_window: 5 }); let m = TimerSm::start(
T::Idle,
TData {
enters: 0,
st_fires: 0,
named_fires: 0,
st_window: 5,
},
);
m.send(Ev2::Cast(TCast::Arm)).unwrap(); // -> Armed, arms 5ms state-timeout m.send(Ev2::Cast(TCast::Arm)).unwrap(); // -> Armed, arms 5ms state-timeout
smarm::sleep(Duration::from_millis(40)); // let it fire smarm::sleep(Duration::from_millis(40)); // let it fire
*got2.lock().unwrap() = m.call(|r| Ev2::Call(TCall::StFires(r))).unwrap(); *got2.lock().unwrap() = m.call(|r| Ev2::Call(TCall::StFires(r))).unwrap();
@@ -116,13 +124,25 @@ fn state_timeout_auto_resets_on_transition() {
let got2 = got.clone(); let got2 = got.clone();
run(move || { run(move || {
// Long window so the explicit Disarm beats it comfortably. // Long window so the explicit Disarm beats it comfortably.
let m = TimerSm::start(T::Idle, TData { enters: 0, st_fires: 0, named_fires: 0, st_window: 50 }); let m = TimerSm::start(
T::Idle,
TData {
enters: 0,
st_fires: 0,
named_fires: 0,
st_window: 50,
},
);
m.send(Ev2::Cast(TCast::Arm)).unwrap(); // -> Armed, arms 50ms state-timeout m.send(Ev2::Cast(TCast::Arm)).unwrap(); // -> Armed, arms 50ms state-timeout
m.send(Ev2::Cast(TCast::Disarm)).unwrap(); // -> Idle, auto-resets it m.send(Ev2::Cast(TCast::Disarm)).unwrap(); // -> Idle, auto-resets it
smarm::sleep(Duration::from_millis(80)); // past the original window smarm::sleep(Duration::from_millis(80)); // past the original window
*got2.lock().unwrap() = m.call(|r| Ev2::Call(TCall::StFires(r))).unwrap(); *got2.lock().unwrap() = m.call(|r| Ev2::Call(TCall::StFires(r))).unwrap();
}); });
assert_eq!(*got.lock().unwrap(), 0, "auto-reset cancelled the pending state-timeout"); assert_eq!(
*got.lock().unwrap(),
0,
"auto-reset cancelled the pending state-timeout"
);
} }
// A named timeout survives a state change: armed in Idle, it still fires after // A named timeout survives a state change: armed in Idle, it still fires after
@@ -133,14 +153,26 @@ fn named_timeout_survives_transition() {
let got2 = got.clone(); let got2 = got.clone();
run(move || { run(move || {
// Armed's own state-timeout is long so it doesn't interfere. // Armed's own state-timeout is long so it doesn't interfere.
let m = TimerSm::start(T::Idle, TData { enters: 0, st_fires: 0, named_fires: 0, st_window: 200 }); let m = TimerSm::start(
T::Idle,
TData {
enters: 0,
st_fires: 0,
named_fires: 0,
st_window: 200,
},
);
m.send(Ev2::Cast(TCast::Ping(20))).unwrap(); // arm "ping" for 20ms (in Idle) m.send(Ev2::Cast(TCast::Ping(20))).unwrap(); // arm "ping" for 20ms (in Idle)
m.send(Ev2::Cast(TCast::Arm)).unwrap(); // -> Armed (ping must survive this) m.send(Ev2::Cast(TCast::Arm)).unwrap(); // -> Armed (ping must survive this)
m.send(Ev2::Cast(TCast::Disarm)).unwrap(); // -> Idle (and this) m.send(Ev2::Cast(TCast::Disarm)).unwrap(); // -> Idle (and this)
smarm::sleep(Duration::from_millis(60)); // let "ping" fire smarm::sleep(Duration::from_millis(60)); // let "ping" fire
*got2.lock().unwrap() = m.call(|r| Ev2::Call(TCall::NamedFires(r))).unwrap(); *got2.lock().unwrap() = m.call(|r| Ev2::Call(TCall::NamedFires(r))).unwrap();
}); });
assert_eq!(*got.lock().unwrap(), 1, "named timeout fired across the transitions"); assert_eq!(
*got.lock().unwrap(),
1,
"named timeout fired across the transitions"
);
} }
// Cancelling a named timeout before its window prevents the fire. // Cancelling a named timeout before its window prevents the fire.
@@ -149,13 +181,25 @@ fn named_timeout_cancel() {
let got = Arc::new(Mutex::new(99u32)); let got = Arc::new(Mutex::new(99u32));
let got2 = got.clone(); let got2 = got.clone();
run(move || { run(move || {
let m = TimerSm::start(T::Idle, TData { enters: 0, st_fires: 0, named_fires: 0, st_window: 200 }); let m = TimerSm::start(
T::Idle,
TData {
enters: 0,
st_fires: 0,
named_fires: 0,
st_window: 200,
},
);
m.send(Ev2::Cast(TCast::Ping(30))).unwrap(); // arm "ping" for 30ms m.send(Ev2::Cast(TCast::Ping(30))).unwrap(); // arm "ping" for 30ms
m.send(Ev2::Cast(TCast::CancelPing)).unwrap(); // cancel before it fires m.send(Ev2::Cast(TCast::CancelPing)).unwrap(); // cancel before it fires
smarm::sleep(Duration::from_millis(60)); // past the original window smarm::sleep(Duration::from_millis(60)); // past the original window
*got2.lock().unwrap() = m.call(|r| Ev2::Call(TCall::NamedFires(r))).unwrap(); *got2.lock().unwrap() = m.call(|r| Ev2::Call(TCall::NamedFires(r))).unwrap();
}); });
assert_eq!(*got.lock().unwrap(), 0, "cancel prevented the named-timeout fire"); assert_eq!(
*got.lock().unwrap(),
0,
"cancel prevented the named-timeout fire"
);
} }
// =========================================================================== // ===========================================================================
@@ -170,7 +214,15 @@ fn cast_then_call_roundtrip() {
let got2 = got.clone(); let got2 = got.clone();
run(move || { run(move || {
// Long state-timeout window so it never fires during the test. // Long state-timeout window so it never fires during the test.
let m = TimerSm::start(T::Idle, TData { enters: 0, st_fires: 0, named_fires: 0, st_window: 10_000 }); let m = TimerSm::start(
T::Idle,
TData {
enters: 0,
st_fires: 0,
named_fires: 0,
st_window: 10_000,
},
);
m.send(Ev2::Cast(TCast::Arm)).unwrap(); // Idle -> Armed (enter) m.send(Ev2::Cast(TCast::Arm)).unwrap(); // Idle -> Armed (enter)
m.send(Ev2::Cast(TCast::Disarm)).unwrap(); // Armed -> Idle (enter) m.send(Ev2::Cast(TCast::Disarm)).unwrap(); // Armed -> Idle (enter)
m.send(Ev2::Cast(TCast::Arm)).unwrap(); // Idle -> Armed (enter) m.send(Ev2::Cast(TCast::Arm)).unwrap(); // Idle -> Armed (enter)
@@ -178,7 +230,11 @@ fn cast_then_call_roundtrip() {
// enters = 1 (start) + 4 transitions = 5. // enters = 1 (start) + 4 transitions = 5.
*got2.lock().unwrap() = m.call(|r| Ev2::Call(TCall::Enters(r))).unwrap(); *got2.lock().unwrap() = m.call(|r| Ev2::Call(TCall::Enters(r))).unwrap();
}); });
assert_eq!(*got.lock().unwrap(), 5, "one enter on start, one per real transition"); assert_eq!(
*got.lock().unwrap(),
5,
"one enter on start, one per real transition"
);
} }
// `enter` fires once on start and once per *real* transition; a stay (a call // `enter` fires once on start and once per *real* transition; a stay (a call
@@ -188,7 +244,15 @@ fn enter_on_start_and_each_transition_but_not_stay() {
let got = Arc::new(Mutex::new((0u32, 0u32, 0u32))); let got = Arc::new(Mutex::new((0u32, 0u32, 0u32)));
let got2 = got.clone(); let got2 = got.clone();
run(move || { run(move || {
let m = TimerSm::start(T::Idle, TData { enters: 0, st_fires: 0, named_fires: 0, st_window: 10_000 }); // enter -> 1 let m = TimerSm::start(
T::Idle,
TData {
enters: 0,
st_fires: 0,
named_fires: 0,
st_window: 10_000,
},
); // enter -> 1
let after_start = m.call(|r| Ev2::Call(TCall::Enters(r))).unwrap(); let after_start = m.call(|r| Ev2::Call(TCall::Enters(r))).unwrap();
// A stay (a counter read returns `prev`) must not bump enters. // A stay (a counter read returns `prev`) must not bump enters.
let _ = m.call(|r| Ev2::Call(TCall::StFires(r))).unwrap(); let _ = m.call(|r| Ev2::Call(TCall::StFires(r))).unwrap();
@@ -208,7 +272,15 @@ fn call_to_panicking_handler_is_down() {
let got = Arc::new(Mutex::new(None::<Result<u32, CallError>>)); let got = Arc::new(Mutex::new(None::<Result<u32, CallError>>));
let got2 = got.clone(); let got2 = got.clone();
run(move || { run(move || {
let m = TimerSm::start(T::Idle, TData { enters: 0, st_fires: 0, named_fires: 0, st_window: 10_000 }); let m = TimerSm::start(
T::Idle,
TData {
enters: 0,
st_fires: 0,
named_fires: 0,
st_window: 10_000,
},
);
let r = m.call(|rep| Ev2::Call(TCall::Boom(rep))); let r = m.call(|rep| Ev2::Call(TCall::Boom(rep)));
*got2.lock().unwrap() = Some(r); *got2.lock().unwrap() = Some(r);
}); });
@@ -299,7 +371,11 @@ fn postponed_call_answered_after_transition() {
smarm::sleep(Duration::from_millis(20)); // let the child wake with its reply smarm::sleep(Duration::from_millis(20)); // let the child wake with its reply
*g2.lock().unwrap() = *taken.lock().unwrap(); *g2.lock().unwrap() = *taken.lock().unwrap();
}); });
assert_eq!(*got.lock().unwrap(), Some(42), "postponed call answered by the Filled state"); assert_eq!(
*got.lock().unwrap(),
Some(42),
"postponed call answered by the Filled state"
);
} }
#[derive(Clone, Copy, PartialEq, Eq, Debug)] #[derive(Clone, Copy, PartialEq, Eq, Debug)]
+128
View File
@@ -0,0 +1,128 @@
//! gen_statem lifetime parity with gen_server: a machine lives until it
//! stops, is shut down, or is killed — its refs are addresses. And
//! `gen_statem::run_named` runs a machine inline as the current actor, so it
//! is a direct `ChildSpec` child addressed by name.
use smarm::gen_statem::{self, GenStatemName, Reply};
use smarm::supervisor::{ChildSpec, OneForOne, Restart, Shutdown};
use smarm::{monitor, request_shutdown, run, sleep, spawn, DownReason};
use std::sync::{Arc, Mutex};
use std::time::Duration;
#[derive(Default, Clone)]
struct Log(Arc<Mutex<Vec<&'static str>>>);
impl Log {
fn push(&self, e: &'static str) {
self.0.lock().unwrap().push(e);
}
fn get(&self) -> Vec<&'static str> {
self.0.lock().unwrap().clone()
}
}
#[derive(Clone, Copy, PartialEq, Eq, Debug)]
enum S {
On,
}
struct D {
log: Log,
trap: bool,
n: u64,
}
enum Cast {
Inc,
StopNow,
}
enum Call {
Get(Reply<u64>),
}
smarm::gen_statem! {
machine: Sm { state: S, data: D };
event: Ev { cast: Cast, call: Call, info: () };
context(data, prev, cx);
enter {
S::On => { data.log.push("enter"); if data.trap { cx.trap_exit() } },
}
on S::On => {
cast Cast::Inc => { data.n += 1; prev },
cast Cast::StopNow => stop,
call Call::Get(r) => { r.reply(data.n); prev },
shutdown => { data.log.push("shutdown"); cx.stop(); prev },
state_timeout => unhandled,
timeout _ => unhandled,
}
terminate { data.log.push("terminate"); }
}
fn d(log: &Log, trap: bool) -> D {
D {
log: log.clone(),
trap,
n: 0,
}
}
#[test]
fn dropping_last_ref_does_not_end_machine() {
let log = Log::default();
let l = log.clone();
run(move || {
let m = Sm::start(S::On, d(&l, true));
let pid = m.pid();
m.send(Ev::Cast(Cast::Inc)).unwrap();
assert_eq!(m.call(|r| Ev::Call(Call::Get(r))).unwrap(), 1);
let mon = monitor(pid);
drop(m);
sleep(Duration::from_millis(30));
assert!(
mon.rx.try_recv().unwrap().is_none(),
"machine must outlive its refs"
);
assert_eq!(l.get(), vec!["enter"]);
request_shutdown(pid);
assert_eq!(mon.rx.recv().unwrap().reason, DownReason::Exit);
});
assert_eq!(log.get(), vec!["enter", "shutdown", "terminate"]);
}
const SM: GenStatemName<Sm> = GenStatemName::new("lifetime-sm");
#[test]
fn run_named_is_a_direct_supervised_child() {
let log = Log::default();
let l = log.clone();
run(move || {
let l2 = l.clone();
let sup = spawn(move || {
let l3 = l2.clone();
OneForOne::new()
.child(
ChildSpec::new(Restart::Permanent, move || {
gen_statem::run_named(SM, Sm::new(S::On, d(&l3, true))).expect("name free");
})
.shutdown(Shutdown::Infinity),
)
.run();
});
sleep(Duration::from_millis(10));
gen_statem::send(SM, Ev::Cast(Cast::Inc)).unwrap();
assert_eq!(gen_statem::call(SM, |r| Ev::Call(Call::Get(r))).unwrap(), 1);
// Normal self-exit → Permanent restart → fresh data, same name.
gen_statem::send(SM, Ev::Cast(Cast::StopNow)).unwrap();
sleep(Duration::from_millis(30));
assert_eq!(gen_statem::call(SM, |r| Ev::Call(Call::Get(r))).unwrap(), 0);
request_shutdown(sup.pid());
sup.join().unwrap();
assert!(gen_statem::whereis_machine(SM).is_none());
});
assert_eq!(
log.get(),
vec!["enter", "terminate", "enter", "shutdown", "terminate"]
);
}
+219
View File
@@ -0,0 +1,219 @@
//! gen_statem graceful shutdown — the gen_server surface, in state-machine
//! clothes. Where gen_server routes a shutdown request to a `handle_shutdown`
//! method, a gen_statem gets it as an **event** so it can be routed by state:
//!
//! - A machine that does not opt in (`cx.trap_exit()` in the initial `enter`)
//! is stopped outright by `request_shutdown`, exactly as by `request_stop`.
//! - A trapping machine sees the request as a `shutdown` row (a unit event
//! like `state_timeout`). The macro's default, when a state writes no
//! `shutdown` row, is `stop` — the loop breaks and `terminate` runs on the
//! normal path. A row may instead transition (e.g. into a Draining state)
//! and `stop` later from any row via the `stop` tail keyword.
//! - Linked-peer deaths reach a trapping machine as `exit <pat>` rows; an
//! unmatched exit is silently dropped, like an unmatched info.
//! - `terminate { … }` is an optional macro block, run on every exit path.
use smarm::gen_statem;
use smarm::gen_statem::{GenStatemRef, Reply};
use smarm::{link, monitor, request_shutdown, run, sleep, spawn, DownReason, ExitSignal};
use std::sync::{Arc, Mutex};
use std::time::Duration;
#[derive(Default, Clone)]
struct Log(Arc<Mutex<Vec<&'static str>>>);
impl Log {
fn push(&self, e: &'static str) {
self.0.lock().unwrap().push(e);
}
fn get(&self) -> Vec<&'static str> {
self.0.lock().unwrap().clone()
}
}
#[derive(Clone, Copy, PartialEq, Eq, Debug)]
enum S {
Idle,
Draining,
}
struct D {
log: Log,
trap: bool,
exits: Vec<ExitSignal>,
}
enum Cast {
Note(&'static str),
StopNow,
}
enum Call {
Exits(Reply<usize>),
}
gen_statem! {
machine: Sm { state: S, data: D };
event: Ev { cast: Cast, call: Call, info: () };
context(data, prev, cx);
enter {
S::Idle => if data.trap { cx.trap_exit() },
S::Draining => { data.log.push("draining"); cx.state_timeout(Duration::from_millis(30)); },
}
on S::Idle => {
// Shutdown in Idle: go drain first, stop later.
shutdown => S::Draining,
cast Cast::StopNow => stop,
state_timeout => unhandled,
}
on S::Draining => {
// Drained: end the machine normally.
state_timeout => { data.log.push("drained"); cx.stop(); prev },
// A second request while draining is ignored.
shutdown => unhandled,
cast Cast::StopNow => stop,
}
on _ => {
cast Cast::Note(s) => { data.log.push(s); prev },
call Call::Exits(r) => { r.reply(data.exits.len()); prev },
exit sig => { data.log.push("exit"); data.exits.push(sig); prev },
timeout _ => unhandled,
}
terminate {
data.log.push("terminate");
}
}
/// A machine with no `shutdown` rows at all: the macro default applies.
#[derive(Clone, Copy, PartialEq, Eq, Debug)]
enum P {
On,
}
struct PD {
log: Log,
}
enum PCast {}
enum PCall {}
gen_statem! {
machine: Plain { state: P, data: PD };
event: PEv { cast: PCast, call: PCall, info: () };
context(data, prev, cx);
enter { P::On => cx.trap_exit(), }
on P::On => {
cast _ => unhandled,
call _ => unhandled,
state_timeout => unhandled,
timeout _ => unhandled,
}
terminate { data.log.push("terminate"); }
}
fn settled(log: &Log, trap: bool) -> GenStatemRef<Sm> {
let r = Sm::start(
S::Idle,
D {
log: log.clone(),
trap,
exits: Vec::new(),
},
);
sleep(Duration::from_millis(20)); // let on_start (trap_exit) run
r
}
#[test]
fn non_trapping_machine_is_stopped_outright() {
let log = Log::default();
let l = log.clone();
run(move || {
let r = settled(&l, false);
let mon = monitor(r.pid());
request_shutdown(r.pid());
assert_eq!(mon.rx.recv().unwrap().reason, DownReason::Stopped);
});
assert_eq!(log.get(), vec!["terminate"]);
}
#[test]
fn shutdown_row_routes_by_state_and_stop_tail_exits_normally() {
let log = Log::default();
let l = log.clone();
run(move || {
let r = settled(&l, true);
let mon = monitor(r.pid());
request_shutdown(r.pid());
// The second request lands in Draining and is `unhandled` (ignored).
sleep(Duration::from_millis(5));
request_shutdown(r.pid());
assert_eq!(mon.rx.recv().unwrap().reason, DownReason::Exit);
});
assert_eq!(log.get(), vec!["draining", "drained", "terminate"]);
}
#[test]
fn default_shutdown_is_stop() {
let log = Log::default();
let l = log.clone();
run(move || {
let r = Plain::start(P::On, PD { log: l });
sleep(Duration::from_millis(20));
let mon = monitor(r.pid());
request_shutdown(r.pid());
assert_eq!(mon.rx.recv().unwrap().reason, DownReason::Exit);
});
assert_eq!(log.get(), vec!["terminate"]);
}
#[test]
fn stop_tail_from_a_cast_is_a_normal_exit() {
let log = Log::default();
let l = log.clone();
run(move || {
let r = settled(&l, false);
let mon = monitor(r.pid());
r.send(Ev::Cast(Cast::Note("a"))).unwrap();
r.send(Ev::Cast(Cast::StopNow)).unwrap();
r.send(Ev::Cast(Cast::Note("after-stop"))).unwrap(); // never dispatched
assert_eq!(mon.rx.recv().unwrap().reason, DownReason::Exit);
});
assert_eq!(log.get(), vec!["a", "terminate"]);
}
#[test]
fn linked_peer_death_reaches_exit_row() {
let log = Log::default();
let l = log.clone();
run(move || {
let r = settled(&l, true);
let pid = r.pid();
let peer = spawn(move || {
link(pid);
panic!("peer dies");
});
let _ = peer.join();
sleep(Duration::from_millis(20));
r.send(Ev::Cast(Cast::Note("still-running"))).unwrap();
assert_eq!(r.call(|r| Ev::Call(Call::Exits(r))).unwrap(), 1);
r.shutdown();
});
assert_eq!(
log.get(),
vec!["exit", "still-running", "draining", "drained", "terminate"]
);
}
#[test]
fn ref_shutdown_is_graceful_and_waits() {
let log = Log::default();
let l = log.clone();
run(move || {
let r = settled(&l, true);
r.shutdown();
// terminate has run by the time shutdown() returns.
assert_eq!(l.get(), vec!["draining", "drained", "terminate"]);
});
}
+162 -5
View File
@@ -56,7 +56,11 @@ fn snapshot_lists_actors_with_parent_edge() {
// The root itself is on-CPU (it's running this code) and rooted under // The root itself is on-CPU (it's running this code) and rooted under
// the forest sentinel. // the forest sentinel.
let root = snap.actors.iter().find(|a| a.pid == me).expect("root present"); let root = snap
.actors
.iter()
.find(|a| a.pid == me)
.expect("root present");
assert_eq!(root.state, ActorState::Running); assert_eq!(root.state, ActorState::Running);
assert_eq!(root.supervisor, smarm::Pid::new(u32::MAX, u32::MAX)); assert_eq!(root.supervisor, smarm::Pid::new(u32::MAX, u32::MAX));
@@ -200,7 +204,11 @@ fn tree_places_child_under_its_spawner() {
// The root is parented at the forest sentinel, so it's a genuine root, // The root is parented at the forest sentinel, so it's a genuine root,
// and the worker it spawned hangs beneath it. // and the worker it spawned hangs beneath it.
let root = t.roots.iter().find(|n| n.info.pid == me).expect("root in forest"); let root = t
.roots
.iter()
.find(|n| n.info.pid == me)
.expect("root in forest");
assert!(!root.orphaned); assert!(!root.orphaned);
assert!( assert!(
root.children.iter().any(|c| c.info.pid == h.pid()), root.children.iter().any(|c| c.info.pid == h.pid()),
@@ -237,6 +245,13 @@ fn tree_from_nests_children_and_reroots_orphans() {
overruns: 0, overruns: 0,
messages_received: 0, messages_received: 0,
budget_cycles: 0, budget_cycles: 0,
stack: smarm::StackInfo {
reserve: 0,
guard: 0,
depth_high_water: 0,
parks_since_shrink: 0,
shrinks: 0,
},
}; };
let snap = RuntimeSnapshot { let snap = RuntimeSnapshot {
@@ -251,14 +266,25 @@ fn tree_from_nests_children_and_reroots_orphans() {
let t = tree_from(snap); let t = tree_from(snap);
assert_eq!(t.roots.len(), 2); assert_eq!(t.roots.len(), 2);
let root = t.roots.iter().find(|n| n.info.pid == root_pid).expect("root present"); let root = t
.roots
.iter()
.find(|n| n.info.pid == root_pid)
.expect("root present");
assert!(!root.orphaned); assert!(!root.orphaned);
assert_eq!(root.children.len(), 1); assert_eq!(root.children.len(), 1);
assert_eq!(root.children[0].info.pid, child); assert_eq!(root.children[0].info.pid, child);
assert!(!root.children[0].orphaned); assert!(!root.children[0].orphaned);
let o = t.roots.iter().find(|n| n.info.pid == orphan).expect("orphan re-rooted"); let o = t
assert!(o.orphaned, "an actor whose parent is absent must be flagged orphaned"); .roots
.iter()
.find(|n| n.info.pid == orphan)
.expect("orphan re-rooted");
assert!(
o.orphaned,
"an actor whose parent is absent must be flagged orphaned"
);
assert!(o.children.is_empty()); assert!(o.children.is_empty());
} }
@@ -352,3 +378,134 @@ fn budget_cycles_accumulate_when_enabled() {
h.join().unwrap(); h.join().unwrap();
}); });
} }
// ---------------------------------------------------------------------------
// RFC 019 §8 — the stack introspection surface.
// ---------------------------------------------------------------------------
/// Burn ~`frames` × 4 KiB of stack with a yield at max depth, so the context
/// save samples the high-water there (RFC 019 §2: hwm is SAMPLED at
/// deschedule, not tracked continuously).
#[inline(never)]
fn burn_stack_yielding(frames: usize) -> u64 {
let mut local = [0u8; 4096];
local[0] = frames as u8;
let below = if frames == 0 {
smarm::yield_now();
0
} else {
burn_stack_yielding(frames - 1)
};
std::hint::black_box(&mut local);
below.wrapping_add(local[0] as u64)
}
#[test]
fn stack_info_reports_defaults_and_sampled_depth() {
run(|| {
let (ready_tx, ready_rx) = channel::<()>();
let (gate_tx, gate_rx) = channel::<()>();
let h = spawn(move || {
// ~32 KiB deep with a yield at the bottom: the sample point.
std::hint::black_box(burn_stack_yielding(8));
ready_tx.send(()).unwrap();
gate_rx.recv().unwrap();
});
ready_rx.recv().unwrap();
let info = spin_until(h.pid(), |a| a.state == ActorState::Parked);
let s = info.stack;
assert_eq!(s.reserve, 64 * 1024, "default reserve");
assert_eq!(
s.guard,
1024 * 1024,
"default guard (kernel stack_guard_gap convention)"
);
assert!(
s.depth_high_water >= 8 * 4096,
"hwm sampled at the deep yield: expected ≥ 32 KiB, got {}",
s.depth_high_water
);
assert!(
s.depth_high_water < s.reserve,
"depth {} cannot exceed the reserve {}",
s.depth_high_water,
s.reserve
);
// Parked at the gate right now, never shrunk (64 KiB reserve cannot
// cross the shrink threshold).
assert!(s.parks_since_shrink >= 1, "the gate park must be counted");
assert_eq!(s.shrinks, 0);
gate_tx.send(()).unwrap();
h.join().unwrap();
});
}
#[test]
fn stack_info_shrink_counters_are_live() {
use smarm::runtime::{Config, SHRINK_COOLDOWN, SHRINK_THRESHOLD};
use smarm::{spawn_with, SpawnOpts};
let rt = smarm::runtime::init(Config::exact(1));
rt.run(|| {
let (park_tx, park_rx) = channel::<()>();
let spike = 768 * 4096;
assert!(spike > SHRINK_THRESHOLD);
let worker = spawn_with(
SpawnOpts {
stack_reserve: Some(8 * 1024 * 1024),
..SpawnOpts::default()
},
move || {
std::hint::black_box(burn_stack_yielding(768));
for _ in 0..(SHRINK_COOLDOWN + 8) {
park_rx.recv().unwrap();
}
},
);
let wpid = worker.pid();
// Before any parks complete: the spike depth is visible.
let info = spin_until(wpid, |a| a.state == ActorState::Parked);
assert!(
info.stack.depth_high_water >= spike,
"spike should be sampled: {} < {spike}",
info.stack.depth_high_water
);
// Cross the cooldown, then read the counters live while the worker
// is parked waiting for the remaining rounds (post-join the slot is
// reclaimed and the generation check correctly hides it).
for _ in 0..(SHRINK_COOLDOWN + 2) {
spin_until(wpid, |a| a.state == ActorState::Parked);
park_tx.send(()).unwrap();
}
let info = spin_until(wpid, |a| {
a.state == ActorState::Parked && a.stack.shrinks >= 1
});
let s = info.stack;
assert!(
s.shrinks >= 1,
"cooldown was crossed with a spike above threshold"
);
assert!(
s.parks_since_shrink < SHRINK_COOLDOWN,
"counter must reset at shrink: {}",
s.parks_since_shrink
);
assert!(
s.depth_high_water < spike,
"hwm resets to the shallow park sp at shrink; got {}",
s.depth_high_water
);
for _ in 0..6 {
spin_until(wpid, |a| a.state == ActorState::Parked);
park_tx.send(()).unwrap();
}
worker.join().unwrap();
});
}
+13 -3
View File
@@ -56,8 +56,16 @@ fn other_actors_run_while_block_on_io_is_in_flight() {
let pos_2 = v.iter().position(|&x| x == 2).unwrap(); let pos_2 = v.iter().position(|&x| x == 2).unwrap();
let pos_3 = v.iter().position(|&x| x == 3).unwrap(); let pos_3 = v.iter().position(|&x| x == 3).unwrap();
let pos_4 = v.iter().position(|&x| x == 4).unwrap(); let pos_4 = v.iter().position(|&x| x == 4).unwrap();
assert!(pos_2 < pos_4, "B's first step ran after A resumed: {:?}", *v); assert!(
assert!(pos_3 < pos_4, "B's second step ran after A resumed: {:?}", *v); pos_2 < pos_4,
"B's first step ran after A resumed: {:?}",
*v
);
assert!(
pos_3 < pos_4,
"B's second step ran after A resumed: {:?}",
*v
);
} }
#[test] #[test]
@@ -76,7 +84,9 @@ fn many_concurrent_block_on_io_calls_all_complete() {
cc.fetch_add(n, Ordering::SeqCst); cc.fetch_add(n, Ordering::SeqCst);
})); }));
} }
for h in handles { h.join().unwrap(); } for h in handles {
h.join().unwrap();
}
}); });
assert_eq!(counter.load(Ordering::SeqCst), 10); assert_eq!(counter.load(Ordering::SeqCst), 10);
} }
+11 -4
View File
@@ -144,8 +144,7 @@ fn write_sugar_sends_bytes_to_pipe() {
// Pipe is empty + has buffer space, so this returns immediately // Pipe is empty + has buffer space, so this returns immediately
// after wait_writable wakes (which happens fast because the // after wait_writable wakes (which happens fast because the
// kernel marks an empty pipe as immediately writable). // kernel marks an empty pipe as immediately writable).
let n = smarm::scheduler::write(p_writer.write, b"smarm") let n = smarm::scheduler::write(p_writer.write, b"smarm").expect("write failed");
.expect("write failed");
assert_eq!(n, 5); assert_eq!(n, 5);
c.fetch_add(1, Ordering::SeqCst); c.fetch_add(1, Ordering::SeqCst);
}); });
@@ -209,10 +208,18 @@ fn other_actors_run_while_one_is_parked_on_wait_readable() {
let pos_lit_a = v.iter().position(|&c| c == b'a').unwrap(); let pos_lit_a = v.iter().position(|&c| c == b'a').unwrap();
let big_b_count = v.iter().filter(|&&c| c == b'B').count(); let big_b_count = v.iter().filter(|&&c| c == b'B').count();
assert_eq!(big_b_count, 3, "B should have made 3 steps: {:?}", *v); assert_eq!(big_b_count, 3, "B should have made 3 steps: {:?}", *v);
assert!(pos_big_a < pos_lit_a, "A pre-park before A post-park: {:?}", *v); assert!(
pos_big_a < pos_lit_a,
"A pre-park before A post-park: {:?}",
*v
);
// At least the last B step should be before A resumes. // At least the last B step should be before A resumes.
let last_big_b = v.iter().rposition(|&c| c == b'B').unwrap(); let last_big_b = v.iter().rposition(|&c| c == b'B').unwrap();
assert!(last_big_b < pos_lit_a, "B should finish before A resumes: {:?}", *v); assert!(
last_big_b < pos_lit_a,
"B should finish before A resumes: {:?}",
*v
);
} }
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
+8 -2
View File
@@ -57,7 +57,10 @@ fn linked_pair_one_panics_other_is_stopped() {
panic!("boom"); panic!("boom");
}); });
let dn = down_b.rx.recv().expect("monitor channel closed before Down"); let dn = down_b
.rx
.recv()
.expect("monitor channel closed before Down");
assert_eq!(dn.pid, b, "Down reported the wrong pid"); assert_eq!(dn.pid, b, "Down reported the wrong pid");
if matches!(dn.reason, DownReason::Stopped) { if matches!(dn.reason, DownReason::Stopped) {
s.store(true, Ordering::SeqCst); s.store(true, Ordering::SeqCst);
@@ -152,7 +155,10 @@ fn link_to_dead_pid_stops_a_nontrapping_caller() {
}); });
let b = hb.pid(); let b = hb.pid();
let down_b = monitor(b); let down_b = monitor(b);
let dn = down_b.rx.recv().expect("monitor channel closed before Down"); let dn = down_b
.rx
.recv()
.expect("monitor channel closed before Down");
if matches!(dn.reason, DownReason::Stopped) { if matches!(dn.reason, DownReason::Stopped) {
s.store(true, Ordering::SeqCst); s.store(true, Ordering::SeqCst);
} }
+27 -6
View File
@@ -67,7 +67,10 @@ fn monitor_already_dead_target_is_noproc() {
// and its generation bumped, so `pid` is now stale. // and its generation bumped, so `pid` is now stale.
h.join().unwrap(); h.join().unwrap();
let down = monitor(pid); let down = monitor(pid);
let d = down.rx.recv().expect("NoProc Down should be delivered immediately"); let d = down
.rx
.recv()
.expect("NoProc Down should be delivered immediately");
assert_eq!(d.pid, pid); assert_eq!(d.pid, pid);
if matches!(d.reason, DownReason::NoProc) { if matches!(d.reason, DownReason::NoProc) {
o.store(true, Ordering::SeqCst); o.store(true, Ordering::SeqCst);
@@ -91,7 +94,11 @@ fn multiple_monitors_all_notified() {
} }
} }
}); });
assert_eq!(count.load(Ordering::SeqCst), 3, "every monitor should see the Down"); assert_eq!(
count.load(Ordering::SeqCst),
3,
"every monitor should see the Down"
);
} }
#[test] #[test]
@@ -103,8 +110,15 @@ fn demonitor_stops_delivery() {
let h = spawn(|| {}); let h = spawn(|| {});
let pid = h.pid(); let pid = h.pid();
let m = monitor(pid); let m = monitor(pid);
assert_eq!(demonitor(&m), Some(m.id), "live registration should be removed"); assert_eq!(
assert!(m.rx.recv().is_err(), "no Down should arrive after demonitor"); demonitor(&m),
Some(m.id),
"live registration should be removed"
);
assert!(
m.rx.recv().is_err(),
"no Down should arrive after demonitor"
);
let _ = h.join(); let _ = h.join();
}); });
} }
@@ -122,7 +136,10 @@ fn demonitor_one_of_many() {
let _ = h.join(); let _ = h.join();
assert!(matches!(ms[0].rx.recv().unwrap().reason, DownReason::Exit)); assert!(matches!(ms[0].rx.recv().unwrap().reason, DownReason::Exit));
assert!(matches!(ms[2].rx.recv().unwrap().reason, DownReason::Exit)); assert!(matches!(ms[2].rx.recv().unwrap().reason, DownReason::Exit));
assert!(ms[1].rx.recv().is_err(), "demonitored channel should be closed"); assert!(
ms[1].rx.recv().is_err(),
"demonitored channel should be closed"
);
}); });
} }
@@ -136,7 +153,11 @@ fn demonitor_after_fire_is_none() {
let m = monitor(pid); let m = monitor(pid);
let d = m.rx.recv().expect("Down before close"); let d = m.rx.recv().expect("Down before close");
assert!(matches!(d.reason, DownReason::Exit)); assert!(matches!(d.reason, DownReason::Exit));
assert_eq!(demonitor(&m), None, "already-fired monitor has nothing to remove"); assert_eq!(
demonitor(&m),
None,
"already-fired monitor has nothing to remove"
);
let _ = h.join(); let _ = h.join();
}); });
} }
+16 -4
View File
@@ -3,9 +3,9 @@
//! needs to be able to park. //! needs to be able to park.
use smarm::{run, spawn, yield_now, LockTimeout, Mutex}; use smarm::{run, spawn, yield_now, LockTimeout, Mutex};
use std::sync::atomic::{AtomicU32, Ordering};
use std::sync::Arc; use std::sync::Arc;
use std::sync::Mutex as StdMutex; use std::sync::Mutex as StdMutex;
use std::sync::atomic::{AtomicU32, Ordering};
use std::time::{Duration, Instant}; use std::time::{Duration, Instant};
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -111,8 +111,16 @@ fn contended_lock_parks_until_holder_releases() {
let pos_b_locked = v.iter().position(|s| *s == "B_locked").unwrap(); let pos_b_locked = v.iter().position(|s| *s == "B_locked").unwrap();
assert!(pos_a_locked < pos_b_try, "log: {:?}", *v); assert!(pos_a_locked < pos_b_try, "log: {:?}", *v);
assert!(pos_b_try < pos_a_dropped, "B should attempt before A drops: {:?}", *v); assert!(
assert!(pos_a_dropped < pos_b_locked, "B should lock only after A drops: {:?}", *v); pos_b_try < pos_a_dropped,
"B should attempt before A drops: {:?}",
*v
);
assert!(
pos_a_dropped < pos_b_locked,
"B should lock only after A drops: {:?}",
*v
);
} }
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -209,7 +217,11 @@ fn waiters_are_granted_the_lock_in_fifo_order() {
}); });
let v = order.lock().unwrap().clone(); let v = order.lock().unwrap().clone();
assert_eq!(v, vec![1, 2, 3, 4], "waiters should acquire in arrival order"); assert_eq!(
v,
vec![1, 2, 3, 4],
"waiters should acquire in arrival order"
);
} }
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
+5 -3
View File
@@ -77,8 +77,7 @@ fn observer_reports_none_for_a_forged_pid() {
// An index that is not in the slab at all — the verb relays the // An index that is not in the slab at all — the verb relays the
// primitive's `None` faithfully. // primitive's `None` faithfully.
let forged = smarm::Pid::new(u32::MAX - 1, 0); let forged = smarm::Pid::new(u32::MAX - 1, 0);
let ObserverReply::ActorInfo(none) = let ObserverReply::ActorInfo(none) = obs.call(ObserverRequest::ActorInfo(forged)).unwrap()
obs.call(ObserverRequest::ActorInfo(forged)).unwrap()
else { else {
panic!("ActorInfo verb must reply ActorInfo"); panic!("ActorInfo verb must reply ActorInfo");
}; };
@@ -113,7 +112,10 @@ fn observer_sees_a_parked_actor_as_parked() {
} }
smarm::yield_now(); smarm::yield_now();
} }
assert!(parked, "observer should eventually report the worker as Parked"); assert!(
parked,
"observer should eventually report the worker as Parked"
);
gate_tx.send(()).unwrap(); gate_tx.send(()).unwrap();
worker.join().unwrap(); worker.join().unwrap();
+69
View File
@@ -0,0 +1,69 @@
//! RFC 018 scheduler park/wake — observable-behavior guards.
//!
//! These pin the two timer-latency properties the park/wake swap must
//! preserve or introduce:
//!
//! - `sleep_fires_under_saturation`: due timers fire even when every
//! scheduler is busy (nobody parked ⇒ no timekeeper) — the busy-path
//! due-check, ratified design point (a). The old drain phase gave this
//! for free (timers drained every loop iteration); the new design must
//! not lose it.
//! - `submillisecond_sleep_is_prompt`: a sub-ms sleep completes promptly.
//! Under the old wake pipe, `poll_wake`'s `as_millis` truncation turned
//! sub-ms deadlines into 0ms busy-polls (correct wall time, pathological
//! CPU); under park/wake the futex timespec carries full nanosecond
//! precision.
use std::sync::atomic::{AtomicBool, Ordering};
use std::sync::Arc;
use std::time::{Duration, Instant};
#[test]
fn sleep_fires_under_saturation() {
let rt = smarm::runtime::init(smarm::runtime::Config::exact(4));
rt.run(|| {
let stop = Arc::new(AtomicBool::new(false));
let mut spinners = Vec::new();
// 8 spinners over 4 schedulers: the run queue never empties, so no
// scheduler ever parks and no timekeeper exists. Only the busy-path
// due-check can fire the sleeper's timer before the spinners quit.
for _ in 0..8 {
let stop = stop.clone();
spinners.push(smarm::spawn(move || {
let t0 = Instant::now();
while !stop.load(Ordering::Relaxed) && t0.elapsed() < Duration::from_secs(5) {
smarm::yield_now();
}
}));
}
let t0 = Instant::now();
smarm::sleep(Duration::from_millis(10));
let dt = t0.elapsed();
stop.store(true, Ordering::Relaxed);
for s in spinners {
let _ = s.join();
}
assert!(
dt < Duration::from_millis(500),
"10ms sleep took {dt:?} under scheduler saturation — busy-path \
timer firing is broken (timekeeper-only firing stalls under load)"
);
});
}
#[test]
fn submillisecond_sleep_is_prompt() {
let rt = smarm::runtime::init(smarm::runtime::Config::exact(2));
rt.run(|| {
// Warm one iteration, then measure.
smarm::sleep(Duration::from_micros(500));
let t0 = Instant::now();
smarm::sleep(Duration::from_micros(500));
let dt = t0.elapsed();
assert!(dt >= Duration::from_micros(400), "woke early: {dt:?}");
assert!(
dt < Duration::from_millis(100),
"500µs sleep took {dt:?} — sub-ms deadline handling is broken"
);
});
}
+13 -3
View File
@@ -44,7 +44,10 @@ fn a_dead_actor_vanishes_from_every_group_it_joined() {
// Drain-on-contact: touching g1 detects the death and sweeps the pid // Drain-on-contact: touching g1 detects the death and sweeps the pid
// out of every group (g2 included), not just g1. // out of every group (g2 included), not just g1.
assert!(members("g1").is_empty(), "evicted from the touched group"); assert!(members("g1").is_empty(), "evicted from the touched group");
assert!(members("g2").is_empty(), "and swept from the untouched group"); assert!(
members("g2").is_empty(),
"and swept from the untouched group"
);
assert_eq!(pick("g1"), None); assert_eq!(pick("g1"), None);
}); });
} }
@@ -83,7 +86,11 @@ fn live_members_survive_a_peers_death() {
tx_a.send(()).unwrap(); tx_a.send(()).unwrap();
a.join().unwrap(); a.join().unwrap();
assert_eq!(members("svc"), vec![b.pid()], "only the dead peer is reaped"); assert_eq!(
members("svc"),
vec![b.pid()],
"only the dead peer is reaped"
);
assert_eq!(pick("svc"), Some(b.pid())); assert_eq!(pick("svc"), Some(b.pid()));
tx_b.send(()).unwrap(); tx_b.send(()).unwrap();
@@ -125,7 +132,10 @@ fn joining_an_already_dead_pid_is_evicted_on_next_contact() {
// monitor() on a gone pid queues a NoProc Down immediately, so the // monitor() on a gone pid queues a NoProc Down immediately, so the
// membership is reaped the next time the group is touched. // membership is reaped the next time the group is touched.
join("late", pid); join("late", pid);
assert!(members("late").is_empty(), "dead-at-join member is reaped on read"); assert!(
members("late").is_empty(),
"dead-at-join member is reaped on read"
);
assert_eq!(pick("late"), None); assert_eq!(pick("late"), None);
}); });
} }
+10 -2
View File
@@ -41,7 +41,11 @@ fn stop_storm_does_not_poison_runtime() {
} }
c.fetch_add(1, Ordering::SeqCst); c.fetch_add(1, Ordering::SeqCst);
}); });
assert_eq!(completed.load(Ordering::SeqCst), 1, "root completed cleanly"); assert_eq!(
completed.load(Ordering::SeqCst),
1,
"root completed cleanly"
);
} }
/// The sharper repro: a stop-flagged actor whose *next allocation* is the /// The sharper repro: a stop-flagged actor whose *next allocation* is the
@@ -85,5 +89,9 @@ fn self_stop_during_spawn_does_not_poison_shared_mutex() {
} }
c.fetch_add(1, Ordering::SeqCst); c.fetch_add(1, Ordering::SeqCst);
}); });
assert_eq!(completed.load(Ordering::SeqCst), 1, "root completed cleanly"); assert_eq!(
completed.load(Ordering::SeqCst),
1,
"root completed cleanly"
);
} }
+15 -4
View File
@@ -43,10 +43,21 @@ fn check_yields_when_timeslice_expired() {
let pos_big_b = v.iter().position(|&c| c == b'B').unwrap(); let pos_big_b = v.iter().position(|&c| c == b'B').unwrap();
let pos_lit_a = v.iter().position(|&c| c == b'a').unwrap(); let pos_lit_a = v.iter().position(|&c| c == b'a').unwrap();
let pos_lit_b = v.iter().position(|&c| c == b'b').unwrap(); let pos_lit_b = v.iter().position(|&c| c == b'b').unwrap();
assert!(pos_big_a < pos_lit_a, "A's tail ran before B's head: {:?}", *v); assert!(
assert!(pos_big_b < pos_lit_b, "B's tail ran before A's head: {:?}", *v); pos_big_a < pos_lit_a,
assert!(pos_big_a.max(pos_big_b) < pos_lit_a.min(pos_lit_b), "A's tail ran before B's head: {:?}",
"preemption didn't interleave: {:?}", *v); *v
);
assert!(
pos_big_b < pos_lit_b,
"B's tail ran before A's head: {:?}",
*v
);
assert!(
pos_big_a.max(pos_big_b) < pos_lit_a.min(pos_lit_b),
"preemption didn't interleave: {:?}",
*v
);
} }
#[test] #[test]
+12 -3
View File
@@ -65,7 +65,10 @@ fn name_held_by_live_actor_is_taken() {
ready_rx.recv().unwrap(); ready_rx.recv().unwrap();
// Root tries to claim a live actor's name for itself -> NameTaken. // Root tries to claim a live actor's name for itself -> NameTaken.
let (tx_b, _rx_b) = channel::<u64>(); let (tx_b, _rx_b) = channel::<u64>();
assert_eq!(register(SVC, tx_b), Err(RegisterError::NameTaken { holder: a.pid() })); assert_eq!(
register(SVC, tx_b),
Err(RegisterError::NameTaken { holder: a.pid() })
);
send(SVC, 0).unwrap(); // release a (delivers to the holder, a) send(SVC, 0).unwrap(); // release a (delivers to the holder, a)
a.join().unwrap(); a.join().unwrap();
}); });
@@ -105,7 +108,10 @@ fn dead_holder_is_pruned_and_name_taken_over() {
fn send_errors_unresolved_and_no_channel() { fn send_errors_unresolved_and_no_channel() {
run(|| { run(|| {
// No actor at all. // No actor at all.
assert!(matches!(send(Name::<u64>::new("ghost"), 1u64), Err(SendError::Unresolved(_)))); assert!(matches!(
send(Name::<u64>::new("ghost"), 1u64),
Err(SendError::Unresolved(_))
));
let (ready_tx, ready_rx) = channel::<()>(); let (ready_tx, ready_rx) = channel::<()>();
let (tx, rx) = channel::<u64>(); let (tx, rx) = channel::<u64>();
@@ -228,7 +234,10 @@ fn send_dyn_delivers_and_reports_wrong_type() {
let p = h.pid(); // a bare Pid<Erased>, as if recovered off a Down let p = h.pid(); // a bare Pid<Erased>, as if recovered off a Down
send_dyn::<u64>(p, 3u64).unwrap(); // right type: delivered send_dyn::<u64>(p, 3u64).unwrap(); // right type: delivered
// Live actor, but it has no channel for &str — the genuinely-fallible case. // Live actor, but it has no channel for &str — the genuinely-fallible case.
assert!(matches!(send_dyn::<&'static str>(p, "nope"), Err(SendError::NoChannel(_)))); assert!(matches!(
send_dyn::<&'static str>(p, "nope"),
Err(SendError::NoChannel(_))
));
done_tx.send(()).unwrap(); done_tx.send(()).unwrap();
h.join().unwrap(); h.join().unwrap();
}); });
+288
View File
@@ -0,0 +1,288 @@
//! Root exit — the run's initial actor returning means "the program is done".
//!
//! When the root finalizes, the runtime delivers `request_shutdown` to every
//! **forest root**: each live actor whose parent is the run itself (a plain
//! `spawn` from the root closure) or is already dead. Nothing below a live
//! parent is touched directly — a supervisor gets one request and runs its
//! own ordered shutdown per child `Shutdown` policy.
//!
//! - Non-trapping actors are stopped outright, exactly as by
//! `request_shutdown` — a `spawn(|| { sleep(..); work() })` the root did
//! not `join` does NOT get to finish. Join it, supervise it, or trap.
//! - Trapping actors get `handle_shutdown` / an `ExitSignal{Shutdown}` and
//! may keep running (`Continue`, drain, then stop themselves) — timers and
//! all; the run ends when they do. There is no second, forcing sweep.
//! - A periodic-timer daemon (the classic wedge) never blocks `run()`.
use smarm::gen_server::{start, GenServer, GenServerCtx, ShutdownAction, StopHandle, TimerHandle};
use smarm::supervisor::{ChildSpec, OneForOne, Restart, Shutdown};
use smarm::{run, sleep, spawn, trap_exit, DownReason};
use std::sync::atomic::{AtomicBool, AtomicUsize, Ordering};
use std::sync::{Arc, Mutex};
use std::time::{Duration, Instant};
fn assert_prompt(start: Instant, what: &str) {
assert!(
start.elapsed() < Duration::from_secs(2),
"{what}: run() took {:?}",
start.elapsed()
);
}
// ---------------------------------------------------------------------------
// Bare (non-gen_server) actors
// ---------------------------------------------------------------------------
/// A non-trapping sleeper the root did not join is stopped, not waited for.
#[test]
fn unjoined_non_trapping_sleeper_is_stopped() {
let finished = Arc::new(AtomicBool::new(false));
let f = finished.clone();
let t = Instant::now();
run(move || {
spawn(move || {
sleep(Duration::from_secs(5));
f.store(true, Ordering::SeqCst);
});
});
assert_prompt(t, "sleeper");
assert!(
!finished.load(Ordering::SeqCst),
"sleeper should have been stopped"
);
}
/// A trapping bare actor sees `Shutdown` from the root's exit and may keep
/// working — here it sleeps (a timer!) after the signal, then returns. The
/// run waits for it: no forcing sweep.
#[test]
fn trapping_actor_may_finish_after_shutdown_signal() {
let finished = Arc::new(AtomicBool::new(false));
let f = finished.clone();
run(move || {
spawn(move || {
let inbox = trap_exit();
let sig = inbox.recv().expect("shutdown signal");
assert_eq!(sig.reason, DownReason::Shutdown);
sleep(Duration::from_millis(100));
f.store(true, Ordering::SeqCst);
});
// Trapping is a runtime opt-in: give the actor a chance to run
// `trap_exit()`; a not-yet-run actor is non-trapping and is stopped.
sleep(Duration::from_millis(20));
});
assert!(
finished.load(Ordering::SeqCst),
"trapping actor must be allowed to finish"
);
}
/// The classic wedge: a lazily spawned daemon that never returns on its own.
#[test]
fn parked_forever_daemon_does_not_block_run() {
let t = Instant::now();
run(|| {
let (_tx, rx) = smarm::channel::<()>();
spawn(move || {
let _ = rx.recv(); // parked forever: sender is held by the root, which returns
});
// Leak the sender into the daemon's own scope so nothing else drops it.
std::mem::forget(_tx);
});
assert_prompt(t, "daemon");
}
// ---------------------------------------------------------------------------
// gen_servers
// ---------------------------------------------------------------------------
/// A non-trapping ticker with a periodic timer: the timer wheel is never empty,
/// and root exit must still end the run.
struct Ticker {
ticks: Arc<AtomicUsize>,
}
impl GenServer for Ticker {
type Call = ();
type Reply = ();
type Cast = ();
type Info = ();
type Timer = ();
fn init(&mut self, ctx: &GenServerCtx<Self>) {
ctx.timer().tick_every(Duration::from_millis(5), ());
}
fn handle_call(&mut self, _: ()) {}
fn handle_cast(&mut self, _: ()) {}
fn handle_timer(&mut self, _: ()) {
self.ticks.fetch_add(1, Ordering::SeqCst);
}
}
#[test]
fn periodic_timer_daemon_does_not_block_run() {
let ticks = Arc::new(AtomicUsize::new(0));
let tk = ticks.clone();
let t = Instant::now();
run(move || {
let _r = start(Ticker { ticks: tk });
sleep(Duration::from_millis(50));
});
assert_prompt(t, "ticker");
assert!(
ticks.load(Ordering::SeqCst) >= 3,
"ticker should have ticked"
);
}
/// A trapping server that answers `Continue`, keeps ticking on its own timer
/// (draining), and stops itself later. Root exit must not cut it short.
struct Drainer {
log: Arc<Mutex<Vec<&'static str>>>,
shutdowns: Arc<AtomicUsize>,
ticks_after_shutdown: usize,
stop: Option<StopHandle<Drainer>>,
timer: Option<TimerHandle<Drainer>>,
draining: bool,
}
impl GenServer for Drainer {
type Call = ();
type Reply = ();
type Cast = ();
type Info = ();
type Timer = ();
fn init(&mut self, ctx: &GenServerCtx<Self>) {
ctx.trap_exit();
self.stop = Some(ctx.stop_handle());
self.timer = Some(ctx.timer());
}
fn handle_call(&mut self, _: ()) {}
fn handle_cast(&mut self, _: ()) {}
fn handle_shutdown(&mut self) -> ShutdownAction {
self.shutdowns.fetch_add(1, Ordering::SeqCst);
self.log.lock().unwrap().push("handle_shutdown");
self.draining = true;
self.timer
.as_ref()
.unwrap()
.tick_every(Duration::from_millis(10), ());
ShutdownAction::Continue
}
fn handle_timer(&mut self, _: ()) {
if !self.draining {
return;
}
self.ticks_after_shutdown += 1;
if self.ticks_after_shutdown == 3 {
self.log.lock().unwrap().push("drained");
self.stop.as_ref().unwrap().stop();
}
}
fn terminate(&mut self) {
self.log.lock().unwrap().push("terminate");
}
}
fn drainer(log: &Arc<Mutex<Vec<&'static str>>>, shutdowns: &Arc<AtomicUsize>) -> Drainer {
Drainer {
log: log.clone(),
shutdowns: shutdowns.clone(),
ticks_after_shutdown: 0,
stop: None,
timer: None,
draining: false,
}
}
#[test]
fn trapping_server_drains_with_timers_after_root_exit() {
let log = Arc::new(Mutex::new(Vec::new()));
let shutdowns = Arc::new(AtomicUsize::new(0));
let (l, s) = (log.clone(), shutdowns.clone());
run(move || {
let r = start(drainer(&l, &s));
// A gen_server's lifetime is governed by its refs: dropping the last
// one closes the inbox and ends the loop cleanly, which would cut the
// drain short for a reason unrelated to root exit. Pin it the way a
// registered name would.
std::mem::forget(r);
sleep(Duration::from_millis(20)); // let init (trap_exit) run
});
assert_eq!(
*log.lock().unwrap(),
vec!["handle_shutdown", "drained", "terminate"]
);
assert_eq!(shutdowns.load(Ordering::SeqCst), 1);
}
// ---------------------------------------------------------------------------
// Supervision trees: only forest roots are addressed
// ---------------------------------------------------------------------------
/// The supervisor gets ONE request and runs its ordered shutdown; a trapping
/// child under it sees exactly one `Shutdown` — from the supervisor, not a
/// second one from the runtime — and is allowed to finish its drain (a sleep,
/// i.e. a timer) under `Shutdown::Infinity`.
#[test]
fn supervised_children_are_shut_down_only_via_their_supervisor() {
let signals = Arc::new(AtomicUsize::new(0));
let drained = Arc::new(AtomicBool::new(false));
let sup_returned = Arc::new(AtomicBool::new(false));
let (sg, dr, sr) = (signals.clone(), drained.clone(), sup_returned.clone());
run(move || {
spawn(move || {
OneForOne::new()
.child(
ChildSpec::new(Restart::Permanent, move || {
let inbox = trap_exit();
while let Ok(sig) = inbox.recv() {
if sig.reason == DownReason::Shutdown {
sg.fetch_add(1, Ordering::SeqCst);
sleep(Duration::from_millis(100));
// A late second signal would land here.
while let Ok(Some(sig)) = inbox.try_recv() {
if sig.reason == DownReason::Shutdown {
sg.fetch_add(1, Ordering::SeqCst);
}
}
dr.store(true, Ordering::SeqCst);
return;
}
}
})
.shutdown(Shutdown::Infinity),
)
.run();
sr.store(true, Ordering::SeqCst);
});
sleep(Duration::from_millis(30)); // let the tree settle
});
assert!(
sup_returned.load(Ordering::SeqCst),
"supervisor should return normally"
);
assert!(
drained.load(Ordering::SeqCst),
"child should finish its drain"
);
assert_eq!(signals.load(Ordering::SeqCst), 1);
}
/// A supervised non-trapping child under `Shutdown::Timeout` is stopped by
/// the supervisor's policy, and the run ends promptly.
#[test]
fn supervisor_tree_is_torn_down_promptly_on_root_exit() {
let t = Instant::now();
run(|| {
spawn(|| {
OneForOne::new()
.child(
ChildSpec::new(Restart::Permanent, || loop {
sleep(Duration::from_millis(5));
})
.shutdown(Shutdown::Timeout(Duration::from_millis(50))),
)
.run();
});
sleep(Duration::from_millis(30));
});
assert_prompt(t, "tree");
}
+36
View File
@@ -0,0 +1,36 @@
//! Under `smarm-trace`, every actor the root-exit sweep reaches is recorded
//! as a `root_sweep` event — the way to *see* unsupervised leftovers. One test
//! per binary: the trace file is process-global.
#![cfg(feature = "smarm-trace")]
use smarm::gen_server::{self, GenServer, GenServerCtx};
use smarm::run;
struct Quiet;
impl GenServer for Quiet {
type Call = ();
type Reply = ();
type Cast = ();
type Info = ();
type Timer = ();
fn init(&mut self, _: &GenServerCtx<Self>) {}
fn handle_call(&mut self, _: ()) {}
fn handle_cast(&mut self, _: ()) {}
}
#[test]
fn forgotten_server_shows_up_as_root_sweep() {
let path = std::env::temp_dir().join(format!("smarm_root_sweep_{}.json", std::process::id()));
std::env::set_var("SMARM_TRACE_FILE", &path);
run(|| {
let srv = gen_server::start(Quiet);
srv.call(()).unwrap();
drop(srv); // forgotten: nobody supervises it, nobody holds it
});
let trace = std::fs::read_to_string(&path).expect("trace file written");
let _ = std::fs::remove_file(&path);
assert!(
trace.contains("root_sweep stopped"),
"expected a root_sweep line for the non-trapping leftover; got:\n{trace}"
);
}
+108 -23
View File
@@ -14,10 +14,17 @@
//! - No slot leaks under high spawn/join churn //! - No slot leaks under high spawn/join churn
//! - Panic on one scheduler thread doesn't kill others //! - Panic on one scheduler thread doesn't kill others
use smarm::{channel, runtime::{Config, Runtime}, spawn, yield_now, JoinHandle}; use smarm::{
use std::sync::{atomic::{AtomicBool, AtomicU64, Ordering}, Arc}; channel,
use std::time::Duration; runtime::{Config, Runtime},
spawn, yield_now, JoinHandle,
};
use std::collections::HashSet; use std::collections::HashSet;
use std::sync::{
atomic::{AtomicBool, AtomicU64, Ordering},
Arc,
};
use std::time::Duration;
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// Helpers // Helpers
@@ -29,7 +36,9 @@ fn rt(n: usize) -> Runtime {
} }
/// Convenient single-threaded runtime (regression guard). /// Convenient single-threaded runtime (regression guard).
fn rt1() -> Runtime { rt(1) } fn rt1() -> Runtime {
rt(1)
}
/// Multi-threaded runtime using all available parallelism. /// Multi-threaded runtime using all available parallelism.
fn rt_par() -> Runtime { fn rt_par() -> Runtime {
@@ -79,7 +88,9 @@ fn config_min_1_max_1_is_single_threaded() {
fn runtime_run_executes_closure() { fn runtime_run_executes_closure() {
let flag = Arc::new(AtomicBool::new(false)); let flag = Arc::new(AtomicBool::new(false));
let f = flag.clone(); let f = flag.clone();
rt(1).run(move || { f.store(true, Ordering::SeqCst); }); rt(1).run(move || {
f.store(true, Ordering::SeqCst);
});
assert!(flag.load(Ordering::SeqCst)); assert!(flag.load(Ordering::SeqCst));
} }
@@ -111,8 +122,12 @@ fn runtime_can_be_used_multiple_times_sequentially() {
let b = Arc::new(AtomicU64::new(0)); let b = Arc::new(AtomicU64::new(0));
let ac = a.clone(); let ac = a.clone();
let bc = b.clone(); let bc = b.clone();
r.run(move || { ac.fetch_add(1, Ordering::SeqCst); }); r.run(move || {
r.run(move || { bc.fetch_add(1, Ordering::SeqCst); }); ac.fetch_add(1, Ordering::SeqCst);
});
r.run(move || {
bc.fetch_add(1, Ordering::SeqCst);
});
assert_eq!(a.load(Ordering::SeqCst), 1); assert_eq!(a.load(Ordering::SeqCst), 1);
assert_eq!(b.load(Ordering::SeqCst), 1); assert_eq!(b.load(Ordering::SeqCst), 1);
} }
@@ -126,7 +141,9 @@ fn exact_1_spawn_join_works() {
let v = Arc::new(AtomicU64::new(0)); let v = Arc::new(AtomicU64::new(0));
let vc = v.clone(); let vc = v.clone();
rt1().run(move || { rt1().run(move || {
let h = spawn(move || { vc.store(42, Ordering::SeqCst); }); let h = spawn(move || {
vc.store(42, Ordering::SeqCst);
});
h.join().unwrap(); h.join().unwrap();
}); });
assert_eq!(v.load(Ordering::SeqCst), 42); assert_eq!(v.load(Ordering::SeqCst), 42);
@@ -155,7 +172,9 @@ fn exact_1_panic_captured() {
let s = saw_err.clone(); let s = saw_err.clone();
rt1().run(move || { rt1().run(move || {
let h = spawn(|| panic!("oops")); let h = spawn(|| panic!("oops"));
if h.join().is_err() { s.store(true, Ordering::SeqCst); } if h.join().is_err() {
s.store(true, Ordering::SeqCst);
}
}); });
assert!(saw_err.load(Ordering::SeqCst)); assert!(saw_err.load(Ordering::SeqCst));
} }
@@ -176,7 +195,9 @@ fn multi_thread_all_actors_complete() {
cc.fetch_add(1, Ordering::SeqCst); cc.fetch_add(1, Ordering::SeqCst);
})); }));
} }
for h in handles { h.join().unwrap(); } for h in handles {
h.join().unwrap();
}
}); });
assert_eq!(counter.load(Ordering::SeqCst), 100); assert_eq!(counter.load(Ordering::SeqCst), 100);
} }
@@ -221,7 +242,9 @@ fn multi_thread_many_channels_no_lost_wakeups() {
tx.send(1).unwrap(); tx.send(1).unwrap();
})); }));
} }
for h in handles { h.join().unwrap(); } for h in handles {
h.join().unwrap();
}
}); });
assert_eq!(count.load(Ordering::SeqCst), PAIRS as u64); assert_eq!(count.load(Ordering::SeqCst), PAIRS as u64);
} }
@@ -247,7 +270,9 @@ fn multi_thread_mutex_contention_no_deadlock() {
} }
})); }));
} }
for h in handles { h.join().unwrap(); } for h in handles {
h.join().unwrap();
}
let g = m.lock_timeout(Duration::from_secs(1)).unwrap(); let g = m.lock_timeout(Duration::from_secs(1)).unwrap();
t.store(*g, Ordering::SeqCst); t.store(*g, Ordering::SeqCst);
}); });
@@ -262,7 +287,9 @@ fn multi_thread_join_across_threads() {
rt_par().run(move || { rt_par().run(move || {
let h = spawn(move || { let h = spawn(move || {
// Do some work to make scheduling interesting. // Do some work to make scheduling interesting.
for _ in 0..10 { yield_now(); } for _ in 0..10 {
yield_now();
}
vc.store(1, Ordering::SeqCst); vc.store(1, Ordering::SeqCst);
}); });
h.join().unwrap(); h.join().unwrap();
@@ -279,8 +306,7 @@ fn multi_thread_join_across_threads() {
#[test] #[test]
fn actors_run_on_multiple_os_threads() { fn actors_run_on_multiple_os_threads() {
let thread_ids: Arc<smarm::Mutex<HashSet<u64>>> = let thread_ids: Arc<smarm::Mutex<HashSet<u64>>> = Arc::new(smarm::Mutex::new(HashSet::new()));
Arc::new(smarm::Mutex::new(HashSet::new()));
rt_par().run({ rt_par().run({
let ids = thread_ids.clone(); let ids = thread_ids.clone();
@@ -294,11 +320,15 @@ fn actors_run_on_multiple_os_threads() {
g.insert(tid); g.insert(tid);
})); }));
} }
for h in handles { h.join().unwrap(); } for h in handles {
h.join().unwrap();
}
} }
}); });
let n = std::thread::available_parallelism().map(|n| n.get()).unwrap_or(1); let n = std::thread::available_parallelism()
.map(|n| n.get())
.unwrap_or(1);
let ids = thread_ids.lock_timeout(Duration::from_secs(1)).unwrap(); let ids = thread_ids.lock_timeout(Duration::from_secs(1)).unwrap();
// If we have >1 scheduler threads, we expect >1 OS thread IDs. // If we have >1 scheduler threads, we expect >1 OS thread IDs.
@@ -326,11 +356,17 @@ fn scheduler_stats_run_queue_len_is_observable() {
// run() completes (queue len == 0 at quiescence). // run() completes (queue len == 0 at quiescence).
let r = rt_par(); let r = rt_par();
r.run(|| { r.run(|| {
for _ in 0..10 { spawn(|| {}); } for _ in 0..10 {
spawn(|| {});
}
// Don't join — let them drain naturally. // Don't join — let them drain naturally.
}); });
let stats = r.stats(); let stats = r.stats();
assert_eq!(stats.total_run_queue_len(), 0, "queue should be empty after run()"); assert_eq!(
stats.total_run_queue_len(),
0,
"queue should be empty after run()"
);
} }
#[test] #[test]
@@ -359,7 +395,9 @@ fn panic_in_actor_does_not_kill_runtime() {
})); }));
} }
let _ = bad.join(); // expect Err let _ = bad.join(); // expect Err
for h in good_handles { h.join().unwrap(); } for h in good_handles {
h.join().unwrap();
}
}); });
assert_eq!(completed.load(Ordering::SeqCst), 10); assert_eq!(completed.load(Ordering::SeqCst), 10);
} }
@@ -379,7 +417,9 @@ fn no_slot_leak_under_churn() {
rt_par().run(move || { rt_par().run(move || {
for _ in 0..500 { for _ in 0..500 {
let cc = c.clone(); let cc = c.clone();
spawn(move || { cc.fetch_add(1, Ordering::SeqCst); }) spawn(move || {
cc.fetch_add(1, Ordering::SeqCst);
})
.join() .join()
.unwrap(); .unwrap();
} }
@@ -474,7 +514,11 @@ fn multi_thread_timer_only_no_pipe_contention() {
} }
}); });
assert_eq!(count.load(Ordering::SeqCst), ACTORS as u64, "not all actors completed"); assert_eq!(
count.load(Ordering::SeqCst),
ACTORS as u64,
"not all actors completed"
);
let elapsed = start.elapsed(); let elapsed = start.elapsed();
assert!( assert!(
@@ -515,5 +559,46 @@ fn runtime_reusable_after_root_panic() {
let ran = Arc::new(AtomicBool::new(false)); let ran = Arc::new(AtomicBool::new(false));
let ran_t = ran.clone(); let ran_t = ran.clone();
r.run(move || ran_t.store(true, Ordering::Relaxed)); r.run(move || ran_t.store(true, Ordering::Relaxed));
assert!(ran.load(Ordering::Relaxed), "runtime unusable after root panic"); assert!(
ran.load(Ordering::Relaxed),
"runtime unusable after root panic"
);
}
// ---------------------------------------------------------------------------
// RFC 019 — Config stack knobs
// ---------------------------------------------------------------------------
/// Burn ~`frames` × 4 KiB of stack; probestack touches pages in order so
/// exceeding the reserve would hit the guard and SIGSEGV the process.
#[inline(never)]
fn burn_stack(frames: usize) -> u64 {
let mut local = [0u8; 4096];
local[0] = frames as u8;
let below = if frames == 0 {
0
} else {
burn_stack(frames - 1)
};
std::hint::black_box(&mut local);
below.wrapping_add(local[0] as u64)
}
#[test]
fn config_stack_reserve_permits_deep_recursion() {
// ~256 KiB of frames: four times the old fixed 64 KiB reserve. With
// Config::stack_reserve raised this must complete; before RFC 019 it
// could only segfault.
let rt = smarm::runtime::init(Config::exact(1).stack_reserve(1024 * 1024));
let done = Arc::new(AtomicBool::new(false));
let done2 = done.clone();
rt.run(move || {
spawn(move || {
std::hint::black_box(burn_stack(64));
done2.store(true, Ordering::SeqCst);
})
.join()
.unwrap();
});
assert!(done.load(Ordering::SeqCst));
} }
+7 -4
View File
@@ -14,7 +14,9 @@ use std::sync::Arc;
fn root_actor_runs() { fn root_actor_runs() {
let captured = Arc::new(AtomicI64::new(0)); let captured = Arc::new(AtomicI64::new(0));
let c = captured.clone(); let c = captured.clone();
run(move || { c.store(99, Ordering::SeqCst); }); run(move || {
c.store(99, Ordering::SeqCst);
});
assert_eq!(captured.load(Ordering::SeqCst), 99); assert_eq!(captured.load(Ordering::SeqCst), 99);
} }
@@ -27,7 +29,9 @@ fn spawn_and_join_returns_exit() {
let captured = Arc::new(AtomicI64::new(0)); let captured = Arc::new(AtomicI64::new(0));
let c = captured.clone(); let c = captured.clone();
run(move || { run(move || {
let h = spawn(move || { c.store(7, Ordering::SeqCst); }); let h = spawn(move || {
c.store(7, Ordering::SeqCst);
});
let res = h.join(); let res = h.join();
assert!(res.is_ok(), "join returned {:?}", res); assert!(res.is_ok(), "join returned {:?}", res);
}); });
@@ -68,8 +72,7 @@ fn yield_now_interleaves_actors() {
#[test] #[test]
fn self_pid_is_stable_within_an_actor() { fn self_pid_is_stable_within_an_actor() {
let pid_cell: Arc<std::sync::Mutex<Option<smarm::Pid>>> = let pid_cell: Arc<std::sync::Mutex<Option<smarm::Pid>>> = Arc::new(std::sync::Mutex::new(None));
Arc::new(std::sync::Mutex::new(None));
let p2 = pid_cell.clone(); let p2 = pid_cell.clone();
run(move || { run(move || {
let h = spawn(move || { let h = spawn(move || {
+14 -3
View File
@@ -19,7 +19,12 @@ fn ready_arm_returns_immediately_without_parking() {
txa.send(42).unwrap(); txa.send(42).unwrap();
let i = select(&[&rxb, &rxa]); let i = select(&[&rxb, &rxa]);
assert_eq!(i, 1); assert_eq!(i, 1);
out2.store(rxa.try_recv().unwrap().expect("ready arm must hold a message"), Ordering::SeqCst); out2.store(
rxa.try_recv()
.unwrap()
.expect("ready arm must hold a message"),
Ordering::SeqCst,
);
}); });
assert_eq!(out.load(Ordering::SeqCst), 42); assert_eq!(out.load(Ordering::SeqCst), 42);
} }
@@ -276,7 +281,10 @@ fn select_timeout_ready_arm_wins_without_arming_a_timer() {
let (txa, rxa) = channel::<i64>(); let (txa, rxa) = channel::<i64>();
let (_keep_b, rxb) = channel::<i64>(); let (_keep_b, rxb) = channel::<i64>();
txa.send(5).unwrap(); txa.send(5).unwrap();
assert_eq!(select_timeout(&[&rxb, &rxa], Duration::from_millis(500)), Some(1)); assert_eq!(
select_timeout(&[&rxb, &rxa], Duration::from_millis(500)),
Some(1)
);
assert_eq!(rxa.try_recv().unwrap(), Some(5)); assert_eq!(rxa.try_recv().unwrap(), Some(5));
}); });
} }
@@ -342,7 +350,10 @@ fn select_timeout_closed_arm_is_ready_not_a_timeout() {
let (_keep_a, rxa) = channel::<i64>(); let (_keep_a, rxa) = channel::<i64>();
let (txb, rxb) = channel::<i64>(); let (txb, rxb) = channel::<i64>();
drop(txb); drop(txb);
assert_eq!(select_timeout(&[&rxa, &rxb], Duration::from_millis(200)), Some(1)); assert_eq!(
select_timeout(&[&rxa, &rxb], Duration::from_millis(200)),
Some(1)
);
assert!(rxb.try_recv().is_err()); assert!(rxb.try_recv().is_err());
}); });
} }
+122
View File
@@ -0,0 +1,122 @@
//! Graceful shutdown — `request_shutdown` (OTP `exit(Pid, shutdown)`).
//!
//! `request_stop` is `exit(Pid, kill)`: an uncatchable unwind at the target's
//! next observation point. `request_shutdown` is the polite form:
//! - a target that is NOT trapping exits is stopped exactly as by
//! `request_stop` (OTP's rule: don't trap, you die);
//! - a target that IS trapping receives an `ExitSignal { reason: Shutdown }`
//! on its trap inbox and keeps running — it is expected to wind down and
//! exit normally on its own.
use smarm::{monitor, request_shutdown, run, self_pid, sleep, spawn, trap_exit, DownReason, Pid};
use std::sync::atomic::{AtomicBool, Ordering};
use std::sync::{mpsc, Arc};
use std::thread;
use std::time::Duration;
const WATCHDOG: Duration = Duration::from_secs(10);
#[test]
fn request_shutdown_stops_a_non_trapping_actor() {
run(|| {
let h = spawn(|| sleep(Duration::from_secs(3600)));
let mon = monitor(h.pid());
request_shutdown(h.pid());
let down = mon.rx.recv().expect("down");
assert_eq!(down.reason, DownReason::Stopped);
});
}
#[test]
fn request_shutdown_is_a_message_to_a_trapping_actor() {
let unwound = Arc::new(AtomicBool::new(false));
let u = unwound.clone();
run(move || {
struct Unwound(Arc<AtomicBool>);
impl Drop for Unwound {
fn drop(&mut self) {
if std::thread::panicking() {
self.0.store(true, Ordering::SeqCst);
}
}
}
let (tx, rx) = smarm::channel::<(Pid, DownReason)>();
let (ready_tx, ready_rx) = smarm::channel::<()>();
let h = spawn(move || {
let _g = Unwound(u);
let inbox = trap_exit();
let _ = ready_tx.send(());
let sig = inbox.recv().expect("exit signal");
let _ = tx.send((sig.from, sig.reason));
// Keep doing work after the request: shutdown is advisory.
sleep(Duration::from_millis(20));
});
// Trapping is set by the target itself; a request that beats it is a
// plain stop (same window as OTP's exit-before-process_flag).
ready_rx.recv().expect("ready");
let me = self_pid();
let mon = monitor(h.pid());
request_shutdown(h.pid());
let (from, reason) = rx.recv().expect("relayed");
assert_eq!(from, me);
assert_eq!(reason, DownReason::Shutdown);
let down = mon.rx.recv().expect("down");
assert_eq!(
down.reason,
DownReason::Exit,
"target exited normally, not stopped"
);
});
assert!(
!unwound.load(Ordering::SeqCst),
"trapping target must not be unwound"
);
}
#[test]
fn request_shutdown_on_dead_pid_is_a_no_op() {
run(|| {
let h = spawn(|| {});
let pid = h.pid();
let _ = h.join();
request_shutdown(pid); // must not panic
});
}
#[test]
fn handle_request_shutdown_from_foreign_thread() {
let rt = smarm::init(smarm::Config::exact(2));
let handle = rt.handle();
let (pid_tx, pid_rx) = mpsc::channel::<Pid>();
let requester = thread::spawn(move || {
let pid = pid_rx.recv().expect("pid");
thread::sleep(Duration::from_millis(50));
handle.request_shutdown(pid);
});
let (done_tx, done_rx) = mpsc::channel();
thread::spawn(move || {
rt.run(move || {
let (tx, rx) = smarm::channel::<DownReason>();
let (ready_tx, ready_rx) = smarm::channel::<()>();
let h = spawn(move || {
let inbox = trap_exit();
let _ = ready_tx.send(());
let sig = inbox.recv().expect("exit signal");
let _ = tx.send(sig.reason);
});
ready_rx.recv().expect("ready");
pid_tx.send(h.pid()).expect("send pid");
let reason = rx.recv().expect("relayed");
assert_eq!(reason, DownReason::Shutdown);
let _ = h.join();
});
let _ = done_tx.send(());
});
done_rx
.recv_timeout(WATCHDOG)
.expect("run did not return: foreign-thread request_shutdown never reached the target");
requester.join().expect("requester thread");
}
+235
View File
@@ -0,0 +1,235 @@
//! RFC 019 commit 2 — the `SpawnOpts` surface.
//!
//! Covers: per-spawn stack shape overrides on every spawn surface, the
//! `None ⇒ Config default` resolution, the pool rule from the outside
//! (obligation 4: a custom-shaped stack never enters the pool), and that a
//! big reserve behaviorally takes effect (deep recursion completes).
use smarm::runtime::{Config, DEFAULT_STACK_GUARD, DEFAULT_STACK_RESERVE};
use smarm::{self_pid, spawn, spawn_under_with, spawn_with, GenServerBuilder, SpawnOpts};
use std::sync::atomic::{AtomicBool, Ordering};
use std::sync::Arc;
fn rt1() -> smarm::runtime::Runtime {
smarm::runtime::init(Config::exact(1))
}
#[test]
fn default_spawn_has_default_shape() {
rt1().run(|| {
let h = spawn(|| {
let shape = smarm::introspect::stack_shape(self_pid()).unwrap();
assert_eq!(shape, (DEFAULT_STACK_RESERVE, DEFAULT_STACK_GUARD));
});
h.join().unwrap();
});
}
#[test]
fn spawn_with_overrides_reserve_and_guard() {
rt1().run(|| {
let opts = SpawnOpts {
stack_reserve: Some(1024 * 1024),
guard_size: Some(256 * 1024),
};
let h = spawn_with(opts, || {
let shape = smarm::introspect::stack_shape(self_pid()).unwrap();
assert_eq!(shape, (1024 * 1024, 256 * 1024));
});
h.join().unwrap();
});
}
#[test]
fn spawn_with_partial_override_keeps_config_default_for_the_rest() {
rt1().run(|| {
let opts = SpawnOpts {
stack_reserve: Some(1024 * 1024),
..SpawnOpts::default()
};
let h = spawn_with(opts, || {
let shape = smarm::introspect::stack_shape(self_pid()).unwrap();
assert_eq!(shape, (1024 * 1024, DEFAULT_STACK_GUARD));
});
h.join().unwrap();
});
}
#[test]
fn spawn_with_rounds_to_pages() {
rt1().run(|| {
let opts = SpawnOpts {
stack_reserve: Some(64 * 1024 + 1),
guard_size: Some(4097),
};
let h = spawn_with(opts, || {
let (reserve, guard) = smarm::introspect::stack_shape(self_pid()).unwrap();
assert_eq!(reserve % 4096, 0);
assert_eq!(guard % 4096, 0);
assert!(reserve >= 64 * 1024 + 1);
assert!(guard >= 4097);
});
h.join().unwrap();
});
}
#[test]
fn spawn_under_with_takes_opts() {
rt1().run(|| {
let me = self_pid();
let opts = SpawnOpts {
stack_reserve: Some(128 * 1024),
..SpawnOpts::default()
};
let h = spawn_under_with(me, opts, || {
let (reserve, _) = smarm::introspect::stack_shape(self_pid()).unwrap();
assert_eq!(reserve, 128 * 1024);
});
h.join().unwrap();
});
}
/// Obligation 4, from the outside: a dead custom stack must not be handed to
/// the next default spawn. The pool is LIFO, so if the custom stack had been
/// (wrongly) pushed at death, the very next default-shaped spawn on this
/// single-threaded runtime would pop it and report a custom shape.
#[test]
fn custom_stack_never_enters_the_pool() {
rt1().run(|| {
spawn_with(
SpawnOpts {
stack_reserve: Some(512 * 1024),
guard_size: Some(128 * 1024),
},
|| {},
)
.join()
.unwrap();
let h = spawn(|| {
let shape = smarm::introspect::stack_shape(self_pid()).unwrap();
assert_eq!(shape, (DEFAULT_STACK_RESERVE, DEFAULT_STACK_GUARD));
});
h.join().unwrap();
});
}
/// The reverse direction of the pool rule: a default-shaped stack IS pooled
/// and reused (cap = threads × 4 ≥ 1 here, pool empty at start).
#[test]
fn default_stack_is_recycled() {
rt1().run(|| {
spawn(|| {}).join().unwrap();
let h = spawn(|| {
let shape = smarm::introspect::stack_shape(self_pid()).unwrap();
assert_eq!(shape, (DEFAULT_STACK_RESERVE, DEFAULT_STACK_GUARD));
});
h.join().unwrap();
});
}
/// Burn ~`frames` × 4 KiB of stack (see tests/runtime.rs twin).
#[inline(never)]
fn burn_stack(frames: usize) -> u64 {
let mut local = [0u8; 4096];
local[0] = frames as u8;
let below = if frames == 0 {
0
} else {
burn_stack(frames - 1)
};
std::hint::black_box(&mut local);
below.wrapping_add(local[0] as u64)
}
#[test]
fn big_reserve_behaviorally_takes_effect() {
// ~1 MiB deep on an 8 MiB per-spawn reserve, runtime default untouched.
rt1().run(|| {
let done = Arc::new(AtomicBool::new(false));
let done2 = done.clone();
spawn_with(
SpawnOpts {
stack_reserve: Some(8 * 1024 * 1024),
..SpawnOpts::default()
},
move || {
std::hint::black_box(burn_stack(256));
done2.store(true, Ordering::SeqCst);
},
)
.join()
.unwrap();
assert!(done.load(Ordering::SeqCst));
});
}
// ---------------------------------------------------------------------------
// Builder surfaces
// ---------------------------------------------------------------------------
struct Echo;
impl smarm::GenServer for Echo {
type Call = ();
type Reply = (usize, usize);
type Cast = ();
type Info = ();
type Timer = ();
fn handle_call(&mut self, _c: ()) -> (usize, usize) {
smarm::introspect::stack_shape(self_pid()).unwrap()
}
fn handle_cast(&mut self, _c: ()) {}
}
#[test]
fn gen_server_builder_stack_opts() {
rt1().run(|| {
let server = GenServerBuilder::new(Echo)
.stack_opts(SpawnOpts {
stack_reserve: Some(256 * 1024),
..SpawnOpts::default()
})
.start();
let (reserve, guard) = server.call(()).unwrap();
assert_eq!(reserve, 256 * 1024);
assert_eq!(guard, DEFAULT_STACK_GUARD);
server.shutdown();
});
}
struct Probe;
impl smarm::Machine for Probe {
type Ev = smarm::channel::Sender<(usize, usize)>;
fn state_timeout_ev() -> Self::Ev {
unreachable!("no timers in this test")
}
fn timeout_ev(_name: &'static str) -> Self::Ev {
unreachable!("no timers in this test")
}
fn on_start(&mut self, _cx: &mut smarm::Cx<Self::Ev>) {}
fn handle(
&mut self,
ev: Self::Ev,
_cx: &mut smarm::Cx<Self::Ev>,
) -> smarm::gen_statem::Step<Self::Ev> {
let _ = ev.send(smarm::introspect::stack_shape(self_pid()).unwrap());
smarm::gen_statem::Step::Stayed
}
}
#[test]
fn gen_statem_spawn_with_stack_opts() {
rt1().run(|| {
let m = smarm::gen_statem::spawn_with(
SpawnOpts {
stack_reserve: Some(256 * 1024),
..SpawnOpts::default()
},
Probe,
);
let (tx, rx) = smarm::channel::channel();
m.send(tx).unwrap();
let (reserve, guard) = rx.recv().unwrap();
assert_eq!(reserve, 256 * 1024);
assert_eq!(guard, DEFAULT_STACK_GUARD);
});
}
+103 -11
View File
@@ -7,13 +7,13 @@ use smarm::stack::Stack;
#[test] #[test]
fn top_is_16_byte_aligned() { fn top_is_16_byte_aligned() {
let s = Stack::new(64 * 1024).unwrap(); let s = Stack::new(64 * 1024, 4096).unwrap();
assert_eq!(s.top() as usize % 16, 0); assert_eq!(s.top() as usize % 16, 0);
} }
#[test] #[test]
fn top_is_within_allocation() { fn top_is_within_allocation() {
let s = Stack::new(64 * 1024).unwrap(); let s = Stack::new(64 * 1024, 4096).unwrap();
let top = s.top() as usize; let top = s.top() as usize;
let base = s.usable_base() as usize; let base = s.usable_base() as usize;
assert!(top > base); assert!(top > base);
@@ -22,7 +22,7 @@ fn top_is_within_allocation() {
#[test] #[test]
fn write_and_read_top_of_stack() { fn write_and_read_top_of_stack() {
let s = Stack::new(64 * 1024).unwrap(); let s = Stack::new(64 * 1024, 4096).unwrap();
let sentinel: u64 = 0xDEAD_BEEF_CAFE_1234; let sentinel: u64 = 0xDEAD_BEEF_CAFE_1234;
unsafe { unsafe {
let ptr = s.top().sub(8) as *mut u64; let ptr = s.top().sub(8) as *mut u64;
@@ -33,7 +33,7 @@ fn write_and_read_top_of_stack() {
#[test] #[test]
fn write_and_read_bottom_of_usable_region() { fn write_and_read_bottom_of_usable_region() {
let s = Stack::new(64 * 1024).unwrap(); let s = Stack::new(64 * 1024, 4096).unwrap();
let sentinel: u64 = 0x0102_0304_0506_0708; let sentinel: u64 = 0x0102_0304_0506_0708;
unsafe { unsafe {
let ptr = s.usable_base() as *mut u64; let ptr = s.usable_base() as *mut u64;
@@ -44,17 +44,17 @@ fn write_and_read_bottom_of_usable_region() {
#[test] #[test]
fn small_stack_allocates() { fn small_stack_allocates() {
assert!(Stack::new(4096).is_ok()); assert!(Stack::new(4096, 4096).is_ok());
} }
#[test] #[test]
fn large_stack_allocates() { fn large_stack_allocates() {
assert!(Stack::new(8 * 1024 * 1024).is_ok()); assert!(Stack::new(8 * 1024 * 1024, 4096).is_ok());
} }
#[test] #[test]
fn stack_size_at_least_requested() { fn stack_size_at_least_requested() {
let s = Stack::new(64 * 1024).unwrap(); let s = Stack::new(64 * 1024, 4096).unwrap();
assert!(s.stack_size() >= 64 * 1024); assert!(s.stack_size() >= 64 * 1024);
} }
@@ -68,15 +68,32 @@ use std::process::Command;
fn run_as_child_if_requested() { fn run_as_child_if_requested() {
match env::var("SMARM_SUBTEST").as_deref() { match env::var("SMARM_SUBTEST").as_deref() {
Ok("guard_page_direct") => { Ok("guard_page_direct") => {
let s = Stack::new(64 * 1024).unwrap(); let s = Stack::new(64 * 1024, 4096).unwrap();
unsafe { unsafe {
let guard_ptr = s.usable_base().sub(1); let guard_ptr = s.usable_base().sub(1);
guard_ptr.write_volatile(0xAB); guard_ptr.write_volatile(0xAB);
} }
std::process::exit(0); std::process::exit(0);
} }
Ok("wide_guard_top") => {
// One byte below the usable region, 64 KiB guard: must fault.
let s = Stack::new(64 * 1024, 64 * 1024).unwrap();
unsafe {
s.usable_base().sub(1).write_volatile(0xAB);
}
std::process::exit(0);
}
Ok("wide_guard_bottom") => {
// The very bottom page of a 64 KiB guard: an unprobed C-style
// leap over a small guard lands here — must still fault.
let s = Stack::new(64 * 1024, 64 * 1024).unwrap();
unsafe {
s.usable_base().sub(64 * 1024).write_volatile(0xAB);
}
std::process::exit(0);
}
Ok("stack_overflow") => { Ok("stack_overflow") => {
let s = Stack::new(64 * 1024).unwrap(); let s = Stack::new(64 * 1024, 4096).unwrap();
unsafe { unsafe {
let mut ptr = s.top().sub(1); let mut ptr = s.top().sub(1);
let stop = s.usable_base().sub(1); let stop = s.usable_base().sub(1);
@@ -107,7 +124,12 @@ fn guard_page_causes_sigsegv() {
#[cfg(unix)] #[cfg(unix)]
{ {
use std::os::unix::process::ExitStatusExt; use std::os::unix::process::ExitStatusExt;
assert_eq!(status.signal(), Some(11), "expected SIGSEGV, got: {:?}", status); assert_eq!(
status.signal(),
Some(11),
"expected SIGSEGV, got: {:?}",
status
);
} }
} }
@@ -118,6 +140,76 @@ fn stack_overflow_causes_sigsegv() {
#[cfg(unix)] #[cfg(unix)]
{ {
use std::os::unix::process::ExitStatusExt; use std::os::unix::process::ExitStatusExt;
assert_eq!(status.signal(), Some(11), "expected SIGSEGV, got: {:?}", status); assert_eq!(
status.signal(),
Some(11),
"expected SIGSEGV, got: {:?}",
status
);
}
}
// ---------------------------------------------------------------------------
// RFC 019 — explicit shape: rounding, guard accessor, wide-guard coverage.
// ---------------------------------------------------------------------------
#[test]
fn sizes_round_up_to_page() {
let s = Stack::new(64 * 1024 + 1, 4096 + 1).unwrap();
assert_eq!(s.stack_size() % 4096, 0);
assert_eq!(s.guard_size() % 4096, 0);
assert!(s.stack_size() >= 64 * 1024 + 1);
assert!(s.guard_size() >= 4096 + 1);
}
#[test]
fn shape_reports_rounded_sizes() {
let s = Stack::new(64 * 1024, 64 * 1024).unwrap();
assert_eq!(s.shape(), (64 * 1024, 64 * 1024));
}
#[test]
fn usable_base_sits_above_guard() {
let s = Stack::new(64 * 1024, 64 * 1024).unwrap();
// The usable region must start exactly guard_size above the mapping
// base: a write at usable_base is legal, one byte below is not (the
// subprocess tests below prove the "not").
let sentinel: u64 = 0x1111_2222_3333_4444;
unsafe {
let ptr = s.usable_base() as *mut u64;
ptr.write_volatile(sentinel);
assert_eq!(ptr.read_volatile(), sentinel);
}
}
#[test]
fn wide_guard_faults_at_top() {
run_as_child_if_requested();
let status = spawn_subtest("wide_guard_top");
#[cfg(unix)]
{
use std::os::unix::process::ExitStatusExt;
assert_eq!(
status.signal(),
Some(11),
"expected SIGSEGV, got: {:?}",
status
);
}
}
#[test]
fn wide_guard_faults_at_bottom() {
run_as_child_if_requested();
let status = spawn_subtest("wide_guard_bottom");
#[cfg(unix)]
{
use std::os::unix::process::ExitStatusExt;
assert_eq!(
status.signal(),
Some(11),
"expected SIGSEGV, got: {:?}",
status
);
} }
} }
+153
View File
@@ -0,0 +1,153 @@
//! RFC 019 §7 — overflow diagnostics, observed from outside via subprocess
//! (mirrors tests/stack.rs's harness, plus stderr capture).
//!
//! Four cases:
//! - Rust recursion at defaults: probed frames walk into the guard →
//! tier-1 definitive message, death by SIGSEGV.
//! - FFI canary (96 KiB unprobed C local) at defaults: first touch lands
//! inside the 1 MiB guard → tier-1 message.
//! - FFI canary with the guard shrunk to 4 KiB: the frame steps over it
//! into unmapped VA below → tier-2 "stepped over" message. This is the
//! RFC's motivating incident (cargo-vendored gz build) reproduced.
//! - FFI canary with reserve raised to 256 KiB: fits, runs clean, exits 0 —
//! the §1 knob is the fix, proven by the same frame.
use std::env;
use std::process::Command;
unsafe extern "C" {
fn smarm_canary_burn();
}
/// Unbounded probed recursion; each frame dirties 4 KiB. black_box defeats
/// tail-call elision so the walk is real.
#[inline(never)]
#[allow(unconditional_recursion)]
fn recurse_forever(depth: u64) -> u64 {
let mut local = [0u8; 4096];
local[0] = depth as u8;
std::hint::black_box(&mut local);
recurse_forever(depth + 1).wrapping_add(local[0] as u64)
}
fn run_as_child_if_requested() {
let mode = match env::var("SMARM_DIAG_SUBTEST") {
Ok(m) => m,
Err(_) => return,
};
use smarm::runtime::Config;
use smarm::{spawn_with, SpawnOpts};
let rt = smarm::runtime::init(Config::exact(1));
rt.run(move || {
let opts = match mode.as_str() {
"rust_overflow" | "ffi_tier1" => SpawnOpts::default(),
// Small guard: the canary's 96 KiB displacement clears it.
"ffi_tier2" => SpawnOpts {
guard_size: Some(4096),
..SpawnOpts::default()
},
// Enough reserve: the same frame simply fits.
"ffi_clean" => SpawnOpts {
stack_reserve: Some(256 * 1024),
..SpawnOpts::default()
},
other => panic!("unknown subtest {other}"),
};
let is_rust = mode == "rust_overflow";
spawn_with(opts, move || {
if is_rust {
std::hint::black_box(recurse_forever(0));
} else {
unsafe { smarm_canary_burn() };
}
})
.join()
.unwrap();
});
std::process::exit(0);
}
fn spawn_subtest(name: &str) -> std::process::Output {
let exe = env::current_exe().unwrap();
Command::new(exe)
.env("SMARM_DIAG_SUBTEST", name)
.args(["--test-threads=1", "--quiet"])
.output()
.expect("failed to spawn subprocess")
}
#[cfg(unix)]
fn assert_died_sigsegv(out: &std::process::Output) {
use std::os::unix::process::ExitStatusExt;
assert_eq!(
out.status.signal(),
Some(11),
"expected death by SIGSEGV, got {:?}; stderr:\n{}",
out.status,
String::from_utf8_lossy(&out.stderr)
);
}
#[test]
fn rust_overflow_dies_with_tier1_message() {
run_as_child_if_requested();
let out = spawn_subtest("rust_overflow");
assert_died_sigsegv(&out);
let err = String::from_utf8_lossy(&out.stderr);
assert!(
err.contains("overflowed its stack") && err.contains("in the guard region"),
"missing tier-1 diagnostic; stderr:\n{err}"
);
assert!(
err.contains("reserve=65536"),
"wrong reserve in message:\n{err}"
);
assert!(
err.contains("guard=1048576"),
"wrong guard in message:\n{err}"
);
}
#[test]
fn ffi_canary_at_defaults_dies_with_tier1_message() {
run_as_child_if_requested();
let out = spawn_subtest("ffi_tier1");
assert_died_sigsegv(&out);
let err = String::from_utf8_lossy(&out.stderr);
// 96 KiB displacement from a 64 KiB reserve lands ~32 KiB into the
// 1 MiB guard: definitively classified.
assert!(
err.contains("in the guard region"),
"wide guard should catch the unprobed frame in tier 1; stderr:\n{err}"
);
}
#[test]
fn ffi_canary_over_small_guard_dies_with_tier2_message() {
run_as_child_if_requested();
let out = spawn_subtest("ffi_tier2");
assert_died_sigsegv(&out);
let err = String::from_utf8_lossy(&out.stderr);
assert!(
err.contains("stepped over it") && err.contains("below the guard"),
"expected tier-2 overshoot attribution; stderr:\n{err}"
);
assert!(err.contains("guard=4096"), "wrong guard in message:\n{err}");
}
#[test]
fn ffi_canary_with_enough_reserve_runs_clean() {
run_as_child_if_requested();
let out = spawn_subtest("ffi_clean");
assert!(
out.status.success(),
"canary should fit in 256 KiB reserve, got {:?}; stderr:\n{}",
out.status,
String::from_utf8_lossy(&out.stderr)
);
let err = String::from_utf8_lossy(&out.stderr);
assert!(
!err.contains("smarm: actor"),
"no diagnostic expected on the clean path; stderr:\n{err}"
);
}
+133
View File
@@ -0,0 +1,133 @@
//! RFC 019 commit 5 — pool recycle zaps a dead stack down to its retained
//! entry end, observed from the outside.
//!
//! A default-shaped stack that spiked deep and then died must not carry its
//! spike into the pool as resident RSS: `recycle_stack` DONTNEEDs everything
//! below the top `RECYCLE_RETAIN` bytes before pushing. The zap is
//! synchronous on the death path, so the drop is immediate — but the death
//! path itself races the observer's `join` return, hence the brief poll.
//!
//! Residency is measured with `mincore`, not smaps: a neighboring rw anon
//! mapping can land flush against the stack top and the kernel merges the
//! VMAs (observed under the full test run), so per-mapping smaps fields
//! over-count. The PROT_NONE guard below can never merge, so the usable
//! base is exactly the anchor VMA's start, and `mincore` counts pages
//! within [usable_base, usable_base + reserve) regardless of merging.
use smarm::runtime::{Config, RECYCLE_RETAIN};
use smarm::{channel, spawn, yield_now};
const RESERVE: usize = 4 * 1024 * 1024;
/// Burn ~`frames` × 4 KiB of stack, dirtying every frame.
#[inline(never)]
fn burn_stack(frames: usize) -> u64 {
let mut local = [0u8; 4096];
local[0] = frames as u8;
let below = if frames == 0 {
0
} else {
burn_stack(frames - 1)
};
std::hint::black_box(&mut local);
below.wrapping_add(local[0] as u64)
}
/// Resident-page count over [lo, lo + len) via mincore (len page-aligned).
fn resident_pages(lo: usize, len: usize) -> usize {
let page = 4096;
let mut vec = vec![0u8; len / page];
let ret = unsafe { libc::mincore(lo as *mut libc::c_void, len, vec.as_mut_ptr()) };
assert_eq!(
ret,
0,
"mincore failed: {}",
std::io::Error::last_os_error()
);
vec.iter().filter(|&&b| b & 1 != 0).count()
}
/// The [start, end) of the VMA containing `addr`.
fn vma_containing(addr: usize) -> (usize, usize) {
let maps = std::fs::read_to_string("/proc/self/maps").unwrap();
for line in maps.lines() {
if let Some((range, _)) = line.split_once(' ') {
if let Some((a, b)) = range.split_once('-') {
if let (Ok(start), Ok(end)) =
(usize::from_str_radix(a, 16), usize::from_str_radix(b, 16))
{
if start <= addr && addr < end {
return (start, end);
}
}
}
}
}
panic!("no VMA contains {addr:#x}");
}
fn vma_exists(addr: usize) -> bool {
let maps = std::fs::read_to_string("/proc/self/maps").unwrap();
for line in maps.lines() {
if let Some((range, _)) = line.split_once(' ') {
if let Some((a, b)) = range.split_once('-') {
if let (Ok(start), Ok(end)) =
(usize::from_str_radix(a, 16), usize::from_str_radix(b, 16))
{
if start <= addr && addr < end {
return true;
}
}
}
}
}
false
}
#[test]
fn recycle_zaps_dead_stack_down_to_retain() {
// Default reserve raised so the pool holds big stacks (default-shaped ⇒
// pooled) and the zap has something to bite; single scheduler.
let rt = smarm::runtime::init(Config::exact(1).stack_reserve(RESERVE));
rt.run(|| {
let (tx, rx) = channel::<usize>();
let h = spawn(move || {
let probe = 0u8;
let anchor = &probe as *const u8 as usize;
// The guard below is PROT_NONE and can never merge with the
// usable region, so the anchor VMA's start IS the usable base.
let (vlo, _) = vma_containing(anchor);
// Dirty ~3 MiB of the 4 MiB reserve, then die.
std::hint::black_box(burn_stack(768));
tx.send(vlo).unwrap();
});
let usable_base = rx.recv().unwrap();
h.join().unwrap();
// The zap span is everything below the retained entry end. DONTNEED
// on private anon discards synchronously and unconditionally, so
// this must go to exactly zero resident pages; the poll only covers
// the death path racing join's return.
let zap_len = RESERVE - RECYCLE_RETAIN;
let mut resident = usize::MAX;
for _ in 0..10_000 {
resident = resident_pages(usable_base, zap_len);
if resident == 0 {
break;
}
yield_now();
}
assert_eq!(
resident, 0,
"recycled stack's zap span still resident: {resident} pages in \
[{usable_base:#x}, +{zap_len:#x})"
);
// Pooled, not munmapped: the mapping must still be there.
assert!(
vma_exists(usable_base),
"default-shaped stack was unmapped instead of pooled"
);
});
}
+161
View File
@@ -0,0 +1,161 @@
//! RFC 019 commit 3 — park-path stack shrink, observed from the outside.
//!
//! The one integration-level claim of the shrink machinery: an actor that
//! spikes deep, returns shallow, and then parks past the cooldown gets its
//! dead span MADV_FREE'd — visible as `LazyFree` in `/proc/self/smaps`
//! within the stack's address range — while everything live survives.
//!
//! The high-water mark is *sampled* at context-save, so the spike yields
//! once at max depth to guarantee a sample there (in production, preemption
//! provides the quasi-random samples; a test must not rely on luck).
use smarm::runtime::{Config, SHRINK_COOLDOWN, SHRINK_THRESHOLD};
use smarm::{actor_info, channel, spawn, spawn_with, yield_now, ActorState, SpawnOpts};
/// Burn ~`frames` × 4 KiB of stack, yielding once at the bottom so the
/// context-save samples `sp` at max depth.
#[inline(never)]
fn burn_stack_yielding(frames: usize) -> u64 {
let mut local = [0u8; 4096];
local[0] = frames as u8;
let below = if frames == 0 {
yield_now();
0
} else {
burn_stack_yielding(frames - 1)
};
std::hint::black_box(&mut local);
below.wrapping_add(local[0] as u64)
}
/// Sum the `LazyFree:` kB of every smaps mapping intersecting [lo, hi).
fn lazy_free_bytes_in(lo: usize, hi: usize) -> usize {
let smaps = std::fs::read_to_string("/proc/self/smaps").unwrap();
let mut total_kb = 0usize;
let mut in_range = false;
for line in smaps.lines() {
if let Some((range, _)) = line.split_once(' ') {
if let Some((a, b)) = range.split_once('-') {
if let (Ok(start), Ok(end)) =
(usize::from_str_radix(a, 16), usize::from_str_radix(b, 16))
{
in_range = start < hi && end > lo;
continue;
}
}
}
if in_range {
if let Some(rest) = line.strip_prefix("LazyFree:") {
let kb: usize = rest.trim().trim_end_matches(" kB").trim().parse().unwrap();
total_kb += kb;
}
}
}
total_kb * 1024
}
#[test]
fn spike_then_parks_marks_lazyfree_and_keeps_live_data() {
// Single scheduler: the controller can gate on the worker being Parked.
let rt = smarm::runtime::init(Config::exact(1));
rt.run(|| {
let (park_tx, park_rx) = channel::<()>();
let (done_tx, done_rx) = channel::<(usize, u64)>();
let spike = 768 * 4096; // ~3 MiB, well past SHRINK_THRESHOLD
assert!(spike > SHRINK_THRESHOLD);
let worker = spawn_with(
SpawnOpts {
stack_reserve: Some(8 * 1024 * 1024),
..SpawnOpts::default()
},
move || {
// Live data that must survive the shrink, and an anchor
// address inside the stack for the smaps scan.
let live = [0xA5u8; 64];
let anchor = live.as_ptr() as usize;
// Spike: ~3 MiB deep, sampled at the bottom, unwound.
std::hint::black_box(burn_stack_yielding(768));
// Park past the cooldown. Each recv on the drained inbox is
// one park; the controller sends only when it sees us Parked.
for _ in 0..(SHRINK_COOLDOWN + 8) {
park_rx.recv().unwrap();
}
// Measure from inside: the stack spans ≤ 8 MiB below anchor.
let lazy = lazy_free_bytes_in(anchor - 8 * 1024 * 1024, anchor + 4096);
let checksum = live.iter().map(|&b| b as u64).sum();
done_tx.send((lazy, checksum)).unwrap();
},
);
let wpid = worker.pid();
for _ in 0..(SHRINK_COOLDOWN + 8) {
// Gate: send only once the worker is genuinely parked so every
// round is a real park-on-empty-mailbox.
loop {
match actor_info(wpid) {
Some(info) if info.state == ActorState::Parked => break,
Some(_) => yield_now(),
None => panic!("worker died early"),
}
}
park_tx.send(()).unwrap();
}
let (lazy, checksum) = done_rx.recv().unwrap();
// The spike was ~3 MiB; demand at least 2 MiB marked to leave slack
// for the redzone, rounding, and pages the unwind re-dirtied.
assert!(
lazy >= 2 * 1024 * 1024,
"expected ≥ 2 MiB LazyFree in the stack range, got {} bytes",
lazy
);
assert_eq!(
checksum,
64 * 0xA5u64,
"live stack data corrupted by shrink"
);
worker.join().unwrap();
});
}
/// Steady-state actors must never pay the syscall: an actor that parks a lot
/// but never spikes past the threshold ends with zero LazyFree in its stack.
#[test]
fn shallow_actor_never_shrinks() {
let rt = smarm::runtime::init(Config::exact(1));
rt.run(|| {
let (park_tx, park_rx) = channel::<()>();
let (done_tx, done_rx) = channel::<usize>();
let worker = spawn(move || {
let probe = 0u8;
let anchor = &probe as *const u8 as usize;
for _ in 0..(SHRINK_COOLDOWN + 8) {
park_rx.recv().unwrap();
}
done_tx
.send(lazy_free_bytes_in(anchor - 64 * 1024, anchor + 4096))
.unwrap();
});
let wpid = worker.pid();
for _ in 0..(SHRINK_COOLDOWN + 8) {
loop {
match actor_info(wpid) {
Some(info) if info.state == ActorState::Parked => break,
Some(_) => yield_now(),
None => panic!("worker died early"),
}
}
park_tx.send(()).unwrap();
}
assert_eq!(done_rx.recv().unwrap(), 0, "steady-state actor was shrunk");
worker.join().unwrap();
});
}
+5 -3
View File
@@ -18,8 +18,8 @@
//! registry entry guarantees for every named server. //! registry entry guarantees for every named server.
use smarm::{ use smarm::{
call, channel, init, request_stop, spawn, Config, GenServer, GenServerBuilder, GenServerName, call, channel, init, request_stop, spawn, CallError, Config, GenServer, GenServerBuilder,
CallError, Receiver, RecvTimeoutError, GenServerName, Receiver, RecvTimeoutError,
}; };
use std::sync::{Arc, Mutex}; use std::sync::{Arc, Mutex};
use std::time::Duration; use std::time::Duration;
@@ -73,7 +73,9 @@ fn named_server_request_stop_releases_queued_caller_with_server_down() {
let (res_tx, res_rx) = channel::<Result<(), CallError>>(); let (res_tx, res_rx) = channel::<Result<(), CallError>>();
// 1. Start the named server and keep its ref alive. // 1. Start the named server and keep its ref alive.
let server = GenServerBuilder::new(Blocker { gate: Some(gate_rx) }) let server = GenServerBuilder::new(Blocker {
gate: Some(gate_rx),
})
.named(BLOCKER) .named(BLOCKER)
.start() .start()
.expect("name should be free"); .expect("name should be free");
+16 -8
View File
@@ -10,7 +10,11 @@
//! out rather than produce a false pass — run with `cargo test -- --timeout` //! out rather than produce a false pass — run with `cargo test -- --timeout`
//! or under a CI timeout. //! or under a CI timeout.
use smarm::{channel, runtime::{Config, Runtime}, spawn, yield_now, JoinHandle}; use smarm::{
channel,
runtime::{Config, Runtime},
spawn, yield_now, JoinHandle,
};
use std::sync::{ use std::sync::{
atomic::{AtomicU64, AtomicUsize, Ordering}, atomic::{AtomicU64, AtomicUsize, Ordering},
Arc, Arc,
@@ -199,7 +203,9 @@ fn thundering_herd_all_wake() {
} }
// Let all receivers park before we send. // Let all receivers park before we send.
for _ in 0..4 { yield_now(); } for _ in 0..4 {
yield_now();
}
// Coordinator blasts all channels. // Coordinator blasts all channels.
handles.push(spawn(move || { handles.push(spawn(move || {
@@ -240,8 +246,7 @@ fn concurrent_spawn_join_churn() {
for _ in 0..PARENTS { for _ in 0..PARENTS {
let tc = t.clone(); let tc = t.clone();
parent_handles.push(spawn(move || { parent_handles.push(spawn(move || {
let mut child_handles: Vec<JoinHandle> = let mut child_handles: Vec<JoinHandle> = Vec::with_capacity(CHILDREN_PER_PARENT);
Vec::with_capacity(CHILDREN_PER_PARENT);
for _ in 0..CHILDREN_PER_PARENT { for _ in 0..CHILDREN_PER_PARENT {
let tcc = tc.clone(); let tcc = tc.clone();
@@ -292,7 +297,9 @@ fn join_race_child_finishes_first() {
} }
// Yield enough to let children run to completion before we join. // Yield enough to let children run to completion before we join.
for _ in 0..8 { yield_now(); } for _ in 0..8 {
yield_now();
}
for h in handles { for h in handles {
// If child already finished, join must return immediately with Ok. // If child already finished, join must return immediately with Ok.
@@ -374,8 +381,7 @@ fn panic_storm_does_not_corrupt_scheduler() {
fn pid_generation_increments_on_reuse() { fn pid_generation_increments_on_reuse() {
use smarm::self_pid; use smarm::self_pid;
let pids: Arc<smarm::Mutex<Vec<smarm::Pid>>> = let pids: Arc<smarm::Mutex<Vec<smarm::Pid>>> = Arc::new(smarm::Mutex::new(Vec::new()));
Arc::new(smarm::Mutex::new(Vec::new()));
let p = pids.clone(); let p = pids.clone();
rt(1).run(move || { rt(1).run(move || {
@@ -392,7 +398,9 @@ fn pid_generation_increments_on_reuse() {
} }
}); });
let g = pids.lock_timeout(std::time::Duration::from_secs(1)).unwrap(); let g = pids
.lock_timeout(std::time::Duration::from_secs(1))
.unwrap();
// Any two PIDs that share an index must have different generations. // Any two PIDs that share an index must have different generations.
for i in 0..g.len() { for i in 0..g.len() {
for j in (i + 1)..g.len() { for j in (i + 1)..g.len() {
+10 -2
View File
@@ -51,7 +51,11 @@ fn transient_child_is_restarted_on_panic_then_settles() {
}); });
sup.join().unwrap(); sup.join().unwrap();
}); });
assert_eq!(runs.load(Ordering::SeqCst), 3, "two restarts then a clean exit"); assert_eq!(
runs.load(Ordering::SeqCst),
3,
"two restarts then a clean exit"
);
} }
#[test] #[test]
@@ -167,7 +171,11 @@ fn one_for_all_restarts_a_normally_exited_sibling() {
sup.join().unwrap(); sup.join().unwrap();
}); });
assert_eq!(a.load(Ordering::SeqCst), 2, "A: crash then clean run"); assert_eq!(a.load(Ordering::SeqCst), 2, "A: crash then clean run");
assert_eq!(b.load(Ordering::SeqCst), 2, "B cycled with the group despite a clean exit"); assert_eq!(
b.load(Ordering::SeqCst),
2,
"B cycled with the group despite a clean exit"
);
} }
#[test] #[test]
+303
View File
@@ -0,0 +1,303 @@
//! Supervisor shutdown — the OTP child-spec `shutdown` policy.
//!
//! A supervisor traps exits. A `request_shutdown` reaching it (from its parent
//! supervisor, or from the app via `request_shutdown`/`RuntimeHandle`) runs
//! the ordered shutdown: children are stopped in reverse start order, each
//! per its `Shutdown` policy — `request_shutdown`, wait up to the timeout for
//! its termination signal, `request_stop` if it overstays — and then `run()`
//! returns normally. Every supervisor-initiated child stop (ordered shutdown,
//! OneForAll/RestForOne sibling cycling) goes through the same policy.
//!
//! A *hard* `request_stop` on a supervisor unwinds it; a drop guard then
//! hard-stops its live children so the subtree is never orphaned.
use smarm::supervisor::{ChildSpec, OneForOne, Restart, Shutdown, Strategy};
use smarm::{
monitor, request_shutdown, request_stop, run, sleep, spawn, trap_exit, DownReason, JoinHandle,
};
use std::sync::atomic::{AtomicBool, AtomicUsize, Ordering};
use std::sync::{Arc, Mutex};
use std::time::{Duration, Instant};
/// A child that traps exits, records the order it was shut down in, and exits
/// normally on the request (after `delay`). Ignores the request if `comply`
/// is false — a straggler that must be hard-stopped.
fn polite_child(
tag: usize,
log: &Arc<Mutex<Vec<usize>>>,
delay: Duration,
comply: bool,
) -> impl Fn() + Send + Sync + 'static {
let log = log.clone();
move || {
let inbox = trap_exit();
loop {
let sig = match inbox.recv() {
Ok(s) => s,
Err(_) => return,
};
if sig.reason == DownReason::Shutdown {
log.lock().unwrap().push(tag);
if comply {
sleep(delay);
return;
}
// Not complying: keep running until hard-stopped.
loop {
sleep(Duration::from_millis(5));
}
}
}
}
}
/// Spawn `sup`, let its children reach `trap_exit`, return the handle.
fn spawn_settled(sup: OneForOne) -> JoinHandle {
let h = spawn(move || sup.run());
sleep(Duration::from_millis(30));
h
}
#[test]
fn shutdown_stops_children_in_reverse_order_and_returns_normally() {
let log = Arc::new(Mutex::new(Vec::new()));
let l = log.clone();
run(move || {
let sup = OneForOne::new()
.child(ChildSpec::new(
Restart::Permanent,
polite_child(1, &l, Duration::ZERO, true),
))
.child(ChildSpec::new(
Restart::Permanent,
polite_child(2, &l, Duration::ZERO, true),
))
.child(ChildSpec::new(
Restart::Permanent,
polite_child(3, &l, Duration::ZERO, true),
));
let h = spawn_settled(sup);
let mon = monitor(h.pid());
request_shutdown(h.pid());
let down = mon.rx.recv().expect("down");
assert_eq!(
down.reason,
DownReason::Exit,
"supervisor exits normally after shutdown"
);
});
assert_eq!(*log.lock().unwrap(), vec![3, 2, 1]);
}
#[test]
fn non_trapping_child_is_simply_stopped() {
let dropped = Arc::new(AtomicBool::new(false));
let d = dropped.clone();
run(move || {
struct G(Arc<AtomicBool>);
impl Drop for G {
fn drop(&mut self) {
self.0.store(true, Ordering::SeqCst);
}
}
let sup = OneForOne::new().child(ChildSpec::new(Restart::Permanent, move || {
let _g = G(d.clone());
loop {
sleep(Duration::from_millis(5));
}
}));
let h = spawn_settled(sup);
request_shutdown(h.pid());
h.join().expect("sup");
});
assert!(dropped.load(Ordering::SeqCst));
}
#[test]
fn straggler_is_hard_stopped_after_timeout() {
let log = Arc::new(Mutex::new(Vec::new()));
let l = log.clone();
run(move || {
let sup = OneForOne::new().child(
ChildSpec::new(
Restart::Permanent,
polite_child(1, &l, Duration::ZERO, false),
)
.shutdown(Shutdown::Timeout(Duration::from_millis(50))),
);
let h = spawn_settled(sup);
let t0 = Instant::now();
request_shutdown(h.pid());
h.join().expect("sup");
let took = t0.elapsed();
assert!(
took >= Duration::from_millis(50),
"returned before the grace period: {took:?}"
);
assert!(
took < Duration::from_secs(2),
"did not fall back to a hard stop: {took:?}"
);
});
assert_eq!(
*log.lock().unwrap(),
vec![1],
"the straggler did receive the request"
);
}
#[test]
fn infinity_waits_for_a_slow_but_compliant_child() {
let log = Arc::new(Mutex::new(Vec::new()));
let l = log.clone();
let finished = Arc::new(AtomicBool::new(false));
let f = finished.clone();
run(move || {
let f2 = f.clone();
let l2 = l.clone();
let sup = OneForOne::new().child(
ChildSpec::new(Restart::Permanent, move || {
let inbox = trap_exit();
let _ = inbox.recv();
l2.lock().unwrap().push(1);
sleep(Duration::from_millis(150));
f2.store(true, Ordering::SeqCst); // only reached if not hard-stopped
})
.shutdown(Shutdown::Infinity),
);
let h = spawn_settled(sup);
request_shutdown(h.pid());
h.join().expect("sup");
});
assert!(
finished.load(Ordering::SeqCst),
"Infinity must not hard-stop a compliant child"
);
}
#[test]
fn brutal_kill_skips_the_request() {
let log = Arc::new(Mutex::new(Vec::new()));
let l = log.clone();
run(move || {
let sup = OneForOne::new().child(
ChildSpec::new(
Restart::Permanent,
polite_child(1, &l, Duration::ZERO, true),
)
.shutdown(Shutdown::BrutalKill),
);
let h = spawn_settled(sup);
request_shutdown(h.pid());
h.join().expect("sup");
});
assert!(
log.lock().unwrap().is_empty(),
"a BrutalKill child never sees the request"
);
}
#[test]
fn hard_stop_of_supervisor_does_not_orphan_children() {
let alive = Arc::new(AtomicUsize::new(0));
let a = alive.clone();
run(move || {
struct Alive(Arc<AtomicUsize>);
impl Drop for Alive {
fn drop(&mut self) {
self.0.fetch_sub(1, Ordering::SeqCst);
}
}
let mk = |a: Arc<AtomicUsize>| {
move || {
a.fetch_add(1, Ordering::SeqCst);
let _g = Alive(a.clone());
loop {
sleep(Duration::from_millis(5));
}
}
};
let sup = OneForOne::new()
.child(ChildSpec::new(Restart::Permanent, mk(a.clone())))
.child(ChildSpec::new(Restart::Permanent, mk(a.clone())));
let h = spawn_settled(sup);
assert_eq!(a.load(Ordering::SeqCst), 2);
let mon = monitor(h.pid());
request_stop(h.pid());
let _ = mon.rx.recv();
sleep(Duration::from_millis(50));
assert_eq!(
a.load(Ordering::SeqCst),
0,
"children orphaned by a hard supervisor stop"
);
});
}
#[test]
fn nested_shutdown_reaches_grandchildren() {
let log = Arc::new(Mutex::new(Vec::new()));
let l = log.clone();
run(move || {
let l_inner = l.clone();
let inner = move || {
OneForOne::new()
.child(ChildSpec::new(
Restart::Permanent,
polite_child(10, &l_inner, Duration::ZERO, true),
))
.child(ChildSpec::new(
Restart::Permanent,
polite_child(11, &l_inner, Duration::ZERO, true),
))
.run()
};
let sup = OneForOne::new()
.child(ChildSpec::new(
Restart::Permanent,
polite_child(1, &l, Duration::ZERO, true),
))
.child(ChildSpec::new(Restart::Permanent, inner).shutdown(Shutdown::Infinity));
let h = spawn_settled(sup);
request_shutdown(h.pid());
h.join().expect("sup");
});
assert_eq!(*log.lock().unwrap(), vec![11, 10, 1]);
}
#[test]
fn sibling_cycling_uses_graceful_shutdown() {
// OneForAll: when child A dies, sibling B (trapping) must receive a
// Shutdown request rather than a bare stop.
let log = Arc::new(Mutex::new(Vec::new()));
let l = log.clone();
let a_runs = Arc::new(AtomicUsize::new(0));
let ar = a_runs.clone();
run(move || {
let ar2 = ar.clone();
let sup = OneForOne::new()
.strategy(Strategy::OneForAll)
.intensity(5, Duration::from_secs(60))
.child(ChildSpec::new(Restart::Transient, move || {
let n = ar2.fetch_add(1, Ordering::SeqCst) + 1;
sleep(Duration::from_millis(30));
if n == 1 {
panic!("first run dies");
}
// Second run: park until shut down.
let inbox = trap_exit();
let _ = inbox.recv();
}))
.child(ChildSpec::new(
Restart::Permanent,
polite_child(2, &l, Duration::ZERO, true),
));
let h = spawn(move || sup.run());
sleep(Duration::from_millis(150));
request_shutdown(h.pid());
h.join().expect("sup");
});
// B was shut down once by the cycle and once by the final shutdown.
assert_eq!(*log.lock().unwrap(), vec![2, 2]);
assert_eq!(a_runs.load(Ordering::SeqCst), 2);
}
+276
View File
@@ -0,0 +1,276 @@
//! The terminal-record contract (bridge soak signature 4): a watch installed
//! *after* its target's death — the async-install race the bridge's proxies
//! live with — must be able to recover the real down reason instead of a
//! blanket `NoProc`. Two primitives carry it:
//!
//! - `finalize_actor` stamps the slot with `(generation, DownReason)`; the
//! record survives reclaim, registry pruning, and the next tenant's
//! install, and is overwritten only by the slot's next death.
//! [`terminal_reason`] reads it generation-matched.
//! - [`resolve_name`] is `whereis` with the corpse kept: the dead-holder arm
//! returns the stored pid it prunes ([`NameResolution::Corpse`]) instead
//! of discarding the only evidence of *who* died. `Unbound` stays the
//! Erlang-shaped `noproc` for names that were never (or are no longer)
//! bound.
//!
//! `monitor()` of a stale pid still queues plain `NoProc` — the upgrade is a
//! caller's deliberate act, not a semantics change.
use smarm::{
init, mark_watchable, request_stop, resolve_name, terminal_reason, CallError, Config,
DownReason, GenServer, GenServerBuilder, GenServerName, NameResolution,
};
use std::sync::{Arc, Mutex};
use std::time::Duration;
const TARGET: GenServerName<Target> = GenServerName::new("terminal_target");
/// Named server that panics on cast — the sig-4 death.
struct Target;
impl GenServer for Target {
type Call = ();
type Reply = ();
type Cast = ();
type Info = ();
type Timer = ();
fn handle_call(&mut self, _req: ()) {}
fn handle_cast(&mut self, _op: ()) {
panic!("terminal_target: induced panic");
}
}
/// Slot filler for the re-tenancy phase (distinct type, held alive).
struct Filler;
impl GenServer for Filler {
type Call = ();
type Reply = ();
type Cast = ();
type Info = ();
type Timer = ();
fn handle_call(&mut self, _req: ()) {}
fn handle_cast(&mut self, _op: ()) {}
}
#[derive(Debug)]
struct Observed {
exit_reason: Option<DownReason>,
anon_reason: Option<DownReason>,
/// Anonymous but export-marked while alive — must stamp (sig 5).
marked_reason: Option<DownReason>,
/// Marked only after death — must remain unknowable.
marked_late_reason: Option<DownReason>,
panic_reason: Option<DownReason>,
stopped_reason: Option<DownReason>,
live_reason: Option<DownReason>,
live_resolution_is_live: bool,
unknown_resolution: NameResolution,
/// First resolve after the named target's panic — must be Corpse(old pid).
corpse_resolution_matches: bool,
/// Second resolve — the Corpse arm pruned, so the name has healed.
resolution_after_prune: NameResolution,
/// Read AFTER the prune above: the record is slot-side, not registry-side.
corpse_reason_after_prune: Option<DownReason>,
/// Record survives the slot being re-tenanted (new tenant still alive).
corpse_reason_after_reuse: Option<DownReason>,
/// ... and dies with the next tenancy's death (overwritten).
corpse_reason_after_tenant_death: Option<DownReason>,
tenant_reason: Option<DownReason>,
}
#[test]
fn terminal_record_recovers_the_reason_a_raced_watch_lost() {
let out: Arc<Mutex<Option<Observed>>> = Arc::new(Mutex::new(None));
let out_w = out.clone();
// Tiny slab: prompt slot recycling for the re-tenancy phase.
init(Config::exact(2).max_actors(32)).run(move || {
// --- Registered plain actors: one record per way of dying. The
// record is named-tenancy-only, so each actor self-registers a
// throwaway channel before dying; the anonymous control below pins
// the complement.
let h = smarm::spawn(|| {
let (tx, _rx) = smarm::channel::<()>();
let _ = smarm::register(smarm::Name::<()>::new("terminal_probe_exit"), tx);
});
let pid_exit = h.pid();
let _ = h.join();
let exit_reason = terminal_reason(pid_exit);
let h = smarm::spawn(|| {
let (tx, _rx) = smarm::channel::<()>();
let _ = smarm::register(smarm::Name::<()>::new("terminal_probe_panic"), tx);
panic!("induced");
});
let pid_panic = h.pid();
let _ = h.join();
let panic_reason = terminal_reason(pid_panic);
let h = smarm::spawn(|| {
let (tx, _rx) = smarm::channel::<()>();
let _ = smarm::register(smarm::Name::<()>::new("terminal_probe_stop"), tx);
loop {
smarm::sleep(Duration::from_millis(2));
}
});
let pid_stop = h.pid();
request_stop(pid_stop);
let _ = h.join();
let stopped_reason = terminal_reason(pid_stop);
// --- Anonymous control: an unregistered death must NOT stamp (nor
// evict) — the free list is LIFO, so green-thread churn would
// otherwise overwrite a watchable record faster than any race
// window this exists to cover.
let h = smarm::spawn(|| panic!("anonymous"));
let pid_anon = h.pid();
let _ = h.join();
let anon_reason = terminal_reason(pid_anon);
// --- mark_watchable: the bridge's export-seam eligibility (sig 5).
// An anonymous actor marked while alive stamps like a named one ...
let h = smarm::spawn(|| loop {
smarm::sleep(Duration::from_millis(2));
});
let pid_marked = h.pid();
mark_watchable(pid_marked);
request_stop(pid_marked);
let _ = h.join();
let marked_reason = terminal_reason(pid_marked);
// ... while marking a pid whose tenancy already ended is a no-op:
// the history is honestly unknowable, not retroactively invented.
mark_watchable(pid_anon);
let marked_late_reason = terminal_reason(pid_anon);
// --- The named target: live readings first. -----------------------
let target = GenServerBuilder::new(Target)
.named(TARGET)
.start()
.expect("name free at test start");
let old_pid = target.pid();
let live_reason = terminal_reason(old_pid);
let live_resolution_is_live =
resolve_name(TARGET.as_str()) == NameResolution::Live(old_pid.erase());
let unknown_resolution = resolve_name("terminal_never_bound");
// --- Kill it by panic; confirm death via the ref, NEVER the name
// (any name reader would take the prune arm and destroy the corpse
// precondition — the same trap stale_name_slot_reuse.rs documents).
let _ = target.cast(());
loop {
match target.call(()) {
Err(CallError::ServerDown) => break,
Ok(()) => smarm::sleep(Duration::from_millis(2)),
}
}
let corpse_resolution_matches =
resolve_name(TARGET.as_str()) == NameResolution::Corpse(old_pid.erase());
let resolution_after_prune = resolve_name(TARGET.as_str());
let corpse_reason_after_prune = terminal_reason(old_pid);
// --- Re-tenant the freed slot; the record must outlive the install
// and die only with the next tenancy's death.
let mut fillers = Vec::new();
let mut tenant = None;
for i in 0..24 {
let name: &'static str = Box::leak(format!("terminal_filler_{i}").into_boxed_str());
let f = GenServerBuilder::new(Filler)
.named(GenServerName::<Filler>::new(name))
.start()
.expect("filler names are fresh");
let fp = f.pid();
let landed = fp.index() == old_pid.index();
fillers.push(f);
if landed {
tenant = Some((fillers.len() - 1, fp));
break;
}
}
let (tenant_at, tenant_pid) = tenant.expect(
"precondition: the freed slot must be re-tenanted within the tiny slab \
(slots are recycled; every filler is held alive)",
);
let corpse_reason_after_reuse = terminal_reason(old_pid);
request_stop(tenant_pid);
loop {
match fillers[tenant_at].call(()) {
Err(CallError::ServerDown) => break,
Ok(()) => smarm::sleep(Duration::from_millis(2)),
}
}
let corpse_reason_after_tenant_death = terminal_reason(old_pid);
let tenant_reason = terminal_reason(tenant_pid);
*out_w.lock().unwrap() = Some(Observed {
exit_reason,
anon_reason,
panic_reason,
stopped_reason,
live_reason,
live_resolution_is_live,
unknown_resolution,
corpse_resolution_matches,
resolution_after_prune,
marked_reason,
marked_late_reason,
corpse_reason_after_prune,
corpse_reason_after_reuse,
corpse_reason_after_tenant_death,
tenant_reason,
});
});
let o = out.lock().unwrap().take().expect("runtime body completed");
assert_eq!(o.exit_reason, Some(DownReason::Exit), "{o:?}");
assert_eq!(
o.anon_reason, None,
"anonymous deaths must not stamp: {o:?}"
);
assert_eq!(o.panic_reason, Some(DownReason::Panic), "{o:?}");
assert_eq!(
o.marked_reason,
Some(DownReason::Stopped),
"mark_watchable while alive must make the death stamp: {o:?}"
);
assert_eq!(
o.marked_late_reason, None,
"marking a dead tenancy must not invent history: {o:?}"
);
assert_eq!(o.stopped_reason, Some(DownReason::Stopped), "{o:?}");
assert_eq!(
o.live_reason, None,
"live tenancy must have no record: {o:?}"
);
assert!(o.live_resolution_is_live, "{o:?}");
assert_eq!(o.unknown_resolution, NameResolution::Unbound, "{o:?}");
assert!(
o.corpse_resolution_matches,
"first post-death resolve must carry the corpse: {o:?}"
);
assert_eq!(
o.resolution_after_prune,
NameResolution::Unbound,
"the Corpse arm prunes — the name heals: {o:?}"
);
assert_eq!(
o.corpse_reason_after_prune,
Some(DownReason::Panic),
"the record is slot-side; registry pruning must not touch it: {o:?}"
);
assert_eq!(
o.corpse_reason_after_reuse,
Some(DownReason::Panic),
"a new tenant's install must leave the previous tenancy's record: {o:?}"
);
assert_eq!(
o.corpse_reason_after_tenant_death, None,
"the next death overwrites — the old generation no longer matches: {o:?}"
);
assert_eq!(o.tenant_reason, Some(DownReason::Stopped), "{o:?}");
}
+7 -4
View File
@@ -35,7 +35,10 @@ impl PipePair {
let mut fds: [libc::c_int; 2] = [0; 2]; let mut fds: [libc::c_int; 2] = [0; 2];
let r = unsafe { libc::pipe2(fds.as_mut_ptr(), libc::O_CLOEXEC | libc::O_NONBLOCK) }; let r = unsafe { libc::pipe2(fds.as_mut_ptr(), libc::O_CLOEXEC | libc::O_NONBLOCK) };
assert_eq!(r, 0, "pipe2 failed"); assert_eq!(r, 0, "pipe2 failed");
PipePair { read: fds[0], write: fds[1] } PipePair {
read: fds[0],
write: fds[1],
}
} }
} }
@@ -67,9 +70,9 @@ fn run_with_watchdog(limit: Duration, body: impl FnOnce() + Send + 'static) {
rt.run(body); rt.run(body);
let _ = done_tx.send(()); let _ = done_tx.send(());
}); });
done_rx done_rx.recv_timeout(limit).expect(
.recv_timeout(limit) "Runtime::run did not return: idle scheduler thread was never woken at termination",
.expect("Runtime::run did not return: idle scheduler thread was never woken at termination"); );
} }
/// Permanent-hang variant: sibling blocked in `poll_wake(wake_fd, None)` /// Permanent-hang variant: sibling blocked in `poll_wake(wake_fd, None)`
+26 -6
View File
@@ -166,14 +166,19 @@ fn timers_only_pop_entries_whose_deadline_has_passed() {
#[test] #[test]
fn timers_mix_sleep_and_wait_timeout_reasons() { fn timers_mix_sleep_and_wait_timeout_reasons() {
let mut t = Timers::new(); let mut t = Timers::new();
let target = Arc::new(RecordingTarget { calls: Mutex::new(Vec::new()) }); let target = Arc::new(RecordingTarget {
calls: Mutex::new(Vec::new()),
});
let now = Instant::now(); let now = Instant::now();
t.insert_sleep(now + Duration::from_millis(5), Pid::new(0, 0), 1); t.insert_sleep(now + Duration::from_millis(5), Pid::new(0, 0), 1);
t.insert( t.insert(
now + Duration::from_millis(10), now + Duration::from_millis(10),
Pid::new(1, 0), Pid::new(1, 0),
Reason::WaitTimeout { target: target.clone(), epoch: 42 }, Reason::WaitTimeout {
target: target.clone(),
epoch: 42,
},
); );
let due = t.pop_due(now + Duration::from_millis(20)); let due = t.pop_due(now + Duration::from_millis(20));
@@ -238,7 +243,10 @@ fn armed_send_timer_is_returned_and_fires() {
let mut due = t.pop_due(now + Duration::from_millis(20)); let mut due = t.pop_due(now + Duration::from_millis(20));
assert_eq!(due.len(), 1, "an armed send timer should pop when due"); assert_eq!(due.len(), 1, "an armed send timer should pop when due");
assert!(!fired.load(Ordering::SeqCst), "pop must not fire on its own"); assert!(
!fired.load(Ordering::SeqCst),
"pop must not fire on its own"
);
run_fire(due.pop().unwrap()); run_fire(due.pop().unwrap());
assert!(fired.load(Ordering::SeqCst), "running the thunk delivers"); assert!(fired.load(Ordering::SeqCst), "running the thunk delivers");
assert!(t.is_empty()); assert!(t.is_empty());
@@ -282,7 +290,11 @@ fn cancel_after_fire_returns_false() {
fn cancel_unknown_id_returns_false() { fn cancel_unknown_id_returns_false() {
let mut t = Timers::new(); let mut t = Timers::new();
let now = Instant::now(); let now = Instant::now();
let id = t.insert_send(now + Duration::from_millis(5), Pid::new(0, 0), Box::new(|| {})); let id = t.insert_send(
now + Duration::from_millis(5),
Pid::new(0, 0),
Box::new(|| {}),
);
assert!(t.cancel(id)); assert!(t.cancel(id));
// Second cancel of the same id: already gone. // Second cancel of the same id: already gone.
assert!(!t.cancel(id)); assert!(!t.cancel(id));
@@ -293,7 +305,11 @@ fn send_timers_interleave_with_sleep_in_deadline_order() {
let mut t = Timers::new(); let mut t = Timers::new();
let now = Instant::now(); let now = Instant::now();
t.insert_sleep(now + Duration::from_millis(30), Pid::new(0, 0), 1); t.insert_sleep(now + Duration::from_millis(30), Pid::new(0, 0), 1);
let _id = t.insert_send(now + Duration::from_millis(10), Pid::new(1, 0), Box::new(|| {})); let _id = t.insert_send(
now + Duration::from_millis(10),
Pid::new(1, 0),
Box::new(|| {}),
);
t.insert_sleep(now + Duration::from_millis(20), Pid::new(2, 0), 1); t.insert_sleep(now + Duration::from_millis(20), Pid::new(2, 0), 1);
let due = t.pop_due(now + Duration::from_millis(50)); let due = t.pop_due(now + Duration::from_millis(50));
@@ -308,7 +324,11 @@ fn send_timers_interleave_with_sleep_in_deadline_order() {
fn clear_drops_armed_send_timers() { fn clear_drops_armed_send_timers() {
let mut t = Timers::new(); let mut t = Timers::new();
let now = Instant::now(); let now = Instant::now();
let id = t.insert_send(now + Duration::from_millis(10), Pid::new(0, 0), Box::new(|| {})); let id = t.insert_send(
now + Duration::from_millis(10),
Pid::new(0, 0),
Box::new(|| {}),
);
t.clear(); t.clear();
assert!(t.is_empty()); assert!(t.is_empty());
// The arm record is gone too: cancelling reports nothing to cancel. // The arm record is gone too: cancelling reports nothing to cancel.
+185
View File
@@ -0,0 +1,185 @@
//! Non-panicking spawn at slab capacity (`try_spawn`).
//!
//! Covers: parity with `spawn` when slots are free; `Err(AtCapacity)` instead
//! of a panic on a full slab (the spawning actor survives — the crash-loop
//! from the motivating slowloris incident cannot start); self-heal (a freed
//! slot makes the next `try_spawn` succeed); and exact claim-or-report
//! accounting under a multi-thread race for the last slots (no TOCTOU
//! overshoot, no panic).
use smarm::runtime::Config;
use smarm::{spawn, try_spawn, try_spawn_under_with, yield_now, SpawnError, SpawnOpts};
use std::sync::atomic::{AtomicBool, AtomicUsize, Ordering};
use std::sync::Arc;
/// A child that holds its slot until `release` flips, without parking
/// machinery: busy-yield keeps the scheduler moving and the slot occupied.
fn holder(release: Arc<AtomicBool>) -> impl FnOnce() + Send + 'static {
move || {
while !release.load(Ordering::Acquire) {
yield_now();
}
}
}
#[test]
fn try_spawn_is_spawn_when_slots_free() {
smarm::runtime::init(Config::exact(1)).run(|| {
let ran = Arc::new(AtomicBool::new(false));
let flag = ran.clone();
let h = try_spawn(move || flag.store(true, Ordering::Release))
.expect("slots free — must behave exactly like spawn");
h.join().unwrap();
assert!(ran.load(Ordering::Acquire));
});
}
#[test]
fn at_capacity_is_err_not_panic_and_accounting_is_exact() {
const MAX: usize = 8;
smarm::runtime::init(Config::exact(1).max_actors(MAX)).run(|| {
let release = Arc::new(AtomicBool::new(false));
// Fill the slab from the initial actor: slots are claimed at spawn
// time, so children need not have run yet. Count until refusal.
let mut held = Vec::new();
loop {
match try_spawn(holder(release.clone())) {
Ok(h) => held.push(h),
Err(e) => {
assert_eq!(e, SpawnError::AtCapacity);
break;
}
}
}
// Initial actor occupies one slot; the rest were spawnable.
assert_eq!(held.len(), MAX - 1, "slab accounting must be exact");
// Still refusing (and still not panicking) on repeat.
assert!(matches!(try_spawn(|| ()), Err(SpawnError::AtCapacity)));
// The `_with` surface refuses identically — a custom shape must not
// reach stack allocation when there is no slot for it.
let opts = SpawnOpts {
stack_reserve: Some(1024 * 1024),
..SpawnOpts::default()
};
assert!(matches!(
try_spawn_under_with(smarm::self_pid(), opts, || ()),
Err(SpawnError::AtCapacity)
));
// Self-heal: free the slots, join, and the next try_spawn succeeds.
release.store(true, Ordering::Release);
for h in held {
h.join().unwrap();
}
let h = try_spawn(|| ()).expect("slots freed — must succeed again");
h.join().unwrap();
});
}
#[test]
fn plain_spawn_still_panics_at_capacity() {
// The existing invariant-check semantics of `spawn` are untouched: at a
// full slab it panics, the panic is caught at the actor isolation
// boundary, and it surfaces as a join error — exactly as before. The
// bomb actor is spawned into the LAST slot (so the slab is full only
// once the bomb itself is live) and the panic lands inside the bomb,
// not the initial actor.
const MAX: usize = 6;
smarm::runtime::init(Config::exact(1).max_actors(MAX)).run(|| {
let release = Arc::new(AtomicBool::new(false));
let mut held = Vec::new();
for _ in 0..MAX - 2 {
held.push(spawn(holder(release.clone())));
}
let armed = Arc::new(AtomicBool::new(false));
let armed2 = armed.clone();
let bomb = spawn(move || {
armed2.store(true, Ordering::Release);
// Slab is now full (initial + MAX−2 holders + this actor); the
// plain spawn must panic this actor.
let _ = spawn(|| ());
unreachable!("allocate_slot must have panicked");
});
let err = bomb
.join()
.expect_err("bomb must die by panic, not run through");
assert!(armed.load(Ordering::Acquire), "bomb must have actually run");
// The panic message is a formatted String (panic! with args).
let msg = err
.payload
.downcast_ref::<String>()
.cloned()
.unwrap_or_else(|| "<non-string payload>".into());
assert!(
msg.contains("slot table exhausted"),
"panic must be the slab-exhaustion invariant message, got: {msg}"
);
release.store(true, Ordering::Release);
for h in held {
h.join().unwrap();
}
});
}
#[test]
fn racing_try_spawns_claim_exactly_the_free_slots() {
// 4 scheduler threads, 4 spawner actors hammering try_spawn for a small
// pool of remaining slots. Claim-or-report must hand out exactly the
// free slots across all racers — no overshoot (TOCTOU), no panic.
const MAX: usize = 32;
const SPAWNERS: usize = 4;
smarm::runtime::init(Config::exact(4).max_actors(MAX)).run(|| {
let release = Arc::new(AtomicBool::new(false));
let won = Arc::new(AtomicUsize::new(0));
let done = Arc::new(AtomicUsize::new(0));
// Occupy some slots up front so the racers fight over a remainder.
let mut pre = Vec::new();
for _ in 0..8 {
pre.push(spawn(holder(release.clone())));
}
// Free slots now: MAX − 1 (initial) − 8 (pre) − SPAWNERS.
let up_for_grabs = MAX - 1 - 8 - SPAWNERS;
let mut spawners = Vec::new();
for _ in 0..SPAWNERS {
let release = release.clone();
let won = won.clone();
let done = done.clone();
spawners.push(spawn(move || {
loop {
match try_spawn(holder(release.clone())) {
Ok(h) => {
won.fetch_add(1, Ordering::AcqRel);
drop(h); // detached; slot held by the holder
}
Err(SpawnError::AtCapacity) => break,
Err(_) => unreachable!("non_exhaustive future-proofing"),
}
}
done.fetch_add(1, Ordering::AcqRel);
}));
}
// Wait for every racer to hit AtCapacity.
while done.load(Ordering::Acquire) < SPAWNERS {
yield_now();
}
assert_eq!(won.load(Ordering::Acquire), up_for_grabs);
release.store(true, Ordering::Release);
for h in pre.into_iter().chain(spawners) {
h.join().unwrap();
}
});
}
#[test]
fn spawn_error_is_a_real_error() {
let e = SpawnError::AtCapacity;
let msg = format!("{e}");
assert!(
msg.contains("capacity"),
"Display should name the condition: {msg}"
);
let _: &dyn std::error::Error = &e;
}