42 Commits
Author SHA1 Message Date
claude-asm-audit 8d2fed10d6 release: v0.8.0
RFC 010 clustering behind `--features cluster` (c1–c16 + Phase 6), plus
the 16 perf-audit commits on v0.7.0. Public changes in the default build:
`DownReason::Disconnected` removed from core (cluster carries it as
`cluster::RemoteDownReason::Disconnected`).
2026-09-12 05:26:14 +00:00
claude-asm-audit 0f10415fb3 style: cargo fmt sweep under rustc 1.98.1 (toolchain reformat, no semantic change)
Four files were already fmt-dirty on 8fd724a; the merge added none.
2026-09-12 05:26:07 +00:00
claude-asm-audit 69a52a5578 Merge branch 'rfc010-cluster': RFC 010 clustering (c1–c16 + Phase 6) onto the v0.7.0 + perf-audit tree
Both lines branched from ca1c983 (v0.6.1 + try_spawn). master carried
graceful shutdown, gen_server/gen_statem lifetime, v0.7.0 and the 16
perf-audit commits; rfc010-cluster carried clustering behind
`--features cluster`. Two resolutions beyond the automatic merge:

- tests/channel.rs: both sides deflaked the spawn-then-monitor race in
  channel_ops_interleaved_with_monitor_churn_multi_thread. Kept master's
  `spawn_monitor` (monitor registered before publish) over the cluster
  side's `go`-gated spawn; same intent, API-level fix.
- src/cluster/envelope.rs: master added `DownReason::Shutdown`, which
  made the `Frame::Down` reason codec non-exhaustive. New wire tag
  DR_SHUTDOWN = 6, encoded and decoded symmetrically. `Shutdown` never
  rides in a `Down` by contract (a target that honours the request
  exits normally); the tag exists so the codec stays total. Tags 1–5
  unchanged.

Gates on the merged tree (rustc 1.98.1): default 405/0, cluster 492/0
(0 ignored beyond the 11 pre-existing `ignore` doctests), clippy --lib
-D warnings on both configs, doctests. cluster_disconnect still ~1.8s
(SMARM_FAST_TIMING plumbing intact).
2026-09-12 05:25:50 +00:00
claude-asm-audit 8fd724a6b6 bench: regenerate baseline.json on the reconciled tree (job 3a1af71f)
Rebased onto upstream v0.7.0, with the Weak-per-park fix (93bc83a) and the
root-sweep lock-free guard (3b0e06e) both in. 20-core box, rq-mpmc default,
wake slot on. Key order kept, so the diff is values only.

Not a flat baseline, and the regress pass in outputs/f21_job_3a1af71f.log
records why. Against 09ad759 the short 1-thread sections carry ~30-45 µs of
residual fixed cost (the root-exit sweep still walks all 16_384 slot words even
though it no longer locks them), and ping_pong_steady is +25% 1T / +23% 20T with
only ~30 µs of that explained — ~28 ns/roundtrip on the channel park chain that
nothing in the upstream diff accounts for. Several 20-thread sections moved
10-18% alongside it. That is finding 20c, open, and regenerating here means
`regress` can no longer see it: the comparison lives in history.md and the job
log instead.

switch_cost 126-128 cyc/roundtrip, unchanged, so the yield path is not involved.
2026-08-21 14:04:19 +00:00
claude-asm-audit 3b0e06ef13 perf(runtime): reject non-Live slots lock-free in the root-exit sweep
`shutdown_forest_roots` locked every slot's `cold` to discover `actor == None`.
The slab is `max_actors` entries (16_384 by default) and is almost entirely
vacant at root exit, so the sweep cost one uncontended mutex round-trip per
slot: a fixed ~152 µs per run on the 5900X, visible as deep_recursion 1T
289 -> 441 µs and the same +152 on chained_spawn 1T, fan_out_compute 1T and
deep_recursion 20T. `general.rs` times `init` and `run` together, so it landed
in every smarm bench number while no tokio section paid it.

`is_live_for` is a lock-free snapshot of the slot word; a non-Live slot has no
actor to shut down, and the ones that survive the check take the lock below as
before. The sweep is not weakened: the snapshot can go Live just after we pass
it, but that was already true of a spawn landing after the scan finished, and
there is no second sweep either way.

Suite green under reltest, including upstream's root_exit / shutdown /
supervisor_shutdown / root_sweep_trace (the last also under --features
smarm-trace, which is what asserts the sweep actually visits its targets).
2026-08-21 13:53:17 +00:00
claude-asm-audit 93bc83a5b3 perf(channel): capture the receiver's runtime Weak once per channel, not once per park
Upstream's off-runtime wake (1002777) captured a Weak<RuntimeInner> in the
parked_receiver tuple on every park, so every park/unpark round-trip paid an
Arc::downgrade plus the matching Weak drop — a locked RMW pair on one globally
shared counter, on the hot path of every channel workload.

The receiver never migrates between runtimes, so one capture is enough: the
Weak moves into Inner<T>, filled in on first park (a branch on an Option
thereafter), and unpark_at_via takes a closure so the fallback only re-locks
and clones off a scheduler thread, where the in-runtime fast path has already
declined. runtime_weak() returns Option and no longer panics off-runtime.

Sandbox 1T channel steady round-trip: 290 -> 270 ns (upstream's shape measured
427 -> 445 on its own base). Full suite green under reltest, including
upstream's tests/cross_thread_wake.rs.
2026-08-21 12:30:18 +00:00
claude-asm-audit a9341e2d82 fix(tls): actor-side thread-local accessors are #[inline(never)] + fence — LLVM caches TLS addresses across context switches (finding 19)
Without LTO (any downstream `cargo build --release`), LLVM keeps `%fs:0` in a
callee-saved register across `switch_to_scheduler`; after the actor migrates,
`ACTOR_DONE`/`PREEMPTION_ENABLED`/`CURRENT_PID` etc. hit the OLD thread's TLS.
8 multi-thread tests aborted with "scheduler resumed a done actor" (error 5).
Thin LTO only worked because it picked the local-exec model.

- context.rs: module doc "Thread-locals and migration" (the rule), tls_fence()
- every TLS accessor reachable from an actor stack is out of line:
  actor::{current_pid, publish_outcome (new)}, preempt::{maybe_preempt,
  check_cancelled, current_slot_ptr, note_*, preemption_swap/enabled (new;
  replaces direct PREEMPTION_ENABLED.with in NoPreempt/RawMutex/with_runtime/
  trace/debug asserts)}, context::{get,set}_scheduler_sp, runtime::{sched_slot,
  slot_push, set_yield_intent}, causal, trace
- raw_mutex order checks: out of line only under debug_assertions
- Cargo.toml: [profile.reltest] = release without LTO; 41 test bins link in
  ~1.5 s each instead of ~12 s (suite 11 min -> 1.5 min) AND it is the
  regression oracle for this bug: `cargo test --profile reltest`
Both profiles: 41/41 + doctests green.
2026-08-21 12:23:46 +00:00
claude-asm-audit d300a9d536 docs: fix 10 rustdoc link warnings, deny broken/private/redundant intra-doc links, doctests deny(warnings) (3 doctests had unused vars/imports) 2026-08-21 12:23:46 +00:00
claude-asm-audit bc0a5e8656 scheduler: spawn_monitor / spawn_monitor_with — monitor registered on the child's slot before publish, so the Down always carries the real reason (spawn-then-monitor could race to NoProc); tests; channel test uses it 2026-08-21 12:23:46 +00:00
claude-asm-audit ea2b222cff runtime: wake_slot default ON (RFC 005 accepted, finding 18); supervisor stops survivors sequentially so reverse-order teardown is guaranteed; deflake gen_server/channel tests; baseline.json regenerated slot-on; ROADMAP notes 2026-08-21 12:23:33 +00:00
claude-asm-audit 7001f04b65 diag(runtime): wake-path counters in SchedulerStats + RuntimeStats::wake_diag; RQDIAG/DIAG lines in rq_runtime and general (target 5, finding 17) 2026-08-21 12:22:46 +00:00
claude-asm-audit 77938cd31d bench: split ping_pong_oneshot — spawn_pair_control + ping_pong_steady; refresh baseline (job e982b5b5)
general.rs sections 5/6: spawn_pair_control is ping_pong_oneshot with the
messages removed (2 spawn + 2 join per round); ping_pong_steady is one
persistent pair × 10k roundtrips over unbounded MPSC (smarm::channel vs
tokio::sync::mpsc::unbounded_channel). Box (5900X, 3729fff, rq-mpmc):
oneshot 804/control 633 → 79% spawn+join; steady 142 ns vs tokio 128
(0.90×) at 1T, 1.8 µs/roundtrip at 20T. history.md finding 16.

baseline.json regenerated from the same run (20T labels, 14 benches).
2026-08-21 12:22:46 +00:00
claude-asm-audit 5ddd122711 perf(preempt): rdtsc unserialised by default; causal attribution opts into lfence
reset_timeslice paid an lfence pipeline drain on every resume via the
shared rdtsc() helper. Nothing in preempt.rs needs it: the timeslice
arm/expiry compare against a ~1e5-cycle slice, and an early stamp only
makes the slice look more used. The consumer that does need it — causal
site attribution, where a speculative early read misattributes a site's
tail — now calls rdtsc_serialising() explicitly (cold_check sample,
SiteGuard enter/exit).

Measured vs 31dc26a (baseline commit), 24-core box, rq-mpmc default:
- switch_cost, taskset -c 2, 9 interleaved old/new pairs: mean_cyc
  149 -> 117 per roundtrip (-32, -21%); mean_ns 43.2 -> 34.9.
- sweep.py regress + run, 20T, two runs agree:
  yield_in_hot_loop 1T 40171 -> 30277/30497 us (-24%)
  yield_many        1T 12513 ->  9965/10013 us (-20%)
  ping_pong_oneshot unchanged (RFC 005 wake-slot handoffs skip
  reset_timeslice, so that path never paid the lfence).
  All other rows within the box's ~+-15% noise floor.
- 1-core sandbox: 218 -> 206 cyc (understates by ~3x).
- bench binary lfence count 24 -> 4 (survivors = the bench's own
  rdtscp;lfence bracket). Tests green.
Baseline (benches/baseline.json) not re-saved in this commit.
2026-08-21 12:22:46 +00:00
claude-asm-audit 343e53e17b bench: baseline under the fixed rq-mpmc default (job 6e71e9b1)
20-core box (taskset 0-19), 5 sets, at 0438f12. Replaces the rq-mutex
baseline (df41ff3) so sweep.py regress compares like with like. The
multi-thread rows show what the mutex contention was costing:
yield_many 20T 156.7ms -> 44.1ms (-72%), spawn_storm_busy 20T -72%,
chained_spawn/catch_unwind 20T -20%. Sweep ran clean (0 panics) —
the finding-13 fix holding under the full suite.
Note: mpsc_contention 1T is ~2x the mutex-default row (3155 -> 6123);
backend characteristic to investigate, not a fix regression (the
shootout shows the fix itself is noise-neutral).
2026-08-21 12:22:46 +00:00
claude-asm-audit 9b215573de fix(run_queue): make the Vyukov rings preemption-tolerant (finding 13)
A consumer OS-preempted between its dequeue_pos claim and its seq
release freezes one cell; once traffic laps the ring (~cap ops ≈ 1-2ms
at yield-storm throughput ≈ one scheduling quantum under load),
try_push reads the stale seq and the original algorithm's 'lap behind
=> full' inference misfires. The old push assert then converted that
liveness stall into an abort blaming a double enqueue that never
happened (soak: occupancy 174-181 of cap 16384 at every failure), and
the dead scheduler threads stranded actors => the observed hangs.
Pristine 5504ef3 failed 14/14 under an 8-spinner soak on the 20-core
box; a retry prototype passed 13/14, its one failure being a stall
that outlived a fixed 1M-spin bound — pure spinning starves the
descheduled consumer, so the wait must yield.

Fix, following crossbeam ArrayQueue's shape (fence + opposite-counter
check; cells hand off, COUNTERS give verdicts):
- push: on a lap-behind cell, fence(SeqCst) + occupancy check.
  occ < cap => transient stall => spin-then-yield backoff and retry;
  occ >= cap => the REAL at-most-once-enqueued violation => panic with
  a truthful message and the counters. enqueue_pos is loaded before
  dequeue_pos so racing pops only underestimate occupancy (no spurious
  panic).
- pop: symmetric counter check before an empty verdict; on a
  mid-publish producer, bounded wait then None — deliberate deviation
  from crossbeam's unbounded retry (spurious None is benign: RFC 018's
  enqueue-wake self-heals; StripedRing's probe must not hang on one
  stripe).
- StripedRing::push probe: yield-escalating backoff after a full
  refused lap (was a bare spin_loop).
- Hand-rolled Backoff (spin 2^n to 64, then yield_now); under loom
  every wait is a yield so models explore the stalled peer's progress.

New loom model mpmc_lap_onto_stalled_consumer_completes reproduces the
old panic in the first explored interleavings (verified FAILED against
5504ef3) and passes with the fix. Lib 60 + integration 41 + all 4 ring
loom models pass. rq-striped inherits the fix (stripes are MpmcRings).
2026-08-21 12:22:46 +00:00
claude-asm-audit bf55cef3e3 perf(run_queue): flip default backend rq-mutex -> rq-mpmc
Evidence (history.md session 5, findings 10-11, 20-core jobrunner run):
rq-mutex collapses with thread count on queue-heavy load (yield-storm
6.9x slower than striped at 20T; attributed cause of the baseline's
multi-thread regression on yield_many/chained_spawn), while rq-mpmc is
best-or-close everywhere: best 1-thread, best slot-on ping-pong
(1237us vs mutex 6754us at 20T), -23%/-41 cyc per yield roundtrip on
real hardware (single-core switch_cost, interleaved). rq-striped
remains the churn-heavy many-core option, selectable per build.

Docs + compile_error hints in run_queue.rs updated to name rq-mpmc as
the default. slot_state.rs untouched, no loom-relevant changes; lib
(60) + scheduler/channel/supervisor/park_wake/wake_slot/preempt (41)
pass under the new default.
2026-08-21 12:22:46 +00:00
claude-asm-audit 2f88264426 bench: refresh baseline.json — 20-core box, fe85197, rq-mutex
Re-measured via jobrunner (taskset -c 0-19 of 24, rust:1.97-slim,
sweep.py run --save-baseline, 5 sets). Supersedes the old 24-thread
baseline; multi-thread labels are now 'smarm 20-thread'. Taken under
the rq-mutex default *before* the backend flip — the yield_many /
chained_spawn multi-thread regression reproduces here and is attributed
to rq-mutex contention (history.md session 5, finding 10).
2026-08-21 12:22:45 +00:00
claude-asm-audit 6172f4231d perf(runtime): skip take_closure's locked swap after first resume
Every resume paid an unconditional AtomicPtr::swap (lock xchg, full
barrier) to check for a first-resume closure that is null on all
resumes after the first. A Relaxed null-load fast path is sound:
store_closure runs only before publish_queued, whose Release pairing
with try_claim's Acquire orders it before this call, so no writer can
race the load within an occupancy.

Measured on switch_cost (1-core sandbox, rq-mutex, cycles): mean
roundtrip 350-355 -> 324-328, ~7.5%. All lib + scheduler/channel/
supervisor tests pass.
2026-08-21 12:22:45 +00:00
claude-asm-audit d7082eb266 perf(context): pass actor sp through registers, not TLS
switch_to_actor takes the target sp in rdi and returns the actor's
next saved sp in rax (handed over by switch_to_scheduler's shim).
Deletes the ACTOR_SP thread-local and halves the helper calls per
one-way switch (2 -> 1); the scheduler loop also drops its
set_actor_sp/get_actor_sp TLS round-trips. SCHEDULER_SP stays: a
yielding actor at arbitrary call depth has no argument channel back.

asm before/after in outputs/history.md session 1-2. Tests: 354 pass.
2026-08-21 12:22:45 +00:00
Claude (sandbox) 741c10337b release: v0.7.0
Bump crate version to 0.7.0.
2026-08-21 13:10:25 +02:00
Claude (sandbox) e570138da5 docs(roadmap): supervisor start order is not start readiness
Filed from the urus v0.3 endpoint work. start_child spawns and moves on,
so a later sibling can whereis an earlier named child before that child's
actor has run. Notes why blocking spawn is not the fix ('has begun
executing' != 'has bound its name', plus a per-accept round-trip tax and
every spawn becoming a context-switch point), that OTP has the same async
spawn and synchronises one level up in gen_server:start_link, the
readiness-ack shape if scheduled, and the structural workaround urus uses
today (registrar spawns its own consumers).
2026-08-20 13:20:42 +00:00
Claude (sandbox) 415effb2e9 feat(gen_server,gen_statem): lifetime is the actor's — refs are addresses; inline named run
Root cause behind the "pin the endpoint" gotcha and the trapping-wrapper
pattern: a gen_server had two lifetime authorities — its refs (last one
dropped → inbox closes → exit) and, when supervised, its supervisor. OTP has
one: a process lives until it stops, is shut down, or is killed; a pid is an
address. Root exit now shutting down every forest root removes the reason the
ref-governed idiom existed (a forgotten server no longer hangs the run), so
adopt the one rule:

- The server/machine loop holds one inbox sender for its life; the inbox
  never closes. GenServerRef / GenStatemRef are addresses. Explicit close is
  `shutdown()`; a forgotten one is swept at root exit.
- `NamedGenServerBuilder::run()` / `gen_statem::run_named(name, m)` run the
  loop inline as the current actor: a server is a direct ChildSpec child,
  gets the supervisor's shutdown as handle_shutdown / a shutdown row, re-binds
  its name on restart, and is addressed by name. The wrapper in
  examples/graceful_shutdown.rs is gone.
- gen_statem gains GenStatemName + whereis_machine/send/call/shutdown by name
  (parity with gen_server); the macro gets `Sm::new`.
- Root-exit sweep records `Event::RootSweep { target, trapping }` under
  smarm-trace ("root_sweep shutdown|stopped"): unsupervised leftovers are
  visible rather than silently owned-by-refs.
- Named start() name-clash path stops the spawned actor instead of relying
  on ref drop.

Tests: tests/gen_server_lifetime.rs, tests/gen_statem_lifetime.rs,
tests/root_sweep_trace.rs (feature-gated); three existing tests that used
drop-closes-inbox now use shutdown(). Docs/README/ROADMAP/Deep Dive updated.
2026-08-19 17:47:31 +00:00
Claude (sandbox) 849a424c8e docs,examples: graceful shutdown — new examples/graceful_shutdown.rs, README 'Stopping actors', named_genserver uses shutdown(), Deep Dive terminate note, ROADMAP open items 2026-08-19 16:28:42 +00:00
Claude (sandbox) 6ceb138f5f feat(gen_statem): graceful-shutdown parity — trap_exit, shutdown/exit rows, stop, terminate
Mirrors the gen_server surface in gen_statem's event model:
- Cx::trap_exit() (in the initial enter): a shutdown request then arrives
  as the Shutdown event, routed by state through `shutdown` rows (default
  for a state with no row: stop); linked-peer deaths as `exit <pat>` rows
  (default: drop). Non-trapping machines are stopped outright, as before.
- Cx::stop(): normal self-exit after the current event; `stop` tail keyword
  is sugar for { cx.stop(); prev }.
- Machine::terminate (optional `terminate { … }` macro block), run from a
  Drop guard on every exit path; the guard also drains armed timers.
- Machine::shutdown_ev / exit_ev (defaults None) so hand-written machines
  keep compiling; GenStatemRef::shutdown() is graceful and waits.
- Loop selects exits > timers > inbox.

Tests: tests/gen_statem_shutdown.rs.
2026-08-19 16:25:15 +00:00
Claude (sandbox) 250f31265b feat(runtime): root exit is graceful shutdown of the forest roots
The RFC 014 root-exit sweep hard-stopped every live slot once nothing was
runnable. That deferral privileged queued work over parked-with-a-pending-
wake work (a sleeper was killed, a queued cast was drained) and any attempt
to widen the notion of pending wake (timers, fd readiness) re-wedges the
run on the periodic-timer daemon the sweep exists to end.

Root exit now means "the program is done": finalize_actor delivers
request_shutdown to every forest root — each live actor whose parent is
the run (ROOT_PID) or is dead — synchronously, before the live-count
decrement. Supervisors cascade per child Shutdown policy; trapping actors
may Continue/drain with working timers and end the run when they stop
themselves; non-trapping actors are stopped outright. No forcing sweep.

Removes root_exited/root_swept, Pop::RootDrain and the idle-verdict
condition; adds tests/root_exit.rs.
2026-08-19 07:11:46 +00:00
claude 9c8f59ca53 feat(scheduler,supervisor,gen_server): graceful shutdown — request_shutdown, child Shutdown policy, handle_shutdown
Lift OTP's `exit(Pid, shutdown)` + child-spec `shutdown` wholesale.

scheduler / runtime
- `request_shutdown(pid)`: the polite stop. A target trapping exits gets an
  `ExitSignal { reason: DownReason::Shutdown }` on its trap inbox and keeps
  running; a non-trapping target is stopped as by `request_stop`, which is
  now documented as the hard stop (`exit(Pid, kill)`). Dead pid: no-op.
- `RuntimeHandle::request_shutdown` for the off-runtime (signal thread) path;
  `from == ROOT_PID` there.
- `DownReason::Shutdown` — appears only in ExitSignal, never in Down (a
  complying target exits *normally*).

supervisor
- `ChildSpec::shutdown(Shutdown::{BrutalKill, Timeout(d), Infinity})`,
  default Timeout(5s). Every supervisor-initiated stop (ordered shutdown and
  OneForAll/RestForOne sibling cycling) is: request_shutdown → await the
  child's Signal up to the grace → request_stop → await. Sequential, reverse
  start order.
- The supervisor traps exits; a Shutdown ExitSignal runs the ordered
  shutdown and `run()` returns normally, so `request_shutdown(root_sup)`
  tears a whole tree down top-down with each child's grace period.
- FIX: a hard `request_stop` on a supervisor previously orphaned its
  children (the ordered shutdown lived after the loop, and the unwind
  skipped it). `Live` (the by_pid map) now carries a drop guard that
  fire-and-forget hard-stops live children when unwinding.

gen_server
- `GenServerCtx::trap_exit()` opt-in in `init`; the trap inbox becomes arm 0
  of the loop's select. Shutdown ExitSignal → `handle_shutdown() ->
  ShutdownAction::{Exit, Continue}` (default Exit: loop breaks, `terminate`
  runs on the normal path and may block). Other ExitSignals →
  `handle_exit(sig)`.
- `GenServerCtx::stop_handle() -> StopHandle`, `stop()` ends the server
  after the current message with a *normal* exit — the missing
  `{stop, normal, State}`; `request_stop(self_pid())` was the only self-exit
  and it is abnormal (Transient restarts it).
- `GenServerRef::shutdown()` / `gen_server::shutdown(name)` now go through
  `request_shutdown`.

Tests: tests/shutdown.rs, tests/supervisor_shutdown.rs,
tests/gen_server_shutdown.rs. Full suite green; fmt + clippy --lib clean.
2026-08-19 06:34:48 +00:00
Claude (sandbox) 1002777ef3 feat(channel,runtime): off-runtime cross-thread wake for parked receivers and stop
Wakes issued from a non-scheduler OS thread were silent no-ops. Every
off-runtime wake primitive (unpark, unpark_at, request_stop) reaches the
runtime through the RUNTIME thread-local, which is unset on any foreign
thread — so a send from a plain std::thread enqueued its message but never
woke the parked receiver, and there was no way to drive a stop into a
runtime from an application thread (e.g. an OS-signal handler). The former
strands a parked recv forever; the latter is why a downstream server must
poll a shutdown flag instead of parking on it.

Generalize RFC 018's rule — a producer reaches the runtime through a Weak it
holds — from the IO backend to channel senders and to a new handle:

- A receiver captures a Weak<RuntimeInner> (provably live at that moment)
  alongside its (pid, epoch) when it parks. send() and the last-sender drop
  wake through scheduler::unpark_at_via, which takes the thread-local path
  when on a scheduler thread (preemption-gated, slot-eligible) and the
  captured Weak otherwise — the same cross-context wake the IO threads do.
- Runtime::handle() returns a Send + Sync RuntimeHandle carrying that Weak;
  RuntimeHandle::request_stop drives a cooperative stop from any thread and
  is a no-op once the runtime is dropped.

The in-runtime wake paths (recv/select timers) are unchanged; only the sites
reachable from a foreign thread route through the Weak. RuntimeHandle exposes
request_stop only: send-wake needs no user-facing handle, and off-runtime
unpark is covered because request_stop drives unpark on the upgraded inner.

tests/cross_thread_wake.rs: a foreign-thread send wakes a parked receiver; a
foreign-thread request_stop wakes and stops a parked actor; a RuntimeHandle
held across and beyond run never blocks all-done and degrades to a no-op.
2026-08-19 05:50:54 +00:00
claude-asm-audit 4c0e42152f feat(cluster): RFC 010 c10–c16, follow-ups and Phase 6 (squash of 16ef583..d9c62a8)
Tree snapshot of d9c62a8 (2026-08-18). The 20 source commits between
16ef583 (c9) and d9c62a8 were never pushed and the clone that held them
was lost; this commit carries their combined tree verbatim so the build
history stays auditable from the c1–c9 commits below it. Original
hashes as recorded in the session handoff:

  c10  f03e94d  pid targeting + auto-serialization (RemotePid, D14 name
                on the wire); Phase 3 gate
  c11  7ef4bad  DownReason::Disconnected, wire tag 5
  c12  d124162  remote monitors (Monitor/Demonitor/Down frames)
  c13  9de967b  connection-loss synthesis (A+B: Monitors::teardown +
                unread-command Disconnected); Phase 4 gate
  c14  7e822b7  eager pg eviction (reaper actor, ReaperInboxes)
       dbe1a22  InboundVerdict::label(), trace::Event::ClusterInbound
       31a9877  tests/channel.rs monitor-churn target gated on `go`
       653559e  Discovery::Withdrawn{name, addr}
  c15  b41d76e  distributed pg: Sync on NodeUp, Join/Leave broadcast,
                NodeDown sweep, members_all; PgMsg wire type
  c16  fafa881  pick_any / dispatch_any; Phase 5 complete
  Phase 6 Tier A:
       195c73e  p4  NodeEvent::NodeDown(NodeInfo)
       48fd766  p1  connector Candidate{name, addr, state}
       ce8cf99  p2+p7 conn.rs select arms as Vec<Arm>; Outbound::Drained
       bf24988  p6  RemotePid::from_local -> Option
       9ae0380  p3  PeerStanding{Free, Claimed, Dialing}
       c7d62a1  p11 cluster::Timing knobs, threaded by value
       46f171d  p11 cluster_disconnect un-ignored on SMARM_FAST_TIMING
  Phase 6 Tier B:
       391a9ae  p5  cluster::RemoteDownReason{Local, Disconnected};
                    DownReason::Disconnected removed from core
       7ddd908  p9  pg ctl channel unconditional, one cfg seam at spawn
       d9c62a8      PeerNameMismatch parks the candidate; ClusterDial trace

Verified at d9c62a8: default 361/0, cluster 448/0, clippy --lib on
default / cluster / cluster+smarm-trace, fmt, 10x flake on
cluster_dial_mismatch, 5x on cluster_pg.
2026-08-18 16:00:00 +02:00
smarm 8f2d513940 README: rewrite overview, add limitations, roadmap, and contribution notes
Expands the intro into an overview/limitations structure, documents
preemption, tracing/causal profiling, URUS, and adds a 'Coming up' and
'A note on open source' section.
2026-08-18 00:25:25 +02:00
Claude 16ef583455 feat(cluster): RFC 010 c9 — remote Name sends: outbound table + the one inbound seam
Outbound (D13, ratified 2026-08-15): a module-private table node → dedicated
Sender<Frame>, populated/torn down by the manager inside the same serialized
handlers that own connection lifetime (Register / Disconnect / reap /
terminate), on RuntimeInner beside the exposure state. remote::send is one
leaf-lock lookup + one channel send: no gen_server on the data plane (rejected
send-through-manager — worse than the c8 argument, it is the data plane), no
published channel a leaked conn pid could inject raw frames into (rejected
send_dyn-at-conn-pid — module privacy cannot fence a published channel).
Ok(()) = handed to the connection's inbox, local knowledge only (RFC §3);
missing entry / closed channel = NotConnected. The outbound sender is a
SEPARATE channel from cmd_tx on purpose: a clone would let the table hold
lifetime authority (D9 violation); its closure is not a stop signal.
Buffering toward a slow peer is unbounded (BEAM busy_dist_port shape);
backpressure out of scope, documented not silently absent.

Inbound: remote::deliver_named is THE resolution seam (RFC v2) — exposed-set
check (unexposed = unreachable, the safety) → hash check against what the
name was exposed with → registry whereis → c8's decode_deliver. Verdicts are
local-only (InboundVerdict, discarded by the conn actor for now; a trace
hook is the place). The conn actor gains a third select arm (outbound
inbox → wire) and interprets SendNamed; Send/Monitor/Down are consumed for
liveness and ignored until c10/c11.

Public surface: RemoteName<M> (node, Name<M>), RemoteSendError, NotConnected,
send, send_remote_raw (untyped escape hatch so tests can put deliberately
wrong frames on the wire — the typed API cannot express a hash mismatch).

CORE (found the hard way this chunk): publishing a second channel of the
same message type on one live actor silently replaced and CLOSED the first,
so a select/recv on it returned 'closed' immediately forever — a hot loop
starving the single-threaded scheduler. publish_channel now asserts when an
existing same-type channel's receiver is still alive AND the new sender is
not a clone of it (Sender::same_channel via Arc::ptr_eq); replacing a
dead-receiver channel stays silent (an actor re-registering after dropping
its inbox is legit). Message names the sanctioned shapes. Three registry
tests pin panic / cloned-sender-ok / dead-receiver-ok. Full default suite
clean.

Harness: wait_line drains stderr before dumping on timeout/EOF (eprintln!
diagnostics no longer vanish); Node::transcript() accessor for
ordering-proof assertions.

tests/cluster_remote_send.rs 1/0, 10/10 flake runs, subprocess harness:
cross-node name-send delivers; unexposed name unreachable (registered
locally, never delivered); wrong hash never misroutes (unknown-hash and
known-hash-wrong-channel flavours); send to an unconnected node =
NotConnected locally; every frame-bearing send is Ok — 'handed to
transport', asserted and documented at the test. Negatives are proven by
STREAM ORDERING (they precede the positive on one in-order connection),
not by sleeping. All cluster suites regression-clean (envelope 15,
handshake 11, transport 11, lifecycle 1, liveness 3, connect 9, two_node 3,
membership 4, mesh 2, expose 5); clippy --lib green both configs; fmt clean;
default build compiles.
2026-08-15 21:22:56 +00:00
Claude a7f98f8d48 feat(cluster): RFC 010 c8 — exposure registry + fixed-seed type hashing
Nothing local is remotely reachable by default (RFC §4). expose(Name<M>)
marks a name remotely addressable and registers M's decoder under
type_hash::<M>(); expose_type::<M>() registers only the decoder (the
reply-to path). exposed_names() is the auditable remote surface.

D3's watchable fold, resolved against the code as it stands (stated in the
module docs): register() ALREADY stamps every named holder watchable ('no
successfully-registered actor can die unflagged', registry.rs), so an
exposed name's holder needs no extra mark — and re-registration after a
holder's death re-stamps the new holder for free, which a per-tenancy mark
taken at expose time could not do. The cluster's own mark_watchable
set-site is therefore the pid crossing the wire (frame serialization, c10)
— the exact analog of the membrane crossing. c8 adds only the name/type
state neither the registry nor slot bits can carry. No new pid registry;
RFC §4 honored.

One-viable calls, flagged:
- State lives on RuntimeInner (the pg pattern: leaf RawMutex field,
  cfg-gated behind cluster, zero-cost-when-off per c1) — c9's inbound
  decode consults it per frame; manager-held state would serialize every
  remote delivery through one gen_server.
- type_hash = FNV-1a 64 (fixed seed: the offset basis) over TypeId: a
  constant of the binary — stable across runs of the same build (the scope
  the build-hash handshake reduces the mesh to), deliberately not across
  builds. Collisions degrade to decode error / refused channel, never a
  misroute (the NoChannel guarantee, RFC §3).
- Decoder = decode-and-deliver-to-pid Arc closure capturing M (the one
  typed site): decode_payload then send_dyn. Wire-name → pid resolution
  stays OUTSIDE — that is c9's single seam, which calls decode_deliver.
  Arc so the call happens with the exposure lock RELEASED: send_dyn takes
  the registry lock, a mutual Leaf (the runtime asserts on nesting — caught
  live by the first test run).
- expose is a name-level fact, valid for an unregistered name (names
  late-bind; c9 resolves per delivery).

tests/cluster_expose.rs 5/0 stable x5, purely local per roadmap:
exposed/unexposed lookup + audit listing; decoder registration and the
delivery contract (happy path into a registered String channel; unknown
hash; corrupt bytes; wrong channel refused — never misrouted); distinct
types distinct hashes; expose/bridge-crossing agreement via the shared
watchable observable (terminal_reason after holder death); hash stability
across runs in the same binary via a c4-harness re-exec. Payload types are
std types — the crate's serde is derive-less by design, user crates bring
their own derive.

All cluster suites regression-clean (envelope 15, handshake 11, transport
11, lifecycle 1, liveness 3, connect 9, two_node 3, membership 4, mesh 2);
clippy --lib green both configs; fmt clean; default build compiles.
2026-08-15 07:25:42 +00:00
Claude 1282c3a08d feat(cluster): RFC 010 c7b — discovery Strategy, static seeds, connector dial loop
Phase 2 gate: 3-node mesh under the subprocess harness, repeatable (10/10).

Strategy (ratified): push-based, spawned as its own actor by the connector —
it emits Discovery events into a channel whenever it learns something and
may run forever; the connector owns all retry/backoff state. StaticSeeds
announces its list once and exits. Discovery is #[non_exhaustive] and
additive-only (candidates announced, never withdrawn) so expiry can land
later without breaking strategies.

One-viable correction to the ratified Discovery shape, flagged: a candidate
is a (name, addr) PAIR, not a bare address. The dial path and the D7
tie-break are keyed by peer name (the dial intent must be registered before
connecting so a crossing inbound Hello sees it), so an anonymous dial would
reintroduce exactly the simultaneous-connect flap D7 exists to prevent.
Discovery mechanisms know names — that is what they discover.

Connector: plain select-loop actor (the c6 shape) folding cmd inbox,
discovery stream, membership stream, and the earliest retry deadline into
one wait. It tracks who is up by SUBSCRIBING TO MEMBERSHIP like any
consumer — first consumer of c7a's snapshot-then-stream surface, no
privileged channel into the manager. Backoff: 250ms doubling to a 5s cap
(the c6c class of one-viable constants), reset on node_up; node_down
schedules a prompt redial with a fresh sequence. A candidate bearing the
local name is parked (that seed is us); every other failure retries — in
particular NameTaken can be our own ghost at the peer, not yet reaped by
its liveness timer, so it must not park. Dials run inline in the loop, the
acceptor's deliberate serialization (each attempt bounded by the connect +
handshake deadlines).

cluster::start(Config {node_name, meta, listen_addr, strategy}) is now the
integrated node start: supervised manager + acceptor + connector. It
completes the node identity: build_hash = cluster::BUILD_HASH (first
consumer, closing the c6d loose end) and incarnation = self_incarnation()
— unix-epoch MILLIS truncated to u32, not seconds: a supervised
crash-and-restart inside one second is routine, and seconds would collide
the ghost with its successor. Cluster handle: local_addr()/local()/
shutdown(); drop stops acceptor+connector loops, manager subtree detaches
(same split as AcceptorHandle alone).

Roadmap-binding, asserted in review: no consumer touches the connection
table — Manager.conns and ConnEntry stay private; the only exposures are
Call::Peers (sorted names, pre-existing) and the membership surface.

tests/cluster_mesh.rs 2/0, 10/10 flake runs: (1) 3-node mesh forms; kill
one (SIGKILL via Drop, per the retractable-state trap: roles park forever)
=> node_down at both survivors; restart same name => new incarnation at
every observer, distinguishable from the ghost; (2) seed unreachable at
start (pre-reserved closed port; accepted micro steal-window, documented)
then arriving later => edge forms via the retry path. All cluster suites
regression-clean (envelope 15, handshake 11, transport 11, lifecycle 1,
liveness 3, connect 9, two_node 3, membership 4); clippy --lib green both
configs; fmt clean; default build compiles.
2026-08-15 07:14:13 +00:00
Claude 160967939b feat(cluster): RFC 010 c7a — membership events + view at the manager
node_up/node_down are derived facts of the manager's own register/remove
events, so the membership state lives in the manager (no cross-actor race
between 'connection exists' and 'node is up'); src/cluster/membership.rs is
the consumer surface: NodeEvent/NodeInfo, subscribe(), view(). The conn
table stays private — no consumer touches it (roadmap-binding).

Ratified semantics: subscribe() is snapshot-then-stream — one NodeUp per
live peer is queued before the subscription joins the list, exact because
gen_server handlers are serialized. Dropped subscribers are pruned on the
next emit (closed channel), no monitor needed.

One-viable call, flagged: NodeId is memoized per (name, incarnation) — a
compact local alias for the wire identity, per pg.rs's framing. A reconnect
blip at the same incarnation keeps its id; a restart (new incarnation) gets
a fresh one, so a ghost and its successor are always distinguishable.
Allocation starts at 1; NodeId(0) stays pg::DEFAULT_NODE_ID (self).

Call::Register now carries the whole handshake Peer (the path already has
it; node_up needs incarnation + meta).

tests/cluster_membership.rs 4/0 stable over 5 runs (live up/down over
localhost TCP with commanded and EOF teardown; late-subscriber snapshot +
view agreement; restart-vs-blip id identity; dead-subscriber pruning). All
cluster suites regression-clean; clippy --lib green both configs; fmt
clean; default build compiles.
2026-08-15 07:07:14 +00:00
Claude 112d6b2e65 feat(cluster): RFC 010 c6d — build_hash derivation
cluster::BUILD_HASH, the derived value for LocalNode::build_hash: a
compile-time const, FNV-1a 64 over the build script's input string
(rustc -V + sorted enabled CARGO_FEATURE_* set) with PROTO_VERSION folded
in as a continuation of the same state — a proto bump moves the hash even
on an identical toolchain. build.rs emits only the raw inputs
(SMARM_BUILD_HASH_INPUTS); the hashing lives in src/cluster.rs next to
PROTO_VERSION rather than build.rs parsing it out of a source file. The
env var is emitted unconditionally — one string costs the default build
nothing, and the const itself is behind the cluster feature with the rest
of the module.

Domain is the flagged lean (one-viable-answer, veto at diff review):
toolchain + declared features + proto version. Tightenable later without a
wire change — it is just a u64 on the Hello.

Unit tests: FNV-1a 64 published vectors (empty/'a'/'foobar'); the proto
fold is a continuation of the same FNV state and moves the hash;
BUILD_HASH is const-evaluable and nonzero. All cluster suites
regression-clean (7 files); clippy --lib green both configs; fmt clean;
default build compiles.
2026-08-15 06:37:57 +00:00
Claude ad4958421f feat(cluster): RFC 010 c6c — heartbeat send + fixed-timeout liveness + teardown
The timeout arm of the connection actor's select: HEARTBEAT_INTERVAL (1s)
paces outbound Frame::Heartbeat (first at spawn, so the peer's window
starts fed) and LIVENESS_TIMEOUT (4s = 4 intervals) declares the peer dead
when no inbound frame arrives inside it — any frame resets the window, so
heartbeats keep an idle connection alive and real traffic (c8+) counts for
free. Fire => close + exit; the manager's monitor reaps the table entry as
on every other exit path. Fixed timeout per RFC v2 §5 (control connection,
heartbeats can't queue behind bulk). Intervals are the one-viable-answer
call flagged for veto at diff review.

The pump was made non-blocking to keep the deadlines honest: a plain recv()
blocks into the socket while the buffer holds a partial frame, parking the
actor past its timers. Two additive FramedConn methods (read_once,
next_buffered): exactly one socket read per level-triggered readable wake
(cannot block, cannot strand — leftovers re-signal), then drain every
complete buffered frame. Liveness resets only on complete frames.

No-fd transports (loopback) still get the command-only loop: no readiness
means no timers, same caveat as recv_deadline.

tests/cluster_conn_liveness.rs 3/0, stable over 5 runs (raw far end over
localhost TCP: heartbeats appear unprompted; mute peer still up at half
the window, gone after it; heartbeat-only peer survives 1.5x the window,
then reaped once silenced). Cluster suites regression-clean; clippy --lib
green both configs; fmt clean.
2026-08-15 06:36:07 +00:00
Claude c8ed858e4c feat(cluster): RFC 010 c6b — handshake on the accept/connect path
Drive the c5 machines as straight-line code on the path (D8): dial_handshake
and accept_handshake do the IO on a shared FramedConn, and a connection actor
is spawned only after a successful handshake. Rejects, tie-break losses (D7),
protocol faults and timeouts are all resolved on the path by closing, so no
actor ever exists for a connection that did not establish. The whole
FramedConn travels into spawn_established, carrying any read-ahead past the
handshake frames.

Handshake deadlines land here rather than in c6c: FramedConn::recv_deadline
enforces them between reads via the connection's fd arm, so a peer that
connects and goes silent cannot wedge the acceptor.

Connection lifetime moves to the manager (pulled forward from c7). The path
registers each established connection and hands over its ConnHandle; the
manager owns it, monitors the actor, and tears the connection down on
Disconnect, on peer close, or at manager shutdown. spawn_established returns
a Pid, so a connection neither outlives nor dies with whichever actor
established it — the ownership that made two-node teardown unorderable.

The manager also tracks in-flight dial intents, monitored so a panicking
dial cannot wedge the tie-break, and answers HelloCtx for the accept path.
2026-08-14 21:09:08 +00:00
Claude bbaaa062e3 feat(cluster): RFC 010 c6a — connection actor + manager subtree
Per-peer connection actor as a single select-loop plain actor owning the
whole FramedConn: one select folds its command inbox and the transport's
readable arm, so reads and control share one execution context — no reader
thread, no read/write split. The handshake is bypassed here (c6b wires it);
the actor is spawned already-established and self-registers with the manager.

Manager gen_server: the peer-name -> conn-pid registry and the uniqueness
source the handshake's NameTaken depends on. It monitors each connection, so
the table self-heals on any exit path. Explicit supervision subtree keeps the
manager up; connections are dynamic and monitored, never restarted (c7
re-dials).

Transport gains an additive Conn::readable_arm -> Option<FdArm> (default None;
TCP returns its fd's arm, loopback stays None). Existing c3 transport tests
unchanged.

Lifecycle test over localhost TCP: up reflected in the table, commanded
shutdown reaps exactly one, peer EOF reaps the other.
2026-08-14 19:23:39 +00:00
claude 9e49038474 feat(cluster): RFC 010 c5 — handshake as a pure state machine
Frames in, actions out — no IO, no clocks, no actors; the c6 connection
actor will drive it. Initiator (dial: emit Hello, interpret the single
response) and Responder (accept: judge the first frame) as consuming-self
machines; check order proto -> hash -> name -> tie-break. Driver-supplied
HelloCtx carries the two facts the pure machine cannot know (name claimed,
own dial in flight). Tie-break ratified as a wire fact: the smaller name's
dial survives; the losing inbound closes silently (both ends compute the
same verdict, no reject frame needed). Peer's own name offered => NameTaken.
build_hash is config-supplied; derivation lands with c6.
2026-08-14 17:17:14 +00:00
Claude 8a9e2b81b1 test(cluster): RFC 010 c4 — subprocess two-node harness
The runtime is a process singleton, so multi-node tests mean multiple
processes. tests/common/mod.rs is the reusable harness (precedent: RFC
019 c6 / tests/stack_diag.rs self-re-exec, extended to live tailing):
re-execs the current test binary as named roles, tails stdout/stderr on
reader threads, waits on protocol-visible lines with bounded timeouts
(panic dumps carry the full transcript), and Drop SIGKILLs+reaps so a
panicking test leaves no orphan or zombie. Children run with
--test-threads=1 --quiet --nocapture; the last flag is load-bearing —
libtest's capture would otherwise swallow role output.

Port assignment is race-free by construction: children bind port 0 and
announce the concrete address (LISTENING <addr>); the parent never
pre-picks.

Smoke suite per roadmap: two real nodes, handshake-less TCP connect
through the real framed codec (one Heartbeat across, clean close seen on
both sides, both exit 0), plus reap-on-drop proven via ESRCH and
nonzero-exit surfacing. Flake budget stated in the module doc: 10 s
bound per wait, 10/10 clean at authoring, >1/100 failures = regression.
2026-08-14 14:46:28 +00:00
Claude 39ab92871e feat(cluster): RFC 010 c3 — transport trait, framed codec, TCP + loopback impls
The control-connection abstraction (RFC v2 §5): object-safe Transport/
Listener/Conn over opaque pre-resolved addresses (resolution stays the c9
seam), with FramedConn as the single shared byte->Frame codec feeding
Frame::decode's incremental contract. Nothing forecloses additional
per-peer connections for the jarred bulk plane; the membrane is not a
transport (D2).

TCP parks the calling actor via scheduler fd readiness (MSG_NOSIGNAL
writes, EINPROGRESS dial resolved through SO_ERROR). Loopback is the
shipped in-memory test transport: OS-thread-blocking condvar pipes with
TCP-shaped close semantics, per-instance address registry.

Conformance suite runs the same codec over both impls: roundtrips both
directions, framing across split writes, coalesced frames, peer-close
mid-frame as TruncatedByPeer (not EOF), clean close as Ok(None). Plus
impl-specific establishment/error cases and a 4 MiB cross-buffer TCP
frame under real backpressure.
2026-08-14 14:31:39 +00:00
Claude 3850f6099b feat(cluster): RFC 010 c2 — owned wire envelope
Frame enum per the RFC inventory; hand-rolled encode/decode with u32 LE
length prefix + u8 tag; strings u16-prefixed, payload blobs u32-prefixed;
MAX_FRAME_LEN cap (control plane never carries bulk, §5). Streaming decode:
Ok(None) = need more bytes, every Err = corruption. postcard confined to
encode_payload/decode_payload — the single codec seam (§2); postcard gains
the alloc feature for to_allocvec (still no_std-aligned, no default
features). DownReason/RejectReason travel as single tag bytes; Down tag 5
is reserved for c11's Disconnected.

Tests: per-frame roundtrip, back-to-back frames, golden heartbeat bytes,
zero-length payload, every-prefix incomplete, unknown frame/enum tags,
length prefix lying long (with and without bytes present) and short,
truncation mid-string, adversarial lengths (u32::MAX, cap+1, zero), UTF-8
corruption, payload seam roundtrip through a real Send frame.
2026-08-14 13:23:50 +00:00
Claude 58a2fe3046 feat(cluster): RFC 010 c1 — cluster feature flag + optional serde/postcard deps
Off by default; the default build stays libc-only. serde (payload contract)
and postcard (payload codec) are optional, no default features. src/cluster.rs
is an intentionally empty cfg-gated stub so the flag's default-build
invariance is reviewable in isolation.
2026-08-14 13:20:20 +00:00
79 changed files with 14941 additions and 1037 deletions
+21 -2
View File
@@ -1,6 +1,6 @@
[package] [package]
name = "smarm" name = "smarm"
version = "0.6.1" version = "0.8.0"
edition = "2021" edition = "2021"
rust-version = "1.95" rust-version = "1.95"
@@ -17,7 +17,7 @@ unwrap_used = "deny"
expect_used = "deny" expect_used = "deny"
[features] [features]
default = ["rq-mutex"] default = ["rq-mpmc"]
smarm-trace = [] smarm-trace = []
# RFC 007: native causal profiling. Zero cost when off (cf. smarm-trace): the # RFC 007: native causal profiling. Zero cost when off (cf. smarm-trace): the
# hook in `maybe_preempt` and the resume-path fast-forward compile away; the # hook in `maybe_preempt` and the resume-path fast-forward compile away; the
@@ -33,6 +33,11 @@ budget-accounting = []
# and unflagged; only the optional gen_server transport sits behind this, so a # and unflagged; only the optional gen_server transport sits behind this, so a
# release build pays nothing for an observer it never starts. # release build pays nothing for an observer it never starts.
observer = [] observer = []
# RFC 010 c1: clustering. Off by default — the default build stays libc-only,
# byte-for-byte (gate checked per phase). serde is the payload contract,
# postcard the payload codec; both minimal (no default features). Everything
# cluster-shaped lives behind this flag.
cluster = ["dep:serde", "dep:postcard"]
# Run-queue selection: exactly one, compile-time (see src/run_queue.rs). # Run-queue selection: exactly one, compile-time (see src/run_queue.rs).
# Non-default variants need --no-default-features (features are additive). # Non-default variants need --no-default-features (features are additive).
rq-mutex = [] rq-mutex = []
@@ -44,12 +49,18 @@ cc = "1"
[dependencies] [dependencies]
libc = "0.2" libc = "0.2"
# RFC 010 §2 — only compiled under `--features cluster`.
serde = { version = "1", default-features = false, optional = true }
# `alloc` (not `std`): the seam serializes to Vec; postcard stays no_std-aligned.
postcard = { version = "1", default-features = false, features = ["alloc"], optional = true }
[target.'cfg(loom)'.dependencies] [target.'cfg(loom)'.dependencies]
loom = "0.7" loom = "0.7"
[dev-dependencies] [dev-dependencies]
libc = "0.2" libc = "0.2"
# derive + std for cluster envelope tests only; the lib itself never needs them
serde = { version = "1", features = ["derive"] }
tokio = { version = "1", features = ["rt", "rt-multi-thread", "macros", "sync", "time"] } tokio = { version = "1", features = ["rt", "rt-multi-thread", "macros", "sync", "time"] }
[profile.dev] [profile.dev]
@@ -60,6 +71,14 @@ panic = "unwind"
lto = "thin" lto = "thin"
codegen-units = 1 codegen-units = 1
# `cargo test --profile reltest`: release codegen for the crate (same opt-level,
# same panic strategy) but no LTO at the final link. Thin LTO is what makes each
# of the ~40 test binaries cost ~12 s to link instead of ~2 s; the tests don't
# need cross-crate LTO, the benches do (they keep using `release`).
[profile.reltest]
inherits = "release"
lto = false
[[bench]] [[bench]]
name = "primes" name = "primes"
harness = false harness = false
+69 -43
View File
@@ -1,35 +1,38 @@
# smarm # smarm
> SMARM — Smarm, Marks Actor Runtime Machinery. A proof-of-concept green-thread actor runtime for Rust. > SMARM: Smarm, Marks Actor Runtime Machinery. A proof-of-concept green-thread actor runtime for Rust.
Implements the core ideas in [`Achitecture.md`](.docs/Architecture.md): green-thread actors on a SMARM is my attempt to implement the erlang/OTP philosophy in the Rust programming language. This has yielded a fault-tolerant, fast, and scalable runtime. This runtime allows the creation of asynchronous applications in Rust without the function coloring associated with the async/await system. It encourages the writing of simple, synchronous code, and largely elides the need for lifetime annotations.
shared heap, scheduled cooperatively, communicating only by `Send` messages.
Erlang's isolation model without Erlang's copying GC, Rust's zero-copy
ownership transfers without async's function colouring.
The scheduler is multi-threaded — one OS thread per available CPU, all drawing
from a shared run queue. The single-threaded `run()` entry point is kept as a
convenience wrapper around `runtime::init(Config::exact(1)).run(f)`.
## What's here ## Overview
SMARM implements green-thread actors on a shared heap, communicating only by `Send` messages. By sharing the heap, SMARM avoids the copying overhead of Erlang, which is safe to do due to Rust's borrow checker.
On top of the core runtime mechanics, SMARM also provides a library of primitives for making applications closely inspired by erlang/OTP. This includes generic servers (gen_servers), generic state machines (gen_statem), and supervision trees.
Supervision trees are the core primitive to allow your application to survive an unexpected panic. Supervisors are processes dedicated to monitoring other processes, which can restart these should they fail. This means that when set up properly an application may 'self-heal' when encountering unforeseen circumstances.
SMARM is not cooperatively scheduled; it uses preemption. This means a heavy task will not starve out other lighter tasks. Everything will make steady progress, which translates to very beneficial behaviour under (over)load: average latency goes up, but tail latency does not blow up.
To help diagnose these unforeseen circumstances, smarm may be compiled with its `tracing` feature, which emits a full trace using [Perfetto](https://perfetto.dev/).
Should you want to optimize your application, SMARM is unusually well poised to help. As the runtime functionally controls time, SMARM comes with a built in causal profiler, under the `causal` feature.
I also built a Phoenix-Framework inspired HTTP 1.1 library on top of SMARM called [URUS](https://git.kalsbeek.dev/Markk116/urus), which implements Pub/Sub, Channels, and basic amenities like Websockets and Server Sent Events.
## Limitations
This runtime requires naked assembly to function, and has thus far only been implemented for x86-64 assembly. It expects an operating system that supports virtual address space, and is therefore not (yet) suited for embedded targets. The IO implementation is currently based around the Linux kernel's `epoll` mechanism, meaning it requires a (GNU+)Linux distribution to run.
The preemption mechanism works by wrapping the memory allocator and checking how many CPU cycles you have used compared to your timeslice budget. This allows preemption to fire in most normal code, but tight zero-allocation loops do not get caught and require manual insertion of `check!()` if you want preemption to function.
This library is still in its early stages, and while I try my best with loom and tests, stable operation cannot be guaranteed. Therefore it is not (yet) recommended for production use.
At this moment, stack memory for each green thread is capped. Uncapping this may lead to performance benefits for deeply recursive algorithms that in a traditional async runtime might require pointer-chases through the heap. This is as yet unrealised.
At this stage, the codebase is largely LLM-generated, which is obvious if you start to read through the internals. While I did the design, and I keep the LLM under tight rein, the codebase is not in a state that I am very happy with. This also goes for the documentation.
| Module | What it does |
|--------------|------------------------------------------------------------------------|
| `stack` | `mmap`'d growable stack with guard page; SIGSEGV on overflow |
| `context` | `#[naked]` x86-64 context-switch shims, callee-saved regs only |
| `preempt` | Allocator-driven preemption; `check!()` macro for no-alloc loops |
| `pid` | `(index, generation)` PIDs; stale handles are detectable, not silent |
| `actor` | Trampoline + `catch_unwind` boundary at the actor entry point |
| `scheduler` | Run queue, slot table, spawn/join, parking, idle path |
| `channel` | Unbounded MPSC channel; `recv` parks the actor; `recv_timeout` bounds it; `select`/`select_timeout` park on many receivers at once (ready-index, priority order) |
| `mutex` | `Mutex<T>` with mandatory timeout; FIFO waiters; parks the green thread |
| `timer` | Min-heap of `(deadline, reason)`; `Sleep` and `WaitTimeout` reasons |
| `io` | `block_on_io` for blocking work; `wait_readable`/`wait_writable` + `read`/`write` via epoll |
| `supervisor` | `Signal::Exit`/`Panic`/`Stopped` funnelled to a parent; `OneForOne`/`OneForAll`/`RestForOne` strategies + restart-intensity cap |
| `monitor` | `monitor(pid)` → `Monitor { id, target, rx }`; one-shot `Down` via `rx`; `demonitor(&m)` tears one registration down; unidirectional death notice |
| `link` | bidirectional `link`/`unlink`; abnormal death propagates (cooperative stop, or an `ExitSignal` message under `trap_exit`) |
| `gen_server` | `call`/`call_timeout` (sync request-reply) / `cast` (async) over one inbox; `handle_info` over static info arms + `handle_down` via `Watcher`-fed monitors, selected ahead of the inbox; `ServerRef`/`ServerBuilder` + `init`/`terminate` hooks; server-down via channel closure |
| `registry` | `register`/`whereis`/`name_of`: name ↔ pid bimap; lazy generation-checked cleanup |
## Quick taste ## Quick taste
@@ -51,6 +54,29 @@ run(|| {
}); });
``` ```
## Stopping actors
Two strengths, as in OTP. `request_stop(pid)` is `exit(Pid, kill)`: a cooperative
hard stop, unwinding at the actor's next observation point. `request_shutdown(pid)`
is `exit(Pid, shutdown)`: an actor that traps exits (`trap_exit()`, or
`ctx.trap_exit()` in a gen_server / `cx.trap_exit()` in a gen_statem) receives it
as a signal — `handle_shutdown` / a `shutdown` row — and may drain before stopping
itself; one that does not trap is stopped outright. Supervisors trap:
`request_shutdown(sup)` tears the tree down top-down, each child per its
`ChildSpec` `Shutdown` policy (`Timeout(d)`, `Infinity`, `BrutalKill`). The run's
root actor returning means "the program is done": every top-level actor gets a
`request_shutdown`, and `run()` returns when they are gone. From outside the
runtime (a signal thread), `Runtime::handle().request_shutdown(pid)` does the
same. `examples/graceful_shutdown.rs` shows all of it.
A gen_server or gen_statem lives until it stops, is shut down, or is killed; its
refs are addresses — dropping them never ends it (a forgotten one is swept at
root exit; with `--features smarm-trace` each such sweep is a `root_sweep`
trace line). The supervised shape is `GenServerBuilder::named(N).run()` /
`gen_statem::run_named(N, m)`: the server runs inline as the `ChildSpec` child
itself, so the supervisor's shutdown reaches it directly, a restart re-binds the
name, and the program addresses it by name.
## Layout ## Layout
``` ```
@@ -67,12 +93,9 @@ benches/
## Building and running ## Building and running
Standard Cargo. Requires Rust 1.95 or newer (the `#[naked]` attribute went stable Standard Cargo. Requires Rust 1.95 or newer (the `#[naked]` attribute went stable in 1.88; we use a few unrelated post-1.88 features). I have worked hard to keep this library as dependency-free as possible. `master` is x86-64 Linux only. An experimental, **untested** aarch64 context-switch backend lives on the `arm-port` branch (extracted into a `target_arch`-gated `src/arch/`); it has not been validated on hardware yet. macOS remains on the deferred list because of the epoll dependency.
in 1.88; we use a few unrelated post-1.88 features). `master` is x86-64 Linux
only. An experimental, **untested** aarch64 context-switch backend lives on the
`arm-port` branch (extracted into a `target_arch`-gated `src/arch/`); it has not
been validated on hardware yet. macOS remains on the deferred list because of the
epoll dependency.
```sh ```sh
cargo test # all tests cargo test # all tests
@@ -80,26 +103,29 @@ cargo test --test mutex # one module
cargo bench # primes benchmark vs tokio cargo bench # primes benchmark vs tokio
``` ```
## What's not here
See the **Defer** section of `Architecture.md`.
`join!` for handle groups, stack growth via remap,
hierarchical timer wheel, fd-wait timeouts, `Signal::Timeout`. Each is
mechanism we know how to add; none belongs in this iteration.
## Docs ## Docs
| Document | What it covers | | Document | What it covers |
|---|---| |---|---|
| [`Architecture.md`](./docs/Architecture.md) | Design intent, runtime model, and deferred work | | [`Architecture.md`](./docs/Architecture.md) | Design intent, runtime model, and deferred work |
| [`smarm - Deep Dive.html`](./docs/smarm%20-%20Deep%20Dive.html) | Generated walkthrough of the system; good starting point | | [`smarm - Deep Dive.html`](./docs/smarm%20-%20Deep%20Dive.html) | Generated walkthrough of the system; good starting point if you want to learn about the internals |
| [`BENCHMARKS_AND_TUNING.md`](./docs/BENCHMARKS_AND_TUNING.md) | Where smarm wins and loses vs tokio, preemption knob recommendations | | [`BENCHMARKS_AND_TUNING.md`](./docs/BENCHMARKS_AND_TUNING.md) | Where smarm wins and loses vs tokio, preemption knob recommendations |
| [`benchmarks.md`](./docs/benchmarks.md) | Raw benchmark results, methodology, and tuning experiment log | | [`benchmarks.md`](./docs/benchmarks.md) | Raw benchmark results, methodology, and tuning experiment log |
## Coming up
Clustering: clustering multiple SMARM nodes together is in the pipeline.
SMARM-BEAM Interop: Running SMARM as a supervised node under the BEAM via a Rustler NIF works, including message passing and supervision trees that span the runtimes. However, this library is still too unstable to release.
SMARM is an interesting platform for implementing a 'dataflow' library, but work on this has not yet started.
## Contributing ## Contributing
This is a personal proof-of-concept. There's no PR workflow. If you fork it and do something interesting, just send me an email. If it's nice, I'll upstream the changes. This started as a personal proof-of-concept, but it is starting to outgrow that name. If you want to contribute, please get in contact to discuss what you want to work on. Code without prior communication is not welcome.
## A note on open source
An open source project is a gift, and by giving it, it is no longer mine. I highly enourage you to fork it, to make it your own. This repository, however, is still mine.
--- ---
+62 -5
View File
@@ -28,6 +28,15 @@ and **excised** (not worth the code cost; preserved on branch
`rfc-004-spinning`). Also a false-sharing fix (`align(64)` on `SchedulerStats`) `rfc-004-spinning`). Also a false-sharing fix (`align(64)` on `SchedulerStats`)
and a termination wake for idle siblings. and a termination wake for idle siblings.
Commits `2708042`, `37d9319`, `eddf3fe`. Commits `2708042`, `37d9319`, `eddf3fe`.
**Default flipped ON 2026-08-18** (history.md findings 17/18): the slot had
shipped default-off "until the shootout accepts it" and the flip was never
made, so every general.rs number since was slot-off. Acceptance sweep
(rq_runtime, 1/2/4/8/20 schedulers, rq-mpmc, 7 runs): ping-pong-pairs
−7% at 1T, 3.6×/5.3×/15.6×/10× faster at 2/4/8/20T, 100% slot hits, 0
displaced; yield-storm and spawn-storm within ±10% noise both ways.
general.rs on the flip: ping_pong_steady 20T 18485→1549 µs, 1T −15%,
mpsc_contention 1T −51%, ping_pong_oneshot 20T −18%; no smarm regressions.
`benches/baseline.json` regenerated slot-on.
--- ---
@@ -77,10 +86,10 @@ Delivered surface:
`ServerBuilder::start` untouched; free `call` / `cast` / `whereis_server`; `ServerBuilder::start` untouched; free `call` / `cast` / `whereis_server`;
`ServerRef::shutdown` + free `shutdown` as the sys-style synchronous stop. `ServerRef::shutdown` + free `shutdown` as the sys-style synchronous stop.
- **Root-exit teardown** (final phase): the run's initial actor is the root; - **Root-exit teardown** (final phase): the run's initial actor is the root;
when it exits, the scheduler's idle verdict stops the parked-forever remainder when it exits the run winds down. *(Reworked with the graceful-shutdown work:
(deferred past the queue drain, so actors with in-flight work finish rather root exit now delivers `request_shutdown` to every forest root — see
than unwinding on the stop). Closes the "app actor blocks AllDone" stall — see "Root exit" below and `tests/root_exit.rs`.)* Closes the "app actor blocks
Look into, below. AllDone" stall — see Look into, below.
Extends — does not retire — the "select exists; a unified per-process mailbox Extends — does not retire — the "select exists; a unified per-process mailbox
still does not" invariant: 014 adds addressable *delivery*, not a unified inbox; still does not" invariant: 014 adds addressable *delivery*, not a unified inbox;
@@ -146,7 +155,11 @@ Needs an RFC.
#### Per-switch cost (context shims, epoch protocol) #### Per-switch cost (context shims, epoch protocol)
The shootout's residual: per-wake latency is 0.16–0.18 µs at N=1 and The shootout's residual: per-wake latency is 0.16–0.18 µs at N=1 and
0.8–1.2 µs at N=8+, dominated by the context-switch shims and the epoch 0.8–1.2 µs at N=8+, dominated by the context-switch shims and the epoch
protocol, not the queue. On current evidence this is the larger constant — protocol, not the queue. **Premise corrected 2026-08-18 (finding 17): the
N=8+ figure was the slot-off futex path (one futex_wake per landed wake,
woken worker steals the pair); slot-on it is ~1.3× N=1. The shims are ~5
cycles (finding 3) and the epoch CASes ~117 cycles/roundtrip (finding 15).
Re-measure slot-on before spending anything here.** On current evidence this is the larger constant —
"the whole game" alongside the v0.9 work — but there is no spec yet. Needs a "the whole game" alongside the v0.9 work — but there is no spec yet. Needs a
profiling spike (where do the cycles actually go per park/unpark round-trip) profiling spike (where do the cycles actually go per park/unpark round-trip)
and then an RFC before it can be scheduled. and then an RFC before it can be scheduled.
@@ -260,8 +273,52 @@ path the atomic-bool workaround stood in for. Re-check the urus crud repro to
confirm the workaround can be retired (the teardown is cooperative — an actor in confirm the workaround can be retired (the teardown is cooperative — an actor in
a tight loop with no observation point still can't be stopped). a tight loop with no observation point still can't be stopped).
**Update (graceful shutdown):** the RFC 014 sweep was a hard `request_stop` of
every live slot, deferred until nothing was runnable — which killed a sleeping
actor (timer pending) but drained a queued one, for no principled reason. It is
now the OTP semantics: root exit = "the program is done" = `request_shutdown`
to every **forest root** (live actor whose parent is the run or is dead), run
synchronously on the root's finalize path. Supervisors cascade with their child
`Shutdown` policies; trapping actors may `Continue`/drain (timers keep working)
and end the run when they stop themselves; non-trapping actors are stopped
outright — `join` what you need finished. No forcing sweep follows.
--- ---
### Open items from the graceful-shutdown work (not scheduled)
- ~~gen_server / gen_statem as a direct supervised child.~~ Done: lifetime is
the actor's (refs are addresses, the loop holds an inbox sender);
`NamedGenServerBuilder::run` / `gen_statem::run_named` run the loop inline as
the `ChildSpec` child; the root-exit sweep traces each leftover as
`root_sweep` under `smarm-trace`.
- Supervisor `Live` drop-guard sweep is `request_stop` (kill propagates as
kill); OTP would deliver a trappable `killed`. Chosen for boundedness.
- A root-exit shutdown reaches only actors live *at that instant*; a
non-trapping forest root that spawns before it unwinds leaves that spawn
to itself (Erlang: an unlinked spawn is nobody's child).
- **Supervisor start *order* is not start *readiness*.** `start_child`
spawns and moves straight on, so an earlier child is merely *scheduled*,
not initialised, when a later sibling starts. A later child that resolves
an earlier one by name (`whereis_server`) can therefore miss it — the
classic "named registry sibling, then its consumers" tree. Ordered
`OneForOne`/`RestForOne` shutdown is unaffected (reverse order is honoured
and each stop *is* awaited); this is a start-side gap only.
Making `spawn` itself block does NOT fix it — it would only shrink the
window to "child has begun executing", while the property callers need is
"child has bound its name / opened its socket", which only the child can
declare. It would also tax the hot path (one round-trip per accepted
connection) and turn every spawn into a context-switch point. OTP has the
same async `spawn` and puts the synchronisation one level up:
`gen_server:start_link` blocks the caller until `init/1` returns.
Fix shape when scheduled: a readiness ack in the supervisor's child-start
path (`ChildSpec` variant whose factory receives a ready-signal;
`NamedGenServerBuilder::run` acks after its name bind, gen_server default
acks after `init`; plain closures ack at spawn as today, i.e. opt-in with
no cost to existing children). Until then the workaround is structural:
have the registrar spawn its own consumers so the ordering is program
order inside one actor, not a cross-actor guarantee (urus v0.3 endpoint
does exactly this).
## Invariants & gotchas (respect these across all cycles) ## Invariants & gotchas (respect these across all cycles)
- **Shared mutex is non-reentrant.** `Sender::send` can call `unpark` → - **Shared mutex is non-reentrant.** `Sender::send` can call `unpark` →
+229
View File
@@ -0,0 +1,229 @@
# urus / smarm handoff — updated 2026-08-19 (session 3)
## TL;DR for the next session
**smarm is done for now** (5 unpushed commits on local `master`, see below).
**Next = urus v0.3 endpoint refactor.** You should NOT need to read smarm
scheduler internals; the contract you build on is fully described here and in
`smarm_full/examples/graceful_shutdown.rs` (read that file first — it is the
exact shape urus's tree will take) plus `smarm_full/tests/root_exit.rs`.
### The smarm contract urus builds on (all on local master, verified by tests)
- `request_stop(pid)` = kill (cooperative hard stop). `request_shutdown(pid)` =
polite: trapping target gets `ExitSignal{reason: Shutdown}`, non-trapping is
stopped outright. `RuntimeHandle::{request_stop,request_shutdown}` do the same
from any OS thread (signal handler); grab `rt.handle()` before `rt.run`.
- Supervisor traps; `request_shutdown(sup)` = ordered reverse-start shutdown,
per-child `ChildSpec::shutdown(Shutdown::{Timeout(d)|Infinity|BrutalKill})`
(default Timeout(5s)); sup then returns normally. `request_stop(sup)`
hard-stops children too (no orphans).
- gen_server: `ctx.trap_exit()` in init; `handle_shutdown() -> Exit|Continue`;
`handle_exit(sig)`; `ctx.stop_handle().stop()` = normal self-exit;
`terminate()` may block only on the graceful path (Exit / stop / inbox close).
`GenServerRef::shutdown()` is graceful and waits.
- gen_statem: same in event clothes — `cx.trap_exit()` in initial enter,
`shutdown` rows (default `stop`), `exit sig` rows, `cx.stop()` / `stop` tail,
optional `terminate { }` block. `GenStatemRef::shutdown()`.
- **Root exit = program done**: when the root actor returns, the runtime
`request_shutdown`s every *forest root* (live actor whose parent is the run
or dead). Supervisors cascade; trapping actors may drain (timers keep
working) and end the run when they stop; non-trapping are stopped; **no
forcing sweep** (`join` what must finish). The old "wait until nothing
runnable then kill all" deferral is gone.
- **Gotcha for urus:** a gen_server's lifetime is governed by its refs — drop
the last `GenServerRef` and the inbox closes → clean exit *even mid-drain*.
The endpoint must be pinned (named, or its ref held by the supervisor
wrapper) or it will terminate the moment the root drops its ref.
- **Known gap (ROADMAP open item):** a gen_server can't be a direct `ChildSpec`
child; use the trapping wrapper pattern in `examples/graceful_shutdown.rs::
drainer_child` (starts `under(self_pid())`, forwards shutdown, waits). Doing
an inline `GenServerBuilder::run()` first may be worth a short smarm detour —
decide with Markk.
### smarm commits this session (local master, NOT pushed, NOT tagged)
`250f312` root-exit = graceful shutdown of forest roots (tests/root_exit.rs)
`6ceb138` gen_statem shutdown parity (tests/gen_statem_shutdown.rs)
`849a424` docs + examples/graceful_shutdown.rs + README "Stopping actors"
On top of `1002777` (cross-thread wake) and `9c8f59c` (graceful shutdown).
Cargo.toml still `0.6.1`. Release cut (push, tag — v0.7 is justified by the
API surface — version bump) is Markk's. Full suite, doc tests, examples,
`cargo fmt`, `cargo clippy --lib` all clean. (`clippy --tests` has pre-existing
unwrap lints in tests/fd_select.rs, untouched.)
### Decisions taken this session (Markk)
- Root exit means "program done" (Go/tokio/OTP), not "wait for pending work";
the previously agreed "sleep(50ms) must finish" test was dropped as encoding
the wrong contract (a timer-wheel gate would re-wedge periodic-timer daemons).
- No behaviour-preserving deferral, no forcing second sweep.
- Examples/docs done in the same session; urus next session.
---
# Previous handoff (still accurate where not superseded above)
## Next-session goal
Phase 1 is **done and committed**; Phase 2 is next:
1. **smarm v0.6.2** — cross-thread wake root fix. **DONE**, committed on `master`
as `1002777`. Not yet tagged, not yet version-bumped (Cargo.toml still reads
`0.6.1`), and **not yet pushed to origin** — it exists only in the delivered
snapshot zip and the local sandbox clone. Cutting the release (push + tag
`v0.6.2` + bump `0.6.1`→`0.6.2`) is Markk's step.
2. **urus v0.3** — endpoint refactor. Working against `smarm = { path = "../smarm_full" }`
with `git update-index --skip-worktree Cargo.toml` (Markk approved); release commit
swaps back to the tag once Markk cuts it. Phase 2 plan below is STALE where it says
drain-in-terminate; the endpoint is a trapping GenServer: `handle_shutdown` →
`Continue`, enter Draining, `StopHandle::stop()` when the conn set empties.
Decisions below are locked unless marked *(confirm)*.
## Reconstruction (the sandbox resets between sessions)
A fresh sandbox has an empty home and **no Rust toolchain**. To restore:
- Install rustup/cargo. smarm reformats under **rustc 1.97.1**; urus `rust-version`
is 1.95. Use 1.97.1.
- urus: `git clone https://git.kalsbeek.dev/Markk116/urus` — `origin` is registered
and public-read. master `8bdec97` = the v0.2.x line. (Zips in outputs are stale;
prefer the remote now.)
- smarm: `git clone https://git.kalsbeek.dev/Markk116/smarm`. Latest tag **v0.6.1**
(`ca1c983`). The cross-thread wake fix is committed as `1002777` on top of the
post-v0.6.1 README commit `8f2d513` (= origin/master). **It is NOT on origin
yet** — a fresh clone won't have it until Markk pushes. Restore it from the
snapshot zip if working before the push. **v0.6.2 is not yet tagged.**
- urus pins smarm by git **tag** in `Cargo.toml` (currently `v0.6.0`). A trivial
first commit bumps it to `v0.6.1` (also picks up `try_spawn` + monitor
terminal-outcome fixes).
## Why (context — the finding that drives the plan)
urus's shutdown machinery (the `AtomicBool` listener flag + the `SHUTDOWN_POLL`
loop in `serve.rs`) is scaffolding around two smarm properties. Their statuses
differ, which is the whole point:
- **Issue A — lossy stop vs a QUEUED actor: ALREADY FIXED in smarm.** Commit
`7bab4d2` added an entry-side `check_cancelled()` in `park_current`. A
`request_stop` against a listener parked in `wait_readable_timeout` now unwinds
cleanly (it parks via `try_select_timeout → park_current`). urus's flag + its
stale "smarm's lossy stop-while-QUEUED window" comment can be deleted.
- **Issue B — foreign-thread wake is a no-op: FIXED in `1002777` (was present
through v0.6.1).** The gap: `unpark`/`unpark_at`/`request_stop` all route through
`try_with_runtime`, which reads a thread-local that is `None` on any non-scheduler
thread, so no cross-thread wake worked — a signal handler / OS thread could not
wake *or* stop a parked actor, which is why `serve.rs` polls the shutdown signal
instead of parking on it. Now closed (see Phase 1 below): urus's `SHUTDOWN_POLL`
loop can be deleted and its `Handle::shutdown` can park on a handle-driven stop.
## Phase 1 — smarm cross-thread wake (root fix) — DONE (`1002777`)
Shipped as one commit generalizing RFC 018 (a producer reaches the runtime through
a `Weak` it holds) from the IO backend to channel senders and a new handle:
- **`Runtime::handle() -> RuntimeHandle`** (`Send + Sync`), holding a
`Weak<RuntimeInner>`. Grab it before `rt.run` and hand it to the signal thread.
- **`RuntimeHandle::request_stop<A>(Pid<A>)`** — upgrades the Weak and calls
`request_stop_inner` on the inner; no-op if the runtime is gone. This is the
signal-handler-drives-shutdown path; it cascades the ordered stop down the tree
exactly like an in-runtime `request_stop`.
- **Send-wake:** the receiver captures `scheduler::runtime_weak()` into its
`parked_receiver` tuple **at park time** (not at channel creation — the resolved
sub-decision; a parked receiver is a live actor so the Weak is provably upgradable,
and it scopes the capture to when a wake is possible). `send()` and last-sender
`drop` wake via `scheduler::unpark_at_via(pid, epoch, &weak)`: thread-local path
when on a scheduler thread (preempt-gated, slot-eligible), captured Weak otherwise.
In-runtime timer wakes (recv/select) were left on `scheduler::unpark_at`.
**API scope decision (signed off):** `RuntimeHandle` exposes **`request_stop` only**.
No public `unpark`/`unpark_at` on the handle — send-wake needs no user-facing handle,
and "unpark off-runtime" is covered because `request_stop` drives `unpark` on the
upgraded inner. No `is_alive()`. Both are one-line additions if a consumer appears.
**No RFC written** — pattern was already established (RFC 018), agreed not needed.
Tests: `tests/cross_thread_wake.rs` (foreign-thread send wakes a parked receiver;
foreign-thread `request_stop` wakes+stops a parked actor; a lingering handle never
blocks all-done and degrades to a no-op once the runtime drops). Full suite green;
`cargo fmt` + `cargo clippy --lib` clean.
**Remaining release step (Markk):** push `master`, tag `v0.6.2`, bump Cargo.toml
`0.6.1`→`0.6.2`. Left paired with the tag as the release cut, not done in `1002777`.
## Phase 1b — smarm graceful shutdown (OTP lift) — DONE (`9c8f59c`, on top of `1002777`)
Decided this session (Markk): B — fix at the smarm level rather than a two-stop
split in urus. No RFC (Markk: "just implement it"). Shipped, tested, committed on
the local `master`, **not pushed, not tagged**. It should ship as the same
release as 1002777 (v0.6.2, or v0.7 given the API surface — Markk's call).
- `request_shutdown(pid)` / `RuntimeHandle::request_shutdown` = `exit(Pid, shutdown)`;
`request_stop` = `exit(Pid, kill)`. Trapping target gets `ExitSignal{reason:
DownReason::Shutdown}`; non-trapping is stopped outright.
- `ChildSpec::shutdown(Shutdown::{BrutalKill, Timeout(d), Infinity})`, default 5s.
Supervisor traps exits; `request_shutdown(sup)` = ordered top-down shutdown,
returns normally. **Also fixed**: `request_stop(sup)` used to ORPHAN children
(probe-verified; the handoff's "cascade" claim was wrong) — `Live` drop guard now
hard-stops them.
- gen_server: `ctx.trap_exit()`, `handle_shutdown() -> ShutdownAction::{Exit,
Continue}`, `handle_exit(ExitSignal)`, `ctx.stop_handle().stop()` = normal
self-exit (`{stop, normal}`; previously impossible — only abnormal `Stopped`).
`GenServerRef::shutdown()` is graceful now.
- Root finding that forced this: gen_server `terminate()` runs from a Drop guard,
mid-unwind on the stop path; any park in it = double panic = abort. So
"drain-in-terminate()" (the old Phase 2 plan) was never viable.
### ~~Next-session smarm work~~ DONE this session (see TL;DR)
1. **Root-exit sweep**: make `Pop::RootDrain` also require an empty timer wheel
(and it already requires nothing runnable; io_out is only checked for AllDone —
check whether it should gate RootDrain too). TDD: an actor in `sleep(50ms)` when
the root returns must finish, not be swept. Then a `Reservoir`-style test that a
*truly* parked-forever daemon still gets swept.
2. **gen_statem parity**: `ctx.trap_exit()`, `handle_shutdown -> ShutdownAction`,
`handle_exit`, stop handle. Mirror gen_server; mechanical.
3. **Examples review**: `examples/*.rs` predate all of this. Rework where they show
shutdown/teardown to use `request_shutdown`, `Shutdown` policies, and
`StopHandle`; `named_genserver.rs` first (uses `shutdown`). Also
`docs/smarm - Deep Dive.html` says terminate() must be non-blocking — now only
true on the unwind paths; and README could use a "Stopping actors" paragraph
(request_stop = kill, request_shutdown = shutdown, Shutdown policy).
### Open smarm items found on the way (noted, not scheduled)
- (root-exit sweep and gen_statem parity moved up to the scheduled list.)
- Sweep in the supervisor `Live` drop guard is `request_stop` (kill propagates as
kill); OTP would deliver a trappable `killed`. Chosen for boundedness.
## Phase 2 — urus v0.3: endpoint refactor (after v0.6.2 is tagged)
Target = the spec's original shape (`urus-spec.md` §2.1/§6: `listener_sup` under the
**user's** root supervisor). Deviation to unwind: `serve` owning `rt.run`.
- App owns the runtime: `smarm::init(cfg).run(|| root_sup.run())`, root e.g.
`RestForOne[ app actors…, urus::endpoint(config, pipeline) ]`. This is what kills
the `Arc<OnceLock>` idiom for the right reason (app state born in-runtime as a
supervised, ordered child).
- `urus::endpoint` = one GenServer child owning the registry + an **internal**
listener sub-supervisor + drain-in-`terminate()`. Listeners stay internal, not
app-visible peers.
- Shutdown = `request_stop` the root supervisor (or via the runtime handle from a
signal thread) → cascades down → `endpoint.terminate()` runs the drain
(`drain_timeout`, force-stop sweep).
- **DELETE:** the `AtomicBool` listener flag (A fixed) and the `SHUTDOWN_POLL` loop
+ its apologetic comment (B fixed → park, don't poll).
- Nuance: `request_stop` → `Signal::Stopped` is *abnormal* → `Transient` restarts.
Stop-without-restart = stop the **supervisor**, not the children.
- *(confirm)* Keep `serve`/`serve_with`/`serve_with_shutdown` as thin wrappers that
build the one-child tree internally, so the simple case stays one line.
- *(confirm)* Keep `Handle`/`ShutdownSignal`? Now that cross-thread wake works,
`Handle::shutdown` can map to handle-driven `request_stop` on the endpoint.
- Breaking → cut **urus v0.3**; bump the smarm pin to `v0.6.2` here.
## Working norms
- Every bash call: `export PATH=$HOME/.cargo/bin:$PATH` (once the toolchain's in).
- **TDD**: failing test first, then implement; keep suites green.
- **Hammer ritual** for ANY connection-lifecycle change (the urus #2 shutdown work
qualifies): 35× subset (`shutdown timeout reaped slowloris streaming chunked sse
stalled ws_ channels session`) + 3× full + 1× trace. `scripts/hammer.sh` does NOT
pass feature flags — loop manually with `--features phoenix`. Subset filter must
NOT use `--test integration` (session tests live in lib).
- Example smoke tests: hold the server's stdin open (`mkfifo` + `sleep > fifo`) or
the Enter-to-shutdown thread fires on EOF instantly.
- Background procs are reaped BETWEEN bash calls; `pkill -f` matches your own shell.
- Artefact store (specs): `curl -H "Authorization: Bearer sk-llmingest-2e45d80c63db24c6781e761eb2a9a58e83d9f48ef77a42185bad311d07c80e68" https://artefacts.kalsbeek.dev/artifacts/<name>`
— `urus-spec.md`, `urus-bench-spec.md`, `rfc_008-implementation-notes.md`, …
- smarm feature flags: `smarm-trace`, `smarm-causal` (urus re-exports both).
## Local cross-repo testing — KEEP OUT OF COMMITS
To test urus #2 against un-tagged smarm 0.6.2, point urus's `Cargo.toml` smarm dep
at a local path (`smarm = { path = "../smarm" }`) instead of the git tag.
- Must NOT land in commits. Guard: `git update-index --skip-worktree Cargo.toml`
after editing (undo with `--no-skip-worktree`), or stash before committing.
- The committed `Cargo.toml` stays pinned to the git tag; restore the tag (bumped to
`v0.6.2`) for the release commit.
+317 -265
View File
@@ -1,314 +1,366 @@
{ {
"catch_unwind_panics": {
"smarm 1-thread": {
"result": 10000,
"median": 122512,
"min": 113735,
"max": 124564
},
"smarm 20-thread": {
"result": 10000,
"median": 20569,
"min": 19671,
"max": 21174
},
"tokio current_thread": {
"result": 10000,
"median": 10134,
"min": 9081,
"max": 10524
},
"tokio multi-thread": {
"result": 10000,
"median": 2461,
"min": 2351,
"max": 2715
}
},
"chained_spawn": { "chained_spawn": {
"smarm 1-thread": { "smarm 1-thread": {
"result": 1000, "result": 1000,
"median": 413, "median": 498,
"min": 410, "min": 491,
"max": 439 "max": 503
}, },
"smarm 24-thread": { "smarm 20-thread": {
"result": 1000, "result": 1000,
"median": 909, "median": 2135,
"min": 888, "min": 1881,
"max": 951 "max": 2284
}, },
"tokio current_thread": { "tokio current_thread": {
"result": 1000, "result": 1000,
"median": 62, "median": 114,
"min": 61, "min": 106,
"max": 62 "max": 119
}, },
"tokio multi-thread": { "tokio multi-thread": {
"result": 1000, "result": 1000,
"median": 197, "median": 164,
"min": 194, "min": 160,
"max": 210 "max": 186
}
},
"yield_many": {
"smarm 1-thread": {
"result": 200000,
"median": 16475,
"min": 16393,
"max": 16732
},
"smarm 24-thread": {
"result": 200000,
"median": 148708,
"min": 111213,
"max": 156462
},
"tokio current_thread": {
"result": 200000,
"median": 4751,
"min": 4740,
"max": 5259
},
"tokio multi-thread": {
"result": 200000,
"median": 8320,
"min": 7862,
"max": 8882
}
},
"fan_out_compute": {
"smarm 1-thread": {
"result": 33860,
"median": 13453,
"min": 13305,
"max": 15077
},
"smarm 24-thread": {
"result": 33860,
"median": 2451,
"min": 2330,
"max": 2520
},
"tokio current_thread": {
"result": 33860,
"median": 14019,
"min": 12339,
"max": 14045
},
"tokio multi-thread": {
"result": 33860,
"median": 1500,
"min": 1426,
"max": 1600
}
},
"ping_pong_oneshot": {
"smarm 1-thread": {
"result": 1000,
"median": 898,
"min": 782,
"max": 920
},
"smarm 24-thread": {
"result": 1000,
"median": 1494,
"min": 1489,
"max": 1546
},
"tokio current_thread": {
"result": 1000,
"median": 396,
"min": 389,
"max": 409
},
"tokio multi-thread": {
"result": 1000,
"median": 10382,
"min": 9559,
"max": 10924
}
},
"spawn_storm_busy": {
"smarm 1-thread": {
"result": 10000,
"median": 106232,
"min": 105627,
"max": 107693
},
"smarm 24-thread": {
"result": 10000,
"median": 49198,
"min": 48894,
"max": 52606
},
"tokio current_thread": {
"result": 10000,
"median": 1140,
"min": 1018,
"max": 1169
},
"tokio multi-thread": {
"result": 10000,
"median": 18650,
"min": 16486,
"max": 19396
}
},
"mpsc_contention": {
"smarm 1-thread": {
"result": 320000,
"median": 6173,
"min": 5446,
"max": 6226
},
"smarm 24-thread": {
"result": 320000,
"median": 36329,
"min": 35121,
"max": 37590
},
"tokio current_thread": {
"result": 320000,
"median": 5563,
"min": 5527,
"max": 6239
},
"tokio multi-thread": {
"result": 320000,
"median": 63966,
"min": 59972,
"max": 67534
}
},
"many_timers": {
"smarm 1-thread": {
"result": 10000,
"median": 107826,
"min": 107093,
"max": 119034
},
"smarm 24-thread": {
"result": 10000,
"median": 96539,
"min": 95413,
"max": 97804
},
"tokio current_thread": {
"result": 10000,
"median": 12584,
"min": 12539,
"max": 12627
},
"tokio multi-thread": {
"result": 10000,
"median": 16183,
"min": 16024,
"max": 16541
}
},
"multi_thread_scaling": {
"smarm 1-thread": {
"result": 33860,
"median": 15083,
"min": 15071,
"max": 15283
},
"smarm 2-thread": {
"result": 33860,
"median": 8070,
"min": 8003,
"max": 8096
},
"smarm 4-thread": {
"result": 33860,
"median": 4460,
"min": 4454,
"max": 4516
},
"smarm 24-thread": {
"result": 33860,
"median": 2333,
"min": 2294,
"max": 2348
},
"tokio multi 1-thread": {
"result": 33860,
"median": 14504,
"min": 14193,
"max": 14562
},
"tokio multi 2-thread": {
"result": 33860,
"median": 7297,
"min": 7289,
"max": 7398
},
"tokio multi 4-thread": {
"result": 33860,
"median": 3795,
"min": 3756,
"max": 3799
},
"tokio multi 24-thread": {
"result": 33860,
"median": 1544,
"min": 1520,
"max": 1610
} }
}, },
"deep_recursion": { "deep_recursion": {
"smarm 1-thread": { "smarm 1-thread": {
"result": 1, "result": 1,
"median": 226, "median": 316,
"min": 222, "min": 314,
"max": 243 "max": 322
}, },
"smarm 24-thread": { "smarm 20-thread": {
"result": 1, "result": 1,
"median": 745, "median": 758,
"min": 744, "min": 725,
"max": 776 "max": 1222
}, },
"tokio current_thread": { "tokio current_thread": {
"result": 1, "result": 1,
"median": 11, "median": 11,
"min": 10, "min": 11,
"max": 13 "max": 14
}, },
"tokio multi-thread": { "tokio multi-thread": {
"result": 1, "result": 1,
"median": 53, "median": 52,
"min": 53, "min": 51,
"max": 57 "max": 55
} }
}, },
"yield_in_hot_loop": { "fan_out_compute": {
"smarm 1-thread": { "smarm 1-thread": {
"result": 1000000, "result": 33860,
"median": 64849, "median": 15074,
"min": 64396, "min": 15064,
"max": 65283 "max": 15081
},
"smarm 20-thread": {
"result": 33860,
"median": 2786,
"min": 2539,
"max": 2913
}, },
"tokio current_thread": { "tokio current_thread": {
"result": 1000000, "result": 33860,
"median": 68507, "median": 14040,
"min": 62018, "min": 13749,
"max": 72341 "max": 14056
},
"tokio multi-thread": {
"result": 33860,
"median": 2404,
"min": 2136,
"max": 2436
}
},
"many_timers": {
"smarm 1-thread": {
"result": 10000,
"median": 117779,
"min": 107641,
"max": 120785
},
"smarm 20-thread": {
"result": 10000,
"median": 54031,
"min": 53646,
"max": 54841
},
"tokio current_thread": {
"result": 10000,
"median": 12417,
"min": 12383,
"max": 12479
},
"tokio multi-thread": {
"result": 10000,
"median": 13880,
"min": 13405,
"max": 13953
}
},
"mpsc_contention": {
"smarm 1-thread": {
"result": 320000,
"median": 3260,
"min": 2896,
"max": 3427
},
"smarm 20-thread": {
"result": 320000,
"median": 41206,
"min": 40562,
"max": 41754
},
"tokio current_thread": {
"result": 320000,
"median": 5895,
"min": 5272,
"max": 5902
},
"tokio multi-thread": {
"result": 320000,
"median": 71069,
"min": 60076,
"max": 80026
}
},
"multi_thread_scaling": {
"smarm 1-thread": {
"result": 33860,
"median": 15163,
"min": 15125,
"max": 15172
},
"smarm 2-thread": {
"result": 33860,
"median": 7980,
"min": 7947,
"max": 8103
},
"smarm 4-thread": {
"result": 33860,
"median": 4324,
"min": 4300,
"max": 4364
},
"smarm 20-thread": {
"result": 33860,
"median": 2704,
"min": 2391,
"max": 2841
},
"tokio multi 1-thread": {
"result": 33860,
"median": 14271,
"min": 14251,
"max": 14459
},
"tokio multi 2-thread": {
"result": 33860,
"median": 7395,
"min": 7195,
"max": 7443
},
"tokio multi 4-thread": {
"result": 33860,
"median": 3740,
"min": 3709,
"max": 3752
},
"tokio multi 20-thread": {
"result": 33860,
"median": 2373,
"min": 2176,
"max": 2380
}
},
"ping_pong_oneshot": {
"smarm 1-thread": {
"result": 1000,
"median": 935,
"min": 910,
"max": 941
},
"smarm 20-thread": {
"result": 1000,
"median": 6168,
"min": 6029,
"max": 6550
},
"tokio current_thread": {
"result": 1000,
"median": 407,
"min": 401,
"max": 430
},
"tokio multi-thread": {
"result": 1000,
"median": 9395,
"min": 8795,
"max": 9812
}
},
"ping_pong_steady": {
"smarm 1-thread": {
"result": 10000,
"median": 1517,
"min": 1501,
"max": 1527
},
"smarm 20-thread": {
"result": 10000,
"median": 1910,
"min": 1848,
"max": 1947
},
"tokio current_thread": {
"result": 10000,
"median": 1284,
"min": 1278,
"max": 1296
},
"tokio multi-thread": {
"result": 10000,
"median": 82015,
"min": 81833,
"max": 82993
}
},
"spawn_pair_control": {
"smarm 1-thread": {
"result": 1000,
"median": 637,
"min": 632,
"max": 661
},
"smarm 20-thread": {
"result": 1000,
"median": 5118,
"min": 5089,
"max": 5157
},
"tokio current_thread": {
"result": 1000,
"median": 288,
"min": 249,
"max": 300
},
"tokio multi-thread": {
"result": 1000,
"median": 8617,
"min": 8534,
"max": 8808
}
},
"spawn_storm_busy": {
"smarm 1-thread": {
"result": 10000,
"median": 105937,
"min": 101927,
"max": 107104
},
"smarm 20-thread": {
"result": 10000,
"median": 15928,
"min": 15918,
"max": 16581
},
"tokio current_thread": {
"result": 10000,
"median": 1248,
"min": 1121,
"max": 1253
},
"tokio multi-thread": {
"result": 10000,
"median": 12313,
"min": 11641,
"max": 16848
} }
}, },
"uncontended_channel": { "uncontended_channel": {
"smarm 1-thread": { "smarm 1-thread": {
"result": 1000000, "result": 1000000,
"median": 11949, "median": 12452,
"min": 11928, "min": 12315,
"max": 13596 "max": 13700
}, },
"tokio current_thread": { "tokio current_thread": {
"result": 1000000, "result": 1000000,
"median": 15083, "median": 15091,
"min": 15038, "min": 15086,
"max": 16994 "max": 16944
} }
}, },
"catch_unwind_panics": { "yield_in_hot_loop": {
"smarm 1-thread": { "smarm 1-thread": {
"result": 10000, "result": 1000000,
"median": 110932, "median": 33187,
"min": 110182, "min": 33000,
"max": 124147 "max": 33325
},
"smarm 24-thread": {
"result": 10000,
"median": 13665,
"min": 13172,
"max": 13784
}, },
"tokio current_thread": { "tokio current_thread": {
"result": 10000, "result": 1000000,
"median": 10431, "median": 74444,
"min": 9346, "min": 67009,
"max": 10915 "max": 75753
}
},
"yield_many": {
"smarm 1-thread": {
"result": 200000,
"median": 10921,
"min": 10828,
"max": 11041
},
"smarm 20-thread": {
"result": 200000,
"median": 44422,
"min": 43715,
"max": 44595
},
"tokio current_thread": {
"result": 200000,
"median": 5355,
"min": 5327,
"max": 5426
}, },
"tokio multi-thread": { "tokio multi-thread": {
"result": 10000, "result": 200000,
"median": 6171, "median": 6170,
"min": 5626, "min": 5982,
"max": 6555 "max": 6898
} }
} }
} }
+163 -2
View File
@@ -13,7 +13,18 @@
//! completeness. //! completeness.
//! 4. ping_pong_oneshot — N rounds of (spawn pair, send oneshot, await). //! 4. ping_pong_oneshot — N rounds of (spawn pair, send oneshot, await).
//! Closer to a request/response workload than channel //! Closer to a request/response workload than channel
//! ping-pong. //! ping-pong. NOTE: 2 spawns + 2 channel allocs + 2
//! joins per round; the message path is a minority
//! of it. Sections 5 and 6 split it apart.
//! 5. spawn_pair_control — section 4 with the messages removed: same
//! spawn/join shape, actors return immediately.
//! (4 − 5) ≈ per-round message-path cost.
//! 6. ping_pong_steady — ONE persistent pair, N roundtrips over unbounded
//! MPSC channels (smarm::channel vs
//! tokio::sync::mpsc::unbounded_channel — both
//! unbounded, non-blocking send). Steady-state
//! park/unpark cost per roundtrip, no spawn in the
//! loop.
use std::sync::atomic::{AtomicU64, Ordering}; use std::sync::atomic::{AtomicU64, Ordering};
use std::sync::Arc; use std::sync::Arc;
@@ -418,6 +429,135 @@ fn bench_pp_tokio_multi() -> (u64, u128) {
(PP_ROUNDS, start.elapsed().as_micros()) (PP_ROUNDS, start.elapsed().as_micros())
} }
// ---------------------------------------------------------------------------
// 5. spawn_pair_control — section 4 minus the messages
// ---------------------------------------------------------------------------
/// Target-5 instrumentation: emit the runtime's wake-path counters for one
/// run when SMARM_WAKE_DIAG is set. Off by default so sweep.py output is
/// unchanged.
fn wake_diag(section: &str, threads: usize, rt: &smarm::runtime::Runtime, us: u128) {
if std::env::var_os("SMARM_WAKE_DIAG").is_some() {
println!("DIAG,{section},{threads},{us},{}", rt.stats().wake_diag());
}
}
fn bench_ctl_smarm(threads: usize) -> (u64, u128) {
let start = Instant::now();
let rt = smarm::runtime::init(bench_cfg(threads));
rt.run(|| {
for _ in 0..PP_ROUNDS {
let hb = smarm::spawn(|| {});
let ha = smarm::spawn(|| {});
ha.join().unwrap();
hb.join().unwrap();
}
});
let us = start.elapsed().as_micros();
wake_diag("spawn_pair_control", threads, &rt, us);
(PP_ROUNDS, us)
}
fn bench_ctl_tokio_current() -> (u64, u128) {
let rt = tokio::runtime::Builder::new_current_thread()
.build()
.unwrap();
let start = Instant::now();
let local = tokio::task::LocalSet::new();
local.block_on(&rt, async move {
for _ in 0..PP_ROUNDS {
let hb = tokio::task::spawn_local(async {});
let ha = tokio::task::spawn_local(async {});
let _ = ha.await;
let _ = hb.await;
}
});
(PP_ROUNDS, start.elapsed().as_micros())
}
fn bench_ctl_tokio_multi() -> (u64, u128) {
let rt = tokio::runtime::Builder::new_multi_thread()
.worker_threads(available_threads())
.build()
.unwrap();
let start = Instant::now();
rt.block_on(async move {
for _ in 0..PP_ROUNDS {
let hb = tokio::spawn(async {});
let ha = tokio::spawn(async {});
let _ = ha.await;
let _ = hb.await;
}
});
(PP_ROUNDS, start.elapsed().as_micros())
}
// ---------------------------------------------------------------------------
// 6. ping_pong_steady — one persistent pair, PP_STEADY roundtrips
// ---------------------------------------------------------------------------
const PP_STEADY: u64 = 10_000;
fn bench_steady_smarm(threads: usize) -> (u64, u128) {
let start = Instant::now();
let rt = smarm::runtime::init(bench_cfg(threads));
rt.run(|| {
let (tx_ab, rx_ab) = smarm::channel::<u64>();
let (tx_ba, rx_ba) = smarm::channel::<u64>();
let echo = smarm::spawn(move || {
for _ in 0..PP_STEADY {
let v = rx_ab.recv().unwrap();
tx_ba.send(v + 1).unwrap();
}
});
for i in 0..PP_STEADY {
tx_ab.send(i).unwrap();
let v = rx_ba.recv().unwrap();
assert_eq!(v, i + 1);
}
echo.join().unwrap();
});
let us = start.elapsed().as_micros();
wake_diag("ping_pong_steady", threads, &rt, us);
(PP_STEADY, us)
}
async fn steady_tokio_body() {
let (tx_ab, mut rx_ab) = tokio::sync::mpsc::unbounded_channel::<u64>();
let (tx_ba, mut rx_ba) = tokio::sync::mpsc::unbounded_channel::<u64>();
let echo = tokio::spawn(async move {
for _ in 0..PP_STEADY {
let v = rx_ab.recv().await.unwrap();
tx_ba.send(v + 1).unwrap();
}
});
for i in 0..PP_STEADY {
tx_ab.send(i).unwrap();
let v = rx_ba.recv().await.unwrap();
assert_eq!(v, i + 1);
}
echo.await.unwrap();
}
fn bench_steady_tokio_current() -> (u64, u128) {
let rt = tokio::runtime::Builder::new_current_thread()
.build()
.unwrap();
let start = Instant::now();
rt.block_on(steady_tokio_body());
(PP_STEADY, start.elapsed().as_micros())
}
fn bench_steady_tokio_multi() -> (u64, u128) {
let rt = tokio::runtime::Builder::new_multi_thread()
.worker_threads(available_threads())
.build()
.unwrap();
let start = Instant::now();
rt.block_on(steady_tokio_body());
(PP_STEADY, start.elapsed().as_micros())
}
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// main // main
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -453,7 +593,8 @@ fn main() {
); );
println!( println!(
"CHAIN_DEPTH={CHAIN_DEPTH}, YIELD_TASKS={YIELD_TASKS}×{YIELD_ROUNDS}, \ "CHAIN_DEPTH={CHAIN_DEPTH}, YIELD_TASKS={YIELD_TASKS}×{YIELD_ROUNDS}, \
PRIME_N={PRIME_N}/{PRIME_WORKERS} workers, PP_ROUNDS={PP_ROUNDS}" PRIME_N={PRIME_N}/{PRIME_WORKERS} workers, PP_ROUNDS={PP_ROUNDS}, \
PP_STEADY={PP_STEADY}"
); );
// ---- 1. chained_spawn ---- // ---- 1. chained_spawn ----
@@ -491,4 +632,24 @@ fn main() {
run_n(&format!("smarm {n}-thread"), ITERS, || bench_pp_smarm(n)); run_n(&format!("smarm {n}-thread"), ITERS, || bench_pp_smarm(n));
run_n("tokio current_thread", ITERS, bench_pp_tokio_current); run_n("tokio current_thread", ITERS, bench_pp_tokio_current);
run_n("tokio multi-thread", ITERS, bench_pp_tokio_multi); run_n("tokio multi-thread", ITERS, bench_pp_tokio_multi);
// ---- 5. spawn_pair_control ----
print_header(&format!(
"spawn_pair_control: {PP_ROUNDS} rounds, no messages"
));
run_n("smarm 1-thread", ITERS, || bench_ctl_smarm(1));
run_n(&format!("smarm {n}-thread"), ITERS, || bench_ctl_smarm(n));
run_n("tokio current_thread", ITERS, bench_ctl_tokio_current);
run_n("tokio multi-thread", ITERS, bench_ctl_tokio_multi);
// ---- 6. ping_pong_steady ----
print_header(&format!(
"ping_pong_steady: 1 pair × {PP_STEADY} roundtrips"
));
run_n("smarm 1-thread", ITERS, || bench_steady_smarm(1));
run_n(&format!("smarm {n}-thread"), ITERS, || {
bench_steady_smarm(n)
});
run_n("tokio current_thread", ITERS, bench_steady_tokio_current);
run_n("tokio multi-thread", ITERS, bench_steady_tokio_multi);
} }
+12
View File
@@ -90,6 +90,7 @@ struct Sample {
us: u128, us: u128,
hits: u64, hits: u64,
displacements: u64, displacements: u64,
diag: String,
} }
fn yield_storm(threads: usize, slot: bool, actors: usize, yields: usize) -> Sample { fn yield_storm(threads: usize, slot: bool, actors: usize, yields: usize) -> Sample {
@@ -116,6 +117,7 @@ fn yield_storm(threads: usize, slot: bool, actors: usize, yields: usize) -> Samp
us, us,
hits: stats.slot_hits(), hits: stats.slot_hits(),
displacements: stats.slot_displacements(), displacements: stats.slot_displacements(),
diag: stats.wake_diag(),
} }
} }
@@ -159,6 +161,7 @@ fn ping_pong_pairs(threads: usize, slot: bool, pairs: usize, roundtrips: usize)
us, us,
hits: stats.slot_hits(), hits: stats.slot_hits(),
displacements: stats.slot_displacements(), displacements: stats.slot_displacements(),
diag: stats.wake_diag(),
} }
} }
@@ -185,6 +188,7 @@ fn spawn_storm(threads: usize, slot: bool, spawns: usize) -> Sample {
us, us,
hits: stats.slot_hits(), hits: stats.slot_hits(),
displacements: stats.slot_displacements(), displacements: stats.slot_displacements(),
diag: stats.wake_diag(),
} }
} }
@@ -253,6 +257,14 @@ fn main() {
mid.us, mid.us,
per_s per_s
); );
println!(
"RQDIAG,{},{},{},{},{}",
variant(),
slot_str,
name,
t,
mid.diag
);
if slot { if slot {
println!( println!(
"RQSLOT,{},{},{},{},{}", "RQSLOT,{},{},{},{},{}",
+25
View File
@@ -8,4 +8,29 @@ fn main() {
.flag_if_supported("-fno-stack-clash-protection") .flag_if_supported("-fno-stack-clash-protection")
.compile("smarm_canary"); .compile("smarm_canary");
println!("cargo:rerun-if-changed=canary/canary.c"); println!("cargo:rerun-if-changed=canary/canary.c");
// RFC 010 c6d — build_hash inputs. The compile-time facts a peer must
// share for a mesh link: the exact toolchain and the declared (enabled)
// feature set. Emitted as a plain string; the hashing (FNV-1a folded
// with PROTO_VERSION) happens in src/cluster.rs where the protocol
// version actually lives — parsing it out of a source file here would
// be a second, fragile copy. Always emitted, even for non-cluster
// builds: one env var costs the default build nothing.
let rustc = std::env::var("RUSTC").unwrap_or_else(|_| "rustc".to_string());
let version = std::process::Command::new(&rustc)
.arg("-V")
.output()
.ok()
.map(|o| String::from_utf8_lossy(&o.stdout).trim().to_string())
.filter(|v| !v.is_empty())
.unwrap_or_else(|| "rustc-unknown".to_string());
let mut feats: Vec<String> = std::env::vars()
.filter_map(|(k, _)| k.strip_prefix("CARGO_FEATURE_").map(str::to_string))
.collect();
feats.sort();
println!(
"cargo:rustc-env=SMARM_BUILD_HASH_INPUTS={version};features={}",
feats.join(",")
);
println!("cargo:rerun-if-env-changed=RUSTC");
} }
+4 -3
View File
@@ -1620,8 +1620,9 @@
wait: <code>select</code> priority is <strong>Down arms › Watcher arm › info channels (declaration order) › wait: <code>select</code> priority is <strong>Down arms › Watcher arm › info channels (declaration order) ›
inbox</strong>, rebuilt each turn. A hot inbox can't starve a death notice or a system message; inbox</strong>, rebuilt each turn. A hot inbox can't starve a death notice or a system message;
conversely a hot info channel <em>can</em> starve the inbox — deliberately. A closed info arm is conversely a hot info channel <em>can</em> starve the inbox — deliberately. A closed info arm is
silently dropped from the set; a closed <em>inbox</em> (every <code>ServerRef</code> gone) is graceful silently dropped from the set. The inbox never closes — the loop holds one sender for its whole
shutdown.</p> life, so a <code>GenServerRef</code> is an address, not an owner: the server ends only by
<code>StopHandle::stop</code>, a shutdown, a hard stop, or a panic.</p>
<h3>Death needs no monitor</h3> <h3>Death needs no monitor</h3>
<p>Server death detection falls out of channel closure. Already dead → the inbox is closed and <p>Server death detection falls out of channel closure. Already dead → the inbox is closed and
@@ -1700,7 +1701,7 @@
</div> </div>
<div class="module-card"> <div class="module-card">
<div class="module-name" style="color:var(--red)">Panics in <code>terminate()</code></div> <div class="module-name" style="color:var(--red)">Panics in <code>terminate()</code></div>
<p>gen_server's <code>terminate()</code> runs from a drop guard, possibly mid-unwind. A panic inside it during an unwind is a double panic → process abort, no supervision tree to save you. Keep it cheap, non-blocking, non-panicking.</p> <p>gen_server's <code>terminate()</code> runs from a drop guard, possibly mid-unwind. A panic inside it during an unwind is a double panic → process abort, no supervision tree to save you. On the panic and hard-stop paths keep it cheap, non-blocking, non-panicking. Only the graceful path (<code>handle_shutdown → Exit</code>, <code>StopHandle::stop</code>) runs it outside an unwind, where it may do real work.</p>
</div> </div>
<div class="module-card"> <div class="module-card">
<div class="module-name" style="color:var(--yellow)">Cold locks are leaf locks</div> <div class="module-name" style="color:var(--yellow)">Cold locks are leaf locks</div>
+139
View File
@@ -0,0 +1,139 @@
//! Graceful shutdown, end to end: a supervised app tree, a server that
//! drains before it exits, and the two ways the whole thing winds down.
//!
//! Stopping an actor comes in two strengths, as in OTP:
//! - `request_stop(pid)` = `exit(Pid, kill)`: cooperative hard stop,
//! unwinds at the next observation point.
//! - `request_shutdown(pid)` = `exit(Pid, shutdown)`: a trapping target gets
//! an `ExitSignal { reason: Shutdown }` and winds
//! down on its own terms; a non-trapping one is
//! stopped outright.
//!
//! A supervisor traps exits. `request_shutdown(sup)` runs its ordered
//! shutdown — children in reverse start order, each per its `ChildSpec`
//! `Shutdown` policy (`Timeout(d)` default 5s, `Infinity`, `BrutalKill`) —
//! and the supervisor then returns normally.
//!
//! Two triggers are shown:
//! 1. **Root exit.** The run's root actor returning means "the program is
//! done": the runtime delivers `request_shutdown` to every top-level actor
//! (here: the supervisor). Trapping actors may keep running to drain and
//! end the run when they stop themselves; non-trapping ones are stopped.
//! 2. **An outside thread** (e.g. a signal handler) driving it via
//! `RuntimeHandle::request_shutdown` on the supervisor — the root then
//! just waits for the tree to come down.
use smarm::gen_server::{
GenServer, GenServerBuilder, GenServerCtx, GenServerName, ShutdownAction, StopHandle,
TimerHandle,
};
use smarm::supervisor::{ChildSpec, OneForOne, Restart, Shutdown};
use smarm::{sleep, spawn};
use std::thread;
use std::time::Duration;
/// A server with in-flight work: on shutdown it stops accepting, finishes what
/// it has (simulated with a ticking timer), then ends itself.
struct Drainer {
pending: u32,
stop: Option<StopHandle<Drainer>>,
timer: Option<TimerHandle<Drainer>>,
}
impl GenServer for Drainer {
type Call = ();
type Reply = ();
type Cast = ();
type Info = ();
type Timer = ();
fn init(&mut self, ctx: &GenServerCtx<Self>) {
ctx.trap_exit(); // opt in: shutdown arrives as handle_shutdown
self.stop = Some(ctx.stop_handle());
self.timer = Some(ctx.timer());
}
fn handle_call(&mut self, _: ()) {}
fn handle_cast(&mut self, _: ()) {}
fn handle_shutdown(&mut self) -> ShutdownAction {
println!(
"drainer: shutdown requested, {} items pending",
self.pending
);
self.timer
.as_ref()
.unwrap()
.tick_every(Duration::from_millis(20), ());
ShutdownAction::Continue // keep serving until drained
}
fn handle_timer(&mut self, _: ()) {
self.pending -= 1;
if self.pending == 0 {
println!("drainer: drained, stopping");
self.stop.as_ref().unwrap().stop(); // normal exit
}
}
fn terminate(&mut self) {
// Graceful path: this runs on the normal path and may block.
println!("drainer: terminate");
}
}
/// The server's name: how the rest of the app reaches it (and the only handle
/// that survives a restart).
const DRAINER: GenServerName<Drainer> = GenServerName::new("drainer");
fn app_tree() -> OneForOne {
OneForOne::new()
.child(
ChildSpec::new(Restart::Permanent, || {
// A plain worker that does not trap: stopped outright on shutdown.
loop {
sleep(Duration::from_millis(10));
}
})
.shutdown(Shutdown::Timeout(Duration::from_millis(100))),
)
// A gen_server is a direct child: `named(N).run()` runs the loop as
// the child actor itself, so the supervisor's shutdown arrives as
// `handle_shutdown` and a restart re-binds the name.
.child(
ChildSpec::new(Restart::Permanent, || {
GenServerBuilder::new(Drainer {
pending: 3,
stop: None,
timer: None,
})
.named(DRAINER)
.run()
.expect("drainer name is free");
})
.shutdown(Shutdown::Infinity),
)
}
fn main() {
println!("--- 1. root exit drives the shutdown ---");
smarm::run(|| {
spawn(|| app_tree().run());
sleep(Duration::from_millis(50)); // the app "runs" for a while
// Returning here asks the supervisor to shut down; the run ends when
// the tree — drainer included — is gone.
});
println!("--- 2. an outside thread drives the shutdown ---");
let rt = smarm::init(smarm::Config::default());
let handle = rt.handle(); // Send + Sync; grab it before run
rt.run(move || {
let sup = spawn(|| app_tree().run());
let sup_pid = sup.pid();
// Stand-in for a SIGTERM handler thread.
thread::spawn(move || {
thread::sleep(Duration::from_millis(50));
println!("signal thread: requesting shutdown");
handle.request_shutdown(sup_pid);
});
sup.join()
.expect("supervisor returns normally after ordered shutdown");
println!("supervisor down; root returns");
});
}
+6
View File
@@ -66,6 +66,12 @@ fn main() {
let svc: Option<GenServerRef<Counter>> = whereis_server(COUNTER); let svc: Option<GenServerRef<Counter>> = whereis_server(COUNTER);
if let Some(svc) = svc { if let Some(svc) = svc {
let _ = svc.call(Query::Get); let _ = svc.call(Query::Get);
// A named server is pinned alive by the registry, so dropping refs
// does not end it. Stop it explicitly: `shutdown()` asks politely
// (a trapping server drains first; this one is stopped outright)
// and waits until it is gone. Left running, the root's return
// would shut it down the same way — see examples/graceful_shutdown.rs.
svc.shutdown();
} }
}); });
} }
+15 -2
View File
@@ -72,7 +72,10 @@ pub fn clear_current_pid() {
CURRENT_PID.with(|c| c.set(None)); CURRENT_PID.with(|c| c.set(None));
} }
/// Actor-side TLS accessor: `#[inline(never)]` + fence, see `context` docs.
#[inline(never)]
pub fn current_pid() -> Option<Pid> { pub fn current_pid() -> Option<Pid> {
crate::context::tls_fence();
CURRENT_PID.with(|c| c.get()) CURRENT_PID.with(|c| c.get())
} }
@@ -109,8 +112,7 @@ pub extern "C-unwind" fn trampoline() {
} }
}; };
LAST_OUTCOME.with(|r| *r.borrow_mut() = Some(outcome)); publish_outcome(outcome);
ACTOR_DONE.with(|c| c.set(true));
// Hand control back. The scheduler will tear down our slot and never // Hand control back. The scheduler will tear down our slot and never
// resume us again. // resume us again.
@@ -119,6 +121,17 @@ pub extern "C-unwind" fn trampoline() {
unreachable!("scheduler resumed a done actor"); unreachable!("scheduler resumed a done actor");
} }
/// Record the outcome for the scheduler that is about to be switched to. Kept
/// out of line: the actor may have migrated threads while its closure ran, so
/// the TLS base `trampoline` computed at entry must not be reused here (see
/// `context` module docs on thread-locals and migration).
#[inline(never)]
fn publish_outcome(outcome: Outcome) {
crate::context::tls_fence();
LAST_OUTCOME.with(|r| *r.borrow_mut() = Some(outcome));
ACTOR_DONE.with(|c| c.set(true));
}
/// One actor's worth of state. Owned by the scheduler's slot table. /// One actor's worth of state. Owned by the scheduler's slot table.
pub struct Actor { pub struct Actor {
/// The PID this actor was assigned at spawn time. /// The PID this actor was assigned at spawn time.
+12 -4
View File
@@ -274,12 +274,16 @@ mod inner {
/// The experiment-active path, kept out of the inlined fast path. /// The experiment-active path, kept out of the inlined fast path.
#[cold] #[cold]
#[inline(never)]
fn cold_check(exp: u64) { fn cold_check(exp: u64) {
crate::context::tls_fence();
let slot = preempt::current_slot_ptr(); let slot = preempt::current_slot_ptr();
if slot.is_null() { if slot.is_null() {
return; return;
} }
let now = preempt::rdtsc(); // Serialised: `now` closes an interval attributed to a code site; a
// speculative early read would drop that site's tail (RFC 007).
let now = preempt::rdtsc_serialising();
let last = LAST_SAMPLE_TSC.with(|c| c.replace(now)); let last = LAST_SAMPLE_TSC.with(|c| c.replace(now));
let target_site = (exp >> 32) as u32; let target_site = (exp >> 32) as u32;
let pct = exp & 0xffff_ffff; let pct = exp & 0xffff_ffff;
@@ -365,8 +369,9 @@ mod inner {
/// - Entering the target site: re-arm the sample clock, so time spent /// - Entering the target site: re-arm the sample clock, so time spent
/// *before* the site can never be attributed to it by the first /// *before* the site can never be attributed to it by the first
/// in-site check (the symmetric over-attribution). /// in-site check (the symmetric over-attribution).
#[inline] #[inline(never)]
fn site_transition(slot: *const crate::runtime::Slot, old: u32, new: u32) { fn site_transition(slot: *const crate::runtime::Slot, old: u32, new: u32) {
crate::context::tls_fence();
let exp = EXPERIMENT.load(Ordering::Relaxed); let exp = EXPERIMENT.load(Ordering::Relaxed);
if exp == 0 || old == new { if exp == 0 || old == new {
return; return;
@@ -374,7 +379,8 @@ mod inner {
let target = (exp >> 32) as u32; let target = (exp >> 32) as u32;
let pct = exp & 0xffff_ffff; let pct = exp & 0xffff_ffff;
if old == target && new != target { if old == target && new != target {
let now = preempt::rdtsc(); // Serialised: guard exit bounds the site's interval exactly.
let now = preempt::rdtsc_serialising();
let last = LAST_SAMPLE_TSC.with(|c| c.replace(now)); let last = LAST_SAMPLE_TSC.with(|c| c.replace(now));
if pct > 0 { if pct > 0 {
if last != 0 { if last != 0 {
@@ -386,7 +392,9 @@ mod inner {
} }
} }
} else if new == target && old != target { } else if new == target && old != target {
LAST_SAMPLE_TSC.with(|c| c.set(preempt::rdtsc())); // Serialised: an early arm would let pre-site work leak into the
// first in-site interval — the over-attribution this guards.
LAST_SAMPLE_TSC.with(|c| c.set(preempt::rdtsc_serialising()));
} }
} }
+39 -3
View File
@@ -90,8 +90,9 @@
use crate::pid::Pid; use crate::pid::Pid;
use crate::raw_mutex::RawMutex; use crate::raw_mutex::RawMutex;
use crate::runtime::RuntimeInner;
use std::collections::VecDeque; use std::collections::VecDeque;
use std::sync::Arc; use std::sync::{Arc, Weak};
/// Create a new channel and return its `(Sender, Receiver)` halves. /// Create a new channel and return its `(Sender, Receiver)` halves.
/// ///
@@ -101,6 +102,7 @@ pub fn channel<T>() -> (Sender<T>, Receiver<T>) {
let inner = Arc::new(RawMutex::new_channel(Inner { let inner = Arc::new(RawMutex::new_channel(Inner {
queue: VecDeque::new(), queue: VecDeque::new(),
parked_receiver: None, parked_receiver: None,
rt: None,
senders: 1, senders: 1,
receiver_alive: true, receiver_alive: true,
})); }));
@@ -120,10 +122,30 @@ struct Inner<T> {
/// `recv_timeout` whose timer fired after it was already satisfied) is /// `recv_timeout` whose timer fired after it was already satisfied) is
/// inert and does nothing when it fires. /// inert and does nothing when it fires.
parked_receiver: Option<(Pid, u32)>, parked_receiver: Option<(Pid, u32)>,
/// The receiver's runtime, captured the first time it parks (so provably
/// alive then) and kept for the life of the channel: it lets a sender on a
/// foreign OS thread wake the receiver without the `RUNTIME` thread-local,
/// which is unset off a scheduler thread. Captured once rather than per
/// park because `Arc::downgrade` + drop is a locked RMW pair on a shared
/// counter, and parking is the channel hot path. A `Receiver` never
/// migrates between runtimes — it is pinned to its actor — so one capture
/// stays correct for every later park.
rt: Option<Weak<RuntimeInner>>,
senders: usize, senders: usize,
receiver_alive: bool, receiver_alive: bool,
} }
impl<T> Inner<T> {
/// Capture the receiver's runtime if we have not already. Called under the
/// channel lock at every park site; after the first park it is one branch
/// on an `Option`, no atomics.
fn note_runtime(&mut self) {
if self.rt.is_none() {
self.rt = crate::scheduler::runtime_weak();
}
}
}
/// The sending half of a channel, created by [`channel`]. Clonable: every /// The sending half of a channel, created by [`channel`]. Clonable: every
/// clone pushes onto the same queue, and the channel stays open as long as /// clone pushes onto the same queue, and the channel stays open as long as
/// any clone is alive. Dropping the last `Sender` closes the channel, which /// any clone is alive. Dropping the last `Sender` closes the channel, which
@@ -207,7 +229,7 @@ impl<T> Drop for Sender<T> {
} }
}; };
if let Some((pid, epoch)) = unpark { if let Some((pid, epoch)) = unpark {
crate::scheduler::unpark_at(pid, epoch); crate::scheduler::unpark_at_via(pid, epoch, || self.inner.lock().rt.clone());
} }
} }
} }
@@ -241,6 +263,16 @@ impl<T> Sender<T> {
self.inner.lock().queue.len() self.inner.lock().queue.len()
} }
/// Whether the [`Receiver`] is still alive (a send would be accepted).
pub(crate) fn receiver_alive(&self) -> bool {
self.inner.lock().receiver_alive
}
/// Whether `other` is a sender of this very channel (a clone).
pub(crate) fn same_channel(&self, other: &Sender<T>) -> bool {
Arc::ptr_eq(&self.inner, &other.inner)
}
/// Push `value` onto the channel. Succeeds unconditionally as long as /// Push `value` onto the channel. Succeeds unconditionally as long as
/// the [`Receiver`] is still alive: the queue has no capacity limit, so /// the [`Receiver`] is still alive: the queue has no capacity limit, so
/// this never blocks and never fails except when the channel is closed, /// this never blocks and never fails except when the channel is closed,
@@ -260,7 +292,7 @@ impl<T> Sender<T> {
.unwrap_or(crate::pid::Pid::new(u32::MAX, u32::MAX)), .unwrap_or(crate::pid::Pid::new(u32::MAX, u32::MAX)),
receiver: Some(pid) receiver: Some(pid)
}); });
crate::scheduler::unpark_at(pid, epoch); crate::scheduler::unpark_at_via(pid, epoch, || self.inner.lock().rt.clone());
} else { } else {
crate::te!(crate::trace::Event::Send { crate::te!(crate::trace::Event::Send {
sender: crate::actor::current_pid() sender: crate::actor::current_pid()
@@ -299,6 +331,7 @@ impl<T> Receiver<T> {
// begin_wait is lock-free, so it's legal under the Channel lock; // begin_wait is lock-free, so it's legal under the Channel lock;
// registering in the same critical section makes the epoch // registering in the same critical section makes the epoch
// atomic with the senders' view of the registration. // atomic with the senders' view of the registration.
g.note_runtime();
g.parked_receiver = Some((me, crate::scheduler::begin_wait())); g.parked_receiver = Some((me, crate::scheduler::begin_wait()));
crate::te!(crate::trace::Event::RecvPark(me)); crate::te!(crate::trace::Event::RecvPark(me));
} }
@@ -351,6 +384,7 @@ impl<T> Receiver<T> {
"channel has more than one receiver" "channel has more than one receiver"
); );
epoch = crate::scheduler::begin_wait(); epoch = crate::scheduler::begin_wait();
g.note_runtime();
g.parked_receiver = Some((me, epoch)); g.parked_receiver = Some((me, epoch));
crate::te!(crate::trace::Event::RecvPark(me)); crate::te!(crate::trace::Event::RecvPark(me));
} }
@@ -426,6 +460,7 @@ impl<T> Receiver<T> {
g.parked_receiver.is_none_or(|(p, _)| p == me), g.parked_receiver.is_none_or(|(p, _)| p == me),
"channel has more than one receiver" "channel has more than one receiver"
); );
g.note_runtime();
g.parked_receiver = Some((me, crate::scheduler::begin_wait())); g.parked_receiver = Some((me, crate::scheduler::begin_wait()));
crate::te!(crate::trace::Event::RecvPark(me)); crate::te!(crate::trace::Event::RecvPark(me));
} }
@@ -565,6 +600,7 @@ impl<T> Selectable for Receiver<T> {
g.parked_receiver.is_none_or(|(p, _)| p == pid), g.parked_receiver.is_none_or(|(p, _)| p == pid),
"channel has more than one receiver" "channel has more than one receiver"
); );
g.note_runtime();
g.parked_receiver = Some((pid, epoch)); g.parked_receiver = Some((pid, epoch));
Ok(true) Ok(true)
} }
+282
View File
@@ -0,0 +1,282 @@
//! RFC 010 — clustering (smarm⇄smarm, explicit remote boundary).
//!
//! c1: feature flag + optional deps. c2: the owned envelope. c3: the
//! transport trait (control connection), framed codec, and the TCP +
//! loopback impls. c5: the handshake state machine. c6: the connection
//! [`manager`] (registry) and per-peer connection actors ([`conn`]), started
//! as an explicit supervision subtree, plus the handshake on the
//! accept/connect path ([`connect`]). Everything above them lands in later
//! chunks.
pub mod conn;
pub mod connect;
pub mod connector;
pub mod discovery;
pub mod envelope;
pub mod expose;
pub mod handshake;
pub mod manager;
pub mod membership;
pub mod pg;
pub mod remote;
pub mod transport;
use std::io;
use std::time::{Duration, SystemTime, UNIX_EPOCH};
use crate::gen_server::{self, GenServerBuilder};
use crate::monitor::monitor;
use crate::pg::Incarnation;
use crate::scheduler::{sleep, spawn, JoinHandle};
use crate::supervisor::{ChildSpec, OneForOne, Restart};
use envelope::NodeMeta;
use handshake::Local;
use transport::tcp::TcpTransport;
use transport::Transport;
pub use conn::{spawn_established, ConnHandle};
pub use connect::{dial, spawn_acceptor, AcceptorHandle};
pub use connector::{spawn_connector, ConnectorHandle};
pub use discovery::{Discovery, StaticSeeds, Strategy};
pub use envelope::RemoteDownReason;
pub use expose::{expose, expose_type, type_hash, DeliverError};
pub use manager::{Manager, MANAGER};
pub use membership::{subscribe, view, MembershipEvents, NodeEvent, NodeInfo};
pub use pg::{dispatch_any, members_all, pick_any, DispatchAnyError, GroupMember, PgMsg, PG_NAME};
pub use remote::{
demonitor_remote, monitor_remote, send_to_remote, NotConnected, RemoteDown, RemoteMonitor,
RemoteName, RemotePid, RemoteSendError, ToRemoteError,
};
/// c6d — the derived build hash for [`handshake::LocalNode::build_hash`]:
/// two builds may mesh only when this matches, and it is a pure function of
/// the compile-time inputs that define wire compatibility today — the exact
/// toolchain (`rustc -V`), the declared feature set, and
/// [`envelope::PROTO_VERSION`]. FNV-1a 64 over the build-script string, then
/// the proto version folded byte-wise, so a proto bump moves the hash even
/// on an identical toolchain. The domain is deliberately lean and
/// tightenable later without a wire change — it is just a `u64`.
pub const BUILD_HASH: u64 = fold_u32(
fnv1a64(env!("SMARM_BUILD_HASH_INPUTS").as_bytes()),
envelope::PROTO_VERSION,
);
/// FNV-1a 64 (const so [`BUILD_HASH`] is a compile-time fact).
const fn fnv1a64(bytes: &[u8]) -> u64 {
let mut h: u64 = 0xcbf2_9ce4_8422_2325;
let mut i = 0;
while i < bytes.len() {
h ^= bytes[i] as u64;
h = h.wrapping_mul(0x0000_0100_0000_01b3);
i += 1;
}
h
}
/// Continue an FNV-1a state over a `u32`'s little-endian bytes.
const fn fold_u32(mut h: u64, v: u32) -> u64 {
let b = v.to_le_bytes();
let mut i = 0;
while i < b.len() {
h ^= b[i] as u64;
h = h.wrapping_mul(0x0000_0100_0000_01b3);
i += 1;
}
h
}
/// The control-plane timing knobs, all with today's fixed values as
/// defaults ([`Timing::default`]). One struct threaded explicitly to the
/// acceptor, the dial path, every connection actor and the connector — no
/// ambient state, so a test can run a fast mesh without touching globals.
/// Every node in a mesh should agree on `heartbeat_interval` <
/// `liveness_timeout`; nothing enforces it.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub struct Timing {
/// Idle-connection heartbeat pace. Default [`conn::HEARTBEAT_INTERVAL`].
pub heartbeat_interval: Duration,
/// Inbound silence that tears a connection down. Default
/// [`conn::LIVENESS_TIMEOUT`].
pub liveness_timeout: Duration,
/// Per-frame handshake deadline on the accept/dial path. Default
/// [`connect::HANDSHAKE_TIMEOUT`].
pub handshake_timeout: Duration,
/// Connector redial delay after the first failure. Default
/// [`connector::INITIAL_BACKOFF`].
pub initial_backoff: Duration,
/// Connector redial delay cap. Default [`connector::MAX_BACKOFF`].
pub max_backoff: Duration,
}
impl Default for Timing {
fn default() -> Self {
Timing {
heartbeat_interval: conn::HEARTBEAT_INTERVAL,
liveness_timeout: conn::LIVENESS_TIMEOUT,
handshake_timeout: connect::HANDSHAKE_TIMEOUT,
initial_backoff: connector::INITIAL_BACKOFF,
max_backoff: connector::MAX_BACKOFF,
}
}
}
/// How to run this node: its identity and how it finds peers.
pub struct Config {
/// This node's claimed name — the mesh-wide identity peers dial by and
/// the tie-break input. Must be unique across the mesh.
pub node_name: String,
/// Metadata offered in this node's `Hello`.
pub meta: NodeMeta,
/// The control-connection listen address (e.g. `"127.0.0.1:0"`; the
/// concrete bound address is [`Cluster::local_addr`]).
pub listen_addr: String,
/// The peer-discovery strategy — [`StaticSeeds`] until richer ones land.
pub strategy: Box<dyn Strategy>,
/// Heartbeat / liveness / handshake / backoff knobs; [`Timing::default`]
/// is the shipping configuration.
pub timing: Timing,
}
/// A running cluster node: the supervised [`Manager`], the acceptor over the
/// bound listener, and the connector driving its [`Strategy`]. Roles will
/// eventually mount this; until the role mechanism lands it is started by
/// hand (RFC 010 §7).
///
/// Dropping the handle stops the acceptor and connector loops (no new
/// connections in either direction) but detaches the manager subtree, which
/// — with every established connection — keeps running for the life of the
/// runtime, the same split as [`AcceptorHandle`] alone.
pub struct Cluster {
_sup: JoinHandle,
acceptor: AcceptorHandle,
connector: ConnectorHandle,
local: Local,
}
impl Cluster {
/// The concrete bound listen address, dialable as-is.
pub fn local_addr(&self) -> &str {
self.acceptor.local_addr()
}
/// This node's handshake identity (name, incarnation, build hash, meta).
pub fn local(&self) -> &Local {
&self.local
}
/// Stop accepting and dialing. Established connections stay up (they
/// belong to the manager); tear those down via the manager.
pub fn shutdown(&self) {
self.acceptor.shutdown();
self.connector.shutdown();
}
}
/// Start a cluster node: the supervised manager (blocking until it is
/// registered and ready to answer), the acceptor bound per
/// [`Config::listen_addr`], and the connector running [`Config::strategy`].
/// The node's identity is completed here: `incarnation` is
/// [`self_incarnation`] and `build_hash` is [`BUILD_HASH`] — c7 is its first
/// consumer. Errs only if the listener cannot bind.
///
/// The manager is a supervised child (restarted on crash); per-peer
/// connection actors are dynamic and monitored by the manager rather than
/// statically supervised — a lost connection is re-established by the
/// connector's dial loop, never resurrected onto a stale socket.
pub fn start(config: Config) -> io::Result<Cluster> {
let sup = spawn(|| {
OneForOne::new()
.child(ChildSpec::new(Restart::Permanent, manager_child))
.run()
});
while gen_server::whereis_server(MANAGER).is_none() {
sleep(Duration::from_millis(1));
}
let local = Local {
node_name: config.node_name,
incarnation: self_incarnation(),
build_hash: BUILD_HASH,
meta: config.meta,
};
// The wire identity serialized pids are stamped with (c10).
remote::set_local_identity(&local.node_name, local.incarnation);
// The pg actor (Phase 5): subscribes membership, owns the "pg" name.
pg::attach_cluster();
let listener = TcpTransport.listen(&config.listen_addr)?;
let acceptor = spawn_acceptor(listener, local.clone(), config.timing);
let connector = spawn_connector(
Box::new(TcpTransport),
local.clone(),
config.strategy,
config.timing,
);
Ok(Cluster {
_sup: sup,
acceptor,
connector,
local,
})
}
/// This process's incarnation epoch: milliseconds since the Unix epoch,
/// truncated to `u32`. Not a clock — its one job is separating a node from
/// its own restart (two starts of the same name land on the same value only
/// if they happen within the same millisecond modulo ~49.7 days). Seconds
/// would be too coarse: a crash-and-restart inside one second is routine
/// under supervision.
pub fn self_incarnation() -> Incarnation {
let ms = SystemTime::now()
.duration_since(UNIX_EPOCH)
.map(|d| d.as_millis())
.unwrap_or(0);
Incarnation::new(ms as u32)
}
/// The supervised manager child body. It *is* the child actor: it starts the
/// named manager, then parks on the manager's own termination so this actor's
/// lifetime tracks the manager's — the supervisor's restart accounting keys off
/// this actor exiting.
fn manager_child() {
let m = match GenServerBuilder::new(Manager::new()).named(MANAGER).start() {
Ok(m) => m,
// Name still held by a not-yet-reaped prior instance: return and let
// the supervisor retry under its restart policy.
Err(_) => return,
};
let _ = monitor(m.pid()).rx.recv();
}
#[cfg(test)]
mod tests {
use super::*;
/// The hash core against the published FNV-1a 64 test vectors — the
/// contract is "this is FNV-1a", not "whatever the fn does".
#[test]
fn fnv1a64_known_vectors() {
assert_eq!(fnv1a64(b""), 0xcbf2_9ce4_8422_2325);
assert_eq!(fnv1a64(b"a"), 0xaf63_dc4c_8601_ec8c);
assert_eq!(fnv1a64(b"foobar"), 0x85944171f73967e8);
}
/// Folding the proto version continues the same FNV state: identical
/// inputs with a different version must land on a different hash.
#[test]
fn proto_version_moves_the_hash() {
let base = fnv1a64(b"same-toolchain;features=CLUSTER");
assert_ne!(fold_u32(base, 1), fold_u32(base, 2));
// And it equals hashing the bytes in one pass — the fold is a
// continuation, not a second construction.
let mut all = b"same-toolchain;features=CLUSTER".to_vec();
all.extend_from_slice(&1u32.to_le_bytes());
assert_eq!(fold_u32(base, 1), fnv1a64(&all));
}
/// The derived constant exists, is compile-time, and is not degenerate.
#[test]
fn build_hash_is_nonzero() {
const H: u64 = BUILD_HASH;
assert_ne!(H, 0);
}
}
+630
View File
@@ -0,0 +1,630 @@
//! RFC 010 c6 — the per-peer connection actor.
//!
//! One actor per established control connection. It owns the whole
//! [`FramedConn`] and, in a single [`select`](crate::select), waits on two
//! things at once: its command inbox and the connection becoming readable (the
//! [`FdArm`](crate::scheduler::FdArm) the transport hands back). That is why it
//! is a plain select-loop actor rather than a `gen_server` or `gen_statem` —
//! neither of those can fold fd-readiness into its wait, and folding it in is
//! the whole job. The single owner sends and receives on the one `FramedConn`,
//! so no read/write split is needed.
//!
//! The handshake completes *before* this actor exists (on the accept/connect
//! path — c6b) and produces the [`Peer`]; the *path* then registers the
//! connection with the [`manager`](crate::cluster::manager), which takes
//! ownership of its [`ConnHandle`] and monitors the actor, so any exit
//! deregisters the connection. The actor itself holds no authority over its
//! own lifetime: it runs until the manager drops its handle (deregistration,
//! `Disconnect`, or manager shutdown), the connection ends, or liveness
//! expires. Heartbeat send and fixed-timeout liveness are the timeout arm of
//! the same `select` (c6c): [`HEARTBEAT_INTERVAL`] paces outbound
//! [`Frame::Heartbeat`](crate::cluster::envelope::Frame::Heartbeat)s, and a
//! [`LIVENESS_TIMEOUT`] window — reset by any inbound frame — tears the
//! connection down when it empties.
//!
//! c9 adds the third arm — the connection's dedicated **outbound inbox**
//! (`Sender<Frame>` bound in the manager-maintained outbound table, D13),
//! drained onto the wire in the same loop — and inbound *interpretation*:
//! `SendNamed` goes to the one resolution seam,
//! [`remote::deliver_named`](crate::cluster::remote::deliver_named).
//! `Send` goes to the pid seam (c10). The outbound
//! sender is a separate channel from `cmd_tx` on purpose: closing it is not
//! a stop signal — lifetime authority stays with the [`ConnHandle`] (D9).
//!
//! c12 adds the monitor plane, and it lives *here* on purpose. Two tables,
//! both owned by this actor and dying with the connection:
//!
//! - **outstanding** — monitors *this* node holds on actors at the peer:
//! `monitor_id → (target, Sender<RemoteDown>)`. Fed by
//! [`MonCmd`](crate::cluster::remote::MonCmd) from `monitor_remote`; the
//! actor records the id and *then* emits the `Monitor` frame, so a `Down`
//! frame can never race an entry that isn't there yet. An inbound `Down`
//! removes the entry and delivers.
//! - **watched** — monitors the *peer* holds on actors here: `monitor_id →
//! local Monitor`. An inbound `Monitor` is admitted only for a pid that
//! was exposed or crossed the wire (`is_watchable`, D12): a corpse answers
//! with its recorded terminal reason (RFC §6), an unwatchable or unknown
//! pid with `NoProc` — indistinguishable from dead, so nothing leaks. A
//! live watchable pid gets a local monitor whose `rx` is one more arm of
//! the select; its `Down` goes back as a frame.
//!
//! Because both tables are actor state, connection loss (c13) needs no
//! second bookkeeping owner: this actor's exit is the one place that knows
//! every monitor the link was carrying. `Monitors::teardown` runs on every
//! exit path and answers each outstanding monitor with `Disconnected` —
//! the roadmap's "partition vs. death" contrast: an actor that dies sends
//! its true reason over the link, a link that dies says only that.
use std::collections::HashMap;
use std::time::{Duration, Instant};
use crate::channel::{channel, try_select_timeout, Receiver, Selectable, Sender};
use crate::cluster::envelope::{Frame, RemoteDownReason};
use crate::cluster::handshake::Peer;
use crate::cluster::manager::{Call, Registered, Reply, MANAGER};
use crate::cluster::remote::{
deliver_named, deliver_to_pid, InboundVerdict, MonCmd, RemoteDown, RemotePid,
};
use crate::cluster::transport::FramedConn;
use crate::cluster::Timing;
use crate::gen_server;
use crate::monitor::{
demonitor, is_watchable, monitor, terminal_reason, DownReason, Monitor, MonitorId,
};
use crate::pid::{Erased, Pid};
use crate::scheduler::spawn;
/// Commands to a running connection actor.
enum Cmd {
Shutdown,
}
/// The manager's authority over one connection actor: while this handle
/// lives the connection lives, and dropping it stops the actor and closes
/// the socket. Only the [`manager`](crate::cluster::manager) holds one —
/// callers of [`spawn_established`] get a [`Pid`] and no lifetime authority,
/// so a connection can never outlive, or die with, whichever actor happened
/// to establish it.
pub struct ConnHandle {
cmd_tx: Sender<Cmd>,
/// The connection's dedicated outbound inboxes — frames and monitor
/// commands. The manager moves them into the outbound table on
/// `Register` (see [`take_outbound`](ConnHandle::take_outbound)); a
/// `Duplicate` verdict drops them with the handle.
out_tx: Option<(Sender<Frame>, Sender<MonCmd>)>,
}
impl std::fmt::Debug for ConnHandle {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
f.write_str("ConnHandle")
}
}
impl ConnHandle {
/// Ask the connection to close and exit. Idempotent, and a no-op if the
/// actor has already gone. Dropping the handle does the same thing; this
/// exists for the manager's explicit `Disconnect` path.
pub fn shutdown(&self) {
let _ = self.cmd_tx.send(Cmd::Shutdown);
}
/// Manager-only: take the outbound senders to bind into the outbound
/// table. Once, at registration.
pub(crate) fn take_outbound(&mut self) -> Option<(Sender<Frame>, Sender<MonCmd>)> {
self.out_tx.take()
}
}
/// The name was already claimed by a live connection, so this one was
/// refused; its actor has been stopped and its socket closed.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub struct RegisterRefused;
/// Spawn a connection actor for an **already-established** connection (the
/// handshake completed on the path and produced `peer`) and register it with
/// the manager, synchronously, before returning. The manager takes the
/// actor's [`ConnHandle`]; the caller gets only the [`Pid`], because
/// connection lifetime belongs to the table and not to the establishing
/// actor. A refusal has already stopped the actor and closed the socket.
pub fn spawn_established(
framed: FramedConn,
peer: Peer,
timing: Timing,
) -> Result<Pid, RegisterRefused> {
let (cmd_tx, cmd_rx) = channel();
let (out_tx, out_rx) = channel();
let (mon_tx, mon_rx) = channel();
let reg_peer = peer.clone();
let pid = spawn(move || run(framed, peer, timing, cmd_rx, out_rx, mon_rx)).pid();
match gen_server::call(
MANAGER,
Call::Register {
peer: reg_peer,
pid,
handle: ConnHandle {
cmd_tx,
out_tx: Some((out_tx, mon_tx)),
},
},
) {
Ok(Reply::Registered(Registered::Ok)) => Ok(pid),
// Duplicate name, or the manager is unreachable. Either way the
// handle went with the call and is dropped there (or never arrived
// and dropped with it), which stops the actor and closes the socket.
_ => Err(RegisterRefused),
}
}
/// Default for [`Timing::heartbeat_interval`]: how often this end emits
/// [`Frame::Heartbeat`] on an idle connection. The first one goes out
/// immediately at spawn, so the peer's liveness window starts fed.
pub const HEARTBEAT_INTERVAL: Duration = Duration::from_secs(1);
/// Default for [`Timing::liveness_timeout`]: how long the connection may go
/// without a single inbound frame before it is declared dead and torn down. Any inbound frame resets the window —
/// heartbeats keep an idle connection alive, and real traffic (c8+) counts
/// for free. Fixed by design (RFC v2 §5): this is the control connection, a
/// heartbeat can never queue behind bulk traffic, so a fixed timeout is an
/// honest detector.
pub const LIVENESS_TIMEOUT: Duration = Duration::from_secs(4);
fn run(
mut framed: FramedConn,
_peer: Peer,
timing: Timing,
cmd_rx: Receiver<Cmd>,
out_rx: Receiver<Frame>,
mon_rx: Receiver<MonCmd>,
) {
let mut mons = Monitors::default();
match framed.readable_arm() {
Some(arm) => run_live(
&mut framed,
arm,
timing,
&cmd_rx,
&out_rx,
&mon_rx,
&mut mons,
),
None => run_inert(&cmd_rx),
}
framed.close();
mons.teardown(&mon_rx);
}
/// The monitor plane's two tables (module docs). Owned by the actor.
#[derive(Default)]
struct Monitors {
/// Monitors this node holds on peer actors: id → (target, delivery).
outstanding: HashMap<MonitorId, (RemotePid<Erased>, Sender<RemoteDown>)>,
/// Monitors the peer holds on local actors: id → the local monitor.
watched: HashMap<MonitorId, Monitor>,
}
impl Monitors {
/// The connection is gone, whatever the exit path (liveness expiry,
/// EOF, wire failure, commanded stop): release the peer's local
/// monitors, and answer every one of ours with `Disconnected` — nothing
/// more can be known about those actors. Commands still sitting in the
/// inbox are folded in first (a `Monitor` handed to us but never
/// processed gets its notice too; a `Demonitor` still cancels), so the
/// only registration that can miss this is one that lands after the
/// drain and before the inbox drops — the reader side backstops that
/// (`RemoteMonitor`). Entries leave the table as they are answered, and
/// this runs once per actor, so no monitor sees two notices.
fn teardown(&mut self, mon_rx: &Receiver<MonCmd>) {
for (_, m) in self.watched.drain() {
let _ = demonitor(&m);
}
while let Ok(Some(cmd)) = mon_rx.try_recv() {
match cmd {
MonCmd::Monitor { id, target, tx } => {
self.outstanding.insert(id, (target, tx));
}
MonCmd::Demonitor { id } => {
self.outstanding.remove(&id);
}
}
}
for (_, (pid, tx)) in self.outstanding.drain() {
let _ = tx.send(RemoteDown {
pid,
reason: RemoteDownReason::Disconnected,
});
}
}
/// Admit a peer's `Monitor` for local `(index, generation)`. Returns
/// the reason to answer with at once, or `None` if a live monitor was
/// installed. Corpse → recorded terminal reason (RFC §6, and only
/// watchable deaths are recorded); live watchable → monitor; anything
/// else → `NoProc`. The check-then-monitor race (dies in between) is
/// closed on the read side: a `NoProc` from a monitor we installed on a
/// live pid is upgraded through `terminal_reason` in `sweep_watched`.
fn admit(&mut self, id: MonitorId, index: u32, generation: u32) -> Option<DownReason> {
let pid = Pid::new(index, generation);
if let Some(reason) = terminal_reason(pid) {
return Some(reason);
}
if !is_watchable(pid) {
return Some(DownReason::NoProc);
}
let m = monitor(pid);
self.watched.insert(id, m);
None
}
fn cancel(&mut self, id: MonitorId) {
if let Some(m) = self.watched.remove(&id) {
let _ = demonitor(&m);
}
}
/// Collect every local `Down` that has arrived for a peer-held monitor.
fn sweep_watched(&mut self) -> Vec<(MonitorId, DownReason)> {
let mut fired = Vec::new();
for (id, m) in self.watched.iter() {
if let Ok(Some(down)) = m.rx.try_recv() {
let reason = match down.reason {
DownReason::NoProc => terminal_reason(m.target).unwrap_or(DownReason::NoProc),
r => r,
};
fired.push((*id, reason));
}
}
for (id, _) in &fired {
self.watched.remove(id);
}
fired
}
/// The peer reports a monitored actor down: deliver locally.
fn down(&mut self, id: MonitorId, reason: RemoteDownReason) {
if let Some((pid, tx)) = self.outstanding.remove(&id) {
let _ = tx.send(RemoteDown { pid, reason });
}
}
}
/// The steady-state loop over an fd-backed connection: one
/// `select_timeout` folds the command inbox, the outbound inbox, socket
/// readability, and the nearer of the two deadlines (`hb_send`,
/// `liveness`) into a single wait.
fn run_live(
framed: &mut FramedConn,
arm: crate::scheduler::FdArm,
timing: Timing,
cmd_rx: &Receiver<Cmd>,
out_rx: &Receiver<Frame>,
mon_rx: &Receiver<MonCmd>,
mons: &mut Monitors,
) {
let mut next_hb = Instant::now();
let mut live_until = Instant::now() + timing.liveness_timeout;
// The outbound senders live in the manager's table and are dropped on
// unbind; after that these arms would wake forever, so they drop out of
// the select (not a stop signal — see the module docs).
let mut out_open = true;
let mut mon_open = true;
// Which wait each select arm stands for. Built in lockstep with the
// `Selectable` vector each iteration, so a wake is decoded by name and
// never by position.
enum Arm {
Cmd,
Fd,
Out,
Mon,
/// A peer-held local monitor (any of them: firing sweeps them all).
Watched,
}
fn push<'s>(
arms: &mut Vec<&'s dyn Selectable>,
what: &mut Vec<Arm>,
s: &'s dyn Selectable,
a: Arm,
) {
arms.push(s);
what.push(a);
}
loop {
let now = Instant::now();
if now >= live_until {
break; // liveness expired: the peer is dead to us
}
if now >= next_hb {
if framed.send(&Frame::Heartbeat).is_err() {
break;
}
next_hb = now + timing.heartbeat_interval;
}
let wait = next_hb.min(live_until).saturating_duration_since(now);
let mut arms: Vec<&dyn Selectable> = Vec::new();
let mut what: Vec<Arm> = Vec::new();
push(&mut arms, &mut what, cmd_rx, Arm::Cmd);
push(&mut arms, &mut what, &arm, Arm::Fd);
if out_open {
push(&mut arms, &mut what, out_rx, Arm::Out);
}
if mon_open {
push(&mut arms, &mut what, mon_rx, Arm::Mon);
}
for m in mons.watched.values() {
push(&mut arms, &mut what, &m.rx, Arm::Watched);
}
match try_select_timeout(&arms, wait).map(|i| i.map(|i| &what[i])) {
Ok(Some(Arm::Cmd)) => {
if should_stop(cmd_rx) {
break;
}
}
Ok(Some(Arm::Fd)) => match pump_readable(framed, mons) {
Pump::Ended => break,
Pump::Frames(n) => {
if n > 0 {
live_until = Instant::now() + timing.liveness_timeout;
}
}
},
Ok(Some(Arm::Out)) => match pump_outbound(framed, out_rx) {
Outbound::Drained => {}
Outbound::Closed => out_open = false,
Outbound::WireFailed => break,
},
Ok(Some(Arm::Mon)) => match pump_moncmds(framed, mon_rx, mons) {
Outbound::Drained => {}
Outbound::Closed => mon_open = false,
Outbound::WireFailed => break,
},
Ok(Some(Arm::Watched)) => {
for (id, reason) in mons.sweep_watched() {
let frame = Frame::Down {
monitor_id: id.0,
reason: reason.into(),
};
if framed.send(&frame).is_err() {
return;
}
}
}
// A deadline passed; the top of the loop acts on whichever.
Ok(None) => {}
// The fd arm failed to register — the connection is gone.
Err(_) => break,
}
}
}
/// Drain the monitor-command inbox: record, then emit (module docs).
fn pump_moncmds(
framed: &mut FramedConn,
mon_rx: &Receiver<MonCmd>,
mons: &mut Monitors,
) -> Outbound {
loop {
match mon_rx.try_recv() {
Ok(Some(MonCmd::Monitor { id, target, tx })) => {
let frame = Frame::Monitor {
monitor_id: id.0,
index: target.index(),
generation: target.generation(),
};
mons.outstanding.insert(id, (target, tx));
if framed.send(&frame).is_err() {
return Outbound::WireFailed;
}
}
Ok(Some(MonCmd::Demonitor { id })) => {
if mons.outstanding.remove(&id).is_some()
&& framed.send(&Frame::Demonitor { monitor_id: id.0 }).is_err()
{
return Outbound::WireFailed;
}
}
Ok(None) => return Outbound::Drained,
Err(_) => return Outbound::Closed,
}
}
}
/// What one outbound-side wake (frames or monitor commands) yielded.
enum Outbound {
/// Everything queued went onto the wire; the inbox is open and empty.
Drained,
/// The manager unbound this connection's sender; nothing more will come.
Closed,
/// The socket refused a write: the connection is gone.
WireFailed,
}
/// Drain every queued outbound frame onto the wire.
fn pump_outbound(framed: &mut FramedConn, out_rx: &Receiver<Frame>) -> Outbound {
loop {
match out_rx.try_recv() {
Ok(Some(frame)) => {
if framed.send(&frame).is_err() {
return Outbound::WireFailed;
}
}
Ok(None) => return Outbound::Drained,
Err(_) => return Outbound::Closed,
}
}
}
/// No fd to select on (loopback): only a command can end the wait, and
/// neither heartbeats nor liveness run — a transport that can't report
/// readiness can't be timed either (same caveat as
/// [`FramedConn::recv_deadline`]). Loopback is a test transport; every real
/// connection is fd-backed.
fn run_inert(cmd_rx: &Receiver<Cmd>) {
loop {
let arms: [&dyn Selectable; 1] = [cmd_rx];
let _ = crate::channel::select(&arms);
if should_stop(cmd_rx) {
break;
}
}
}
/// Drain the command arm. Returns `true` when the actor should exit — a
/// shutdown was requested, or the last handle was dropped.
fn should_stop(cmd_rx: &Receiver<Cmd>) -> bool {
match cmd_rx.try_recv() {
Ok(Some(Cmd::Shutdown)) => true,
Ok(None) => false, // spurious wake
Err(_) => true, // all senders dropped
}
}
/// What one readable wake yielded.
enum Pump {
/// The connection has ended: EOF (clean or mid-frame) or an
/// unrecoverable stream error.
Ended,
/// Still up; this many complete frames were consumed (possibly zero, if
/// the wake delivered only part of a frame). Any nonzero count resets
/// the liveness window.
Frames(usize),
}
/// Surface an inbound verdict: one `smarm-trace` event, nothing else — it
/// is local knowledge (RFC §3). A no-op without the feature.
fn note_verdict(verdict: InboundVerdict) {
#[cfg(feature = "smarm-trace")]
crate::te!(crate::trace::Event::ClusterInbound(verdict.label()));
#[cfg(not(feature = "smarm-trace"))]
drop(verdict);
}
/// Consume one readable wake: exactly one socket read (which cannot block
/// after a level-triggered readable indication), then drain every complete
/// frame the buffer now holds. A blocking `recv` here would park the actor
/// past its heartbeat and liveness deadlines whenever a frame arrives split.
/// Every consumed frame counts for liveness; `SendNamed` goes to the one
/// name-resolution seam and `Send` to the pid seam. Verdicts are local
/// knowledge only — nothing goes back on the wire (RFC §3) — and surface
/// as one `smarm-trace` `ClusterInbound` event each (zero cost off).
/// `Monitor`/`Demonitor`/`Down` go to the [`Monitors`] tables; a `Monitor`
/// that can be answered at once is answered inline.
fn pump_readable(framed: &mut FramedConn, mons: &mut Monitors) -> Pump {
let eof = match framed.read_once() {
Ok(n) => n == 0,
Err(_) => return Pump::Ended,
};
let mut got = 0;
loop {
match framed.next_buffered() {
Ok(Some(frame)) => {
got += 1;
match frame {
Frame::SendNamed {
name,
type_hash,
payload,
} => {
note_verdict(deliver_named(&name, type_hash, &payload));
}
Frame::Send {
index,
generation,
type_hash,
payload,
} => {
note_verdict(deliver_to_pid(index, generation, type_hash, &payload));
}
Frame::Monitor {
monitor_id,
index,
generation,
} => {
let id = MonitorId(monitor_id);
if let Some(reason) = mons.admit(id, index, generation) {
let frame = Frame::Down {
monitor_id,
reason: reason.into(),
};
if framed.send(&frame).is_err() {
return Pump::Ended;
}
}
}
Frame::Demonitor { monitor_id } => mons.cancel(MonitorId(monitor_id)),
Frame::Down { monitor_id, reason } => mons.down(MonitorId(monitor_id), reason),
// Heartbeat: liveness only. Handshake frames after
// establishment: ignored.
_ => {}
}
}
Ok(None) => break,
Err(_) => return Pump::Ended, // corrupt stream
}
}
if eof {
Pump::Ended
} else {
Pump::Frames(got)
}
}
#[cfg(test)]
mod tests {
//! `Monitors::teardown` in isolation: the actor-side half of c13, pinned
//! separately because from the outside it is indistinguishable from the
//! read-side backstop in `RemoteMonitor` (both yield `Disconnected`).
use super::*;
use crate::pg::Incarnation;
fn pid(index: u32) -> RemotePid<Erased> {
RemotePid::from_parts("peer", Incarnation::new(1), index, 1)
}
#[test]
fn teardown_answers_every_outstanding_and_unread_monitor_once() {
crate::run(|| {
let mut mons = Monitors::default();
let (mon_tx, mon_rx) = channel::<MonCmd>();
// Already registered.
let (tx1, rx1) = channel::<RemoteDown>();
mons.outstanding.insert(MonitorId(1), (pid(1), tx1));
// In the inbox, never processed.
let (tx2, rx2) = channel::<RemoteDown>();
mon_tx
.send(MonCmd::Monitor {
id: MonitorId(2),
target: pid(2),
tx: tx2,
})
.ok()
.unwrap();
// Registered, then cancelled in the inbox: silence.
let (tx3, rx3) = channel::<RemoteDown>();
mons.outstanding.insert(MonitorId(3), (pid(3), tx3));
mon_tx
.send(MonCmd::Demonitor { id: MonitorId(3) })
.ok()
.unwrap();
mons.teardown(&mon_rx);
let d1 = rx1.recv().unwrap();
assert_eq!(
(d1.pid, d1.reason),
(pid(1), RemoteDownReason::Disconnected)
);
let d2 = rx2.recv().unwrap();
assert_eq!(
(d2.pid, d2.reason),
(pid(2), RemoteDownReason::Disconnected)
);
// Cancelled: no notice was sent (its sender is dropped, channel
// closed-empty), and nobody got a second one.
assert!(rx3.try_recv().is_err());
assert!(rx1.try_recv().is_err());
assert!(rx2.try_recv().is_err());
assert!(mons.outstanding.is_empty());
});
}
}
+365
View File
@@ -0,0 +1,365 @@
//! RFC 010 c6b — the handshake on the accept/connect path.
//!
//! Per D8 (re-amended): the c5 machines are driven by **straight-line code
//! on the path**, not by an actor. The dial side runs [`Initiator`]; the
//! acceptor loop runs [`Responder`]. A connection actor is spawned only
//! *after* a successful handshake ([`spawn_established`]); every reject,
//! protocol failure, timeout, and tie-break loss is resolved right here,
//! on the path, by closing — no actor ever exists for a connection that
//! didn't establish.
//!
//! Buffer trap (binding): the path reader and the steady-state actor share
//! ONE [`FramedConn`]. Its decode buffer may hold read-ahead past the
//! handshake frames, so the *whole* `FramedConn` travels into
//! [`spawn_established`] — never a fresh codec over the same socket.
//!
//! Layering: [`dial_handshake`] and [`accept_handshake`] are the bare path
//! steps — IO on a `FramedConn`, no manager, no actors — testable over the
//! loopback transport on plain threads. [`dial`] and [`spawn_acceptor`] are
//! the manager-integrated layer (actor context required): they keep the
//! [`manager`](crate::cluster::manager)'s dial-intent set honest and spawn
//! the connection actor on success.
use std::io;
use std::time::{Duration, Instant};
use crate::channel::{channel, Receiver, Selectable, Sender};
use crate::cluster::conn::spawn_established;
use crate::cluster::envelope::{Frame, RejectReason};
use crate::cluster::handshake::{
Initiator, InitiatorOutcome, Local, Peer, PeerStanding, Responder, ResponderOutcome,
};
use crate::cluster::manager::{Call, Reply, MANAGER};
use crate::cluster::transport::{FramedConn, Listener, RecvError, SendError, Transport};
use crate::cluster::Timing;
use crate::gen_server;
use crate::pid::Pid;
use crate::scheduler::{self, spawn};
/// Default for [`Timing::handshake_timeout`]: how long either side waits for
/// the peer's handshake frame before giving up and closing. Enforced on the path via [`FramedConn::recv_deadline`],
/// so a peer that connects and goes silent cannot wedge the acceptor.
pub const HANDSHAKE_TIMEOUT: Duration = Duration::from_secs(5);
/// Why a handshake did not establish. In every case the connection has
/// already been closed on the path by the time this is returned.
#[derive(Debug)]
pub enum HandshakeError {
/// A `HelloReject` travelled — sent by us (accept side) or received by
/// us (dial side).
Rejected(RejectReason),
/// Accept side only: the inbound dial lost the simultaneous-connect
/// tie-break (D7) and was closed silently, no frame sent.
TieBreakLoss,
/// The peer spoke a valid frame that is wrong here (non-`Hello` first
/// frame; non-response to our `Hello`), or an undecodable byte stream.
Protocol,
/// EOF before the handshake resolved. On the dial side this is also
/// what losing the tie-break looks like: the peer closes silently.
Closed,
/// [`HANDSHAKE_TIMEOUT`] (or the caller's deadline) passed first.
TimedOut,
/// The transport failed mid-handshake.
Transport(io::Error),
}
fn from_send(e: SendError) -> HandshakeError {
match e {
// Handshake frames are small and self-made; an encode failure is a
// protocol-level impossibility, not a transport fault.
SendError::Encode(_) => HandshakeError::Protocol,
SendError::Io(e) => HandshakeError::Transport(e),
}
}
fn from_recv(e: RecvError) -> HandshakeError {
match e {
RecvError::Corrupt(_) => HandshakeError::Protocol,
RecvError::TruncatedByPeer => HandshakeError::Closed,
RecvError::Io(e) => HandshakeError::Transport(e),
RecvError::TimedOut => HandshakeError::TimedOut,
}
}
/// Dial-side path step: send our `Hello`, interpret the one response. On
/// `Ok` the connection is established and `framed` is live (with any
/// read-ahead intact in its buffer); on `Err` the connection is closed.
pub fn dial_handshake(
framed: &mut FramedConn,
local: &Local,
deadline: Instant,
) -> Result<Peer, HandshakeError> {
let (initiator, hello) = Initiator::new(local);
if let Err(e) = framed.send(&hello) {
framed.close();
return Err(from_send(e));
}
let outcome = match framed.recv_deadline(deadline) {
Ok(Some(frame)) => initiator.on_frame(frame),
Ok(None) => {
framed.close();
return Err(HandshakeError::Closed);
}
Err(e) => {
framed.close();
return Err(from_recv(e));
}
};
match outcome {
InitiatorOutcome::Established(peer) => Ok(peer),
InitiatorOutcome::Rejected(reason) => {
framed.close();
Err(HandshakeError::Rejected(reason))
}
InitiatorOutcome::Failed(_) => {
framed.close();
Err(HandshakeError::Protocol)
}
}
}
/// Accept-side path step: read the first frame, judge it, answer or close.
///
/// `standing_of` supplies the [`PeerStanding`] of the *offered* name — knowledge
/// only the frame reveals, which is why it is a callback and not a value
/// (the integrated acceptor asks the manager; loopback tests fabricate).
/// It is not called when the first frame is not a `Hello`.
///
/// On `Ok` the ack has been sent and `framed` is live (read-ahead intact);
/// on `Err` any owed reject has been sent and the connection is closed.
pub fn accept_handshake(
framed: &mut FramedConn,
local: Local,
standing_of: impl FnOnce(&str) -> PeerStanding,
deadline: Instant,
) -> Result<Peer, HandshakeError> {
let frame = match framed.recv_deadline(deadline) {
Ok(Some(frame)) => frame,
Ok(None) => {
framed.close();
return Err(HandshakeError::Closed);
}
Err(e) => {
framed.close();
return Err(from_recv(e));
}
};
let standing = match &frame {
Frame::Hello { node_name, .. } => standing_of(node_name),
_ => PeerStanding::Free,
};
match Responder::new(local).on_frame(frame, standing) {
ResponderOutcome::Accepted { reply, peer } => {
if let Err(e) = framed.send(&reply) {
framed.close();
return Err(from_send(e));
}
Ok(peer)
}
ResponderOutcome::Rejected { reply, reason } => {
// Best effort: the reject is the cross-version compatibility
// anchor, but if the write fails the peer sees a bare close,
// which it must survive anyway.
let _ = framed.send(&reply);
framed.close();
Err(HandshakeError::Rejected(reason))
}
ResponderOutcome::TieBreakLoss => {
// D7: close silently — the peer computes the same verdict.
framed.close();
Err(HandshakeError::TieBreakLoss)
}
ResponderOutcome::Failed(_) => {
framed.close();
Err(HandshakeError::Protocol)
}
}
}
// ---------------------------------------------------------------------------
// Manager-integrated layer
// ---------------------------------------------------------------------------
/// Why an integrated [`dial`] did not produce a connection.
#[derive(Debug)]
pub enum DialError {
/// Another dial to this peer name is already in flight.
AlreadyDialing,
/// The manager is not running (or answered nonsense).
ManagerUnavailable,
/// The transport could not connect.
Connect(io::Error),
/// Connected, but the handshake did not establish.
Handshake(HandshakeError),
/// The peer at `addr` established, but answered as a different name
/// than the one we dialed — the tie-break bookkeeping (keyed by the
/// dialed name) would be unsound, so the connection is closed.
PeerNameMismatch { expected: String, got: String },
/// The handshake established, but the manager refused the registration:
/// a connection to this peer already exists. The loser has been closed.
Duplicate,
}
impl DialError {
/// A short static label per kind, for the `smarm-trace` `ClusterDial`
/// event; the payload (io error, names) is not carried.
pub fn label(&self) -> &'static str {
match self {
DialError::AlreadyDialing => "already_dialing",
DialError::ManagerUnavailable => "manager_unavailable",
DialError::Connect(_) => "connect",
DialError::Handshake(_) => "handshake",
DialError::PeerNameMismatch { .. } => "peer_name_mismatch",
DialError::Duplicate => "duplicate",
}
}
}
/// Dial `peer_name` at `addr` and run the handshake, keeping the manager's
/// dial-intent set honest around it: the intent is registered *before*
/// connecting (so a crossing inbound `Hello` sees it) and cleared the
/// moment the handshake resolves, before the connection actor is spawned.
/// Must run inside an actor. Retrying is the caller's business (c7's dial
/// loop); a lost tie-break surfaces as `Handshake(Closed)` — the peer's
/// accepted connection is already on its way.
pub fn dial(
transport: &dyn Transport,
addr: &str,
peer_name: &str,
local: &Local,
timing: Timing,
) -> Result<Pid, DialError> {
let me = scheduler::self_pid();
match gen_server::call(
MANAGER,
Call::DialBegin {
name: peer_name.to_string(),
pid: me,
},
) {
Ok(Reply::DialBegan(true)) => {}
Ok(Reply::DialBegan(false)) => return Err(DialError::AlreadyDialing),
_ => return Err(DialError::ManagerUnavailable),
}
let result = connect_and_shake(transport, addr, local, timing);
// Cleared immediately on outcome — a stale intent during the established
// window would corrupt later tie-breaks. Synchronous (a call): the
// intent is provably gone before anything else happens.
let _ = gen_server::call(
MANAGER,
Call::DialEnd {
name: peer_name.to_string(),
},
);
let (mut framed, peer) = result?;
if peer.node_name != peer_name {
framed.close();
return Err(DialError::PeerNameMismatch {
expected: peer_name.to_string(),
got: peer.node_name,
});
}
spawn_established(framed, peer, timing).map_err(|_| DialError::Duplicate)
}
fn connect_and_shake(
transport: &dyn Transport,
addr: &str,
local: &Local,
timing: Timing,
) -> Result<(FramedConn, Peer), DialError> {
let conn = transport.dial(addr).map_err(DialError::Connect)?;
let mut framed = FramedConn::new(conn);
let deadline = Instant::now() + timing.handshake_timeout;
let peer = dial_handshake(&mut framed, local, deadline).map_err(DialError::Handshake)?;
Ok((framed, peer))
}
/// A running acceptor. [`shutdown`](AcceptorHandle::shutdown) (or dropping
/// the last handle) stops the accept loop only: connections it established
/// belong to the [`manager`](crate::cluster::manager) and keep running, to
/// be torn down through the table (`Disconnect`, a peer close, or manager
/// shutdown).
pub struct AcceptorHandle {
cmd_tx: Sender<()>,
addr: String,
}
impl AcceptorHandle {
/// Ask the acceptor to stop. Idempotent; a no-op if it already has.
pub fn shutdown(&self) {
let _ = self.cmd_tx.send(());
}
/// The concrete bound address, dialable as-is.
pub fn local_addr(&self) -> &str {
&self.addr
}
}
/// Spawn the acceptor actor over a bound listener. Each inbound connection
/// is handshaken **inline in the loop** (a deliberate serialization: the
/// per-frame deadline bounds how long any one peer can hold the line, and
/// nothing concurrent exists to be starved before c7). The listener must be
/// fd-backed ([`Listener::readable_arm`]); the loopback listener is not,
/// and its acceptor exits immediately — loopback handshakes are driven
/// synchronously through the path fns instead, per D8.
pub fn spawn_acceptor(listener: Box<dyn Listener>, local: Local, timing: Timing) -> AcceptorHandle {
let addr = listener.local_addr();
let (cmd_tx, cmd_rx) = channel();
spawn(move || accept_loop(listener, local, timing, cmd_rx));
AcceptorHandle { cmd_tx, addr }
}
fn accept_loop(
mut listener: Box<dyn Listener>,
local: Local,
timing: Timing,
cmd_rx: Receiver<()>,
) {
loop {
let Some(arm) = listener.readable_arm() else {
return;
};
let arms: [&dyn Selectable; 2] = [&cmd_rx, &arm];
match crate::channel::try_select(&arms) {
Ok(0) => match cmd_rx.try_recv() {
Ok(Some(())) => return,
Ok(None) => continue, // spurious wake
Err(_) => return, // all handles dropped
},
Ok(_) => {
// The listener is readable: accept completes without parking.
let conn = match listener.accept() {
Ok(conn) => conn,
Err(_) => return, // listener itself is broken
};
handle_inbound(FramedConn::new(conn), &local, timing);
}
Err(_) => return, // fd arm failed to register: listener is gone
}
}
}
/// Run the accept-side handshake for one inbound connection, asking the
/// manager for the [`PeerStanding`], and hand the established connection to the
/// manager. Every failure was already resolved on the path (reject sent /
/// closed, or the registration refused and the actor stopped), so there is
/// nothing for the acceptor to carry forward.
fn handle_inbound(mut framed: FramedConn, local: &Local, timing: Timing) {
let deadline = Instant::now() + timing.handshake_timeout;
let standing_of = |name: &str| match gen_server::call(
MANAGER,
Call::Standing {
peer_name: name.to_string(),
},
) {
Ok(Reply::Standing(s)) => s,
// Manager unreachable: nobody could register this connection anyway,
// so claim the name taken and reject rather than accept an orphan.
_ => PeerStanding::Claimed,
};
if let Ok(peer) = accept_handshake(&mut framed, local.clone(), standing_of, deadline) {
let _ = spawn_established(framed, peer, timing);
}
}
+316
View File
@@ -0,0 +1,316 @@
//! RFC 010 c7b — the connector: the dial loop that turns discovered
//! candidates into a full mesh.
//!
//! A plain select-loop actor (the c6 shape). It spawns its [`Strategy`] as a
//! child actor and receives [`Discovery`] events from it; it tracks which
//! peers are up by **subscribing to membership like any other consumer** —
//! no privileged channel into the manager, the same snapshot-then-stream
//! surface c8 will use. One `select` folds the command inbox, the discovery
//! stream, the membership stream, and the earliest retry deadline into a
//! single wait.
//!
//! Per-candidate state: dial on arrival; on failure retry with capped
//! exponential backoff ([`INITIAL_BACKOFF`] doubling to [`MAX_BACKOFF`]);
//! on the peer's `node_up` stop dialing and reset the backoff; on its
//! `node_down` resume immediately (a fresh sequence — the reconnect case is
//! the one backoff exists to pace, but the *first* retry after a death
//! should be prompt). A candidate bearing our own name is parked permanently
//! — that seed is us; so is one whose address answers as a different name
//! (`PeerNameMismatch`: a misconfigured or stale seed — each retry would
//! only blip the peer's membership). Every other failure retries: in
//! particular a `NameTaken` reject can be our own ghost at the peer, not
//! yet reaped by its liveness timer, so it must not park. Each attempt's
//! outcome is one `smarm-trace` `ClusterDial` event. A [`Discovery::Withdrawn`]
//! drops its `(name, addr)` from the dial set — only that: a live
//! connection is membership's, and a re-announce re-adds it fresh.
//!
//! Dials run **inline in the loop** — the same deliberate serialization as
//! the acceptor (c6b): each attempt is bounded by the connect + handshake
//! deadlines, and nothing concurrent exists to be starved. A wall of slow
//! unreachable seeds would stretch the loop's latency; revisit if a real
//! deployment ever hits that shape.
use std::collections::HashSet;
use std::time::{Duration, Instant};
use crate::channel::{channel, select, select_timeout, Receiver, Selectable, Sender};
use crate::cluster::connect::{dial, DialError};
use crate::cluster::discovery::{Discovery, Strategy};
use crate::cluster::handshake::Local;
use crate::cluster::membership::{subscribe, NodeEvent};
use crate::cluster::transport::Transport;
use crate::cluster::Timing;
use crate::scheduler::spawn;
/// Default for [`Timing::initial_backoff`]: first retry delay after a failed
/// dial attempt.
pub const INITIAL_BACKOFF: Duration = Duration::from_millis(250);
/// Default for [`Timing::max_backoff`]: an unreachable seed is retried this
/// often, forever.
pub const MAX_BACKOFF: Duration = Duration::from_secs(5);
enum Cmd {
Shutdown,
}
/// A running connector. `shutdown` (or dropping the last handle) stops the
/// dial loop and its strategy only — established connections belong to the
/// manager, exactly as with the acceptor.
pub struct ConnectorHandle {
cmd_tx: Sender<Cmd>,
}
impl ConnectorHandle {
/// Ask the connector to stop. Idempotent; a no-op if it already has.
pub fn shutdown(&self) {
let _ = self.cmd_tx.send(Cmd::Shutdown);
}
}
/// One discovered `(name, addr)` and our dial intent towards it.
struct Candidate {
name: String,
addr: String,
state: State,
}
/// The connector's *intent* for a candidate. Whether the peer is currently
/// up is a separate, name-keyed membership fact (`up` in [`run`]): a
/// candidate can arrive after its peer's `node_up` (the snapshot lands
/// before the strategy has said anything), so "up" cannot live on the
/// candidate alone — it is a filter over dialing, not a candidate state.
enum State {
/// Never dialed: this seed is the local node itself, or the address
/// answered as a *different* name than the one seeded
/// (`DialError::PeerNameMismatch` — a misconfigured or stale seed;
/// redialing would only blip the peer's membership forever). The way
/// back is the strategy's: `Withdrawn` then a fresh `Candidate`.
Parked,
/// Dial when due; on failure, back off.
Dialing {
/// Delay to apply after the *next* failure.
backoff: Duration,
next_attempt: Instant,
},
}
impl State {
fn fresh(timing: &Timing) -> Self {
State::Dialing {
backoff: timing.initial_backoff,
next_attempt: Instant::now(),
}
}
}
impl Candidate {
/// The retry deadline, if this candidate is dialing at all.
fn due(&self) -> Option<Instant> {
match self.state {
State::Parked => None,
State::Dialing { next_attempt, .. } => Some(next_attempt),
}
}
/// A dial attempt was made: schedule the retry, grow the backoff.
fn attempted(&mut self, timing: &Timing) {
if let State::Dialing {
backoff,
next_attempt,
} = &mut self.state
{
*next_attempt = Instant::now() + *backoff;
*backoff = (*backoff * 2).min(timing.max_backoff);
}
}
/// The peer came up: the next sequence (after a later `node_down`)
/// starts from the initial delay again.
fn peer_up(&mut self, timing: &Timing) {
if let State::Dialing { backoff, .. } = &mut self.state {
*backoff = timing.initial_backoff;
}
}
/// The peer went down: redial promptly, fresh sequence.
fn peer_down(&mut self, timing: &Timing) {
if matches!(self.state, State::Dialing { .. }) {
self.state = State::fresh(timing);
}
}
}
/// Spawn the connector actor. The strategy is spawned as its child; the
/// membership subscription is taken inside the actor. Must be called from
/// inside an actor (the same requirement as `dial`).
pub fn spawn_connector(
transport: Box<dyn Transport>,
local: Local,
strategy: Box<dyn Strategy>,
timing: Timing,
) -> ConnectorHandle {
let (cmd_tx, cmd_rx) = channel();
spawn(move || run(transport, local, strategy, timing, cmd_rx));
ConnectorHandle { cmd_tx }
}
fn run(
transport: Box<dyn Transport>,
local: Local,
strategy: Box<dyn Strategy>,
timing: Timing,
cmd_rx: Receiver<Cmd>,
) {
// Membership is the connector's source of truth for "who is up" — the
// snapshot seeds `up` before any candidate arrives.
let Some(events) = subscribe() else {
return; // no manager, no cluster to connect
};
let (disc_tx, disc_rx) = channel();
spawn(move || strategy.run(disc_tx));
let mut cands: Vec<Candidate> = Vec::new();
let mut up: HashSet<String> = HashSet::new();
let mut strategy_done = false;
loop {
// Drain every input, then act. Order does not matter: acting is
// idempotent against the resulting state.
match drain_cmd(&cmd_rx) {
Drained::Stop => return,
Drained::Open => {}
}
if !strategy_done {
strategy_done = drain_discoveries(&disc_rx, &local, &timing, &mut cands);
}
match drain_events(&events.rx, &timing, &mut up, &mut cands) {
Drained::Stop => return, // manager gone: the cluster is tearing down
Drained::Open => {}
}
// Dial everything due, inline (see the module docs on serialization).
let now = Instant::now();
for c in cands
.iter_mut()
.filter(|c| !up.contains(&c.name) && c.due().is_some_and(|d| d <= now))
{
// On success the manager's node_up is on its way and lands in
// `up` (backing off meanwhile keeps a racing re-attempt from
// spinning); every failure retries — see the module docs —
// except a peer-name mismatch, which parks the candidate.
let outcome = dial(&*transport, &c.addr, &c.name, &local, timing);
note_dial(&outcome);
match outcome {
Err(DialError::PeerNameMismatch { .. }) => c.state = State::Parked,
_ => c.attempted(&timing),
}
}
// Wait: until the earliest retry deadline among actionable
// candidates, or indefinitely if none is pending.
let deadline = cands
.iter()
.filter(|c| !up.contains(&c.name))
.filter_map(Candidate::due)
.min();
let mut arms: Vec<&dyn Selectable> = vec![&cmd_rx, &events.rx];
if !strategy_done {
arms.push(&disc_rx);
}
match deadline {
Some(d) => {
let wait = d.saturating_duration_since(Instant::now());
let _ = select_timeout(&arms, wait);
}
None => {
let _ = select(&arms);
}
}
}
}
enum Drained {
Open,
Stop,
}
fn drain_cmd(rx: &Receiver<Cmd>) -> Drained {
match rx.try_recv() {
Ok(Some(Cmd::Shutdown)) => Drained::Stop,
Ok(None) => Drained::Open,
Err(_) => Drained::Stop, // all handles dropped
}
}
/// Pull every pending discovery into the candidate set (deduplicated by
/// `(name, addr)`; a candidate bearing the local name is parked; a
/// `Withdrawn` removes its pair from the dial set and nothing else — see
/// [`Discovery::Withdrawn`]). Returns `true` once the strategy's channel
/// closes — it has said all it will.
fn drain_discoveries(
rx: &Receiver<Discovery>,
local: &Local,
timing: &Timing,
cands: &mut Vec<Candidate>,
) -> bool {
loop {
match rx.try_recv() {
Ok(Some(Discovery::Withdrawn { name, addr })) => {
cands.retain(|c| !(c.name == name && c.addr == addr));
}
Ok(Some(Discovery::Candidate { name, addr })) => {
if cands.iter().any(|c| c.name == name && c.addr == addr) {
continue;
}
let state = if name == local.node_name {
State::Parked
} else {
State::fresh(timing)
};
cands.push(Candidate { name, addr, state });
}
Ok(None) => return false,
Err(_) => return true, // strategy done; its candidates live on here
}
}
}
/// Fold pending membership events into `up`; each is also a transition on
/// that peer's candidates (see [`Candidate::peer_up`] / [`peer_down`]).
///
/// [`peer_down`]: Candidate::peer_down
fn drain_events(
rx: &Receiver<NodeEvent>,
timing: &Timing,
up: &mut HashSet<String>,
cands: &mut [Candidate],
) -> Drained {
loop {
match rx.try_recv() {
Ok(Some(NodeEvent::NodeUp(info))) => {
cands
.iter_mut()
.filter(|c| c.name == info.name)
.for_each(|c| c.peer_up(timing));
up.insert(info.name);
}
Ok(Some(NodeEvent::NodeDown(info))) => {
up.remove(&info.name);
cands
.iter_mut()
.filter(|c| c.name == info.name)
.for_each(|c| c.peer_down(timing));
}
Ok(None) => return Drained::Open,
Err(_) => return Drained::Stop,
}
}
}
/// Surface a dial outcome: one `smarm-trace` event, nothing else. The
/// connector's bookkeeping is decided by the caller.
fn note_dial(outcome: &Result<crate::pid::Pid, DialError>) {
#[cfg(feature = "smarm-trace")]
crate::te!(crate::trace::Event::ClusterDial(
outcome.as_ref().map_or_else(DialError::label, |_| "ok")
));
#[cfg(not(feature = "smarm-trace"))]
let _ = outcome;
}
+76
View File
@@ -0,0 +1,76 @@
//! RFC 010 c7b — peer discovery: the [`Strategy`] seam and the static-seeds
//! implementation.
//!
//! A strategy is **push-based and runs as its own actor**: the
//! [`connector`](crate::cluster::connector) spawns it with the sending end of
//! a channel, and the strategy emits [`Discovery`] events whenever it learns
//! something — once at startup for a static list, continuously for a future
//! mDNS/DNS strategy — for as long as it cares to run. Returning ends the
//! strategy actor; the candidates it pushed live on in the connector (the
//! connector owns all retry/backoff state, so a strategy never re-announces).
//!
//! A candidate is a **`(node_name, addr)` pair**, not a bare address: the
//! dial path and the D7 tie-break are keyed by peer *name* (the dial intent
//! must be registered before connecting so a crossing inbound `Hello` sees
//! it), so an anonymous dial would reintroduce exactly the
//! simultaneous-connect flap D7 exists to prevent. Discovery mechanisms know
//! names — that is what they discover.
use crate::channel::Sender;
/// A discovery event, as pushed by a [`Strategy`].
///
/// `Candidate` announces, `Withdrawn` retracts — the primitive pair. A
/// strategy that wants TTL semantics builds them on top (track its own
/// last-seen times, emit `Withdrawn` on expiry); the connector deliberately
/// has no clock of its own for candidates (D11: strategies never
/// re-announce, the connector owns retry). `#[non_exhaustive]` so more can
/// land without breaking strategies.
#[derive(Debug, Clone, PartialEq, Eq)]
#[non_exhaustive]
pub enum Discovery {
/// A peer worth dialing: its claimed node name and a dialable address.
Candidate { name: String, addr: String },
/// Stop dialing this `(name, addr)`. Dial-set only: a connection that
/// is already up is membership's business and is left alone; an
/// attempt in flight completes on its own; a later `Candidate` for the
/// same pair re-adds it with fresh backoff. Unknown pairs are ignored.
Withdrawn { name: String, addr: String },
}
/// A source of peers to dial. Implementations are spawned as actors by the
/// connector — see the module docs for the contract.
pub trait Strategy: Send + 'static {
/// Run the strategy: push [`Discovery`] events into `out` as they are
/// learned; return when done discovering (or when `out` reports closed —
/// the connector is gone). Runs inside an actor, so blocking
/// cooperatively is fine.
fn run(self: Box<Self>, out: Sender<Discovery>);
}
/// The static-seeds strategy: a fixed `(name, addr)` list, announced once.
#[derive(Debug, Clone, Default)]
pub struct StaticSeeds {
seeds: Vec<(String, String)>,
}
impl StaticSeeds {
pub fn new(seeds: impl IntoIterator<Item = (impl Into<String>, impl Into<String>)>) -> Self {
StaticSeeds {
seeds: seeds
.into_iter()
.map(|(n, a)| (n.into(), a.into()))
.collect(),
}
}
}
impl Strategy for StaticSeeds {
fn run(self: Box<Self>, out: Sender<Discovery>) {
for (name, addr) in self.seeds {
if out.send(Discovery::Candidate { name, addr }).is_err() {
return; // connector gone; nobody to discover for
}
}
}
}
+549
View File
@@ -0,0 +1,549 @@
//! RFC 010 c2 — the owned wire envelope.
//!
//! Every control-plane frame is `u32` little-endian length prefix (of tag +
//! body), `u8` tag, hand-encoded body. postcard appears in exactly one place:
//! the payload blob inside `Send`/`SendNamed`, via [`encode_payload`] /
//! [`decode_payload`] — the seam where a codec swap would land (RFC 010 §2).
//! Everything else is hand-rolled and wholly owned.
//!
//! Integers are little-endian. Strings are `u16` length + UTF-8 bytes.
//! Payload blobs are `u32` length + bytes. Enum-shaped fields
//! ([`RejectReason`], [`DownReason`]) are a single tag byte.
use crate::monitor::DownReason;
use crate::pg::Incarnation;
/// Wire protocol version, checked in the handshake (c5).
pub const PROTO_VERSION: u32 = 1;
/// Hard cap on the length prefix. The control plane never carries bulk data
/// (RFC 010 §5 — that is the jarred rkyv plane), so anything larger is
/// corruption or an attack, not a legitimate frame.
pub const MAX_FRAME_LEN: usize = 16 * 1024 * 1024;
/// Per-node metadata exchanged in the handshake (RFC 010 §1: not identity).
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct NodeMeta {
pub role: String,
pub region: String,
}
/// Why a `Hello` was rejected.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum RejectReason {
/// Build hashes differ — not the same binary.
HashMismatch,
/// The offered node name is already claimed by a live peer.
NameTaken,
/// Wire protocol version mismatch.
ProtoVersion,
}
/// The control-plane frame inventory (RFC 010, *Implementation details*).
#[derive(Debug, Clone, PartialEq, Eq)]
pub enum Frame {
Hello {
proto_version: u32,
build_hash: u64,
node_name: String,
incarnation: Incarnation,
meta: NodeMeta,
},
HelloAck {
node_name: String,
incarnation: Incarnation,
meta: NodeMeta,
},
HelloReject {
reason: RejectReason,
},
Heartbeat,
Send {
/// Target slot index (node is implicit in the connection, incarnation
/// is bound at handshake — RFC 010 §3).
index: u32,
generation: u32,
type_hash: u64,
payload: Vec<u8>,
},
SendNamed {
name: String,
type_hash: u64,
payload: Vec<u8>,
},
Monitor {
monitor_id: u64,
index: u32,
generation: u32,
},
Demonitor {
monitor_id: u64,
},
Down {
monitor_id: u64,
reason: RemoteDownReason,
},
}
/// Why a remotely-monitored actor is reported down: either the target's own
/// terminal [`DownReason`] as its node recorded it, or the *link* to that
/// node was lost (or absent) — which says nothing about the actor itself.
///
/// This is the cluster-side widening of `DownReason` (p5): `Disconnected`
/// is a fact about a connection, never about a local actor, so it lives
/// here rather than in the core enum — a local `Down` can never carry it,
/// and matches on `DownReason` stay exhaustive over actor outcomes only.
/// On the wire `Local(r)` uses `r`'s tag and `Disconnected` is tag 5,
/// bound since c11; no peer emits it today (a lost link is synthesized
/// locally), but the codec honours it both ways.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum RemoteDownReason {
/// The target itself terminated; the peer reported this reason.
Local(DownReason),
/// The link to the target's node was lost or was never up.
Disconnected,
}
impl RemoteDownReason {
/// The actor's own reason, if this was not a link loss.
pub fn local(self) -> Option<DownReason> {
match self {
RemoteDownReason::Local(r) => Some(r),
RemoteDownReason::Disconnected => None,
}
}
}
impl From<DownReason> for RemoteDownReason {
fn from(r: DownReason) -> Self {
RemoteDownReason::Local(r)
}
}
// Frame tags. 0 is deliberately unassigned so an all-zero buffer never parses.
const TAG_HELLO: u8 = 1;
const TAG_HELLO_ACK: u8 = 2;
const TAG_HELLO_REJECT: u8 = 3;
const TAG_HEARTBEAT: u8 = 4;
const TAG_SEND: u8 = 5;
const TAG_SEND_NAMED: u8 = 6;
const TAG_MONITOR: u8 = 7;
const TAG_DEMONITOR: u8 = 8;
const TAG_DOWN: u8 = 9;
// RejectReason tags.
const REJ_HASH_MISMATCH: u8 = 1;
const REJ_NAME_TAKEN: u8 = 2;
const REJ_PROTO_VERSION: u8 = 3;
// DownReason tags. Do not reuse tags.
const DR_EXIT: u8 = 1;
const DR_PANIC: u8 = 2;
const DR_STOPPED: u8 = 3;
const DR_NOPROC: u8 = 4;
const DR_DISCONNECTED: u8 = 5;
// `Shutdown` never rides in a `Down` by contract (a target that honours the
// request exits normally) — the tag exists so the codec stays total.
const DR_SHUTDOWN: u8 = 6;
/// Frame could not be encoded. The output buffer is left exactly as it was.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum EncodeError {
/// tag + body exceed [`MAX_FRAME_LEN`].
FrameTooLarge { len: usize },
/// A string field exceeds `u16::MAX` bytes.
StringTooLong { len: usize },
}
impl core::fmt::Display for EncodeError {
fn fmt(&self, f: &mut core::fmt::Formatter<'_>) -> core::fmt::Result {
match self {
Self::FrameTooLarge { len } => {
write!(f, "frame body of {len} bytes exceeds MAX_FRAME_LEN")
}
Self::StringTooLong { len } => {
write!(f, "string field of {len} bytes exceeds u16::MAX")
}
}
}
}
impl std::error::Error for EncodeError {}
/// Frame could not be decoded. Everything here is *corruption* — "not enough
/// bytes yet" is the `Ok(None)` streaming case, never an error.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum DecodeError {
/// The length prefix exceeds [`MAX_FRAME_LEN`].
FrameTooLarge { declared: usize },
/// The length prefix is zero — there is no tag byte.
EmptyFrame,
/// Unknown frame tag.
UnknownTag(u8),
/// Unknown tag for an enum-shaped field.
UnknownEnumTag { what: &'static str, tag: u8 },
/// A field ran past the declared frame end (the length prefix lied long,
/// or a length-carrying field inside the body lied).
Truncated,
/// Bytes were left over after the body (the length prefix lied short).
Trailing { extra: usize },
/// A string field was not valid UTF-8.
Utf8,
}
impl core::fmt::Display for DecodeError {
fn fmt(&self, f: &mut core::fmt::Formatter<'_>) -> core::fmt::Result {
match self {
Self::FrameTooLarge { declared } => {
write!(f, "declared frame length {declared} exceeds MAX_FRAME_LEN")
}
Self::EmptyFrame => write!(f, "zero-length frame (no tag byte)"),
Self::UnknownTag(t) => write!(f, "unknown frame tag {t}"),
Self::UnknownEnumTag { what, tag } => write!(f, "unknown {what} tag {tag}"),
Self::Truncated => write!(f, "frame body truncated mid-field"),
Self::Trailing { extra } => write!(f, "{extra} trailing bytes after frame body"),
Self::Utf8 => write!(f, "string field is not valid UTF-8"),
}
}
}
impl std::error::Error for DecodeError {}
impl Frame {
/// Append this frame, length-prefixed, to `out`.
///
/// On error `out` is left untouched.
pub fn encode(&self, out: &mut Vec<u8>) -> Result<(), EncodeError> {
let start = out.len();
out.extend_from_slice(&[0u8; 4]); // length placeholder, patched below
let result = self.encode_body(out);
match result {
Ok(()) => {
let frame_len = out.len() - start - 4;
if frame_len > MAX_FRAME_LEN {
out.truncate(start);
return Err(EncodeError::FrameTooLarge { len: frame_len });
}
// Cast is lossless: MAX_FRAME_LEN < u32::MAX, checked above.
let len32 = frame_len as u32;
out[start..start + 4].copy_from_slice(&len32.to_le_bytes());
Ok(())
}
Err(e) => {
out.truncate(start);
Err(e)
}
}
}
fn encode_body(&self, out: &mut Vec<u8>) -> Result<(), EncodeError> {
match self {
Frame::Hello {
proto_version,
build_hash,
node_name,
incarnation,
meta,
} => {
out.push(TAG_HELLO);
put_u32(out, *proto_version);
put_u64(out, *build_hash);
put_str(out, node_name)?;
put_u32(out, incarnation.get());
put_meta(out, meta)?;
}
Frame::HelloAck {
node_name,
incarnation,
meta,
} => {
out.push(TAG_HELLO_ACK);
put_str(out, node_name)?;
put_u32(out, incarnation.get());
put_meta(out, meta)?;
}
Frame::HelloReject { reason } => {
out.push(TAG_HELLO_REJECT);
out.push(match reason {
RejectReason::HashMismatch => REJ_HASH_MISMATCH,
RejectReason::NameTaken => REJ_NAME_TAKEN,
RejectReason::ProtoVersion => REJ_PROTO_VERSION,
});
}
Frame::Heartbeat => out.push(TAG_HEARTBEAT),
Frame::Send {
index,
generation,
type_hash,
payload,
} => {
out.push(TAG_SEND);
put_u32(out, *index);
put_u32(out, *generation);
put_u64(out, *type_hash);
put_blob(out, payload)?;
}
Frame::SendNamed {
name,
type_hash,
payload,
} => {
out.push(TAG_SEND_NAMED);
put_str(out, name)?;
put_u64(out, *type_hash);
put_blob(out, payload)?;
}
Frame::Monitor {
monitor_id,
index,
generation,
} => {
out.push(TAG_MONITOR);
put_u64(out, *monitor_id);
put_u32(out, *index);
put_u32(out, *generation);
}
Frame::Demonitor { monitor_id } => {
out.push(TAG_DEMONITOR);
put_u64(out, *monitor_id);
}
Frame::Down { monitor_id, reason } => {
out.push(TAG_DOWN);
put_u64(out, *monitor_id);
out.push(match reason {
RemoteDownReason::Local(DownReason::Exit) => DR_EXIT,
RemoteDownReason::Local(DownReason::Panic) => DR_PANIC,
RemoteDownReason::Local(DownReason::Stopped) => DR_STOPPED,
RemoteDownReason::Local(DownReason::NoProc) => DR_NOPROC,
RemoteDownReason::Local(DownReason::Shutdown) => DR_SHUTDOWN,
RemoteDownReason::Disconnected => DR_DISCONNECTED,
});
}
}
Ok(())
}
/// Try to decode one frame from the start of `buf`.
///
/// `Ok(Some((frame, consumed)))` — a full frame; the caller advances by
/// `consumed`. `Ok(None)` — not enough bytes yet (streaming); read more
/// and retry. `Err(_)` — the bytes are corrupt; the connection is dead.
pub fn decode(buf: &[u8]) -> Result<Option<(Frame, usize)>, DecodeError> {
let Some(prefix) = buf.get(0..4) else {
return Ok(None);
};
let mut len4 = [0u8; 4];
len4.copy_from_slice(prefix);
let declared = u32::from_le_bytes(len4) as usize;
if declared > MAX_FRAME_LEN {
return Err(DecodeError::FrameTooLarge { declared });
}
if declared == 0 {
return Err(DecodeError::EmptyFrame);
}
let Some(body) = buf.get(4..4 + declared) else {
return Ok(None);
};
let mut r = Reader { buf: body, pos: 0 };
let frame = Self::decode_body(&mut r)?;
if r.pos != body.len() {
return Err(DecodeError::Trailing {
extra: body.len() - r.pos,
});
}
Ok(Some((frame, 4 + declared)))
}
fn decode_body(r: &mut Reader<'_>) -> Result<Frame, DecodeError> {
let tag = r.u8()?;
let frame = match tag {
TAG_HELLO => Frame::Hello {
proto_version: r.u32()?,
build_hash: r.u64()?,
node_name: r.string()?,
incarnation: Incarnation::new(r.u32()?),
meta: r.meta()?,
},
TAG_HELLO_ACK => Frame::HelloAck {
node_name: r.string()?,
incarnation: Incarnation::new(r.u32()?),
meta: r.meta()?,
},
TAG_HELLO_REJECT => Frame::HelloReject {
reason: match r.u8()? {
REJ_HASH_MISMATCH => RejectReason::HashMismatch,
REJ_NAME_TAKEN => RejectReason::NameTaken,
REJ_PROTO_VERSION => RejectReason::ProtoVersion,
t => {
return Err(DecodeError::UnknownEnumTag {
what: "RejectReason",
tag: t,
})
}
},
},
TAG_HEARTBEAT => Frame::Heartbeat,
TAG_SEND => Frame::Send {
index: r.u32()?,
generation: r.u32()?,
type_hash: r.u64()?,
payload: r.blob()?,
},
TAG_SEND_NAMED => Frame::SendNamed {
name: r.string()?,
type_hash: r.u64()?,
payload: r.blob()?,
},
TAG_MONITOR => Frame::Monitor {
monitor_id: r.u64()?,
index: r.u32()?,
generation: r.u32()?,
},
TAG_DEMONITOR => Frame::Demonitor {
monitor_id: r.u64()?,
},
TAG_DOWN => Frame::Down {
monitor_id: r.u64()?,
reason: match r.u8()? {
DR_EXIT => RemoteDownReason::Local(DownReason::Exit),
DR_PANIC => RemoteDownReason::Local(DownReason::Panic),
DR_STOPPED => RemoteDownReason::Local(DownReason::Stopped),
DR_NOPROC => RemoteDownReason::Local(DownReason::NoProc),
DR_SHUTDOWN => RemoteDownReason::Local(DownReason::Shutdown),
DR_DISCONNECTED => RemoteDownReason::Disconnected,
t => {
return Err(DecodeError::UnknownEnumTag {
what: "RemoteDownReason",
tag: t,
})
}
},
},
t => return Err(DecodeError::UnknownTag(t)),
};
Ok(frame)
}
}
// ---------------------------------------------------------------------------
// Body writers
// ---------------------------------------------------------------------------
fn put_u32(out: &mut Vec<u8>, v: u32) {
out.extend_from_slice(&v.to_le_bytes());
}
fn put_u64(out: &mut Vec<u8>, v: u64) {
out.extend_from_slice(&v.to_le_bytes());
}
fn put_str(out: &mut Vec<u8>, s: &str) -> Result<(), EncodeError> {
let Ok(len) = u16::try_from(s.len()) else {
return Err(EncodeError::StringTooLong { len: s.len() });
};
out.extend_from_slice(&len.to_le_bytes());
out.extend_from_slice(s.as_bytes());
Ok(())
}
fn put_blob(out: &mut Vec<u8>, b: &[u8]) -> Result<(), EncodeError> {
let Ok(len) = u32::try_from(b.len()) else {
return Err(EncodeError::FrameTooLarge { len: b.len() });
};
out.extend_from_slice(&len.to_le_bytes());
out.extend_from_slice(b);
Ok(())
}
fn put_meta(out: &mut Vec<u8>, m: &NodeMeta) -> Result<(), EncodeError> {
put_str(out, &m.role)?;
put_str(out, &m.region)
}
// ---------------------------------------------------------------------------
// Body reader
// ---------------------------------------------------------------------------
struct Reader<'a> {
buf: &'a [u8],
pos: usize,
}
impl Reader<'_> {
fn take(&mut self, n: usize) -> Result<&[u8], DecodeError> {
let end = self.pos.checked_add(n).ok_or(DecodeError::Truncated)?;
let s = self.buf.get(self.pos..end).ok_or(DecodeError::Truncated)?;
self.pos = end;
Ok(s)
}
fn u8(&mut self) -> Result<u8, DecodeError> {
Ok(self.take(1)?[0])
}
fn u16(&mut self) -> Result<u16, DecodeError> {
let mut b = [0u8; 2];
b.copy_from_slice(self.take(2)?);
Ok(u16::from_le_bytes(b))
}
fn u32(&mut self) -> Result<u32, DecodeError> {
let mut b = [0u8; 4];
b.copy_from_slice(self.take(4)?);
Ok(u32::from_le_bytes(b))
}
fn u64(&mut self) -> Result<u64, DecodeError> {
let mut b = [0u8; 8];
b.copy_from_slice(self.take(8)?);
Ok(u64::from_le_bytes(b))
}
fn string(&mut self) -> Result<String, DecodeError> {
let len = self.u16()? as usize;
let bytes = self.take(len)?;
match core::str::from_utf8(bytes) {
Ok(s) => Ok(s.to_owned()),
Err(_) => Err(DecodeError::Utf8),
}
}
fn blob(&mut self) -> Result<Vec<u8>, DecodeError> {
let len = self.u32()? as usize;
Ok(self.take(len)?.to_vec())
}
fn meta(&mut self) -> Result<NodeMeta, DecodeError> {
Ok(NodeMeta {
role: self.string()?,
region: self.string()?,
})
}
}
// ---------------------------------------------------------------------------
// The postcard seam (RFC 010 §2) — the ONLY place payload bytes are produced
// or consumed. A codec swap lands here and nowhere else.
// ---------------------------------------------------------------------------
/// Payload (de)serialization failed at the codec seam.
#[derive(Debug)]
pub struct PayloadError(String);
impl core::fmt::Display for PayloadError {
fn fmt(&self, f: &mut core::fmt::Formatter<'_>) -> core::fmt::Result {
write!(f, "payload codec: {}", self.0)
}
}
impl std::error::Error for PayloadError {}
/// Serialize a payload value to the wire blob.
pub fn encode_payload<T: serde::Serialize + ?Sized>(value: &T) -> Result<Vec<u8>, PayloadError> {
postcard::to_allocvec(value).map_err(|e| PayloadError(e.to_string()))
}
/// Deserialize a payload value from the wire blob.
pub fn decode_payload<T: serde::de::DeserializeOwned>(bytes: &[u8]) -> Result<T, PayloadError> {
postcard::from_bytes(bytes).map_err(|e| PayloadError(e.to_string()))
}
+238
View File
@@ -0,0 +1,238 @@
//! RFC 010 c8 — explicit exposure: the node's remote surface, and the
//! fixed-seed type hash.
//!
//! Nothing local is remotely reachable by default (RFC §4 — "a gun needs a
//! safety"). [`expose`] marks a registered name remotely addressable and
//! registers `M`'s decoder under [`type_hash::<M>()`](type_hash);
//! [`expose_type`] registers only the decoder (the reply-to path: a
//! `RemotePid<A>` received in a message is sendable only if `A::Msg`'s
//! decoder was explicitly registered). The exposed set is the node's
//! visible, auditable remote surface ([`exposed_names`]).
//!
//! ## Where the state lives
//!
//! On `RuntimeInner`, the [`pg`](crate::pg) pattern: a leaf-locked table,
//! cfg-gated behind the `cluster` feature (zero-cost-when-off, per c1).
//! Chosen over manager-held state because c9's inbound decode consults it
//! per frame — a hot path that must not serialize every remote delivery
//! through one gen_server. The state resets with the runtime, like every
//! registry.
//!
//! ## The watchable fold (D3), against the code as it stands
//!
//! RFC §4: the exposed set is not a new registry — it folds into the
//! existing `watchable` machinery, one set, two set-sites (a pid crossing
//! the membrane, and expose). Reading the code: `register` **already
//! stamps every named holder watchable** ("no successfully-registered actor
//! can die unflagged", registry.rs), so an exposed *name*'s holder needs no
//! extra mark here — the guarantee holds by registration, and re-registration
//! after a holder's death re-stamps the new holder for free (a per-tenancy
//! mark taken at expose time could not do that). The cluster's own
//! `mark_watchable` set-site is therefore the **pid crossing the wire** —
//! serialization of a pid into a frame, c10 — the exact analog of the
//! membrane crossing. What lives here is only the name/type-level state
//! neither the registry nor the slot bits can carry: which names are
//! exposed, and how to decode each type hash.
//!
//! ## The hash
//!
//! [`type_hash`] is FNV-1a 64 (fixed seed: the FNV offset basis) over
//! `TypeId`, so it is a constant of the binary: stable across runs of the
//! same build — exactly the scope the build-hash handshake reduces the mesh
//! to — and deliberately *not* stable across builds (scope guard: no
//! cross-version wire compatibility). A collision between two exposed types
//! degrades to a decode error or a refused channel, never a misroute — the
//! local `SendError::NoChannel` guarantee survives the network (RFC §3).
//!
//! ## The decoder contract
//!
//! A decoder is **decode-and-deliver-to-pid**: it captures `M` (the one
//! typed site), decodes the payload, and hands the value to the target's
//! published channel via the registry's own dynamic send. Wire-name →
//! local-pid resolution deliberately stays *outside* — that is c9's single
//! resolution seam, and it calls [`decode_deliver`].
use std::any::TypeId;
use std::collections::HashMap;
use std::hash::{Hash, Hasher};
use crate::cluster::envelope::{decode_payload, PayloadError};
use crate::pid::{Name, Pid};
use crate::registry::{send_dyn, SendError};
use crate::scheduler::with_runtime;
/// The fixed-seed `TypeId` → `u64` hash: FNV-1a 64 over the `TypeId`'s hash
/// bytes, seeded with the FNV offset basis. A constant of the binary — see
/// the module docs for scope.
pub fn type_hash<M: 'static>() -> u64 {
let mut h = Fnv1a64::new();
TypeId::of::<M>().hash(&mut h);
h.finish()
}
/// FNV-1a 64 as a `Hasher`, so `TypeId` (opaque, `Hash`-only) can feed it.
/// Same constants as the const fns in [`crate::cluster`] (BUILD_HASH).
struct Fnv1a64(u64);
impl Fnv1a64 {
fn new() -> Self {
Fnv1a64(0xcbf2_9ce4_8422_2325)
}
}
impl Hasher for Fnv1a64 {
fn write(&mut self, bytes: &[u8]) {
for &b in bytes {
self.0 ^= b as u64;
self.0 = self.0.wrapping_mul(0x0000_0100_0000_01b3);
}
}
fn finish(&self) -> u64 {
self.0
}
}
/// Why a [`decode_deliver`] did not deliver. Payload-free mirror of the
/// registry's `SendError` where relevant — the caller (c9's inbound path)
/// has only bytes to give back, not a typed message.
#[derive(Debug)]
pub enum DeliverError {
/// No decoder is registered under this hash — the type was never
/// exposed here.
UnknownType,
/// The bytes did not decode as the registered type.
Decode(PayloadError),
/// The target actor is dead (or was never alive).
Dead,
/// The target is live but has no channel for this message type, or that
/// channel is closed — the `NoChannel` guarantee: a decoded value is
/// refused, never misrouted.
WrongChannel,
}
/// A registered decoder: decode `bytes` as the captured type and deliver to
/// `pid`'s published channel. `Arc`, so [`decode_deliver`] can clone it out
/// from under the exposure lock and call it lock-free — the decoder's
/// `send_dyn` takes the registry lock, and the two are mutual Leaves that
/// must never nest.
type Decoder = std::sync::Arc<dyn Fn(Pid, &[u8]) -> Result<(), DeliverError> + Send + Sync>;
/// The exposure state, one per runtime (a `RuntimeInner` field, pg-style).
pub(crate) struct ExposureState {
/// The exposed names: registry key → the type hash it expects.
exposed: HashMap<&'static str, u64>,
/// The decoders: type hash → decode-and-deliver.
decoders: HashMap<u64, Decoder>,
}
impl ExposureState {
pub(crate) fn new() -> Self {
ExposureState {
exposed: HashMap::new(),
decoders: HashMap::new(),
}
}
}
/// Mark `name` remotely addressable and register `M`'s decoder under its
/// type hash (so both name-sends and pid-sends of `M` work — RFC §4).
/// Returns the hash.
///
/// Exposure is a **name-level fact**, independent of who currently holds the
/// name (names late-bind: the registry re-resolves on every send, and c9's
/// seam resolves per delivery). Exposing an unregistered name is therefore
/// valid — deliveries fail with "unresolved" until someone registers it.
/// Idempotent. Must run inside [`run`](crate::run).
pub fn expose<M>(name: Name<M>) -> u64
where
M: serde::de::DeserializeOwned + Send + 'static,
{
let h = ensure_decoder::<M>();
with_runtime(|inner| {
inner.exposure.lock().exposed.insert(name.as_str(), h);
});
h
}
/// Register only `M`'s decoder (no name): the reply-to path. Returns the
/// hash. Idempotent. Must run inside [`run`](crate::run).
pub fn expose_type<M>() -> u64
where
M: serde::de::DeserializeOwned + Send + 'static,
{
ensure_decoder::<M>()
}
fn ensure_decoder<M>() -> u64
where
M: serde::de::DeserializeOwned + Send + 'static,
{
let h = type_hash::<M>();
with_runtime(|inner| {
inner
.exposure
.lock()
.decoders
.entry(h)
.or_insert_with(decoder::<M>);
});
h
}
/// The one typed site: decode as `M`, deliver via the registry's dynamic
/// send. See the module docs for the error mapping.
fn decoder<M>() -> Decoder
where
M: serde::de::DeserializeOwned + Send + 'static,
{
std::sync::Arc::new(|pid, bytes| {
let m: M = decode_payload(bytes).map_err(DeliverError::Decode)?;
send_dyn(pid, m).map_err(|e| match e {
SendError::Dead(_) | SendError::Unresolved(_) | SendError::NoMember(_) => {
DeliverError::Dead
}
SendError::NoChannel(_) | SendError::Closed(_) => DeliverError::WrongChannel,
})
})
}
/// The type hash `name` was exposed with, or `None` if it is not exposed.
/// Must run inside [`run`](crate::run).
pub fn exposed_hash(name: &str) -> Option<u64> {
with_runtime(|inner| inner.exposure.lock().exposed.get(name).copied())
}
/// Whether a decoder is registered under `hash`. Must run inside
/// [`run`](crate::run).
pub fn decoder_registered(hash: u64) -> bool {
with_runtime(|inner| inner.exposure.lock().decoders.contains_key(&hash))
}
/// Decode `bytes` under `hash`'s registered decoder and deliver to `pid`.
/// This is the delivery half c9's single resolution seam calls after it has
/// resolved a wire name to a local pid. Must run inside [`run`](crate::run).
pub fn decode_deliver(hash: u64, to: Pid, bytes: &[u8]) -> Result<(), DeliverError> {
// Clone the Arc under the lock, call outside it: the decoder's
// `send_dyn` takes the registry lock — a mutual Leaf with the exposure
// lock (the runtime asserts if Leaves nest). This also keeps unrelated
// deliveries uncoupled from a slow decode.
let d = with_runtime(|inner| inner.exposure.lock().decoders.get(&hash).cloned());
match d {
Some(d) => d(to, bytes),
None => Err(DeliverError::UnknownType),
}
}
/// The auditable remote surface: every exposed name and its type hash,
/// unordered. Must run inside [`run`](crate::run).
pub fn exposed_names() -> Vec<(&'static str, u64)> {
with_runtime(|inner| {
inner
.exposure
.lock()
.exposed
.iter()
.map(|(&n, &h)| (n, h))
.collect()
})
}
+178
View File
@@ -0,0 +1,178 @@
//! RFC 010 c5 — the handshake as a pure state machine.
//!
//! Frames in, actions out — no IO, no clocks, no actors. The c6 connection
//! actor drives these machines and executes their actions; everything
//! time-shaped (handshake deadline, heartbeats) lives there.
use crate::cluster::envelope::{Frame, NodeMeta, RejectReason, PROTO_VERSION};
use crate::pg::Incarnation;
/// This node's identity and metadata, as offered in (or checked against) a
/// `Hello`.
#[derive(Debug, Clone)]
pub struct Local {
pub node_name: String,
pub incarnation: Incarnation,
pub build_hash: u64,
pub meta: NodeMeta,
}
/// The peer identity a successful handshake yields (what c7 feeds `node_up`).
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct Peer {
pub node_name: String,
pub incarnation: Incarnation,
pub meta: NodeMeta,
}
/// Driver-supplied standing of the *offered* name at this node — knowledge
/// the pure machine cannot have (c6 owns the connection table and dial
/// set). One answer, in the responder's own precedence: an established
/// peer under that name outranks an in-flight dial to it.
#[derive(Debug, Clone, Copy, Default, PartialEq, Eq)]
pub enum PeerStanding {
/// Neither connected to nor dialing that name.
#[default]
Free,
/// An established peer already holds that name.
Claimed,
/// We have our own dial in flight to that name.
Dialing,
}
/// Simultaneous-connect tie-break: does the connection dialed by
/// `dialer_name` survive against the reverse dial?
/// The rule (ratified 2026-08-14, a wire-protocol fact): the connection
/// dialed by the lexicographically **smaller** name survives. Both ends know
/// both names, so both compute the same verdict — which is why the losing
/// side may close silently instead of sending a reject.
pub fn dial_wins(dialer_name: &str, acceptor_name: &str) -> bool {
dialer_name < acceptor_name
}
/// Dial side: emits `Hello` at construction, interprets the single response.
#[must_use]
#[derive(Debug)]
pub struct Initiator(());
/// What the dial side's response frame meant.
#[must_use]
#[derive(Debug, PartialEq, Eq)]
pub enum InitiatorOutcome {
Established(Peer),
Rejected(RejectReason),
/// Protocol violation before the ack — close. Carries the offending frame.
Failed(Frame),
}
impl Initiator {
/// Start a dial-side handshake: the returned frame is the `Hello` to
/// send; the returned machine is the right to interpret the response.
pub fn new(local: &Local) -> (Self, Frame) {
let hello = Frame::Hello {
proto_version: PROTO_VERSION,
build_hash: local.build_hash,
node_name: local.node_name.clone(),
incarnation: local.incarnation,
meta: local.meta.clone(),
};
(Initiator(()), hello)
}
/// Interpret the response. The `HelloAck` carries no hash or version —
/// the responder already checked ours against its own, and equality is
/// symmetric, so a one-sided check is sound.
pub fn on_frame(self, frame: Frame) -> InitiatorOutcome {
match frame {
Frame::HelloAck {
node_name,
incarnation,
meta,
} => InitiatorOutcome::Established(Peer {
node_name,
incarnation,
meta,
}),
Frame::HelloReject { reason } => InitiatorOutcome::Rejected(reason),
other => InitiatorOutcome::Failed(other),
}
}
}
/// Accept side: awaits exactly one `Hello`, answers or closes.
#[must_use]
#[derive(Debug)]
pub struct Responder {
local: Local,
}
/// What to do with an inbound connection's first frame.
#[must_use]
#[derive(Debug, PartialEq, Eq)]
pub enum ResponderOutcome {
/// Send the ack; the connection is established.
Accepted { reply: Frame, peer: Peer },
/// Send the reject, then close.
Rejected { reply: Frame, reason: RejectReason },
/// Lost the simultaneous-connect tie-break: close silently, no frame.
TieBreakLoss,
/// Protocol violation before Hello — close, no reply. Carries the frame.
Failed(Frame),
}
impl Responder {
pub fn new(local: Local) -> Self {
Responder { local }
}
/// Judge the connection's first frame. Check order is proto → hash →
/// name → tie-break: validity before identity. `HelloReject` is the
/// cross-version compatibility anchor, so a version-mismatched peer
/// still gets one.
pub fn on_frame(self, frame: Frame, standing: PeerStanding) -> ResponderOutcome {
let Frame::Hello {
proto_version,
build_hash,
node_name,
incarnation,
meta,
} = frame
else {
return ResponderOutcome::Failed(frame);
};
let reject = |reason| ResponderOutcome::Rejected {
reply: Frame::HelloReject { reason },
reason,
};
if proto_version != PROTO_VERSION {
return reject(RejectReason::ProtoVersion);
}
if build_hash != self.local.build_hash {
return reject(RejectReason::HashMismatch);
}
if node_name == self.local.node_name || standing == PeerStanding::Claimed {
return reject(RejectReason::NameTaken);
}
// Simultaneous connect: the inbound frame is the peer's dial. If our
// own in-flight dial wins instead, drop this one silently — the peer
// computes the same verdict (see `dial_wins`).
if standing == PeerStanding::Dialing && !dial_wins(&node_name, &self.local.node_name) {
return ResponderOutcome::TieBreakLoss;
}
ResponderOutcome::Accepted {
reply: Frame::HelloAck {
node_name: self.local.node_name,
incarnation: self.local.incarnation,
meta: self.local.meta,
},
peer: Peer {
node_name,
incarnation,
meta,
},
}
}
}
+295
View File
@@ -0,0 +1,295 @@
//! RFC 010 c6 — the cluster connection manager.
//!
//! One manager per runtime: the single registry of live peer connections and
//! the source of truth for whether a peer name is already claimed. The
//! accept/connect path registers each established connection here, handing
//! over its [`ConnHandle`] — **the manager owns connection lifetime**. A
//! connection lives as long as its table entry, so it neither outlives nor
//! dies with whichever actor happened to establish it. Registered actors are
//! also *monitored*, so the table self-heals on any exit path — a connection
//! that panics, is cancelled, or closes cleanly is removed without
//! cooperation from the dying actor.
//!
//! The manager also holds the **membership state** (c7a): `node_up` fires on
//! a successful registration and `node_down` on removal — they are derived
//! facts of the exact events this table already owns, so holding the view
//! here means no cross-actor race between "connection exists" and "node is
//! up". The consumer surface (event types, [`subscribe`], [`view`],
//! semantics) is [`membership`](crate::cluster::membership); no consumer
//! ever touches the table itself.
//!
//! The connector dial loop is c7b, built on top of both.
use std::collections::HashMap;
use crate::channel::Sender;
use crate::cluster::conn::ConnHandle;
use crate::cluster::handshake::{Peer, PeerStanding};
use crate::cluster::membership::{NodeEvent, NodeInfo};
use crate::cluster::remote::{bind_outbound, unbind_outbound};
use crate::gen_server::{GenServer, GenServerCtx, GenServerName, Watcher};
use crate::monitor::{monitor, Down};
use crate::pg::NodeId;
use crate::pid::Pid;
/// Well-known name of the singleton manager within a runtime. Connection
/// actors reach it by name rather than by a passed-around ref, so a restarted
/// manager is always found at the same key.
pub const MANAGER: GenServerName<Manager> = GenServerName::new("smarm.cluster.manager");
/// One live connection's entry: the actor running it, the handle whose
/// lifetime *is* the connection's (see the module docs), and the peer's
/// membership identity (what `node_up` announced and `node_down` will name).
struct ConnEntry {
pid: Pid,
info: NodeInfo,
_handle: ConnHandle,
}
/// The connection registry: peer name → the connection actor that owns that
/// peer's control connection. Plus the membership state layered on it (c7a):
/// subscribers, and the `(name, incarnation)` → [`NodeId`] memo.
pub struct Manager {
conns: HashMap<String, ConnEntry>,
/// In-flight dial intents: peer name -> the actor performing the dial.
/// Registered *before* connecting so a crossing inbound `Hello` sees it
/// ([`PeerStanding::Dialing`]); cleared the moment
/// the dial resolves, and — because the dialer is monitored — on the
/// dialer's death, so a panicking dial can never wedge the tie-break.
dials: HashMap<String, Pid>,
/// Membership subscribers; a closed channel is pruned on the next emit.
subscribers: Vec<Sender<NodeEvent>>,
/// The [`NodeId`] memo: a reconnect at the same incarnation keeps its id,
/// a restart (new incarnation) allocates a fresh one. Grows one entry per
/// distinct `(name, incarnation)` ever seen — unbounded in principle,
/// bounded in practice by restarts actually happening.
ids: HashMap<(String, u32), NodeId>,
/// Next id to allocate. Starts at 1: id 0 is
/// [`DEFAULT_NODE_ID`](crate::pg::DEFAULT_NODE_ID), the local node.
next_id: u32,
watcher: Option<Watcher<Manager>>,
}
impl Manager {
pub fn new() -> Self {
Manager {
conns: HashMap::new(),
dials: HashMap::new(),
subscribers: Vec::new(),
ids: HashMap::new(),
next_id: 1,
watcher: None,
}
}
/// The memoized id for `(name, incarnation)` — see the field docs.
fn node_id(&mut self, name: &str, incarnation: u32) -> NodeId {
*self
.ids
.entry((name.to_string(), incarnation))
.or_insert_with(|| {
let id = NodeId::new(self.next_id);
self.next_id += 1;
id
})
}
/// Deliver `event` to every live subscriber, pruning the dead: a closed
/// channel means the subscriber dropped its [`MembershipEvents`]
/// (crate::cluster::membership::MembershipEvents).
fn emit(&mut self, event: &NodeEvent) {
self.subscribers.retain(|tx| tx.send(event.clone()).is_ok());
}
}
impl Default for Manager {
fn default() -> Self {
Manager::new()
}
}
/// Outcome of a [`Call::Register`].
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum Registered {
/// The name was free; this connection is now the peer of record.
Ok,
/// Another live connection already holds this name — the caller lost the
/// race (or is a duplicate) and must not run.
Duplicate,
}
/// Requests to the manager.
pub enum Call {
/// The path claims its peer's name for a freshly-established connection,
/// handing the manager the actor's [`ConnHandle`] and the handshake's
/// [`Peer`] (the membership identity `node_up` announces). The manager
/// monitors `pid` and holds the handle for as long as the entry lives; a
/// [`Registered::Duplicate`] verdict drops the handle here, which stops
/// the refused actor.
Register {
peer: Peer,
pid: Pid,
handle: ConnHandle,
},
/// Tear down the connection to `name`: the manager drops its handle, the
/// actor stops, and the monitor removes the entry. A no-op if no such
/// connection is live.
Disconnect { name: String },
/// The current peer names, sorted. For observation and tests.
Peers,
/// A dialer declares an in-flight dial to `name` before connecting. The
/// pid is the dialing actor, monitored so the intent dies with it.
DialBegin { name: String, pid: Pid },
/// The dial to `name` resolved (either way): drop the intent. A call,
/// not a cast, so the intent is provably gone before the dialer moves on.
DialEnd { name: String },
/// The [`PeerStanding`] of an inbound `Hello` offering `peer_name` — the
/// accept path asks this between reading the frame and judging it.
Standing { peer_name: String },
/// Subscribe `tx` to membership events, snapshot-then-stream: one
/// [`NodeEvent::NodeUp`] per live peer is queued into `tx` before this
/// call answers, so the stream is exact from its first event (handlers
/// are serialized — nothing interleaves with the snapshot). Use
/// [`subscribe`](crate::cluster::membership::subscribe).
Subscribe { tx: Sender<NodeEvent> },
/// The current view: every live peer's [`NodeInfo`], unordered. Use
/// [`view`](crate::cluster::membership::view).
View,
}
/// Replies from the manager.
#[derive(Debug)]
pub enum Reply {
Registered(Registered),
Disconnected,
Peers(Vec<String>),
/// `false`: another dial to this name is already in flight — do not dial.
DialBegan(bool),
DialEnded,
Standing(PeerStanding),
Subscribed,
View(Vec<NodeInfo>),
}
impl GenServer for Manager {
type Call = Call;
type Reply = Reply;
type Cast = ();
type Info = ();
type Timer = ();
fn init(&mut self, ctx: &GenServerCtx<Self>) {
self.watcher = Some(ctx.watcher());
}
/// Manager shutdown drops every entry (and with it every ConnHandle);
/// the outbound table must not outlive the connections it names.
fn terminate(&mut self) {
for name in self.conns.keys() {
unbind_outbound(name);
}
}
fn handle_call(&mut self, request: Call) -> Reply {
match request {
Call::Register {
peer,
pid,
mut handle,
} => {
if self.conns.contains_key(&peer.node_name) {
// `handle` drops here: the refused actor stops itself.
return Reply::Registered(Registered::Duplicate);
}
if let Some(w) = &self.watcher {
w.watch(monitor(pid));
}
// The outbound table (c9) is maintained here, inside the same
// serialized handlers that own the connection's lifetime.
if let Some((frames, monitors)) = handle.take_outbound() {
bind_outbound(&peer.node_name, peer.incarnation, frames, monitors);
}
let info = NodeInfo {
node: self.node_id(&peer.node_name, peer.incarnation.get()),
name: peer.node_name.clone(),
incarnation: peer.incarnation,
meta: peer.meta,
};
self.conns.insert(
peer.node_name,
ConnEntry {
pid,
info: info.clone(),
_handle: handle,
},
);
self.emit(&NodeEvent::NodeUp(info));
Reply::Registered(Registered::Ok)
}
Call::Disconnect { name } => {
// Dropping the entry drops the handle, which stops the actor.
if let Some(entry) = self.conns.remove(&name) {
unbind_outbound(&name);
self.emit(&NodeEvent::NodeDown(entry.info));
}
Reply::Disconnected
}
Call::Peers => {
let mut names: Vec<String> = self.conns.keys().cloned().collect();
names.sort();
Reply::Peers(names)
}
Call::DialBegin { name, pid } => {
if self.dials.contains_key(&name) {
return Reply::DialBegan(false);
}
if let Some(w) = &self.watcher {
w.watch(monitor(pid));
}
self.dials.insert(name, pid);
Reply::DialBegan(true)
}
Call::DialEnd { name } => {
self.dials.remove(&name);
Reply::DialEnded
}
Call::Standing { peer_name } => {
Reply::Standing(if self.conns.contains_key(&peer_name) {
PeerStanding::Claimed
} else if self.dials.contains_key(&peer_name) {
PeerStanding::Dialing
} else {
PeerStanding::Free
})
}
Call::Subscribe { tx } => {
// The snapshot: queued before `tx` joins the list, and — the
// handlers being serialized — before any later event.
for entry in self.conns.values() {
let _ = tx.send(NodeEvent::NodeUp(entry.info.clone()));
}
self.subscribers.push(tx);
Reply::Subscribed
}
Call::View => Reply::View(self.conns.values().map(|e| e.info.clone()).collect()),
}
}
fn handle_cast(&mut self, _request: ()) {}
fn handle_down(&mut self, down: Down) {
let mut downs = Vec::new();
self.conns.retain(|name, entry| {
let dead = entry.pid == down.pid;
if dead {
downs.push((name.clone(), entry.info.clone()));
}
!dead
});
for (name, info) in downs {
unbind_outbound(&name);
self.emit(&NodeEvent::NodeDown(info));
}
self.dials.retain(|_, pid| *pid != down.pid);
}
}
+97
View File
@@ -0,0 +1,97 @@
//! RFC 010 c7a — membership: `node_up`/`node_down` events and the view.
//!
//! The membership *state* lives inside the [`manager`](crate::cluster::manager)
//! — `node_up` and `node_down` are derived facts of the exact events the
//! manager already owns (a successful registration; a reap or `Disconnect`),
//! so holding the view anywhere else would only add a cross-actor ordering
//! seam. This module is the consumer surface: the event and view types, and
//! the [`subscribe`]/[`view`] entry points. No consumer ever touches the
//! connection table (roadmap-binding, enforced by module privacy: the table
//! is a private field, and nothing here exposes names→pids).
//!
//! ## Subscription semantics (ratified 2026-08-15)
//!
//! [`subscribe`] is **snapshot-then-stream**: the returned receiver first
//! yields one [`NodeEvent::NodeUp`] per currently-live peer, then live events
//! as they happen. Because the manager is a `gen_server` (handlers are
//! serialized), the snapshot is exact — no event can interleave with it, and
//! per-subscriber ordering matches the manager's processing order. There is
//! no join-race for late subscribers and no separate "get, then diff" dance;
//! [`view`] exists for observation, not for synchronization.
//!
//! A dropped subscriber is pruned on the next emission (its channel reports
//! closed) — no monitor needed, the sender itself tells us.
//!
//! ## NodeId identity
//!
//! A [`NodeId`] is a compact **local alias for the wire identity**
//! `(node_name, incarnation)`, memoized by the manager: a reconnect blip at
//! the same incarnation keeps its id (down, then up, same id), while a
//! restart — a new incarnation — gets a fresh one, so a node's ghost and its
//! successor are always distinguishable. Ids are allocated from 1;
//! [`DEFAULT_NODE_ID`](crate::pg::DEFAULT_NODE_ID) (0) remains the local
//! node, per [`pg`](crate::pg)'s framing.
use crate::channel::{channel, Receiver};
use crate::cluster::envelope::NodeMeta;
use crate::cluster::manager::{Call, Reply, MANAGER};
use crate::gen_server;
use crate::pg::{Incarnation, NodeId};
/// One live remote node, as the view and [`NodeEvent::NodeUp`] describe it.
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct NodeInfo {
/// The local alias for `(name, incarnation)` — see the module docs.
pub node: NodeId,
/// The peer's claimed node name (handshake-verified).
pub name: String,
/// The peer's incarnation epoch, as offered in its `Hello`.
pub incarnation: Incarnation,
/// The peer's `Hello` metadata.
pub meta: NodeMeta,
}
/// A membership change, as delivered to subscribers.
#[derive(Debug, Clone, PartialEq, Eq)]
pub enum NodeEvent {
/// A peer's control connection established and registered.
NodeUp(NodeInfo),
/// That peer's connection ended — reaped, commanded down, or the manager
/// itself shut down. Carries the same [`NodeInfo`] the corresponding
/// `NodeUp` delivered, so consumers need no id→name reverse map.
NodeDown(NodeInfo),
}
/// A live membership subscription: the receiving end of the event stream
/// (the [`Monitor`](crate::monitor::Monitor) shape — read from [`rx`], drop
/// to unsubscribe).
///
/// [rx]: MembershipEvents::rx
pub struct MembershipEvents {
/// The event stream: the snapshot's `NodeUp`s first, then live events.
/// Fold it into a `select` from a plain actor, or pipe it into a
/// `gen_server` via `with_info`.
pub rx: Receiver<NodeEvent>,
}
/// Subscribe to membership events (snapshot-then-stream — see the module
/// docs). `None`: the manager is not running. Must be called from inside an
/// actor.
pub fn subscribe() -> Option<MembershipEvents> {
let (tx, rx) = channel();
match gen_server::call(MANAGER, Call::Subscribe { tx }) {
Ok(Reply::Subscribed) => Some(MembershipEvents { rx }),
_ => None,
}
}
/// The current view: every live peer's [`NodeInfo`], unordered. For
/// observation and tests; consumers that need to *track* the view should
/// [`subscribe`] instead (the snapshot makes the stream self-sufficient).
/// `None`: the manager is not running. Must be called from inside an actor.
pub fn view() -> Option<Vec<NodeInfo>> {
match gen_server::call(MANAGER, Call::View) {
Ok(Reply::View(v)) => Some(v),
_ => None,
}
}
+546
View File
@@ -0,0 +1,546 @@
//! RFC 010 c15 — distributed process groups (Phase 5).
//!
//! The Erlang `pg` shape (D18): every node's group store is the union of
//! its own local members and each peer's *announced* local members. There
//! is one **pg actor** per node — the c14 reaper grown up — and it is the
//! only writer of remote entries and the only sender of announcements:
//!
//! - **Origin owns its members.** Joins are local (`pg::join`), the eager
//! reaper is the liveness authority, and the origin announces every
//! change: `Join`/`Leave` incrementally to every up node, and a full
//! `Sync` of its local groups to a peer the moment that peer comes up
//! (`NodeUp`). Nobody monitors a remote member; a peer's `NodeDown` sweeps
//! every member it announced.
//! - **Transport is a pure consumer** of Phase 3/4: the exposed name
//! [`PG_NAME`] (`"pg"`) carrying [`PgMsg`] over postcard, sent with
//! [`remote::send`]. No new frame, no manager change.
//! - **No anti-entropy.** Per-origin ordering rides the single TCP link:
//! the actor sends `Sync` to a peer *before* it can send that peer any
//! `Join`/`Leave` (both from the same loop, over the same connection), and
//! a reconnect is a fresh `NodeUp` ⇒ fresh `Sync` replacing that peer's
//! set wholesale.
//! - **Local API unchanged.** `members`/`pick`/`dispatch` stay local-only
//! (`get_local_members`); a remote entry in the store carries the peer's
//! `NodeId` and never surfaces there. Cluster-wide reads are the new,
//! additive [`members_all`] over [`GroupMember`] (c16 adds `pick_any` /
//! `dispatch_any`).
//!
//! ## Ordering inside the node
//!
//! `pg::join`/`pg::leave` mutate the store on the caller's thread and then
//! *announce* to the actor's control inbox. Because the store op precedes the
//! announcement and the actor re-reads the store before broadcasting a
//! `Joined`, an announcement that has been overtaken (the member left or died
//! before the actor got to it) is dropped rather than advertised: the wire
//! never sees a `Join` for a member the origin no longer holds. `Leave`
//! broadcasts unconditionally — a spurious `Leave` is a no-op at the peer.
//!
//! Inbound: `NodeUp` is emitted by the manager on the accept/connect path,
//! *before* the peer's connection actor exists, so it is queued on the
//! membership stream before any frame from that peer can reach this inbox.
//! The actor still drains membership before it interprets a `PgMsg` whose
//! sender it does not know, and drops the message if the sender is still not
//! up (a ghost — its next `NodeUp` brings a `Sync`).
use std::collections::HashMap;
use crate::channel::{channel, select, Receiver, Selectable};
use crate::cluster::expose::expose;
use crate::cluster::membership::{subscribe, MembershipEvents, NodeEvent, NodeInfo};
use crate::cluster::remote::{
self, local_identity, send_to_remote, RemoteName, RemotePid, ToRemoteError,
};
use crate::monitor::Down;
use crate::pg::{
live, member_for, reaper_inboxes, sweep_local_death, Incarnation, Member, Membership, PgEvent,
};
use crate::pid::{assert_type, Addressable, Erased, Pid};
use crate::registry::{register, send_to, SendError};
use crate::scheduler::with_runtime;
use crate::Name;
/// The exposed name every node's pg actor answers under.
pub const PG_NAME: Name<PgMsg> = Name::new("pg");
/// The pg wire protocol. Every variant is origin-authored: `from` / the
/// pid's node is the node whose local members are being described.
#[derive(Debug, Clone, PartialEq, Eq)]
pub enum PgMsg {
/// The origin's complete local membership, sent to a peer on `NodeUp`.
/// Replaces whatever the receiver held for that origin.
Sync {
from: String,
groups: Vec<(String, Vec<RemotePid<Erased>>)>,
},
/// The origin added `pid` (its own) to `group`.
Join {
group: String,
pid: RemotePid<Erased>,
},
/// The origin removed `pid` from `group` — voluntary leave or death.
Leave {
group: String,
pid: RemotePid<Erased>,
},
}
// Hand-rolled serde (the crate carries no serde-derive), as a 3-tuple with a
// leading tag: (0, from, groups) | (1, group, pid) | (2, group, pid).
impl serde::Serialize for PgMsg {
fn serialize<S: serde::Serializer>(&self, s: S) -> Result<S::Ok, S::Error> {
use serde::ser::SerializeTuple;
let mut t = s.serialize_tuple(3)?;
match self {
PgMsg::Sync { from, groups } => {
t.serialize_element(&0u8)?;
t.serialize_element(from)?;
t.serialize_element(groups)?;
}
PgMsg::Join { group, pid } => {
t.serialize_element(&1u8)?;
t.serialize_element(group)?;
t.serialize_element(pid)?;
}
PgMsg::Leave { group, pid } => {
t.serialize_element(&2u8)?;
t.serialize_element(group)?;
t.serialize_element(pid)?;
}
}
t.end()
}
}
impl<'de> serde::Deserialize<'de> for PgMsg {
fn deserialize<D: serde::Deserializer<'de>>(d: D) -> Result<Self, D::Error> {
struct V;
impl<'de> serde::de::Visitor<'de> for V {
type Value = PgMsg;
fn expecting(&self, f: &mut std::fmt::Formatter) -> std::fmt::Result {
f.write_str("a pg message tuple")
}
fn visit_seq<A: serde::de::SeqAccess<'de>>(
self,
mut seq: A,
) -> Result<PgMsg, A::Error> {
use serde::de::Error;
let tag: u8 = seq
.next_element()?
.ok_or_else(|| A::Error::custom("pg: missing tag"))?;
let text: String = seq
.next_element()?
.ok_or_else(|| A::Error::custom("pg: missing name"))?;
match tag {
0 => {
let groups = seq
.next_element()?
.ok_or_else(|| A::Error::custom("pg: missing groups"))?;
Ok(PgMsg::Sync { from: text, groups })
}
1 | 2 => {
let pid = seq
.next_element()?
.ok_or_else(|| A::Error::custom("pg: missing pid"))?;
Ok(if tag == 1 {
PgMsg::Join { group: text, pid }
} else {
PgMsg::Leave { group: text, pid }
})
}
t => Err(A::Error::custom(format!("pg: unknown tag {t}"))),
}
}
}
d.deserialize_tuple(3, V)
}
}
/// A member of a group as the cluster sees it: on this node (a plain
/// [`Pid`], sendable locally) or on a peer (a [`RemotePid`], sendable via
/// [`send_to_remote`](remote::send_to_remote)). `Pid` cannot hold a remote
/// (D14), hence the two-variant shape.
#[derive(Debug, Clone, PartialEq, Eq)]
pub enum GroupMember {
Local(Pid),
Remote(RemotePid<Erased>),
}
/// Every member of `group` cluster-wide, in the store's order: local members
/// filtered by the same liveness backstop as [`members`](crate::pg::members),
/// remote members exactly as their origins last announced them. Must run
/// inside [`run`](crate::run).
pub fn members_all(group: &str) -> Vec<GroupMember> {
with_runtime(|inner| {
let me = inner.node_id;
let pg = inner.process_groups.lock();
pg.all_of(group)
.into_iter()
.filter_map(|m| {
if m.node == me {
live(inner, m.pid).then_some(GroupMember::Local(m.pid))
} else {
// A remote entry always has its node's name recorded
// (they land under the same lock); a missing one is a
// node already swept, so it hides rather than misnames.
pg.node_name(m.node).map(|name| {
GroupMember::Remote(RemotePid::from_parts(
name,
m.incarnation,
m.pid.index(),
m.pid.generation(),
))
})
}
})
.collect()
})
}
/// One member of `group` cluster-wide, or `None` if it has none: the first
/// entry in the store's order (this node's members in join order first when
/// they joined first — the same stateless first-live scan as
/// [`pick`](crate::pg::pick), extended over the peers' announced members).
/// Must run inside [`run`](crate::run).
pub fn pick_any(group: &str) -> Option<GroupMember> {
members_all(group).into_iter().next()
}
/// Why [`dispatch_any`] handed `msg` back.
#[derive(Debug)]
pub enum DispatchAnyError<M> {
/// The group has no member anywhere.
NoMember(M),
/// The pick was local and the local typed send failed.
Local(SendError<M>),
/// The pick was remote and the remote send failed at this node.
Remote(ToRemoteError<M>),
}
impl<M> DispatchAnyError<M> {
/// The undelivered message.
pub fn into_inner(self) -> M {
match self {
DispatchAnyError::NoMember(m) => m,
DispatchAnyError::Local(e) => e.into_inner(),
DispatchAnyError::Remote(e) => e.into_inner(),
}
}
}
/// [`pick_any`] and send in one step, returning the member reached: a local
/// pick goes through [`send_to`], a remote one through [`send_to_remote`]
/// (so `Ok` for a remote member means "handed to the connection", RFC 010
/// §3). Homogeneous pool assumed, as for [`dispatch`](crate::pg::dispatch);
/// a wrong `A` degrades to a clean error at the target, never a misroute.
/// Must run inside [`run`](crate::run).
pub fn dispatch_any<A>(group: &str, msg: A::Msg) -> Result<GroupMember, DispatchAnyError<A::Msg>>
where
A: Addressable,
A::Msg: serde::Serialize,
{
match pick_any(group) {
None => Err(DispatchAnyError::NoMember(msg)),
Some(GroupMember::Local(pid)) => send_to(assert_type::<A>(pid), msg)
.map(|()| GroupMember::Local(pid))
.map_err(DispatchAnyError::Local),
Some(GroupMember::Remote(rp)) => send_to_remote(rp.clone().assert_type::<A>(), msg)
.map(|()| GroupMember::Remote(rp))
.map_err(DispatchAnyError::Remote),
}
}
/// Attach the pg actor to the running cluster. Called once by
/// `cluster::start` after the manager is up and the local identity is set;
/// spawns the actor if this run has not joined anything yet.
pub(crate) fn attach_cluster() {
let _ = reaper_inboxes().ctl.send(PgEvent::Attach);
}
/// The attached half of the actor's state: who is up (by name) and the
/// membership stream.
struct Attached {
events: MembershipEvents,
peers: HashMap<String, NodeInfo>,
/// This node's wire identity: `Sync`'s `from`, and the stamp on every
/// pid we ship (attach requires it, so no `None` path exists here).
me: String,
incarnation: Incarnation,
}
/// The pg actor: the c14 reaper (`deaths`), the local API's announcements
/// (`ctl`), and — once attached — the membership stream and the exposed
/// `"pg"` inbox, all in one drain-then-select loop. `deaths`/`ctl` closing
/// is the run tearing down; the membership stream closing is the manager
/// gone (detach, keep reaping).
pub(crate) fn actor(deaths: Receiver<Down>, ctl: Receiver<PgEvent>) {
let (pg_tx, pg_rx) = channel::<PgMsg>();
let mut cl: Option<Attached> = None;
loop {
loop {
match deaths.try_recv() {
Ok(Some(down)) => on_death(cl.as_ref(), down.pid),
Ok(None) => break,
Err(_) => return,
}
}
loop {
match ctl.try_recv() {
Ok(Some(PgEvent::Attach)) => {
// Own the name BEFORE subscribing (which yields to the
// manager): a peer's first frame must find "pg" exposed
// and resolvable, or it is dropped. Idempotent for a
// re-attach: same actor, same channel (the registry
// refuses a *second* live one).
let _ = register(PG_NAME, pg_tx.clone());
expose(PG_NAME);
if let Some(a) = attach() {
cl = Some(a);
}
}
Ok(Some(PgEvent::Joined { group, pid })) => on_joined(cl.as_ref(), &group, pid),
Ok(Some(PgEvent::Left { group, pid })) => on_left(cl.as_ref(), &group, pid),
Ok(None) => break,
Err(_) => return,
}
}
if let Some(a) = cl.as_mut() {
if !drain_events(a) {
cl = None;
continue;
}
loop {
match pg_rx.try_recv() {
Ok(Some(msg)) => on_msg(a, msg),
Ok(None) => break,
Err(_) => return, // our own inbox: only on teardown
}
}
}
// Wait. Control first (attach/teardown must be prompt), then deaths,
// then the cluster arms.
let mut arms: Vec<&dyn Selectable> = vec![&ctl, &deaths];
if let Some(a) = cl.as_ref() {
arms.push(&a.events.rx);
arms.push(&pg_rx);
}
let _ = select(&arms);
}
}
fn attach() -> Option<Attached> {
let events = subscribe()?;
let (me, incarnation) = local_identity()?;
Some(Attached {
events,
peers: HashMap::new(),
me,
incarnation,
})
}
/// Fold pending membership events: `NodeUp` ⇒ record + `Sync` that peer;
/// `NodeDown` ⇒ sweep every member it announced. `false` when the stream
/// has closed.
fn drain_events(a: &mut Attached) -> bool {
loop {
match a.events.rx.try_recv() {
Ok(Some(NodeEvent::NodeUp(info))) => {
let name = info.name.clone();
a.peers.insert(name.clone(), info);
// Snapshot under the store lock, then stamp wire pids
// outside it (`from_local` marks watchable under the slot's
// cold lock — Leaf-on-Leaf nesting is asserted).
let local: Vec<(String, Vec<Pid>)> =
with_runtime(|inner| inner.process_groups.lock().groups_on(inner.node_id));
let groups = local
.into_iter()
.map(|(g, pids)| (g, pids.into_iter().map(|p| wire(a, p)).collect()))
.collect();
let msg = PgMsg::Sync {
from: a.me.clone(),
groups,
};
let _ = remote::send(RemoteName::new(name, PG_NAME), msg);
}
Ok(Some(NodeEvent::NodeDown(info))) => {
a.peers.remove(&info.name);
with_runtime(|inner| {
let mut pg = inner.process_groups.lock();
pg.remove_where(|m| m.node == info.node);
pg.forget_node_name(info.node);
});
}
Ok(None) => return true,
Err(_) => return false,
}
}
}
/// Send `msg` to every up peer. `NotConnected` is ignored: that peer's
/// `NodeDown` is on its way and its next `NodeUp` gets a `Sync`.
fn broadcast(a: &Attached, msg: PgMsg) {
for name in a.peers.keys() {
let _ = remote::send(RemoteName::new(name.clone(), PG_NAME), msg.clone());
}
}
fn on_death(a: Option<&Attached>, pid: Pid) {
let evicted = sweep_local_death(pid);
if let Some(a) = a {
for (group, ms) in evicted {
broadcast(
a,
PgMsg::Leave {
group,
pid: wire(a, ms.member.pid),
},
);
}
}
}
fn on_joined(a: Option<&Attached>, group: &str, pid: Pid) {
let Some(a) = a else { return };
// Re-check: a leave/death may have overtaken the announcement.
let still = with_runtime(|inner| {
let m = member_for(inner, pid);
inner.process_groups.lock().contains(group, &m)
});
if still {
broadcast(
a,
PgMsg::Join {
group: group.to_owned(),
pid: wire(a, pid),
},
);
}
}
fn on_left(a: Option<&Attached>, group: &str, pid: Pid) {
let Some(a) = a else { return };
broadcast(
a,
PgMsg::Leave {
group: group.to_owned(),
pid: wire(a, pid),
},
);
}
/// The wire form of a local member pid, stamped with the identity the
/// actor was attached with (marks watchable, like `from_local`).
fn wire(a: &Attached, pid: Pid) -> RemotePid<Erased> {
RemotePid::from_local_at(pid, a.me.clone(), a.incarnation)
}
/// The named origin's `NodeInfo`, if it is up. A second look at the
/// membership stream covers a `NodeUp` that landed after this loop
/// iteration's drain; anything still unknown is a ghost and is dropped.
fn origin(a: &mut Attached, name: &str) -> Option<NodeInfo> {
if let Some(i) = a.peers.get(name) {
return Some(i.clone());
}
drain_events(a);
a.peers.get(name).cloned()
}
/// `origin`, additionally requiring `pid` to be stamped with the origin's
/// current incarnation — a pid from a previous life of that node is a ghost.
fn origin_of(a: &mut Attached, pid: &RemotePid<Erased>) -> Option<NodeInfo> {
origin(a, pid.node()).filter(|i| i.incarnation == pid.incarnation())
}
fn remote_membership(origin: &NodeInfo, pid: &RemotePid<Erased>) -> Membership {
Membership {
member: Member {
node: origin.node,
incarnation: origin.incarnation,
pid: Pid::new(pid.index(), pid.generation()),
},
monitor: None,
}
}
fn on_msg(a: &mut Attached, msg: PgMsg) {
match msg {
PgMsg::Sync { from, groups } => {
let Some(info) = origin(a, &from) else { return };
with_runtime(|inner| {
let mut pg = inner.process_groups.lock();
pg.remove_where(|m| m.node == info.node);
pg.set_node_name(info.node, info.name.clone());
for (group, pids) in &groups {
// Origin-authored: only its own current-incarnation pids.
for p in pids
.iter()
.filter(|p| p.node() == from && p.incarnation() == info.incarnation)
{
pg.join(group, remote_membership(&info, p));
}
}
});
}
PgMsg::Join { group, pid } => {
let Some(info) = origin_of(a, &pid) else {
return;
};
with_runtime(|inner| {
let mut pg = inner.process_groups.lock();
pg.set_node_name(info.node, info.name.clone());
pg.join(&group, remote_membership(&info, &pid));
});
}
PgMsg::Leave { group, pid } => {
let Some(info) = origin_of(a, &pid) else {
return;
};
with_runtime(|inner| {
let ms = remote_membership(&info, &pid);
inner.process_groups.lock().leave(&group, ms.member);
});
}
}
}
#[cfg(test)]
mod tests {
use super::*;
use crate::cluster::envelope::{decode_payload, encode_payload};
#[test]
fn pg_msg_roundtrips_every_variant() {
let p = RemotePid::<Erased>::from_parts("a", Incarnation::new(9), 3, 1);
for m in [
PgMsg::Sync {
from: "a".into(),
groups: vec![
("g".into(), vec![p.clone(), p.clone()]),
("h".into(), vec![]),
],
},
PgMsg::Sync {
from: "a".into(),
groups: vec![],
},
PgMsg::Join {
group: "g".into(),
pid: p.clone(),
},
PgMsg::Leave {
group: "g".into(),
pid: p.clone(),
},
] {
let bytes = encode_payload(&m).unwrap();
let back: PgMsg = decode_payload(&bytes).unwrap();
assert_eq!(back, m);
}
}
#[test]
fn pg_msg_rejects_unknown_tag() {
let bytes = encode_payload(&(7u8, "x", 0u32)).unwrap();
assert!(decode_payload::<PgMsg>(&bytes).is_err());
}
}
+835
View File
@@ -0,0 +1,835 @@
//! RFC 010 c9 — remote `Name` sends: the outbound path and the single
//! inbound name-resolution seam.
//!
//! ## Outbound (D13, ratified 2026-08-15)
//!
//! A module-private table `node name → Sender<Frame>` — one dedicated
//! outbound channel per live connection, populated and torn down by the
//! manager inside the same serialized handlers that own the connection's
//! lifetime (Register / Disconnect / reap), living on `RuntimeInner` beside
//! the exposure state. [`send`] is one leaf-lock lookup + one channel send:
//! no gen_server on the data plane, no published channel anyone holding a
//! pid could inject raw frames into. `Ok(())` means **handed to the
//! connection's inbox** — local knowledge only, exactly the BEAM contract
//! (RFC §3): a missing entry or a closed channel is
//! [`RemoteSendError::NotConnected`]; delivery confirmation is the monitor's
//! job (c11). The entry-present/actor-dying-mid-send window is *honest*
//! under that contract, not a bug.
//!
//! The outbound sender is deliberately **separate from the conn actor's
//! command channel**: if it were a clone of `cmd_tx`, the manager dropping
//! its `ConnHandle` would no longer close that channel and connection
//! lifetime would leak to whoever holds a sender — a D9 violation.
//!
//! Buffering is unbounded toward a slow peer (the BEAM `busy_dist_port`
//! shape); backpressure is out of c9's scope and noted here rather than
//! silently absent.
//!
//! ## Inbound — the ONE resolution seam (RFC v2)
//!
//! Every wire-name → local-pid resolution goes through [`deliver_named`],
//! and nothing else: the conn actor hands it the three fields of a
//! `SendNamed` and gets back a verdict. It checks the exposed set first (an
//! unexposed name is unreachable — the gun's safety), then the type hash
//! against what the name was exposed with, then resolves the name through
//! the registry and delivers via c8's [`decode_deliver`]. When an owned-name
//! table lands beside the `&'static str` registry, it slots in here without
//! touching call sites. Module privacy enforces the funnel: the exposed and
//! outbound tables are `pub(crate)`, and no other module resolves names for
//! the wire.
//!
//! Refusals are silent to the sender by design (§3: send failure reflects
//! local knowledge only); they are observable locally as the returned
//! [`InboundVerdict`], which the conn actor may log or count.
//!
//! ## Pids (c10, D14)
//!
//! [`RemotePid<A>`] = `(node_name, incarnation, index, generation)` +
//! phantom — identity-bound, dead when that incarnation dies, never
//! redirects. The node travels as its **name** (a global identifier, so a pid
//! forwarded through a third node needs no re-mapping); NodeId is a local
//! alias and never crosses. A local `Pid<A>` serializes *as* a `RemotePid`
//! stamped from the ambient [local identity](set_local_identity); a
//! `RemotePid` deserializes into `Pid<A>` only when it names this node (the
//! collapse), else it is a decode error — fields that may hold a pid from
//! anywhere are typed `RemotePid<A>`.
//!
//! [`send_to_remote`] is the pid-targeted send. A self-node pid short-
//! circuits to the local typed send with the message object itself — no
//! encode, no frame (zero-copy-equivalent). Otherwise the outbound table
//! (widened to carry each node's **current incarnation**) does the RFC v2 §3
//! check at the send site: a pid of a dead incarnation is
//! [`ToRemoteError::DeadIncarnation`] and no frame is emitted. Inbound
//! `Send` frames are delivered by index/generation through c8's
//! [`decode_deliver`]: the target actor's published channel for the exposed
//! type is the only route (the reply-to path requires
//! [`expose_type`](crate::cluster::expose::expose_type) at the receiver).
use std::cell::Cell;
use std::collections::HashMap;
use std::marker::PhantomData;
use crate::channel::{channel, Receiver, RecvError, Selectable, Sender};
use crate::cluster::envelope::{encode_payload, Frame, PayloadError, RemoteDownReason};
use crate::cluster::expose::{decode_deliver, exposed_hash, type_hash, DeliverError};
use crate::monitor::{demonitor, monitor, Monitor, MonitorId};
use crate::pg::Incarnation;
use crate::pid::{Addressable, Erased, Name, Pid};
use crate::registry::{send_to, whereis, SendError};
use crate::scheduler::with_runtime;
/// A name on a specific remote node: `(node_name, Name<M>)`. Sendable via
/// [`send`]; typed, so the payload is `M` and the wire hash is
/// [`type_hash::<M>()`](type_hash).
pub struct RemoteName<M> {
node: String,
name: Name<M>,
_marker: PhantomData<fn() -> M>,
}
impl<M> RemoteName<M> {
pub fn new(node: impl Into<String>, name: Name<M>) -> Self {
RemoteName {
node: node.into(),
name,
_marker: PhantomData,
}
}
pub fn node(&self) -> &str {
&self.node
}
pub fn name(&self) -> Name<M> {
self.name
}
}
impl<M> Clone for RemoteName<M> {
fn clone(&self) -> Self {
RemoteName {
node: self.node.clone(),
name: self.name,
_marker: PhantomData,
}
}
}
impl<M> std::fmt::Debug for RemoteName<M> {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
write!(f, "{}@{}", self.name.as_str(), self.node)
}
}
/// Why a remote send did not leave this node. Local knowledge only.
#[derive(Debug)]
pub enum RemoteSendError<M> {
/// No live connection to that node right now (never connected, or gone
/// and not yet re-dialed). The message is handed back.
NotConnected(M),
/// The payload did not serialize.
Encode(M, PayloadError),
}
impl<M> RemoteSendError<M> {
pub fn into_inner(self) -> M {
match self {
RemoteSendError::NotConnected(m) | RemoteSendError::Encode(m, _) => m,
}
}
}
/// The outbound table, one per runtime (a `RuntimeInner` field): per live
/// node, its current incarnation (the RFC v2 §3 send-site check) and the
/// connection's dedicated outbound sender. Plus this node's own wire
/// identity, which pid serialization stamps.
pub(crate) struct Outbound {
by_node: HashMap<String, Route>,
local: Option<(String, Incarnation)>,
}
/// One live connection as the outbound path sees it: the peer's current
/// incarnation and the two inboxes of its connection actor — frames (c9)
/// and monitor bookkeeping (c12, [`MonCmd`]).
pub(crate) struct Route {
incarnation: Incarnation,
frames: Sender<Frame>,
monitors: Sender<MonCmd>,
}
impl Outbound {
pub(crate) fn new() -> Self {
Outbound {
by_node: HashMap::new(),
local: None,
}
}
}
/// Set this node's wire identity — what serialized pids are stamped with
/// and what a `RemotePid` must name to collapse. `cluster::start` sets it;
/// exposed for local tests. Must run inside [`run`](crate::run).
pub fn set_local_identity(node: &str, incarnation: Incarnation) {
with_runtime(|inner| {
inner.outbound.lock().local = Some((node.to_string(), incarnation));
});
}
/// This node's wire identity, if set. Must run inside [`run`](crate::run).
pub fn local_identity() -> Option<(String, Incarnation)> {
with_runtime(|inner| inner.outbound.lock().local.clone())
}
/// Manager-only: bind `node`'s outbound channels at `incarnation`. Called
/// inside `Register`.
pub(crate) fn bind_outbound(
node: &str,
incarnation: Incarnation,
frames: Sender<Frame>,
monitors: Sender<MonCmd>,
) {
with_runtime(|inner| {
inner.outbound.lock().by_node.insert(
node.to_string(),
Route {
incarnation,
frames,
monitors,
},
);
});
}
/// Test probe: bind an arbitrary sender as `node`'s outbound so a test can
/// assert what frames leave — or don't. Same table, same lookup as the real
/// path (this is how "no frame emitted" is asserted at the frame level).
/// Frames only: there is no connection actor behind a probe, so a
/// [`monitor_remote`] against a probed node reports `Disconnected`.
pub fn bind_outbound_probe(node: &str, incarnation: Incarnation, tx: Sender<Frame>) {
drop(bind_outbound_probe_with_monitors(node, incarnation, tx));
}
/// The monitor half of a probed node's inbox: opaque, held only to be
/// dropped. See [`bind_outbound_probe_with_monitors`].
pub struct MonitorInbox {
_rx: Receiver<MonCmd>,
}
/// Test probe: like [`bind_outbound_probe`], but the monitor-command
/// receiver is handed back instead of dropped, so a test can stage the
/// c13 drain gap — a `Monitor` command that reached the connection's inbox
/// and dies unread when the inbox is dropped. While the inbox lives,
/// [`monitor_remote`] against the probed node is simply in flight.
pub fn bind_outbound_probe_with_monitors(
node: &str,
incarnation: Incarnation,
tx: Sender<Frame>,
) -> MonitorInbox {
let (mon_tx, mon_rx) = channel();
bind_outbound(node, incarnation, tx, mon_tx);
MonitorInbox { _rx: mon_rx }
}
/// Manager-only: unbind `node`'s outbound channel. Called on `Disconnect`,
/// reap, and manager shutdown. Dropping the sender is what closes the conn
/// actor's outbound arm — but that arm's closure is NOT a stop signal (the
/// cmd channel is, per D9); the actor simply stops selecting on it.
pub(crate) fn unbind_outbound(node: &str) {
with_runtime(|inner| {
inner.outbound.lock().by_node.remove(node);
});
}
/// Send `msg` to `target`. `Ok(())` = handed to the connection's inbox, and
/// nothing more — see the module docs. Must run inside
/// [`run`](crate::run).
pub fn send<M>(target: RemoteName<M>, msg: M) -> Result<(), RemoteSendError<M>>
where
M: serde::Serialize + Send + 'static,
{
let payload = match encode_payload(&msg) {
Ok(p) => p,
Err(e) => return Err(RemoteSendError::Encode(msg, e)),
};
let frame = Frame::SendNamed {
name: target.name.as_str().to_string(),
type_hash: type_hash::<M>(),
payload,
};
match hand_to_connection(&target.node, frame) {
Ok(()) => Ok(()),
Err(NotConnected) => Err(RemoteSendError::NotConnected(msg)),
}
}
/// No live connection to the named node — the payload-free form of
/// [`RemoteSendError::NotConnected`], for the raw path.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub struct NotConnected;
/// The untyped escape hatch: send pre-encoded `payload` under an explicit
/// `type_hash`. Exists so tests (and future codecs) can put deliberately
/// wrong frames on the wire; the typed [`send`] cannot express a hash/type
/// mismatch, by design. Same `Ok` semantics as [`send`].
pub fn send_remote_raw(
node: &str,
name: &str,
type_hash: u64,
payload: &[u8],
) -> Result<(), NotConnected> {
hand_to_connection(
node,
Frame::SendNamed {
name: name.to_string(),
type_hash,
payload: payload.to_vec(),
},
)
}
/// One lookup, one send. Clone the sender out under the lock and send
/// outside it (a channel send can unpark the conn actor).
fn hand_to_connection(node: &str, frame: Frame) -> Result<(), NotConnected> {
let tx = with_runtime(|inner| {
inner
.outbound
.lock()
.by_node
.get(node)
.map(|r| r.frames.clone())
});
match tx {
Some(tx) => tx.send(frame).map_err(|_| NotConnected),
None => Err(NotConnected),
}
}
// ---- pids ---------------------------------------------------------------
/// A pid on some node: `(node_name, incarnation, index, generation)` plus
/// the actor type. See the module docs. Serializes as a 4-tuple.
pub struct RemotePid<A> {
node: String,
incarnation: Incarnation,
index: u32,
generation: u32,
_marker: PhantomData<fn() -> A>,
}
impl<A> RemotePid<A> {
/// Build from raw parts (tests, and codecs re-hydrating a pid).
pub fn from_parts(
node: impl Into<String>,
incarnation: Incarnation,
index: u32,
generation: u32,
) -> Self {
RemotePid {
node: node.into(),
incarnation,
index,
generation,
_marker: PhantomData,
}
}
/// The wire form of a local pid, stamped with this node's identity, and
/// **marked watchable** — asking for the wire form *is* the intent to
/// ship the pid, so this is the same D12 set-site as `Pid::serialize`
/// (c12 made it explicit: a peer may monitor exactly the pids that
/// crossed, and a pid handed out via `from_local` in a hand-built reply
/// has crossed). Must run inside [`run`](crate::run).
///
/// `None` when this runtime has no wire identity (no `cluster::start`,
/// no [`set_local_identity`]): such a pid cannot name a node, and a
/// stamped `("", 0)` would be dropped by every peer with no signal.
/// The pid is not marked watchable in that case either.
pub fn from_local(pid: Pid<A>) -> Option<Self> {
let (node, incarnation) = local_identity()?;
Some(Self::from_local_at(pid, node, incarnation))
}
/// `from_local` with the identity supplied by the caller — for a holder
/// that already carries the node's identity (the pg actor) and must not
/// have a `None` path. Marks watchable like `from_local`.
pub(crate) fn from_local_at(pid: Pid<A>, node: String, incarnation: Incarnation) -> Self {
crate::monitor::mark_watchable(pid);
RemotePid::from_parts(node, incarnation, pid.index(), pid.generation())
}
/// The collapse: `Some(local pid)` iff this pid names this very node
/// (name and incarnation). Must run inside [`run`](crate::run).
pub fn local(&self) -> Option<Pid<A>> {
let (n, i) = local_identity()?;
(n == self.node && i == self.incarnation)
.then(|| crate::pid::assert_type::<A>(Pid::new(self.index, self.generation)))
}
/// Drop the actor type: the untyped `RemotePid<Erased>`, the form
/// [`RemoteDown`] and [`RemoteMonitor`] carry (mirrors [`Pid::erase`]).
pub fn erase(self) -> RemotePid<Erased> {
RemotePid::from_parts(self.node, self.incarnation, self.index, self.generation)
}
/// Re-type an erased pid as `RemotePid<B>` — the unchecked mirror of
/// `pid::assert_type`, with the same degradation: a wrong `B` means the
/// target refuses the payload's hash (never a misroute).
pub(crate) fn assert_type<B>(self) -> RemotePid<B> {
RemotePid::from_parts(self.node, self.incarnation, self.index, self.generation)
}
pub fn node(&self) -> &str {
&self.node
}
pub fn incarnation(&self) -> Incarnation {
self.incarnation
}
pub fn index(&self) -> u32 {
self.index
}
pub fn generation(&self) -> u32 {
self.generation
}
}
impl<A> Clone for RemotePid<A> {
fn clone(&self) -> Self {
RemotePid::from_parts(
self.node.clone(),
self.incarnation,
self.index,
self.generation,
)
}
}
impl<A> PartialEq for RemotePid<A> {
fn eq(&self, o: &Self) -> bool {
self.node == o.node
&& self.incarnation == o.incarnation
&& self.index == o.index
&& self.generation == o.generation
}
}
impl<A> Eq for RemotePid<A> {}
impl<A> std::fmt::Debug for RemotePid<A> {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
write!(
f,
"<{}.{}@{}#{}>",
self.index,
self.generation,
self.node,
self.incarnation.get()
)
}
}
impl<A> serde::Serialize for RemotePid<A> {
fn serialize<S: serde::Serializer>(&self, s: S) -> Result<S::Ok, S::Error> {
(
self.node.as_str(),
self.incarnation.get(),
self.index,
self.generation,
)
.serialize(s)
}
}
impl<'de, A> serde::Deserialize<'de> for RemotePid<A> {
fn deserialize<D: serde::Deserializer<'de>>(d: D) -> Result<Self, D::Error> {
let (node, inc, index, generation) = <(String, u32, u32, u32)>::deserialize(d)?;
Ok(RemotePid::from_parts(
node,
Incarnation::new(inc),
index,
generation,
))
}
}
/// Why a pid-targeted send did not leave this node. Local knowledge only.
#[derive(Debug)]
pub enum ToRemoteError<M> {
/// No live connection to the pid's node.
NotConnected(M),
/// The pid's incarnation is not that node's current one (RFC v2 §3): the
/// actor died with its incarnation. Detected at the send site; no frame.
DeadIncarnation(M),
/// The payload did not serialize.
Encode(M, PayloadError),
/// The pid collapsed to a local one and the local typed send failed.
Local(SendError<M>),
}
impl<M> ToRemoteError<M> {
/// The undelivered message.
pub fn into_inner(self) -> M {
match self {
ToRemoteError::NotConnected(m)
| ToRemoteError::DeadIncarnation(m)
| ToRemoteError::Encode(m, _) => m,
ToRemoteError::Local(e) => e.into_inner(),
}
}
}
/// Send `msg` to a pid, wherever it lives. Self-node pids short-circuit to
/// the local typed send with `msg` itself (no encode, no frame); others go
/// out as a `Send` frame after the incarnation check. `Ok(())` for a remote
/// target = handed to the connection's inbox. Must run inside
/// [`run`](crate::run).
pub fn send_to_remote<A>(target: RemotePid<A>, msg: A::Msg) -> Result<(), ToRemoteError<A::Msg>>
where
A: Addressable,
A::Msg: serde::Serialize,
{
if let Some(local) = target.local() {
return send_to(local, msg).map_err(ToRemoteError::Local);
}
let route = with_runtime(|inner| {
inner
.outbound
.lock()
.by_node
.get(&target.node)
.map(|r| (r.incarnation, r.frames.clone()))
});
let (current, tx) = match route {
Some(r) => r,
None => return Err(ToRemoteError::NotConnected(msg)),
};
if current != target.incarnation {
return Err(ToRemoteError::DeadIncarnation(msg));
}
let payload = match encode_payload(&msg) {
Ok(p) => p,
Err(e) => return Err(ToRemoteError::Encode(msg, e)),
};
let frame = Frame::Send {
index: target.index,
generation: target.generation,
type_hash: type_hash::<A::Msg>(),
payload,
};
tx.send(frame).map_err(|_| ToRemoteError::NotConnected(msg))
}
/// The inbound `Send` seam: deliver `payload` under `type_hash` to the local
/// actor `(index, generation)`. Node and incarnation are implicit in the
/// connection (bound at handshake) — the frame carries only the slot
/// identity. Delivery goes through c8's decoder table, so only types the
/// receiver has [`expose_type`](crate::cluster::expose::expose_type)d (or
/// exposed by name) can land; anything else is refused, never misrouted.
pub fn deliver_to_pid(
index: u32,
generation: u32,
type_hash: u64,
payload: &[u8],
) -> InboundVerdict {
let pid = Pid::new(index, generation);
match decode_deliver(type_hash, pid, payload) {
Ok(()) => InboundVerdict::Delivered,
Err(e) => InboundVerdict::Refused(e),
}
}
/// What the inbound seam did with a `SendNamed`. Local observability only;
/// nothing goes back on the wire (RFC §3).
#[derive(Debug)]
pub enum InboundVerdict {
/// Decoded and handed to the name's holder.
Delivered,
/// The name is not in this node's exposed set.
NotExposed,
/// The frame's hash is not the hash the name was exposed with.
HashMismatch { expected: u64, got: u64 },
/// Exposed, but no live holder right now (unbound, or its holder died
/// and the binding is being pruned).
Unresolved,
/// Resolved, but the delivery half refused it (decode failure, or the
/// holder's channel does not accept the exposed type — a local
/// re-registration under a different type; never a misroute).
Refused(DeliverError),
}
impl InboundVerdict {
/// A short static label for tracing/counting (`smarm-trace` records one
/// `ClusterInbound` event per frame with it).
pub fn label(&self) -> &'static str {
match self {
InboundVerdict::Delivered => "delivered",
InboundVerdict::NotExposed => "not_exposed",
InboundVerdict::HashMismatch { .. } => "hash_mismatch",
InboundVerdict::Unresolved => "unresolved",
InboundVerdict::Refused(_) => "refused",
}
}
}
/// THE inbound resolution seam: exposed-set check → hash check → registry
/// resolution → c8 delivery. See the module docs. Must run inside
/// [`run`](crate::run) — the conn actor's context.
pub fn deliver_named(name: &str, type_hash: u64, payload: &[u8]) -> InboundVerdict {
let Some(expected) = exposed_hash(name) else {
return InboundVerdict::NotExposed;
};
if expected != type_hash {
return InboundVerdict::HashMismatch {
expected,
got: type_hash,
};
}
let Some(pid) = whereis(name) else {
return InboundVerdict::Unresolved;
};
match decode_deliver(type_hash, pid, payload) {
Ok(()) => InboundVerdict::Delivered,
Err(e) => InboundVerdict::Refused(e),
}
}
// ---- monitors (c12) -----------------------------------------------------
/// Bookkeeping commands from [`monitor_remote`]/[`demonitor_remote`] to the
/// connection actor that owns the link to the target's node. The actor
/// records the registration and *then* emits the `Monitor` frame itself, so
/// a `Down` can never arrive at a table that does not yet know the id. It
/// lives in the actor (not on `RuntimeInner`) so the bookkeeping dies with
/// the connection — exactly what c13 needs to synthesize `Disconnected`.
pub(crate) enum MonCmd {
Monitor {
id: MonitorId,
target: RemotePid<Erased>,
tx: Sender<RemoteDown>,
},
Demonitor {
id: MonitorId,
},
}
/// A remotely-monitored actor's termination notice — the cluster analog of
/// [`Down`](crate::monitor::Down), with the pid in its wire form because it
/// may name any node.
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct RemoteDown {
/// The pid that was being monitored.
pub pid: RemotePid<Erased>,
/// How it went down. `Disconnected` means the *link* to its node was
/// lost (or absent) — nothing is known about the actor itself.
pub reason: RemoteDownReason,
}
enum Watch {
/// The target collapsed to this node: an ordinary local monitor,
/// translated on read.
Local(Monitor),
/// The target is elsewhere: the connection actor for its node holds the
/// registration and forwards the peer's `Down` frame here. The
/// `RemoteState` is the read-side backstop (c13): a channel that closes
/// while `Live` — the connection died with our command unread — reads
/// as `Disconnected` once; afterwards, and after a cancel, closed is
/// just closed.
Remote(Receiver<RemoteDown>, Cell<RemoteState>),
}
/// Where a remote-watch stands from the reader's side.
#[derive(Clone, Copy, PartialEq, Eq)]
enum RemoteState {
/// No notice yet, not cancelled: a closed channel means `Disconnected`.
Live,
/// The one notice has been read (or synthesized): nothing more is due.
Done,
/// `demonitor_remote` ran: never synthesize.
Cancelled,
}
/// A live remote monitor: read its one [`RemoteDown`] with
/// [`recv`](RemoteMonitor::recv)/[`try_recv`](RemoteMonitor::try_recv), or
/// fold it into a `select` via [`arm`](RemoteMonitor::arm). Distinct from
/// [`Monitor`] on purpose: its target is a [`RemotePid`], its notice a
/// [`RemoteDown`], and it can report `Disconnected` — none of which a local
/// monitor can express. Dropping it discards an unread notice, like the
/// local one; after [`demonitor_remote`] the channel is closed and empty, so
/// `recv` errs rather than parking — also like the local one.
///
/// Exactly one notice is guaranteed even if the connection actor dies with
/// the registration unread (the c13 drain gap): a channel that closes
/// before any notice — and before any cancel — reads as `Disconnected`,
/// once. The next read is the ordinary closed-channel `Err`.
pub struct RemoteMonitor {
/// This registration's process-unique id — minted here, echoed by the
/// peer in its `Down` frame.
pub id: MonitorId,
/// The pid being monitored.
pub target: RemotePid<Erased>,
watch: Watch,
}
impl RemoteMonitor {
/// Block (cooperatively) for the notice.
pub fn recv(&self) -> Result<RemoteDown, RecvError> {
match &self.watch {
Watch::Local(m) => m.rx.recv().map(|d| RemoteDown {
pid: self.target.clone(),
reason: d.reason.into(),
}),
Watch::Remote(rx, st) => match rx.recv() {
Ok(d) => {
st.set(RemoteState::Done);
Ok(d)
}
Err(e) => self.closed(st).ok_or(e),
},
}
}
/// The notice if it has arrived; `Ok(None)` if not yet.
pub fn try_recv(&self) -> Result<Option<RemoteDown>, RecvError> {
match &self.watch {
Watch::Local(m) => m.rx.try_recv().map(|o| {
o.map(|d| RemoteDown {
pid: self.target.clone(),
reason: d.reason.into(),
})
}),
Watch::Remote(rx, st) => match rx.try_recv() {
Ok(Some(d)) => {
st.set(RemoteState::Done);
Ok(Some(d))
}
Ok(None) => Ok(None),
Err(e) => self.closed(st).map(Some).ok_or(e),
},
}
}
/// The channel closed. While `Live` — no notice yet, no cancel — that
/// is the connection having died with our registration unread, so
/// synthesize the one `Disconnected` and mark `Done`; otherwise closed
/// is just closed.
fn closed(&self, st: &Cell<RemoteState>) -> Option<RemoteDown> {
if st.get() != RemoteState::Live {
return None;
}
st.set(RemoteState::Done);
Some(RemoteDown {
pid: self.target.clone(),
reason: RemoteDownReason::Disconnected,
})
}
/// The selectable arm: readiness means [`try_recv`](Self::try_recv)
/// will yield the notice.
pub fn arm(&self) -> &dyn Selectable {
match &self.watch {
Watch::Local(m) => &m.rx,
Watch::Remote(rx, _) => rx,
}
}
}
impl std::fmt::Debug for RemoteMonitor {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
f.debug_struct("RemoteMonitor")
.field("id", &self.id)
.field("target", &self.target)
.finish_non_exhaustive()
}
}
/// Monitor `target`, wherever it lives. Exactly one [`RemoteDown`] arrives:
///
/// - self-node pid ⇒ an ordinary local monitor underneath (same NoProc rule);
/// - no live connection to the pid's node ⇒ `Disconnected`, queued at once
/// (the remote analog of NoProc: nothing can be known);
/// - the pid's incarnation is not the node's current one ⇒ `NoProc`, queued
/// at once — the node is *known* to have restarted, so its actor is a
/// corpse, not a partition (RFC v2 §3);
/// - otherwise the connection actor registers the id and sends `Monitor`;
/// the peer answers with the true terminal reason on exit, or immediately
/// with the recorded reason for a corpse (`terminal_reason`, RFC §6) or
/// `NoProc` for a pid it never exposed and never shipped.
///
/// The connection dropping while the monitor is outstanding delivers
/// `Disconnected` (c13): the connection actor synthesizes it on teardown,
/// and the monitor's own read path backstops the case where the actor died
/// with the registration still unread. Must run inside [`run`](crate::run).
pub fn monitor_remote<A>(target: RemotePid<A>) -> RemoteMonitor {
if let Some(local) = target.local() {
let m = monitor(local);
return RemoteMonitor {
id: m.id,
target: target.erase(),
watch: Watch::Local(m),
};
}
let target = target.erase();
let (id, route) = with_runtime(|inner| {
let id = inner.alloc_monitor_id();
let route = inner
.outbound
.lock()
.by_node
.get(&target.node)
.map(|r| (r.incarnation, r.monitors.clone()));
(id, route)
});
let (tx, rx) = channel::<RemoteDown>();
let immediate = match route {
None => Some(RemoteDownReason::Disconnected),
Some((current, _)) if current != target.incarnation => {
Some(RemoteDownReason::Local(crate::monitor::DownReason::NoProc))
}
Some((_, mon_tx)) => {
let cmd = MonCmd::Monitor {
id,
target: target.clone(),
tx: tx.clone(),
};
match mon_tx.send(cmd) {
Ok(()) => None,
Err(_) => Some(RemoteDownReason::Disconnected), // actor already gone
}
}
};
if let Some(reason) = immediate {
let _ = tx.send(RemoteDown {
pid: target.clone(),
reason,
});
}
RemoteMonitor {
id,
target,
watch: Watch::Remote(rx, Cell::new(RemoteState::Live)),
}
}
/// Cancel `m`. No future notice will be *sent* for it; a notice already in
/// flight from the peer is dropped on arrival, and one already sitting in
/// `m` is discarded when `m` is dropped (same contract as
/// [`demonitor`]). Unlike the local form this returns nothing: the
/// registration is owned by the connection actor, so whether the `Down`
/// beat the cancel is not local knowledge. Must run inside
/// [`run`](crate::run).
pub fn demonitor_remote(m: &RemoteMonitor) {
match &m.watch {
Watch::Local(local) => {
let _ = demonitor(local);
}
Watch::Remote(_, st) => {
// Cancel first: a channel closing after this is closed, not a
// Disconnected notice — the caller asked for silence.
st.set(RemoteState::Cancelled);
let mon_tx = with_runtime(|inner| {
inner
.outbound
.lock()
.by_node
.get(&m.target.node)
.map(|r| r.monitors.clone())
});
if let Some(mon_tx) = mon_tx {
let _ = mon_tx.send(MonCmd::Demonitor { id: m.id });
}
}
}
}
+298
View File
@@ -0,0 +1,298 @@
//! RFC 010 c3 — transport abstraction for the **control** connection.
//!
//! Scope, per RFC 010 v2 §5 and D2:
//!
//! - A "connection" here is the *control* connection: the one carrying this
//! RFC's frame inventory ([`crate::cluster::envelope::Frame`]), whose
//! heartbeats feed failure detection. The trait deliberately says nothing
//! about how many connections a peer pair may hold — the jarred rkyv bulk
//! plane opens **additional per-peer connections** outside this trait, and
//! nothing here may foreclose that.
//! - Homogeneous smarm⇄smarm only. The BEAM membrane is *not* a transport
//! impl and the trait does not accommodate it (D2).
//! - Addresses are opaque, **pre-resolved** strings. Name resolution is a
//! single separate seam (roadmap c9); impls reject unresolved names rather
//! than resolving them.
//!
//! Blocking model: [`Conn`] calls block the caller. The TCP impl parks the
//! calling *actor* (fd readiness via the scheduler); the loopback impl blocks
//! the calling *OS thread* and is a test transport — do not drive it from a
//! scheduler thread.
//!
//! Framing is not part of the trait: [`FramedConn`] is the single shared
//! codec that turns any byte-stream [`Conn`] into a frame pipe, feeding
//! [`Frame::decode`]'s incremental contract. Impls never re-implement
//! framing, and the conformance suite exercises the same codec over every
//! impl.
use std::io;
use crate::cluster::envelope::{DecodeError, EncodeError, Frame};
pub mod loopback;
pub mod tcp;
/// An established control connection: a bidirectional byte stream.
pub trait Conn: Send {
/// Read at least one byte, blocking the caller until data is available,
/// EOF, or error. `Ok(0)` means EOF: the peer closed and all bytes it
/// wrote before closing have been consumed.
fn read(&mut self, buf: &mut [u8]) -> io::Result<usize>;
/// Write the whole buffer, blocking the caller as needed.
fn write_all(&mut self, buf: &[u8]) -> io::Result<()>;
/// Close both directions. Idempotent. Bytes already written remain
/// readable at the peer, which then observes EOF; peer writes after this
/// fail.
fn close(&mut self);
/// Diagnostic label for logs only. Mesh identity comes from the
/// handshake (`Hello`/`HelloAck`), never from the transport.
fn peer_addr(&self) -> String;
/// Readiness as a [`select`](crate::select) arm, for transports backed by
/// a file descriptor. `Some` lets a driver wait on "this connection is
/// readable" alongside an ordinary command inbox in a single `select`, so
/// one actor can interleave reading with control messages without a
/// second thread. The default is `None`: a transport with no fd (the
/// in-memory loopback) cannot be selected on and must be driven another
/// way.
fn readable_arm(&self) -> Option<crate::scheduler::FdArm> {
None
}
}
/// A bound listen point producing inbound [`Conn`]s.
pub trait Listener: Send {
/// Accept the next inbound connection, blocking the caller.
fn accept(&mut self) -> io::Result<Box<dyn Conn>>;
/// The concrete bound address, dialable as-is (e.g. the real port when
/// bound with port 0).
fn local_addr(&self) -> String;
/// Readiness as a [`select`](crate::select) arm, mirroring
/// [`Conn::readable_arm`]: `Some` lets an acceptor wait on "an inbound
/// connection is pending" alongside a command inbox in one `select`, so
/// it can be told to stop without a poll loop. Default `None` (the
/// loopback listener has no fd and must be driven synchronously).
fn readable_arm(&self) -> Option<crate::scheduler::FdArm> {
None
}
}
/// A way of establishing control connections. Object-safe on purpose: the
/// connector and membership layers hold `&dyn Transport` / boxed conns
/// rather than growing a generic parameter.
pub trait Transport: Send + Sync {
/// Connect to a peer's listen address. Blocks the caller until
/// established or failed.
fn dial(&self, addr: &str) -> io::Result<Box<dyn Conn>>;
/// Bind a listen point.
fn listen(&self, addr: &str) -> io::Result<Box<dyn Listener>>;
}
impl std::fmt::Debug for dyn Conn {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
write!(f, "Conn({})", self.peer_addr())
}
}
impl std::fmt::Debug for dyn Listener {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
write!(f, "Listener({})", self.local_addr())
}
}
/// Error surface of [`FramedConn::send`].
#[derive(Debug)]
pub enum SendError {
/// The frame could not be encoded (e.g. a field over its wire limit).
Encode(EncodeError),
/// The transport failed mid-write.
Io(io::Error),
}
impl std::fmt::Display for SendError {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
match self {
SendError::Encode(e) => write!(f, "frame encode failed: {e:?}"),
SendError::Io(e) => write!(f, "transport write failed: {e}"),
}
}
}
impl std::error::Error for SendError {}
/// Error surface of [`FramedConn::recv`].
#[derive(Debug)]
pub enum RecvError {
/// The byte stream is not a valid frame stream (bad tag, lying length,
/// oversized frame, …). The connection is unusable.
Corrupt(DecodeError),
/// The peer closed mid-frame: EOF arrived with a partial frame buffered.
/// Distinct from a clean close, which is `Ok(None)`.
TruncatedByPeer,
/// The transport failed mid-read.
Io(io::Error),
/// The deadline passed before a full frame arrived
/// ([`FramedConn::recv_deadline`] only; plain [`recv`](FramedConn::recv)
/// never returns this).
TimedOut,
}
impl std::fmt::Display for RecvError {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
match self {
RecvError::Corrupt(e) => write!(f, "frame stream corrupt: {e:?}"),
RecvError::TruncatedByPeer => write!(f, "peer closed mid-frame"),
RecvError::Io(e) => write!(f, "transport read failed: {e}"),
RecvError::TimedOut => write!(f, "deadline passed mid-receive"),
}
}
}
impl std::error::Error for RecvError {}
/// How many bytes each blocking read asks the transport for.
const READ_CHUNK: usize = 8 * 1024;
/// The shared framed codec: one of these per control connection, owning the
/// [`Conn`] and the reassembly buffer. Frames may arrive split or coalesced
/// arbitrarily; [`recv`](FramedConn::recv) reassembles either way.
pub struct FramedConn {
conn: Box<dyn Conn>,
rbuf: Vec<u8>,
}
impl FramedConn {
pub fn new(conn: Box<dyn Conn>) -> Self {
FramedConn {
conn,
rbuf: Vec::new(),
}
}
/// Encode and write one frame.
pub fn send(&mut self, frame: &Frame) -> Result<(), SendError> {
let mut out = Vec::new();
frame.encode(&mut out).map_err(SendError::Encode)?;
self.conn.write_all(&out).map_err(SendError::Io)
}
/// Receive the next frame. `Ok(None)` is a clean close: EOF at a frame
/// boundary. EOF mid-frame is [`RecvError::TruncatedByPeer`].
pub fn recv(&mut self) -> Result<Option<Frame>, RecvError> {
loop {
match Frame::decode(&self.rbuf) {
Ok(Some((frame, consumed))) => {
self.rbuf.drain(..consumed);
return Ok(Some(frame));
}
Ok(None) => {}
Err(e) => return Err(RecvError::Corrupt(e)),
}
let mut chunk = [0u8; READ_CHUNK];
let n = self.conn.read(&mut chunk).map_err(RecvError::Io)?;
if n == 0 {
return if self.rbuf.is_empty() {
Ok(None)
} else {
Err(RecvError::TruncatedByPeer)
};
}
self.rbuf.extend_from_slice(&chunk[..n]);
}
}
/// Like [`recv`](FramedConn::recv), but gives up with
/// [`RecvError::TimedOut`] once `deadline` passes without a full frame.
/// The deadline is enforced between reads via the connection's fd arm
/// (so the caller must be an actor); a transport with no fd (loopback)
/// cannot be timed out and this degrades to a plain blocking `recv` —
/// the same caveat as liveness.
pub fn recv_deadline(
&mut self,
deadline: std::time::Instant,
) -> Result<Option<Frame>, RecvError> {
loop {
match Frame::decode(&self.rbuf) {
Ok(Some((frame, consumed))) => {
self.rbuf.drain(..consumed);
return Ok(Some(frame));
}
Ok(None) => {}
Err(e) => return Err(RecvError::Corrupt(e)),
}
if let Some(arm) = self.conn.readable_arm() {
let left = deadline.saturating_duration_since(std::time::Instant::now());
if left.is_zero() {
return Err(RecvError::TimedOut);
}
match crate::channel::try_select_timeout(&[&arm], left) {
Ok(Some(_)) => {}
Ok(None) => return Err(RecvError::TimedOut),
Err(e) => return Err(RecvError::Io(e)),
}
}
let mut chunk = [0u8; READ_CHUNK];
let n = self.conn.read(&mut chunk).map_err(RecvError::Io)?;
if n == 0 {
return if self.rbuf.is_empty() {
Ok(None)
} else {
Err(RecvError::TruncatedByPeer)
};
}
self.rbuf.extend_from_slice(&chunk[..n]);
}
}
/// One socket read, appended to the reassembly buffer. Returns the byte
/// count (`0` = EOF). For select-loop callers that were just told the fd
/// is readable: under the level-triggered IO thread exactly one read per
/// readable wake never blocks and never loses data — leftover socket
/// bytes re-signal on the next select, and complete frames already
/// reassembled are drained with [`next_buffered`](FramedConn::next_buffered).
/// (A plain [`recv`](FramedConn::recv) can block into the socket while
/// the buffer holds a partial frame, which a loop with deadlines to keep
/// cannot afford.)
pub fn read_once(&mut self) -> std::io::Result<usize> {
let mut chunk = [0u8; READ_CHUNK];
let n = self.conn.read(&mut chunk)?;
self.rbuf.extend_from_slice(&chunk[..n]);
Ok(n)
}
/// Decode the next complete frame already sitting in the reassembly
/// buffer, without touching the socket. `Ok(None)` means the buffer
/// holds no complete frame (empty, or a partial awaiting more bytes).
pub fn next_buffered(&mut self) -> Result<Option<Frame>, DecodeError> {
match Frame::decode(&self.rbuf) {
Ok(Some((frame, consumed))) => {
self.rbuf.drain(..consumed);
Ok(Some(frame))
}
Ok(None) => Ok(None),
Err(e) => Err(e),
}
}
/// Close the underlying connection (idempotent, see [`Conn::close`]).
pub fn close(&mut self) {
self.conn.close();
}
/// Diagnostic label of the underlying connection.
pub fn peer_addr(&self) -> String {
self.conn.peer_addr()
}
/// The underlying connection's readiness arm, if it is fd-backed (see
/// [`Conn::readable_arm`]).
pub fn readable_arm(&self) -> Option<crate::scheduler::FdArm> {
self.conn.readable_arm()
}
}
+274
View File
@@ -0,0 +1,274 @@
//! In-memory loopback transport — a shipped **test** transport.
//!
//! Lets Phases 2–4 exercise protocol logic (connector, membership,
//! monitors) through the real transport trait and the real framed codec
//! without sockets or timing flake.
//!
//! Blocking model: calls block the **OS thread** on a condvar. That is the
//! right shape for plain `#[test]`s driving protocol state machines; it is
//! the wrong shape for scheduler threads. Do not drive a loopback conn from
//! inside an actor — use the TCP impl there.
//!
//! Semantics mirror TCP shutdown where it matters for the codec: bytes
//! written before `close` remain readable at the peer, which then sees EOF;
//! writes toward a closed peer fail with `BrokenPipe`. Write buffers are
//! unbounded, so writes never block — backpressure is not simulated.
use std::collections::{HashMap, VecDeque};
use std::io;
use std::sync::{Arc, Condvar, Mutex, MutexGuard};
use super::{Conn, Listener, Transport};
/// Poison-tolerant lock: a panicked holder in a *test* transport must not
/// cascade; the byte-queue state stays consistent under every early return.
fn lock<T>(m: &Mutex<T>) -> MutexGuard<'_, T> {
match m.lock() {
Ok(g) => g,
Err(poisoned) => poisoned.into_inner(),
}
}
// ---------------------------------------------------------------------------
// One direction of a duplex: a byte queue with close flags for both ends
// ---------------------------------------------------------------------------
#[derive(Default)]
struct PipeState {
bytes: VecDeque<u8>,
/// The writing end closed: readers drain remaining bytes, then EOF.
write_closed: bool,
/// The reading end closed: writers fail with `BrokenPipe`.
read_closed: bool,
}
#[derive(Default)]
struct Pipe {
state: Mutex<PipeState>,
cv: Condvar,
}
impl Pipe {
fn write_all(&self, buf: &[u8]) -> io::Result<()> {
let mut st = lock(&self.state);
if st.write_closed {
return Err(io::Error::new(
io::ErrorKind::NotConnected,
"loopback conn closed locally",
));
}
if st.read_closed {
return Err(io::Error::new(
io::ErrorKind::BrokenPipe,
"loopback peer closed",
));
}
st.bytes.extend(buf);
self.cv.notify_all();
Ok(())
}
fn read(&self, buf: &mut [u8]) -> io::Result<usize> {
if buf.is_empty() {
return Ok(0);
}
let mut st = lock(&self.state);
loop {
if !st.bytes.is_empty() {
let n = st.bytes.len().min(buf.len());
for (slot, byte) in buf.iter_mut().zip(st.bytes.drain(..n)) {
*slot = byte;
}
return Ok(n);
}
if st.write_closed || st.read_closed {
return Ok(0); // EOF: peer closed, or our own end closed.
}
st = match self.cv.wait(st) {
Ok(g) => g,
Err(poisoned) => poisoned.into_inner(),
};
}
}
/// Close from the writer side: remaining bytes stay readable, then EOF.
fn close_write(&self) {
lock(&self.state).write_closed = true;
self.cv.notify_all();
}
/// Close from the reader side: peer writes fail from now on.
fn close_read(&self) {
lock(&self.state).read_closed = true;
self.cv.notify_all();
}
}
// ---------------------------------------------------------------------------
// Conn: two pipes, one per direction
// ---------------------------------------------------------------------------
/// One end of an established loopback connection.
pub struct LoopbackConn {
tx: Arc<Pipe>,
rx: Arc<Pipe>,
peer: String,
}
impl Conn for LoopbackConn {
fn read(&mut self, buf: &mut [u8]) -> io::Result<usize> {
self.rx.read(buf)
}
fn write_all(&mut self, buf: &[u8]) -> io::Result<()> {
self.tx.write_all(buf)
}
fn close(&mut self) {
self.tx.close_write();
self.rx.close_read();
}
fn peer_addr(&self) -> String {
self.peer.clone()
}
}
impl Drop for LoopbackConn {
fn drop(&mut self) {
self.close();
}
}
fn conn_pair(listen_addr: &str, conn_no: u64) -> (LoopbackConn, LoopbackConn) {
let a_to_b = Arc::new(Pipe::default());
let b_to_a = Arc::new(Pipe::default());
let dialer = LoopbackConn {
tx: a_to_b.clone(),
rx: b_to_a.clone(),
peer: listen_addr.to_string(),
};
let accepted = LoopbackConn {
tx: b_to_a,
rx: a_to_b,
peer: format!("{listen_addr}#dialer-{conn_no}"),
};
(dialer, accepted)
}
// ---------------------------------------------------------------------------
// Listener + registry
// ---------------------------------------------------------------------------
#[derive(Default)]
struct AcceptState {
pending: VecDeque<LoopbackConn>,
closed: bool,
}
#[derive(Default)]
struct AcceptQueue {
state: Mutex<AcceptState>,
cv: Condvar,
}
/// A bound loopback listen point.
pub struct LoopbackListener {
addr: String,
queue: Arc<AcceptQueue>,
registry: Arc<Mutex<Registry>>,
}
impl Listener for LoopbackListener {
fn accept(&mut self) -> io::Result<Box<dyn Conn>> {
let mut st = lock(&self.queue.state);
loop {
if let Some(conn) = st.pending.pop_front() {
return Ok(Box::new(conn));
}
if st.closed {
return Err(io::Error::new(
io::ErrorKind::NotConnected,
"loopback listener closed",
));
}
st = match self.queue.cv.wait(st) {
Ok(g) => g,
Err(poisoned) => poisoned.into_inner(),
};
}
}
fn local_addr(&self) -> String {
self.addr.clone()
}
}
impl Drop for LoopbackListener {
fn drop(&mut self) {
lock(&self.registry).listeners.remove(&self.addr);
let mut st = lock(&self.queue.state);
st.closed = true;
self.queue.cv.notify_all();
}
}
#[derive(Default)]
struct Registry {
listeners: HashMap<String, Arc<AcceptQueue>>,
dial_count: u64,
}
/// The loopback transport. Addresses are arbitrary strings scoped to one
/// transport instance; distinct instances never see each other's listeners.
#[derive(Default)]
pub struct LoopbackTransport {
registry: Arc<Mutex<Registry>>,
}
impl Transport for LoopbackTransport {
fn dial(&self, addr: &str) -> io::Result<Box<dyn Conn>> {
let (queue, conn_no) = {
let mut reg = lock(&self.registry);
reg.dial_count += 1;
let no = reg.dial_count;
match reg.listeners.get(addr) {
Some(q) => (q.clone(), no),
None => {
return Err(io::Error::new(
io::ErrorKind::ConnectionRefused,
format!("no loopback listener at {addr:?}"),
));
}
}
};
let (dialer, accepted) = conn_pair(addr, conn_no);
let mut st = lock(&queue.state);
if st.closed {
return Err(io::Error::new(
io::ErrorKind::ConnectionRefused,
format!("loopback listener at {addr:?} closed"),
));
}
st.pending.push_back(accepted);
queue.cv.notify_all();
Ok(Box::new(dialer))
}
fn listen(&self, addr: &str) -> io::Result<Box<dyn Listener>> {
let queue = Arc::new(AcceptQueue::default());
let mut reg = lock(&self.registry);
if reg.listeners.contains_key(addr) {
return Err(io::Error::new(
io::ErrorKind::AddrInUse,
format!("loopback listener already bound at {addr:?}"),
));
}
reg.listeners.insert(addr.to_string(), queue.clone());
Ok(Box::new(LoopbackListener {
addr: addr.to_string(),
queue,
registry: self.registry.clone(),
}))
}
}
+285
View File
@@ -0,0 +1,285 @@
//! TCP transport — the production control-plane transport.
//!
//! Blocking model: every blocking point parks the **calling actor** on fd
//! readiness ([`crate::scheduler::wait_readable`] / `wait_writable`); the
//! scheduler thread is never blocked. All conn/listener methods must
//! therefore run inside an actor. `listen` itself only binds (no waiting)
//! and is callable anywhere.
//!
//! Addresses are pre-resolved `ip:port` strings (`SocketAddr` syntax, IPv4
//! or IPv6). Hostnames are rejected with `InvalidInput`: name resolution is
//! the single c9 seam, not something each transport does on the side.
//!
//! Writes use `send(2)` with `MSG_NOSIGNAL` — a peer reset must surface as
//! `BrokenPipe`/`ConnectionReset`, not `SIGPIPE`.
use std::io;
use std::net::{SocketAddr, TcpListener as StdListener, TcpStream};
use std::os::fd::{AsRawFd, RawFd};
use crate::scheduler::{wait_readable, wait_writable};
use super::{Conn, Listener, Transport};
// ---------------------------------------------------------------------------
// sockaddr plumbing
// ---------------------------------------------------------------------------
/// A `sockaddr_in`/`sockaddr_in6` built from a parsed `SocketAddr`, plus its
/// length, ready for `connect(2)`.
union SockAddrUnion {
v4: libc::sockaddr_in,
v6: libc::sockaddr_in6,
}
fn to_sockaddr(sa: &SocketAddr) -> (SockAddrUnion, libc::socklen_t) {
match sa {
SocketAddr::V4(v4) => {
let raw = libc::sockaddr_in {
sin_family: libc::AF_INET as libc::sa_family_t,
sin_port: v4.port().to_be(),
sin_addr: libc::in_addr {
s_addr: u32::from_be_bytes(v4.ip().octets()).to_be(),
},
sin_zero: [0; 8],
};
(
SockAddrUnion { v4: raw },
std::mem::size_of::<libc::sockaddr_in>() as libc::socklen_t,
)
}
SocketAddr::V6(v6) => {
let raw = libc::sockaddr_in6 {
sin6_family: libc::AF_INET6 as libc::sa_family_t,
sin6_port: v6.port().to_be(),
sin6_flowinfo: v6.flowinfo(),
sin6_addr: libc::in6_addr {
s6_addr: v6.ip().octets(),
},
sin6_scope_id: v6.scope_id(),
};
(
SockAddrUnion { v6: raw },
std::mem::size_of::<libc::sockaddr_in6>() as libc::socklen_t,
)
}
}
}
fn parse_addr(addr: &str) -> io::Result<SocketAddr> {
addr.parse().map_err(|_| {
io::Error::new(
io::ErrorKind::InvalidInput,
format!("{addr:?} is not a resolved ip:port — resolution is the c9 seam"),
)
})
}
fn so_error(fd: RawFd) -> io::Result<()> {
let mut err: libc::c_int = 0;
let mut len = std::mem::size_of::<libc::c_int>() as libc::socklen_t;
let rc = unsafe {
libc::getsockopt(
fd,
libc::SOL_SOCKET,
libc::SO_ERROR,
(&mut err) as *mut _ as *mut libc::c_void,
&mut len,
)
};
if rc != 0 {
return Err(io::Error::last_os_error());
}
if err != 0 {
return Err(io::Error::from_raw_os_error(err));
}
Ok(())
}
// ---------------------------------------------------------------------------
// Conn
// ---------------------------------------------------------------------------
/// One established TCP control connection. Owns the socket; drop closes it.
pub struct TcpConn {
stream: TcpStream,
closed: bool,
}
impl TcpConn {
fn fd(&self) -> RawFd {
self.stream.as_raw_fd()
}
}
impl Conn for TcpConn {
fn read(&mut self, buf: &mut [u8]) -> io::Result<usize> {
if self.closed {
return Ok(0);
}
if buf.is_empty() {
return Ok(0);
}
loop {
wait_readable(self.fd())?;
let n = unsafe { libc::read(self.fd(), buf.as_mut_ptr() as *mut _, buf.len()) };
if n >= 0 {
return Ok(n as usize);
}
let e = io::Error::last_os_error();
match e.kind() {
// Spurious readiness or signal: park again.
io::ErrorKind::WouldBlock | io::ErrorKind::Interrupted => continue,
_ => return Err(e),
}
}
}
fn write_all(&mut self, mut buf: &[u8]) -> io::Result<()> {
if self.closed {
return Err(io::Error::new(
io::ErrorKind::NotConnected,
"tcp conn closed locally",
));
}
while !buf.is_empty() {
wait_writable(self.fd())?;
let n = unsafe {
libc::send(
self.fd(),
buf.as_ptr() as *const _,
buf.len(),
libc::MSG_NOSIGNAL,
)
};
if n >= 0 {
buf = &buf[n as usize..];
continue;
}
let e = io::Error::last_os_error();
match e.kind() {
io::ErrorKind::WouldBlock | io::ErrorKind::Interrupted => continue,
_ => return Err(e),
}
}
Ok(())
}
fn close(&mut self) {
if !self.closed {
self.closed = true;
// Best-effort: the peer sees EOF after draining. The fd itself
// is released when the owning stream drops.
let _ = self.stream.shutdown(std::net::Shutdown::Both);
}
}
fn peer_addr(&self) -> String {
match self.stream.peer_addr() {
Ok(sa) => sa.to_string(),
Err(_) => "<disconnected>".to_string(),
}
}
fn readable_arm(&self) -> Option<crate::scheduler::FdArm> {
Some(crate::scheduler::FdArm::readable(self.fd()))
}
}
// ---------------------------------------------------------------------------
// Listener
// ---------------------------------------------------------------------------
/// A bound TCP listen point (non-blocking socket; accept parks the actor).
pub struct TcpListener {
inner: StdListener,
local: SocketAddr,
}
impl Listener for TcpListener {
fn readable_arm(&self) -> Option<crate::scheduler::FdArm> {
Some(crate::scheduler::FdArm::readable(self.inner.as_raw_fd()))
}
fn accept(&mut self) -> io::Result<Box<dyn Conn>> {
loop {
wait_readable(self.inner.as_raw_fd())?;
match self.inner.accept() {
Ok((stream, _peer)) => {
stream.set_nonblocking(true)?;
return Ok(Box::new(TcpConn {
stream,
closed: false,
}));
}
Err(e)
if e.kind() == io::ErrorKind::WouldBlock
|| e.kind() == io::ErrorKind::Interrupted =>
{
continue;
}
Err(e) => return Err(e),
}
}
}
fn local_addr(&self) -> String {
self.local.to_string()
}
}
// ---------------------------------------------------------------------------
// Transport
// ---------------------------------------------------------------------------
/// The TCP transport. Stateless; every call stands alone.
pub struct TcpTransport;
impl Transport for TcpTransport {
fn dial(&self, addr: &str) -> io::Result<Box<dyn Conn>> {
let sa = parse_addr(addr)?;
let family = match sa {
SocketAddr::V4(_) => libc::AF_INET,
SocketAddr::V6(_) => libc::AF_INET6,
};
let fd = unsafe {
libc::socket(
family,
libc::SOCK_STREAM | libc::SOCK_NONBLOCK | libc::SOCK_CLOEXEC,
0,
)
};
if fd < 0 {
return Err(io::Error::last_os_error());
}
// From here the fd is owned by `stream`; any early return drops it.
let stream = unsafe {
use std::os::fd::FromRawFd;
TcpStream::from_raw_fd(fd)
};
let (raw, len) = to_sockaddr(&sa);
let rc = unsafe { libc::connect(fd, (&raw) as *const _ as *const libc::sockaddr, len) };
if rc != 0 {
let e = io::Error::last_os_error();
if e.raw_os_error() != Some(libc::EINPROGRESS) {
return Err(e);
}
// Connect in flight: park until the socket is writable, then the
// verdict is in SO_ERROR.
wait_writable(fd)?;
so_error(fd)?;
}
Ok(Box::new(TcpConn {
stream,
closed: false,
}))
}
fn listen(&self, addr: &str) -> io::Result<Box<dyn Listener>> {
let sa = parse_addr(addr)?;
let inner = StdListener::bind(sa)?;
inner.set_nonblocking(true)?;
let local = inner.local_addr()?;
Ok(Box::new(TcpListener { inner, local }))
}
}
+82 -34
View File
@@ -4,30 +4,66 @@
//! actor running on its own mmap'd stack. The compiler cannot do this; the //! actor running on its own mmap'd stack. The compiler cannot do this; the
//! whole point of `#[unsafe(naked)]` is that we control every instruction. //! whole point of `#[unsafe(naked)]` is that we control every instruction.
//! //!
//! `SCHEDULER_SP` and `ACTOR_SP` are thread-locals holding each side's saved //! The actor's stack pointer travels in registers: `switch_to_actor` takes
//! stack pointer. `init_actor_stack` builds the initial stack so that the //! the target sp as its argument and returns the sp the actor saved when it
//! first `switch_to_actor` lands inside the entry function with `rsp % 16 == 8` //! next yielded (handed back in `rax` by `switch_to_scheduler`'s shim). Only
//! (the x86-64 ABI requirement at function entry). //! the *scheduler* sp lives in a thread-local — the yielding actor sits at
//! arbitrary call depth with no argument channel back to the scheduler, so
//! TLS is the one place it can find the way home. `init_actor_stack` builds
//! the initial stack so that the first `switch_to_actor` lands inside the
//! entry function with `rsp % 16 == 8` (the x86-64 ABI requirement at
//! function entry).
//!
//! # Thread-locals and migration (read before touching any `thread_local!`)
//!
//! An actor may park on scheduler thread A and be resumed on thread B. LLVM
//! treats the address of a thread-local as a loop-invariant, side-effect-free
//! value: it computes `%fs:0 + offset` once per function and happily keeps it
//! in a callee-saved register across calls — including across
//! `switch_to_scheduler`. Any function that touches a scheduler thread-local
//! both before and after a switch (or that gets *inlined* into one that does)
//! therefore reads and writes the *old thread's* TLS after migration. Nothing
//! at the switch can prevent this: it is not a memory clobber problem, the
//! address is not memory-derived in LLVM's model. Under thin LTO the code
//! happened to use the local-exec model (`%fs:imm` operands, nothing to
//! cache) so it worked by luck; a plain `cargo build --release` of a
//! downstream crate broke multi-thread runs (`ACTOR_DONE` written to the wrong
//! thread → "scheduler resumed a done actor").
//!
//! Rule: every function that touches a thread-local and can execute on an
//! actor stack must be `#[inline(never)]` and call [`tls_fence`] first, so the
//! TLS base is recomputed inside a callee that cannot be inlined into a frame
//! spanning a switch, and so LLVM cannot infer the accessor is pure and merge
//! two calls to it. Scheduler-side code (`schedule_loop` and what it calls
//! before/after `switch_to_actor`) never migrates and is exempt. Do not return
//! `&Cell`/pointers into TLS from these accessors; return values.
//! `cargo test --profile reltest` (no LTO) is the regression oracle.
use std::cell::Cell; use std::cell::Cell;
thread_local! { /// Compiler barrier for TLS accessors, see the module docs. Emits no code; a
static SCHEDULER_SP: Cell<usize> = const { Cell::new(0) }; /// side-effecting empty asm keeps LLVM from marking the enclosing
static ACTOR_SP: Cell<usize> = const { Cell::new(0) }; /// `#[inline(never)]` accessor `memory(none)` and merging calls to it.
#[inline(always)]
pub(crate) fn tls_fence() {
// SAFETY: empty asm, no operands, no stack, no flags.
unsafe { core::arch::asm!("", options(nostack, preserves_flags)) }
} }
thread_local! {
static SCHEDULER_SP: Cell<usize> = const { Cell::new(0) };
}
#[inline(never)]
fn get_scheduler_sp() -> usize { fn get_scheduler_sp() -> usize {
tls_fence();
SCHEDULER_SP.with(|c| c.get()) SCHEDULER_SP.with(|c| c.get())
} }
#[inline(never)]
fn set_scheduler_sp(v: usize) { fn set_scheduler_sp(v: usize) {
tls_fence();
SCHEDULER_SP.with(|c| c.set(v)) SCHEDULER_SP.with(|c| c.set(v))
} }
pub fn get_actor_sp() -> usize {
ACTOR_SP.with(|c| c.get())
}
pub fn set_actor_sp(v: usize) {
ACTOR_SP.with(|c| c.set(v))
}
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// Initial stack layout // Initial stack layout
@@ -78,12 +114,24 @@ pub fn init_actor_stack(top: *mut u8, entry: extern "C-unwind" fn()) -> usize {
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// Context switch shims // Context switch shims
// //
// Each shim: // switch_to_actor_asm (rdi = target actor sp, returns rax = the sp the actor
// 1. Pushes the six callee-saved integer registers. // saved when it next yielded):
// 2. Snaps rsp into rdi and calls the Rust helper that stores it. // 1. Pushes the six callee-saved integer registers (scheduler side).
// 3. Calls the Rust helper that returns the *other* side's saved rsp. // 2. Stashes the target sp in rbx — free scratch: the register's live
// 4. Moves that into rsp. // value is on the stack we just pushed to, and the pops below load the
// 5. Pops the six registers and rets. // *other* side's values anyway — then snaps rsp into rdi and calls the
// Rust helper that stores it in SCHEDULER_SP.
// 3. Installs the target sp and pops the actor's registers; `ret` lands
// where the actor yielded (or in `entry` on first resume).
//
// switch_to_scheduler_asm (no args; its "return value" materialises on the
// OTHER stack, as switch_to_actor's rax):
// 1. Pushes the six callee-saved integer registers (actor side).
// 2. Stashes its own rsp in rbx (same free-scratch argument), asks the
// Rust helper for SCHEDULER_SP.
// 3. Installs the scheduler sp, moves the saved actor sp into rax, pops
// the scheduler's registers and rets — completing the scheduler's
// `switch_to_actor(sp)` call with the actor's new sp as its result.
// //
// XMM registers are NOT saved here. We rely on every yield happening through // XMM registers are NOT saved here. We rely on every yield happening through
// a Rust call site, which means the compiler has spilled any live XMM state // a Rust call site, which means the compiler has spilled any live XMM state
@@ -94,31 +142,32 @@ pub fn init_actor_stack(top: *mut u8, entry: extern "C-unwind" fn()) -> usize {
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
#[unsafe(naked)] #[unsafe(naked)]
unsafe extern "C" fn switch_to_actor_asm() { unsafe extern "C" fn switch_to_actor_asm(actor_sp: usize) -> usize {
core::arch::naked_asm!( core::arch::naked_asm!(
"push rbx", "push rbp", "push r12", "push r13", "push r14", "push r15", "push rbx", "push rbp", "push r12", "push r13", "push r14", "push r15",
"mov rbx, rdi",
"mov rdi, rsp", "mov rdi, rsp",
"call {set_sched_sp}", "call {set_sched_sp}",
"call {get_actor_sp}", "mov rsp, rbx",
"mov rsp, rax",
"pop r15", "pop r14", "pop r13", "pop r12", "pop rbp", "pop rbx", "pop r15", "pop r14", "pop r13", "pop r12", "pop rbp", "pop rbx",
"ret", "ret",
set_sched_sp = sym set_scheduler_sp, set_sched_sp = sym set_scheduler_sp,
get_actor_sp = sym get_actor_sp,
); );
} }
/// Resume the actor whose sp is in `ACTOR_SP`. Returns when the actor yields. /// Resume the actor whose saved stack pointer is `actor_sp`. Returns when the
/// actor yields, with the stack pointer the actor saved as it did — store it
/// back into the slot for the next resume.
/// ///
/// # Safety /// # Safety
/// ///
/// The caller must be running on a scheduler thread with a valid actor stack /// `actor_sp` must be a valid saved actor stack pointer — either from
/// pointer installed in `ACTOR_SP` — either by `init_actor_stack` (first /// `init_actor_stack` (first resume) or the value a prior `switch_to_actor`
/// resume) or by a prior `switch_to_scheduler` (subsequent resumes). Resuming /// returned for this actor (subsequent resumes). Resuming with a stale or
/// with an unset or stale `ACTOR_SP` transfers control to an arbitrary address. /// forged sp transfers control to an arbitrary address. Must not be called
/// Must not be called from within an actor (only the scheduler side may resume). /// from within an actor (only the scheduler side may resume).
pub unsafe fn switch_to_actor() { pub unsafe fn switch_to_actor(actor_sp: usize) -> usize {
unsafe { switch_to_actor_asm() }; unsafe { switch_to_actor_asm(actor_sp) }
} }
/// Yield from the running actor back to its scheduler thread. Returns when the /// Yield from the running actor back to its scheduler thread. Returns when the
@@ -135,13 +184,12 @@ pub unsafe fn switch_to_actor() {
pub unsafe extern "C" fn switch_to_scheduler() { pub unsafe extern "C" fn switch_to_scheduler() {
core::arch::naked_asm!( core::arch::naked_asm!(
"push rbx", "push rbp", "push r12", "push r13", "push r14", "push r15", "push rbx", "push rbp", "push r12", "push r13", "push r14", "push r15",
"mov rdi, rsp", "mov rbx, rsp",
"call {set_actor_sp}",
"call {get_sched_sp}", "call {get_sched_sp}",
"mov rsp, rax", "mov rsp, rax",
"mov rax, rbx",
"pop r15", "pop r14", "pop r13", "pop r12", "pop rbp", "pop rbx", "pop r15", "pop r14", "pop r13", "pop r12", "pop rbp", "pop rbx",
"ret", "ret",
set_actor_sp = sym set_actor_sp,
get_sched_sp = sym get_scheduler_sp, get_sched_sp = sym get_scheduler_sp,
); );
} }
+226 -46
View File
@@ -127,16 +127,51 @@
//! - [`GenServer::init`] runs once before the first message. Use it to start //! - [`GenServer::init`] runs once before the first message. Use it to start
//! timers or set up monitors; see the [`GenServerCtx`] it receives. //! timers or set up monitors; see the [`GenServerCtx`] it receives.
//! - [`GenServer::terminate`] runs when the server is about to exit. It fires //! - [`GenServer::terminate`] runs when the server is about to exit. It fires
//! on every exit path (all `GenServerRef`s dropped, a handler panic, or an //! on every exit path (a self-stop, a graceful shutdown, a cooperative hard
//! explicit [`GenServerRef::shutdown`]), not only on clean shutdown. Keep it //! stop, or a handler panic), not only on clean shutdown.
//! short and non-blocking: if `terminate` panics while the server is already //! Keep it non-panicking: on the panic and hard-stop paths it runs
//! unwinding from a handler panic, the process aborts. //! mid-unwind, where a second panic aborts the process and where it must
//! not block (any park re-observes the stop). Only on the graceful path
//! (see below) may it do real work.
//! //!
//! ## When the server stops //! ## When the server stops
//! //!
//! The server runs as long as at least one [`GenServerRef`] exists. When the last //! A server lives until it stops, is shut down, or is killed — as an OTP
//! one is dropped, the inbox closes and the loop exits gracefully. To stop a //! process does. A [`GenServerRef`] is an *address*: cloning and dropping it
//! server explicitly and wait for it to finish, call [`GenServerRef::shutdown`]. //! never changes the server's lifetime, and a ref that nobody holds is not a
//! leak — a forgotten server idles until the run ends, when the root-exit
//! shutdown (see [`Runtime::run`](crate::Runtime::run)) takes it down with
//! every other unsupervised actor. Anything meant to live long should be
//! supervised (see *Supervised servers* below); [`start`] / [`start_under`]
//! are for scripts, tests and short-lived helpers, and the explicit close is
//! [`GenServerRef::shutdown`].
//!
//! A server can end itself: clone a [`StopHandle`] from
//! [`GenServerCtx::stop_handle`] in `init` and call [`StopHandle::stop`] from
//! any handler — the loop breaks after the current message and exits
//! *normally* (OTP's `{stop, normal}`). This is distinct from
//! `request_stop(self_pid())`, which is an abnormal `Stopped` and gets a
//! `Transient` child restarted.
//!
//! ## Graceful shutdown
//!
//! From outside, [`GenServerRef::shutdown`] (or a plain
//! [`request_shutdown`](crate::request_shutdown), which is what a supervisor
//! sends) asks the server to stop. What happens next is the server's choice:
//!
//! - By default a server does not trap exits, and the request stops it
//! outright at its next observation point — `terminate` runs mid-unwind.
//! - A server that calls [`GenServerCtx::trap_exit`] in `init` receives the
//! request as [`GenServer::handle_shutdown`]. Return
//! [`ShutdownAction::Exit`] (the default) to have the loop break and
//! `terminate` run on the normal path, where it may block; return
//! [`ShutdownAction::Continue`] to keep serving — e.g. to drain in-flight
//! work — and end the server later with a [`StopHandle`]. The supervisor's
//! [`Shutdown`](crate::supervisor::Shutdown) policy bounds how long that
//! may take before it falls back to a hard stop.
//!
//! A trapping server also receives the deaths of its linked peers as
//! [`GenServer::handle_exit`] messages instead of dying with them.
//! //!
//! If the server panics inside a handler, the panic unwinds the server thread. //! If the server panics inside a handler, the panic unwinds the server thread.
//! Any caller currently waiting in `call` sees `Err(ServerDown)`: the reply //! Any caller currently waiting in `call` sees `Err(ServerDown)`: the reply
@@ -167,9 +202,23 @@
//! Names and registration: a server can be given a static name so other //! Names and registration: a server can be given a static name so other
//! actors can reach it without holding a `GenServerRef`. Use //! actors can reach it without holding a `GenServerRef`. Use
//! [`GenServerBuilder::named`] to register on start, and the free functions //! [`GenServerBuilder::named`] to register on start, and the free functions
//! [`call`], [`cast`], and [`whereis_server`] to address it by name. Registered //! [`call`], [`cast`], and [`whereis_server`] to address it by name.
//! servers are a natural fit for supervision; see `supervisor` for how to //!
//! build a tree that restarts servers on failure. //! ## Supervised servers
//!
//! The supervised shape is [`NamedGenServerBuilder::run`]: it runs the loop
//! **inline, as the current actor**, so the closure of a
//! [`ChildSpec`](crate::supervisor::ChildSpec) *is* the server — the
//! supervisor's shutdown arrives as [`GenServer::handle_shutdown`], a restart
//! runs the factory again and re-binds the name, and the rest of the program
//! addresses it by name (a held ref would go stale on restart anyway).
//!
//! ```ignore
//! const COUNTER: GenServerName<Counter> = GenServerName::new("counter");
//! OneForOne::new().child(ChildSpec::new(Restart::Permanent, || {
//! GenServerBuilder::new(Counter::default()).named(COUNTER).run().unwrap();
//! }));
//! ```
//! //!
//! ## Limitations //! ## Limitations
//! //!
@@ -181,10 +230,12 @@
use crate::channel::{ use crate::channel::{
channel, select, select_timeout, Receiver, RecvTimeoutError, Selectable, Sender, channel, select, select_timeout, Receiver, RecvTimeoutError, Selectable, Sender,
}; };
use crate::link::ExitSignal;
use crate::monitor::DownReason;
use crate::monitor::{demonitor, monitor, Down, Monitor}; use crate::monitor::{demonitor, monitor, Down, Monitor};
use crate::pid::Pid; use crate::pid::Pid;
use crate::registry::{register_with, resolve_named_sender, RegisterError}; use crate::registry::{register_with, resolve_named_sender, RegisterError};
use crate::scheduler::{cancel_timer, request_stop, send_after_to}; use crate::scheduler::{cancel_timer, request_shutdown, send_after_to};
use crate::timer::TimerId; use crate::timer::TimerId;
use std::cell::Cell; use std::cell::Cell;
use std::collections::HashMap; use std::collections::HashMap;
@@ -250,14 +301,38 @@ pub trait GenServer: Send + 'static {
/// full idle window (set via [`GenServerCtx::idle_after`] in `init`) without /// full idle window (set via [`GenServerCtx::idle_after`] in `init`) without
/// dispatching any message. The window resets automatically after this /// dispatching any message. The window resets automatically after this
/// fires, so it acts as a steady idle detector. To shut down after one idle /// fires, so it acts as a steady idle detector. To shut down after one idle
/// period, call [`request_stop`](crate::scheduler::request_stop) here. /// period, call [`request_stop`] here.
/// Default: no-op. /// Default: no-op.
fn handle_idle(&mut self) {} fn handle_idle(&mut self) {}
/// A graceful shutdown request (a [`request_shutdown`](crate::request_shutdown)
/// reaching this server), delivered only if `init` called
/// [`GenServerCtx::trap_exit`]. Return [`ShutdownAction::Exit`] to stop
/// now (the default), or [`ShutdownAction::Continue`] to keep serving and
/// end the server later with a [`StopHandle`].
fn handle_shutdown(&mut self) -> ShutdownAction {
ShutdownAction::Exit
}
/// A linked peer's abnormal death (an [`ExitSignal`] that is not a
/// shutdown request), delivered only if `init` called
/// [`GenServerCtx::trap_exit`]. Default: drop it.
fn handle_exit(&mut self, _sig: ExitSignal) {}
/// Runs as the server actor exits, on any exit path (see module docs). /// Runs as the server actor exits, on any exit path (see module docs).
fn terminate(&mut self) {} fn terminate(&mut self) {}
} }
/// What a server does with a shutdown request; see [`GenServer::handle_shutdown`].
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum ShutdownAction {
/// Break the loop now. `terminate` runs on the normal path and may block.
Exit,
/// Keep dispatching. The server is expected to end itself with a
/// [`StopHandle`] once it is done winding down.
Continue,
}
/// What travels the server's single inbox channel: a synchronous call (with a /// What travels the server's single inbox channel: a synchronous call (with a
/// reply sender) or an asynchronous cast. Private — callers use [`GenServerRef`]. /// reply sender) or an asynchronous cast. Private — callers use [`GenServerRef`].
enum Envelope<G: GenServer> { enum Envelope<G: GenServer> {
@@ -265,9 +340,9 @@ enum Envelope<G: GenServer> {
Cast(G::Cast), Cast(G::Cast),
} }
/// A clonable handle to a running server. Cloning yields another sender to the /// A clonable handle to a running server: an *address*, not an owner. Cloning
/// same inbox; the server lives until the last `GenServerRef` is dropped, at which /// yields another sender to the same inbox; dropping refs never ends the
/// point its inbox closes and the loop exits normally. /// server (see the module docs, *When the server stops*).
pub struct GenServerRef<G: GenServer> { pub struct GenServerRef<G: GenServer> {
tx: Sender<Envelope<G>>, tx: Sender<Envelope<G>>,
pid: Pid, pid: Pid,
@@ -362,18 +437,22 @@ impl<G: GenServer> GenServerRef<G> {
/// Stop the server and block until it has fully exited. /// Stop the server and block until it has fully exited.
/// ///
/// Sends a cooperative stop signal to the server actor and waits for it to /// Asks the server to shut down (a [`request_shutdown`](crate::request_shutdown))
/// exit, so [`GenServer::terminate`] has run by the time this returns. /// and waits for it to exit, so [`GenServer::terminate`] has run by the
/// Returns immediately if the server is already gone. /// time this returns. A trapping server gets to wind down via
/// [`GenServer::handle_shutdown`]; any other is stopped outright. Returns
/// immediately if the server is already gone. Waits as long as the server
/// takes — the caller, not the server, decides whether that is acceptable;
/// a supervisor uses its child's [`Shutdown`](crate::supervisor::Shutdown)
/// policy to bound it.
/// ///
/// This is the right teardown for a server kept alive by a registered /// This is the explicit close: dropping refs never ends a server. Like
/// [`GenServerName`], where dropping every external `GenServerRef` is not enough /// all cooperative cancellation, it is best-effort:
/// to close the inbox. Like all cooperative cancellation, it is best-effort:
/// a server wedged in a tight loop with no observation point cannot be /// a server wedged in a tight loop with no observation point cannot be
/// stopped this way. Panics if called outside `Runtime::run()`. /// stopped this way. Panics if called outside `Runtime::run()`.
pub fn shutdown(&self) { pub fn shutdown(&self) {
let mon = monitor(self.pid); let mon = monitor(self.pid);
request_stop(self.pid); request_shutdown(self.pid);
// The Down lands when the server finalizes; an already-dead target makes // The Down lands when the server finalizes; an already-dead target makes
// `monitor` deliver NoProc immediately, so this never blocks forever. // `monitor` deliver NoProc immediately, so this never blocks forever.
let _ = mon.rx.recv(); let _ = mon.rx.recv();
@@ -396,6 +475,8 @@ enum Sys<G: GenServer> {
/// the payload factory, dispatches it to [`GenServer::handle_timer`], and /// the payload factory, dispatches it to [`GenServer::handle_timer`], and
/// re-arms the next tick before returning. /// re-arms the next tick before returning.
Tick(crate::timer::TimerId), Tick(crate::timer::TimerId),
/// The state asked to end the server (via [`StopHandle::stop`]).
Stop,
} }
/// The server loop's runtime hook, passed to [`GenServer::init`]. Hands out the /// The server loop's runtime hook, passed to [`GenServer::init`]. Hands out the
@@ -411,9 +492,29 @@ pub struct GenServerCtx<G: GenServer> {
/// because `init` holds only `&ctx`; not `Send`, but `GenServerCtx` is only ever /// because `init` holds only `&ctx`; not `Send`, but `GenServerCtx` is only ever
/// borrowed on the actor's own stack during `init`, never sent. /// borrowed on the actor's own stack during `init`, never sent.
idle: Cell<Option<Duration>>, idle: Cell<Option<Duration>>,
/// Whether the loop should trap exits (set via [`trap_exit`](Self::trap_exit)
/// during `init`, read by the loop after).
trap: Cell<bool>,
} }
impl<G: GenServer> GenServerCtx<G> { impl<G: GenServer> GenServerCtx<G> {
/// Trap exits for the server's lifetime: shutdown requests then arrive as
/// [`GenServer::handle_shutdown`] and linked-peer deaths as
/// [`GenServer::handle_exit`], instead of stopping the server outright.
/// Call this once during `init`.
pub fn trap_exit(&self) {
self.trap.set(true);
}
/// A clonable handle that lets the state end the server from any handler
/// (a normal exit; see the module docs). Store it on the state during
/// `init`.
pub fn stop_handle(&self) -> StopHandle<G> {
StopHandle {
sys_tx: self.sys_tx.clone(),
}
}
/// A clonable handle to the loop's monitor intake. Store it in the state /// A clonable handle to the loop's monitor intake. Store it in the state
/// during `init` to watch monitors from later handlers. /// during `init` to watch monitors from later handlers.
pub fn watcher(&self) -> Watcher<G> { pub fn watcher(&self) -> Watcher<G> {
@@ -452,6 +553,30 @@ impl<G: GenServer> GenServerCtx<G> {
} }
} }
/// Lets a server's state end the server, cloned from
/// [`GenServerCtx::stop_handle`] during `init`. [`stop`](Self::stop) makes the
/// loop break after the current message and exit normally; `terminate` runs on
/// the normal path.
pub struct StopHandle<G: GenServer> {
sys_tx: Sender<Sys<G>>,
}
impl<G: GenServer> Clone for StopHandle<G> {
fn clone(&self) -> Self {
StopHandle {
sys_tx: self.sys_tx.clone(),
}
}
}
impl<G: GenServer> StopHandle<G> {
/// End the server after the current message. Idempotent; a no-op once the
/// server is gone.
pub fn stop(&self) {
let _ = self.sys_tx.send(Sys::Stop);
}
}
/// Per-server timer bookkeeping, shared between the loop and every /// Per-server timer bookkeeping, shared between the loop and every
/// [`TimerHandle`] clone. A gen_server actor is single-threaded — handlers /// [`TimerHandle`] clone. A gen_server actor is single-threaded — handlers
/// and the loop never run concurrently — so this `Mutex` is always /// and the loop never run concurrently — so this `Mutex` is always
@@ -689,7 +814,7 @@ impl<G: GenServer> GenServerBuilder<G> {
self self
} }
/// Spawn the server under an explicit supervisor pid (via [`spawn_under`]) /// Spawn the server under an explicit supervisor pid (via [`spawn_under`](crate::scheduler::spawn_under))
/// so it slots into the supervision tree. /// so it slots into the supervision tree.
pub fn under(mut self, supervisor: Pid) -> Self { pub fn under(mut self, supervisor: Pid) -> Self {
self.supervisor = Some(supervisor); self.supervisor = Some(supervisor);
@@ -704,9 +829,10 @@ impl<G: GenServer> GenServerBuilder<G> {
self self
} }
/// Spawn the server actor and hand back its [`GenServerRef`]. The server's /// Spawn the server actor and hand back its [`GenServerRef`] (an address;
/// lifetime is governed by its refs, not by joining, so the backing join /// the server's lifetime is its own, see the module docs). The backing
/// handle is dropped. /// join handle is dropped. For a supervised server use
/// [`named`](Self::named) + [`NamedGenServerBuilder::run`] instead.
pub fn start(self) -> GenServerRef<G> { pub fn start(self) -> GenServerRef<G> {
self.spawn_server() self.spawn_server()
} }
@@ -734,13 +860,14 @@ impl<G: GenServer> GenServerBuilder<G> {
supervisor, supervisor,
stack_opts, stack_opts,
} = self; } = self;
let keep = tx.clone();
let handle = match supervisor { let handle = match supervisor {
Some(sup) => crate::scheduler::spawn_under_with(sup, stack_opts, move || { Some(sup) => crate::scheduler::spawn_under_with(sup, stack_opts, move || {
server_loop::<G>(rx, state, infos) server_loop::<G>(keep, rx, state, infos)
}),
None => crate::scheduler::spawn_with(stack_opts, move || {
server_loop::<G>(keep, rx, state, infos)
}), }),
None => {
crate::scheduler::spawn_with(stack_opts, move || server_loop::<G>(rx, state, infos))
}
}; };
GenServerRef { GenServerRef {
tx, tx,
@@ -828,19 +955,38 @@ impl<G: GenServer> NamedGenServerBuilder<G> {
/// The inbox sender is published under the name **from the parent side, /// The inbox sender is published under the name **from the parent side,
/// before this returns**, so a by-name `call` / `cast` resolves the instant /// before this returns**, so a by-name `call` / `cast` resolves the instant
/// `start()` returns — no race with the server body. On a name clash the /// `start()` returns — no race with the server body. On a name clash the
/// just-spawned server is wound down (its only ref is dropped, closing the /// just-spawned server is stopped, so a failed bind leaks no actor.
/// inbox), so a failed bind leaks no actor.
pub fn start(self) -> Result<GenServerRef<G>, RegisterError> { pub fn start(self) -> Result<GenServerRef<G>, RegisterError> {
let NamedGenServerBuilder { builder, name } = self; let NamedGenServerBuilder { builder, name } = self;
let server = builder.spawn_server(); let server = builder.spawn_server();
match register_with::<Envelope<G>>(server.pid, name, server.tx.clone()) { match register_with::<Envelope<G>>(server.pid, name, server.tx.clone()) {
Ok(()) => Ok(server), Ok(()) => Ok(server),
Err(e) => { Err(e) => {
drop(server); // inbox closes → loop exits gracefully crate::scheduler::request_stop(server.pid); // never ran init
Err(e) Err(e)
} }
} }
} }
/// Run the server **inline, as the current actor**, bound to its name.
/// This is the supervised shape: the closure of a
/// [`ChildSpec`](crate::supervisor::ChildSpec) *is* the server, so the
/// supervisor's shutdown reaches it as [`GenServer::handle_shutdown`], a
/// restart runs the factory again and re-binds the name, and clients
/// address it by name ([`call`], [`cast`], [`whereis_server`]). Returns
/// when the server exits; [`RegisterError::NameTaken`] (before `init`) if
/// the name is held by a different live actor.
///
/// `under` / `stack_opts` are spawn options and do not apply here — the
/// actor already exists.
pub fn run(self) -> Result<(), RegisterError> {
let NamedGenServerBuilder { builder, name } = self;
let GenServerBuilder { state, infos, .. } = builder;
let (tx, rx) = channel::<Envelope<G>>();
register_with::<Envelope<G>>(crate::scheduler::self_pid(), name, tx.clone())?;
server_loop::<G>(tx, rx, state, infos);
Ok(())
}
} }
/// Resolve a [`GenServerName`] to a [`GenServerRef`] when you want a handle to hold or /// Resolve a [`GenServerName`] to a [`GenServerRef`] when you want a handle to hold or
@@ -886,23 +1032,30 @@ pub fn shutdown<G: GenServer>(name: GenServerName<G>) {
} }
} }
/// Spawn `state` as a server under the current actor (via [`spawn`]). Returns a /// Spawn `state` as a server under the current actor (via [`spawn`](crate::scheduler::spawn)). Returns a
/// [`GenServerRef`]. Shorthand for `GenServerBuilder::new(state).start()`. /// [`GenServerRef`]. Shorthand for `GenServerBuilder::new(state).start()`.
pub fn start<G: GenServer>(state: G) -> GenServerRef<G> { pub fn start<G: GenServer>(state: G) -> GenServerRef<G> {
GenServerBuilder::new(state).start() GenServerBuilder::new(state).start()
} }
/// Like [`start`], but spawns the server under an explicit supervisor pid (via /// Like [`start`], but spawns the server under an explicit supervisor pid (via
/// [`spawn_under`]) so it slots into the supervision tree. /// [`spawn_under`](crate::scheduler::spawn_under)) so it slots into the supervision tree.
pub fn start_under<G: GenServer>(supervisor: Pid, state: G) -> GenServerRef<G> { pub fn start_under<G: GenServer>(supervisor: Pid, state: G) -> GenServerRef<G> {
GenServerBuilder::new(state).under(supervisor).start() GenServerBuilder::new(state).under(supervisor).start()
} }
fn server_loop<G: GenServer>( fn server_loop<G: GenServer>(
keep: Sender<Envelope<G>>,
rx: Receiver<Envelope<G>>, rx: Receiver<Envelope<G>>,
state: G, state: G,
mut infos: Vec<Receiver<G::Info>>, mut infos: Vec<Receiver<G::Info>>,
) { ) {
// The loop holds one inbox sender for its whole life: the inbox never
// closes, so refs are addresses and the server's lifetime is the actor's
// (stop handle, shutdown, stop, panic). The `Disconnected` arms below are
// defensive only.
let _keep = keep;
// Drop guard — owns the server state and the timer registry. // Drop guard — owns the server state and the timer registry.
// //
// Why a guard rather than code after the loop: // Why a guard rather than code after the loop:
@@ -974,9 +1127,14 @@ fn server_loop<G: GenServer>(
sys_tx, sys_tx,
reg: reg.clone(), reg: reg.clone(),
idle: Cell::new(None), idle: Cell::new(None),
trap: Cell::new(false),
}; };
guard.0.init(&ctx); guard.0.init(&ctx);
let idle = ctx.idle.get(); let idle = ctx.idle.get();
// Trapping is opted into during init and fixed for the loop's life. The
// inbox is armed only when set: an untrapped server keeps the fast path,
// and a shutdown request simply stops it as `request_stop` would.
let exits: Option<Receiver<ExitSignal>> = ctx.trap.get().then(crate::link::trap_exit);
drop(ctx); drop(ctx);
let mut monitors: Vec<Monitor> = Vec::new(); let mut monitors: Vec<Monitor> = Vec::new();
@@ -992,7 +1150,7 @@ fn server_loop<G: GenServer>(
}; };
loop { loop {
if monitors.is_empty() && !sys_open && infos.is_empty() { if exits.is_none() && monitors.is_empty() && !sys_open && infos.is_empty() {
// Fast path: no extra arms, no select overhead — park directly on // Fast path: no extra arms, no select overhead — park directly on
// the inbox. Mirrors the inbox arm of the select path below; any // the inbox. Mirrors the inbox arm of the select path below; any
// change there must be applied here too. // change there must be applied here too.
@@ -1009,7 +1167,7 @@ fn server_loop<G: GenServer>(
guard.0.handle_idle(); guard.0.handle_idle();
reset_idle(&mut idle_deadline); reset_idle(&mut idle_deadline);
} }
// All ServerRefs dropped → inbox closed → shutdown. // Defensive: the loop holds a sender, so unreachable.
Err(RecvTimeoutError::Disconnected) => break, Err(RecvTimeoutError::Disconnected) => break,
} }
} }
@@ -1020,16 +1178,21 @@ fn server_loop<G: GenServer>(
} }
} else { } else {
// Slow path: one or more extra arms live — build the arm slice and // Slow path: one or more extra arms live — build the arm slice and
// select. Arm order encodes priority: downs → system → infos → // select. Arm order encodes priority: exits → downs → system →
// inbox. The slice is rebuilt each iteration because the monitor // infos → inbox (a shutdown request is noticed under any load).
// and info sets shrink/grow. Mirrors the fast-path inbox park // The slice is rebuilt each iteration because the monitor and info
// above; keep them in sync. // sets shrink/grow. Mirrors the fast-path inbox park above; keep
let nd = monitors.len(); // monitor band: [0, nd) // them in sync.
let nw = sys_open as usize; // system arm: [nd, nd+nw) let ne = exits.is_some() as usize; // exit arm: [0, ne)
// info band: [nd+nw, nd+nw+ni) let nd = ne + monitors.len(); // monitor band: [ne, nd)
let nw = sys_open as usize; // system arm: [nd, nd+nw)
// info band: [nd+nw, nd+nw+ni)
// inbox arm: [nd+nw+ni] // inbox arm: [nd+nw+ni]
let sel = { let sel = {
let mut arms: Vec<&dyn Selectable> = Vec::with_capacity(nd + nw + infos.len() + 1); let mut arms: Vec<&dyn Selectable> = Vec::with_capacity(nd + nw + infos.len() + 1);
if let Some(e) = &exits {
arms.push(e);
}
for m in &monitors { for m in &monitors {
arms.push(&m.rx); arms.push(&m.rx);
} }
@@ -1055,10 +1218,25 @@ fn server_loop<G: GenServer>(
continue; continue;
} }
}; };
if i < nd { if i < ne {
// Exit arm: a shutdown request or a linked peer's death.
// The inbox lives for the loop's life, so it never closes.
let sig = exits.as_ref().and_then(|e| e.try_recv().ok().flatten());
if let Some(sig) = sig {
if sig.reason == DownReason::Shutdown {
match guard.0.handle_shutdown() {
ShutdownAction::Exit => break,
ShutdownAction::Continue => {}
}
} else {
guard.0.handle_exit(sig);
}
reset_idle(&mut idle_deadline);
}
} else if i < nd {
// Monitor band: a Down retires its arm either way (one-shot) // Monitor band: a Down retires its arm either way (one-shot)
// or closes without delivering (defensive; shouldn't happen). // or closes without delivering (defensive; shouldn't happen).
let m = monitors.remove(i); let m = monitors.remove(i - ne);
if let Ok(Some(down)) = m.rx.try_recv() { if let Ok(Some(down)) = m.rx.try_recv() {
guard.0.handle_down(down); guard.0.handle_down(down);
reset_idle(&mut idle_deadline); reset_idle(&mut idle_deadline);
@@ -1067,6 +1245,8 @@ fn server_loop<G: GenServer>(
match sys_rx.try_recv() { match sys_rx.try_recv() {
// Control intake, not a dispatched message: no idle reset. // Control intake, not a dispatched message: no idle reset.
Ok(Some(Sys::Watch(m))) => monitors.push(m), Ok(Some(Sys::Watch(m))) => monitors.push(m),
// The state ended the server: a normal exit.
Ok(Some(Sys::Stop)) => break,
Ok(Some(Sys::Timer(id, msg))) => { Ok(Some(Sys::Timer(id, msg))) => {
// The one-shot fired: retire its registry entry so the // The one-shot fired: retire its registry entry so the
// live set tracks only still-pending timers, then // live set tracks only still-pending timers, then
+469 -66
View File
@@ -68,10 +68,38 @@
//! inbox or timer event; a replayed event may postpone again (it re-queues for //! inbox or timer event; a replayed event may postpone again (it re-queues for
//! the next transition). See the macro docs for the row surface and [`Step`] //! the next transition). See the macro docs for the row surface and [`Step`]
//! for how a postpone surfaces to the loop. //! for how a postpone surfaces to the loop.
//!
//! ## Stopping, shutdown, and exits
//!
//! A machine lives until it stops, is shut down, or is killed; a
//! [`GenStatemRef`] is an address, and dropping refs never ends it (the
//! gen_server rule — see its *When the server stops*). The supervised shape
//! is [`run_named`], which runs the machine inline as the current actor so it
//! is a direct `ChildSpec` child addressed by [`GenStatemName`]. A machine
//! can end itself: any row
//! body may call [`cx.stop()`](Cx::stop) (or use the `stop` tail keyword,
//! sugar for `{ cx.stop(); prev }`) — the loop breaks after that event and the
//! actor exits *normally* (OTP's `{stop, normal}`; a `Transient` child is not
//! restarted). [`terminate`](Machine::terminate) — the optional `terminate
//! { … }` macro block — runs on every exit path.
//!
//! From outside, [`GenStatemRef::shutdown`] (or a plain
//! [`request_shutdown`](crate::request_shutdown), which is what a supervisor
//! sends) asks the machine to stop. By default a machine does not trap exits
//! and the request stops it outright. A machine that calls
//! [`cx.trap_exit()`](Cx::trap_exit) in its initial `enter` instead receives
//! it as the **`shutdown` event**, routed by state like any other — a
//! `Connected` state may transition into `Draining` and stop later from a
//! timeout row, while a state with no `shutdown` row takes the macro's
//! default, `stop`. Linked-peer deaths reach a trapping machine as
//! `exit <pat>` events; an unmatched one is dropped like an unmatched info.
use crate::channel::{channel, select, Receiver, Sender}; use crate::channel::{channel, select, Receiver, Selectable, Sender};
use crate::link::ExitSignal;
use crate::monitor::{monitor, DownReason};
use crate::pid::Pid; use crate::pid::Pid;
use crate::scheduler::{cancel_timer, send_after_to}; use crate::registry::{register_with, resolve_named_sender, RegisterError};
use crate::scheduler::{cancel_timer, request_shutdown, send_after_to};
use crate::timer::TimerId; use crate::timer::TimerId;
use std::collections::{HashMap, VecDeque}; use std::collections::{HashMap, VecDeque};
use std::marker::PhantomData; use std::marker::PhantomData;
@@ -119,6 +147,33 @@ pub trait Machine: Send + 'static {
/// state's `enter`, and returns [`Step::Transitioned`]; a stay or unmatched /// state's `enter`, and returns [`Step::Transitioned`]; a stay or unmatched
/// event returns [`Step::Stayed`]. /// event returns [`Step::Stayed`].
fn handle(&mut self, ev: Self::Ev, cx: &mut Cx<Self::Ev>) -> Step<Self::Ev>; fn handle(&mut self, ev: Self::Ev, cx: &mut Cx<Self::Ev>) -> Step<Self::Ev>;
/// Wrap a graceful shutdown request into this machine's event, so a
/// trapping machine (see [`Cx::trap_exit`]) can route it **by state**. The
/// macro generates it as `Ev::Shutdown` and matches it in `shutdown` rows;
/// its default for a state that writes no such row is `stop`. A
/// hand-written machine that returns `None` (the default) is simply
/// stopped — the loop breaks and `terminate` runs on the normal path.
fn shutdown_ev() -> Option<Self::Ev> {
None
}
/// Wrap a linked peer's death (an [`ExitSignal`] that is not a shutdown
/// request, delivered only when trapping) into this machine's event. The
/// macro generates it as `Ev::Exit(sig)` and matches it in `exit <pat>`
/// rows; an unmatched exit is silently dropped, like an unmatched info.
/// A hand-written machine that returns `None` (the default) drops it.
fn exit_ev(_sig: ExitSignal) -> Option<Self::Ev> {
None
}
/// Runs as the machine actor exits, on any exit path (a `stop`, a
/// graceful shutdown, a handler panic, a hard stop). Like
/// `gen_server::terminate`: on the panic and hard-stop paths it runs
/// mid-unwind — do not panic or park there; only on the normal path
/// (`stop`, shutdown rows) may it do real work. The macro's optional
/// `terminate { … }` block generates it.
fn terminate(&mut self) {}
} }
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -247,6 +302,11 @@ impl Timers {
pub struct Cx<Ev> { pub struct Cx<Ev> {
sys_tx: Sender<Sys>, sys_tx: Sender<Sys>,
reg: Arc<Mutex<Timers>>, reg: Arc<Mutex<Timers>>,
/// Set by [`trap_exit`](Self::trap_exit) during `on_start`; read once by
/// the loop right after, fixed for the machine's life.
trap: bool,
/// Set by [`stop`](Self::stop); the loop breaks after the current event.
stop: bool,
_ev: PhantomData<fn() -> Ev>, _ev: PhantomData<fn() -> Ev>,
} }
@@ -255,10 +315,30 @@ impl<Ev> Cx<Ev> {
Cx { Cx {
sys_tx, sys_tx,
reg, reg,
trap: false,
stop: false,
_ev: PhantomData, _ev: PhantomData,
} }
} }
/// Trap exits for the machine's lifetime: a shutdown request then arrives
/// as the `shutdown` event (routed by state) and linked-peer deaths as
/// `exit` events, instead of stopping the machine outright. Call it in the
/// initial state's `enter` (i.e. during `on_start`); later calls have no
/// effect.
pub fn trap_exit(&mut self) {
self.trap = true;
}
/// End the machine after the current event: the loop breaks and the actor
/// exits *normally* (OTP's `{stop, normal}`); `terminate` runs on the
/// normal path. Anything still queued or postponed is dropped. The `stop`
/// tail keyword in a macro row is sugar for `{ cx.stop(); prev }`.
/// Mirrors gen_server's [`StopHandle`](crate::gen_server::StopHandle).
pub fn stop(&mut self) {
self.stop = true;
}
/// Arm the **state timeout**: fire a `state_timeout` event after `after` in /// Arm the **state timeout**: fire a `state_timeout` event after `after` in
/// the current state. Auto-reset on any state change (the loop cancels and /// the current state. Auto-reset on any state change (the loop cancels and
/// clears it on every real transition), so it measures quiet time *within* a /// clears it on every real transition), so it measures quiet time *within* a
@@ -386,8 +466,7 @@ pub enum SendError {
} }
/// A clonable handle to a running machine. Cloning yields another sender to the /// A clonable handle to a running machine. Cloning yields another sender to the
/// same inbox; the machine lives until the last `GenStatemRef` is dropped, at which /// same inbox. An address, not an owner: dropping refs never ends the machine.
/// point its inbox closes and the loop exits.
pub struct GenStatemRef<M: Machine> { pub struct GenStatemRef<M: Machine> {
tx: Sender<M::Ev>, tx: Sender<M::Ev>,
pid: Pid, pid: Pid,
@@ -433,6 +512,19 @@ impl<M: Machine> GenStatemRef<M> {
self.send(ev).map_err(|_| CallError::Down)?; self.send(ev).map_err(|_| CallError::Down)?;
rx.recv().map_err(|_| CallError::Down) rx.recv().map_err(|_| CallError::Down)
} }
/// Ask the machine to shut down and block until it has fully exited, so
/// [`Machine::terminate`] has run by the time this returns. A trapping
/// machine winds down through its `shutdown` rows; any other is stopped
/// outright. Returns immediately if the machine is already gone. Waits as
/// long as the machine takes — a supervisor bounds that with its child's
/// [`Shutdown`](crate::supervisor::Shutdown) policy. Mirrors
/// [`GenServerRef::shutdown`](crate::gen_server::GenServerRef::shutdown).
pub fn shutdown(&self) {
let mon = monitor(self.pid);
request_shutdown(self.pid);
let _ = mon.rx.recv();
}
} }
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -441,7 +533,8 @@ impl<M: Machine> GenStatemRef<M> {
/// Spawn `machine` as an actor and hand back its [`GenStatemRef`]. Shape mirrors /// Spawn `machine` as an actor and hand back its [`GenStatemRef`]. Shape mirrors
/// `gen_server::start`: make the inbox, spawn the loop, return the ref; the /// `gen_server::start`: make the inbox, spawn the loop, return the ref; the
/// backing join handle is dropped (lifetime is governed by refs, not joining). /// backing join handle is dropped (the machine's lifetime is its own). For a
/// supervised machine use [`run_named`].
/// ///
/// Panics if called outside `Runtime::run()`. /// Panics if called outside `Runtime::run()`.
pub fn spawn<M: Machine>(machine: M) -> GenStatemRef<M> { pub fn spawn<M: Machine>(machine: M) -> GenStatemRef<M> {
@@ -456,40 +549,188 @@ pub fn spawn<M: Machine>(machine: M) -> GenStatemRef<M> {
/// Panics if called outside `Runtime::run()`. /// Panics if called outside `Runtime::run()`.
pub fn spawn_with<M: Machine>(opts: crate::scheduler::SpawnOpts, machine: M) -> GenStatemRef<M> { pub fn spawn_with<M: Machine>(opts: crate::scheduler::SpawnOpts, machine: M) -> GenStatemRef<M> {
let (tx, rx) = channel::<M::Ev>(); let (tx, rx) = channel::<M::Ev>();
let handle = crate::scheduler::spawn_with(opts, move || statem_loop(rx, machine)); let keep = tx.clone();
let handle = crate::scheduler::spawn_with(opts, move || statem_loop(keep, rx, machine));
GenStatemRef { GenStatemRef {
tx, tx,
pid: handle.pid(), pid: handle.pid(),
} }
} }
/// A typed, static name for a gen_statem, used to address a machine through
/// the registry without holding a [`GenStatemRef`]. Mirrors
/// [`GenServerName`](crate::gen_server::GenServerName): declare it as a
/// constant and bind it with [`run_named`].
pub struct GenStatemName<M> {
name: &'static str,
_marker: PhantomData<fn() -> M>,
}
impl<M> GenStatemName<M> {
/// Bind a static string as a machine name.
#[inline]
pub const fn new(name: &'static str) -> Self {
Self {
name,
_marker: PhantomData,
}
}
/// The underlying registry key.
#[inline]
pub const fn as_str(self) -> &'static str {
self.name
}
}
impl<M> Copy for GenStatemName<M> {}
impl<M> Clone for GenStatemName<M> {
fn clone(&self) -> Self {
*self
}
}
/// Run `machine` **inline, as the current actor**, bound to `name`. The
/// supervised shape: the closure of a
/// [`ChildSpec`](crate::supervisor::ChildSpec) *is* the machine, so the
/// supervisor's shutdown reaches it as a `shutdown` row, a restart runs the
/// factory again and re-binds the name, and clients address it by name
/// ([`send`], [`call`], [`whereis_machine`]). Returns when the machine exits;
/// [`RegisterError::NameTaken`] (before `on_start`) if the name is held by a
/// different live actor. Mirrors
/// [`NamedGenServerBuilder::run`](crate::gen_server::NamedGenServerBuilder::run).
pub fn run_named<M: Machine>(name: GenStatemName<M>, machine: M) -> Result<(), RegisterError> {
let (tx, rx) = channel::<M::Ev>();
register_with::<M::Ev>(crate::scheduler::self_pid(), name.as_str(), tx.clone())?;
statem_loop(tx, rx, machine);
Ok(())
}
/// Resolve a [`GenStatemName`] to a [`GenStatemRef`]; `None` if no live
/// machine holds the name. Panics if called outside `Runtime::run()`.
pub fn whereis_machine<M: Machine>(name: GenStatemName<M>) -> Option<GenStatemRef<M>> {
resolve_named_sender::<M::Ev>(name.as_str()).map(|(pid, tx)| GenStatemRef { tx, pid })
}
/// Push an event to the machine registered under `name`, resolving per send.
/// [`SendError::Down`] if no live machine holds the name.
pub fn send<M: Machine>(name: GenStatemName<M>, ev: M::Ev) -> Result<(), SendError> {
match whereis_machine(name) {
Some(m) => m.send(ev),
None => Err(SendError::Down),
}
}
/// Synchronous request-reply to the machine registered under `name`,
/// resolving per call (a machine restarted under the same name is reached
/// transparently). [`CallError::Down`] if no live machine holds the name.
pub fn call<M, T, F>(name: GenStatemName<M>, make: F) -> Result<T, CallError>
where
M: Machine,
T: Send + 'static,
F: FnOnce(Reply<T>) -> M::Ev,
{
match whereis_machine(name) {
Some(m) => m.call(make),
None => Err(CallError::Down),
}
}
/// Shut down the machine registered under `name` and wait for it (see
/// [`GenStatemRef::shutdown`]). A no-op if no live machine holds the name.
pub fn shutdown<M: Machine>(name: GenStatemName<M>) {
if let Some(m) = whereis_machine(name) {
m.shutdown();
}
}
/// The machine actor body: `on_start`, then one `handle` per event until the /// The machine actor body: `on_start`, then one `handle` per event until the
/// inbox closes (all refs dropped → graceful shutdown). /// row resolves to `stop`, a shutdown row stops it, or the actor is stopped
/// from outside.
/// ///
/// Two intake sources are selected each iteration with the **timer arm above /// Intake arms are selected each iteration in priority order — **exits**
/// the inbox**, so a timeout fire is never starved by inbox traffic: `sys_rx` /// (only when trapping) above **timers** above the **inbox** — so a shutdown
/// carries timer fires armed through `cx`, `rx` is the user inbox. A fire is /// request or a timeout fire is never starved by inbox traffic. `sys_rx`
/// turned into the matching internal event (`state_timeout` / `timeout(name)`) /// carries timer fires armed through `cx`; a fire is turned into the matching
/// and run through the same `handle` dispatch as an inbox event — the /// internal event (`state_timeout` / `timeout(name)`) and run through the same
/// gen_statem model, where timeouts surface as ordinary events. /// `handle` dispatch as an inbox event — the gen_statem model, where timeouts
/// (and, when trapping, shutdown and exits) surface as ordinary events.
/// ///
/// The loop owns the **postpone queue**: a `handle` that defers its event hands /// The loop owns the **postpone queue**: a `handle` that defers its event hands
/// it back ([`Step::Postponed`]) for the queue; a `handle` that transitions /// it back ([`Step::Postponed`]) for the queue; a `handle` that transitions
/// ([`Step::Transitioned`]) triggers a [`replay`] of the queue in the new /// ([`Step::Transitioned`]) triggers a [`replay`] of the queue in the new
/// state, ahead of the next intake. /// state, ahead of the next intake.
fn statem_loop<M: Machine>(rx: Receiver<M::Ev>, mut machine: M) { fn statem_loop<M: Machine>(keep: Sender<M::Ev>, rx: Receiver<M::Ev>, machine: M) {
// One inbox sender lives with the loop: the inbox never closes, refs are
// addresses, the machine's lifetime is the actor's (stop row, shutdown,
// stop, panic). The `Disconnected` inbox arm below is defensive only.
let _keep = keep;
// Drop guard — owns the machine and the timer registry, so `terminate`
// fires on every exit path (clean close, `stop`, a handler panic, a hard
// stop) and the timer drain is sequenced before it. Same shape as
// gen_server's guard; see the rationale there.
struct Terminate<M: Machine>(M, Arc<Mutex<Timers>>);
impl<M: Machine> Drop for Terminate<M> {
fn drop(&mut self) {
{
let mut reg = match self.1.lock() {
Ok(g) => g,
Err(e) => panic!("smarm: gen_statem reg lock poisoned (core corrupt): {e}"),
};
if let Some((_, sub)) = reg.state.take() {
cancel_timer(sub);
}
for (_, (_, sub)) in reg.named.drain() {
cancel_timer(sub);
}
}
self.0.terminate();
}
}
let (sys_tx, sys_rx) = channel::<Sys>(); let (sys_tx, sys_rx) = channel::<Sys>();
let reg = Arc::new(Mutex::new(Timers::new())); let reg = Arc::new(Mutex::new(Timers::new()));
let mut guard = Terminate(machine, reg.clone());
// The loop owns `cx` (and through it a `sys_tx` clone) for its whole life, // The loop owns `cx` (and through it a `sys_tx` clone) for its whole life,
// so the sys arm never closes from under us — no auto-close dance needed. // so the sys arm never closes from under us — no auto-close dance needed.
let mut cx = Cx::new(sys_tx, reg.clone()); let mut cx = Cx::new(sys_tx, reg.clone());
// Events deferred by `postpone` rows, replayed FIFO on the next transition. // Events deferred by `postpone` rows, replayed FIFO on the next transition.
let mut postpone: VecDeque<M::Ev> = VecDeque::new(); let mut postpone: VecDeque<M::Ev> = VecDeque::new();
machine.on_start(&mut cx); guard.0.on_start(&mut cx);
// Trapping is opted into during on_start and fixed for the loop's life.
// The inbox is armed only when set: an untrapped machine keeps the
// two-arm select, and a shutdown request simply stops it as
// `request_stop` would.
let exits: Option<Receiver<ExitSignal>> = cx.trap.then(crate::link::trap_exit);
loop { loop {
// Timer arm first: a ready fire is taken in preference to the inbox. // Arm order encodes priority: exits → timers → inbox.
let i = select(&[&sys_rx, &rx]); let ne = exits.is_some() as usize;
if i == 0 { let i = {
let mut arms: Vec<&dyn Selectable> = Vec::with_capacity(3);
if let Some(e) = &exits {
arms.push(e);
}
arms.push(&sys_rx);
arms.push(&rx);
select(&arms)
};
if i < ne {
// Exit arm: a shutdown request or a linked peer's death. The trap
// inbox lives for the loop's life, so it never closes.
let sig = exits.as_ref().and_then(|e| e.try_recv().ok().flatten());
match sig {
Some(sig) if sig.reason == DownReason::Shutdown => match M::shutdown_ev() {
Some(ev) => dispatch(&mut guard.0, &mut cx, &mut postpone, ev),
None => cx.stop(),
},
Some(sig) => {
if let Some(ev) = M::exit_ev(sig) {
dispatch(&mut guard.0, &mut cx, &mut postpone, ev);
}
}
None => {}
}
} else if i == ne {
match sys_rx.try_recv() { match sys_rx.try_recv() {
Ok(Some(fire)) => { Ok(Some(fire)) => {
// Confirm the fire is still the live one before dispatching: // Confirm the fire is still the live one before dispatching:
@@ -528,7 +769,7 @@ fn statem_loop<M: Machine>(rx: Receiver<M::Ev>, mut machine: M) {
} }
}; };
if let Some(ev) = ev { if let Some(ev) = ev {
dispatch(&mut machine, &mut cx, &mut postpone, ev); dispatch(&mut guard.0, &mut cx, &mut postpone, ev);
} }
} }
// Single-receiver: nothing can drain the arm between select's // Single-receiver: nothing can drain the arm between select's
@@ -540,12 +781,16 @@ fn statem_loop<M: Machine>(rx: Receiver<M::Ev>, mut machine: M) {
} }
} else { } else {
match rx.try_recv() { match rx.try_recv() {
Ok(Some(ev)) => dispatch(&mut machine, &mut cx, &mut postpone, ev), Ok(Some(ev)) => dispatch(&mut guard.0, &mut cx, &mut postpone, ev),
Ok(None) => debug_assert!(false, "ready inbox was empty"), Ok(None) => debug_assert!(false, "ready inbox was empty"),
// All GenStatemRefs dropped → inbox closed → shutdown. // Defensive: the loop holds a sender, so unreachable.
Err(_) => break, Err(_) => break,
} }
} }
// A handler (or a replay) asked to stop: a normal exit.
if cx.stop {
break;
}
// Observation point so a machine fed a hot inbox stays preemptible and // Observation point so a machine fed a hot inbox stays preemptible and
// cancellable. // cancellable.
crate::check!(); crate::check!();
@@ -554,7 +799,8 @@ fn statem_loop<M: Machine>(rx: Receiver<M::Ev>, mut machine: M) {
/// Run one event through `handle` and act on its [`Step`]: stash a deferred /// Run one event through `handle` and act on its [`Step`]: stash a deferred
/// event on the postpone queue, or — on a real transition — [`replay`] the /// event on the postpone queue, or — on a real transition — [`replay`] the
/// queue in the new state. A stay/unmatched event needs nothing further. /// queue in the new state. A stay/unmatched event needs nothing further. A
/// [`Cx::stop`] raised by the handler skips the replay; the loop breaks next.
fn dispatch<M: Machine>( fn dispatch<M: Machine>(
machine: &mut M, machine: &mut M,
cx: &mut Cx<M::Ev>, cx: &mut Cx<M::Ev>,
@@ -564,14 +810,19 @@ fn dispatch<M: Machine>(
match machine.handle(ev, cx) { match machine.handle(ev, cx) {
Step::Postponed(ev) => postpone.push_back(ev), Step::Postponed(ev) => postpone.push_back(ev),
Step::Stayed => {} Step::Stayed => {}
Step::Transitioned => replay(machine, cx, postpone), Step::Transitioned => {
if !cx.stop {
replay(machine, cx, postpone)
}
}
} }
} }
/// Replay deferred events after a real transition: each goes back through /// Replay deferred events after a real transition: each goes back through
/// `handle` in FIFO order, in the now-current state. An event that postpones /// `handle` in FIFO order, in the now-current state. An event that postpones
/// again re-queues (to wait for the *next* transition); one that transitions /// again re-queues (to wait for the *next* transition); one that transitions
/// re-arms the replay, so a later state can in turn drain what is still pending. /// re-arms the replay, so a later state can in turn drain what is still pending;
/// one that raises [`Cx::stop`] ends the replay (and the machine) at once.
/// Subsequent events in a batch already see the post-transition state, since /// Subsequent events in a batch already see the post-transition state, since
/// `handle` reads the live state cell — the outer loop only re-runs to give /// `handle` reads the live state cell — the outer loop only re-runs to give
/// re-queued events another pass once a transition has occurred within a batch. /// re-queued events another pass once a transition has occurred within a batch.
@@ -590,6 +841,9 @@ fn replay<M: Machine>(machine: &mut M, cx: &mut Cx<M::Ev>, postpone: &mut VecDeq
Step::Stayed => {} Step::Stayed => {}
Step::Transitioned => transitioned = true, Step::Transitioned => transitioned = true,
} }
if cx.stop {
return;
}
} }
if !transitioned { if !transitioned {
return; return;
@@ -691,8 +945,9 @@ fn replay<M: Machine>(machine: &mut M, cx: &mut Cx<M::Ev>, postpone: &mut VecDeq
/// // The transition table. Group rows by current state with `on <pat>`. /// // The transition table. Group rows by current state with `on <pat>`.
/// // A row is: <kind> <event-pattern> [if <guard>] => <tail> , /// // A row is: <kind> <event-pattern> [if <guard>] => <tail> ,
/// // where <kind> is one of `cast`, `call`, `info`, `state_timeout` /// // where <kind> is one of `cast`, `call`, `info`, `state_timeout`
/// // (no pattern — it is a unit event), or `timeout <name-pattern>`, /// // (no pattern — it is a unit event), `timeout <name-pattern>`,
/// // and the tail is one of: /// // `shutdown` (unit; trapping machines only) or `exit <sig-pattern>`
/// // (trapping only), and the tail is one of:
/// // * a state tag `Door::Closed` (transition, or "stay" /// // * a state tag `Door::Closed` (transition, or "stay"
/// // if it equals current) /// // if it equals current)
/// // * a block ending in one `{ data.enters += 1; Door::Closed }` /// // * a block ending in one `{ data.enters += 1; Door::Closed }`
@@ -701,6 +956,9 @@ fn replay<M: Machine>(machine: &mut M, cx: &mut Cx<M::Ev>, postpone: &mut VecDeq
/// // * the keyword `postpone` (defer until next /// // * the keyword `postpone` (defer until next
/// // transition; cast/call/ /// // transition; cast/call/
/// // info only) /// // info only)
/// // * the keyword `stop` (end the machine
/// // normally; sugar for
/// // `{ cx.stop(); prev }`)
/// on Door::Open => { /// on Door::Open => {
/// cast DoorCast::Push => Door::Closed, /// cast DoorCast::Push => Door::Closed,
/// // An armed state-timeout surfaces as an ordinary event: /// // An armed state-timeout surfaces as an ordinary event:
@@ -724,6 +982,9 @@ fn replay<M: Machine>(machine: &mut M, cx: &mut Cx<M::Ev>, postpone: &mut VecDeq
/// on _ => { /// on _ => {
/// call DoorCall::GetState(r) => { r.reply(prev); prev }, /// call DoorCall::GetState(r) => { r.reply(prev); prev },
/// } /// }
///
/// // Optional: runs as the machine exits, on every exit path.
/// terminate { data.enters = 0; }
/// } /// }
/// ``` /// ```
/// ///
@@ -748,6 +1009,12 @@ fn replay<M: Machine>(machine: &mut M, cx: &mut Cx<M::Ev>, postpone: &mut VecDeq
/// ignores info writes no `info` rows at all. State-timeouts and named /// ignores info writes no `info` rows at all. State-timeouts and named
/// timeouts have **no** such default: a state that can see one must handle it /// timeouts have **no** such default: a state that can see one must handle it
/// (or `unhandled` it) or the match is non-exhaustive. /// (or `unhandled` it) or the match is non-exhaustive.
/// * **`shutdown` defaults to `stop`, `exit` to a silent drop.** Both reach a
/// machine only if its initial `enter` called `cx.trap_exit()`; a
/// non-trapping machine is stopped outright by a shutdown request. Write
/// `shutdown => …` rows only in the states that want to wind down first
/// (transition into a draining state, `stop` later); write `exit sig => …`
/// rows to react to linked-peer deaths. Neither is postponable.
/// * **Stay** = return the current tag. The `prev` you named in `context` is /// * **Stay** = return the current tag. The `prev` you named in `context` is
/// bound to the pre-handler state for exactly this — handy in any-state /// bound to the pre-handler state for exactly this — handy in any-state
/// (`on _`) rows where there is no single literal tag to write. /// (`on _`) rows where there is no single literal tag to write.
@@ -773,12 +1040,12 @@ fn replay<M: Machine>(machine: &mut M, cx: &mut Cx<M::Ev>, postpone: &mut VecDeq
/// # What it emits /// # What it emits
/// ///
/// The unified `enum $Ev` (the `Cast`/`Call`/`Info` wrappers plus the internal /// The unified `enum $Ev` (the `Cast`/`Call`/`Info` wrappers plus the internal
/// `StateTimeout` / `Timeout` events), `struct $Sm { state, data }`, /// `StateTimeout` / `Timeout` / `Shutdown` / `Exit` events), `struct $Sm {
/// `$Sm::start(init, data) -> GenStatemRef<$Sm>`, and the `Machine` impl: /// state, data }`, `$Sm::start(init, data) -> GenStatemRef<$Sm>`, and the
/// `on_start` runs the initial `enter`; `handle` is the dispatch match plus the /// `Machine` impl: `on_start` runs the initial `enter`; `handle` is the
/// stay/transition/unhandled apply-tail (the cell's sole writer, which also /// dispatch match plus the stay/transition/unhandled apply-tail (the cell's
/// auto-resets the state-timeout on every real transition); and the `enter` /// sole writer, which also auto-resets the state-timeout on every real
/// dispatch. /// transition); the `enter` dispatch; and `terminate` when the block is given.
/// ///
/// # Limitation /// # Limitation
/// ///
@@ -794,9 +1061,10 @@ macro_rules! gen_statem {
context ( $data:ident , $cur:ident , $cx:ident ) ; context ( $data:ident , $cur:ident , $cx:ident ) ;
enter { $( $est:pat => $ebody:expr ),+ $(,)? } enter { $( $est:pat => $ebody:expr ),+ $(,)? }
$( on $st:pat => { $($rows:tt)* } )+ $( on $st:pat => { $($rows:tt)* } )+
$( terminate { $($tbody:tt)* } )?
) => { ) => {
/// Unified inbox payload: the user's `cast`/`call`/`info` enums folded /// Unified inbox payload: the user's `cast`/`call`/`info` enums folded
/// together with the runtime's internal timeout events. /// together with the runtime's internal events.
enum $Ev { enum $Ev {
Cast($Cast), Cast($Cast),
Call($Call), Call($Call),
@@ -807,6 +1075,12 @@ macro_rules! gen_statem {
StateTimeout, StateTimeout,
/// A named timeout fired (matched in `timeout <pat>` rows). /// A named timeout fired (matched in `timeout <pat>` rows).
Timeout(&'static str), Timeout(&'static str),
/// A graceful shutdown request reached this (trapping) machine
/// (matched in `shutdown` rows). A state with no such row `stop`s.
Shutdown,
/// A linked peer died (trapping only; matched in `exit <pat>`
/// rows). An unmatched exit is silently dropped.
Exit($crate::ExitSignal),
} }
struct $sm { struct $sm {
@@ -815,8 +1089,14 @@ macro_rules! gen_statem {
} }
impl $sm { impl $sm {
/// The machine value, for [`gen_statem::run_named`]
/// (`$crate::gen_statem::run_named`) or `spawn`.
fn new(init: $State, data: $Data) -> $sm {
$sm { state: init, data }
}
fn start(init: $State, data: $Data) -> $crate::gen_statem::GenStatemRef<$sm> { fn start(init: $State, data: $Data) -> $crate::gen_statem::GenStatemRef<$sm> {
$crate::gen_statem::spawn($sm { state: init, data }) $crate::gen_statem::spawn($sm::new(init, data))
} }
#[allow(unused_variables)] #[allow(unused_variables)]
@@ -840,11 +1120,27 @@ macro_rules! gen_statem {
$Ev::Timeout(name) $Ev::Timeout(name)
} }
fn shutdown_ev() -> Option<$Ev> {
Some($Ev::Shutdown)
}
fn exit_ev(sig: $crate::ExitSignal) -> Option<$Ev> {
Some($Ev::Exit(sig))
}
fn on_start(&mut self, $cx: &mut $crate::gen_statem::Cx<$Ev>) { fn on_start(&mut self, $cx: &mut $crate::gen_statem::Cx<$Ev>) {
let s = self.state; let s = self.state;
self.enter(s, $cx); self.enter(s, $cx);
} }
$(
#[allow(unused_variables)]
fn terminate(&mut self) {
let $data = &mut self.data;
$($tbody)*
}
)?
#[allow(unused_variables)] #[allow(unused_variables)]
#[deny(unreachable_patterns)] // conflicting rows must fail even though #[deny(unreachable_patterns)] // conflicting rows must fail even though
// this match is external-macro-expanded // this match is external-macro-expanded
@@ -860,7 +1156,7 @@ macro_rules! gen_statem {
// untouched), then the consuming `match (state, event)` whose // untouched), then the consuming `match (state, event)` whose
// value is this `Resolution`. // value is this `Resolution`.
let next: $crate::gen_statem::Resolution<$State> = let next: $crate::gen_statem::Resolution<$State> =
$crate::gen_statem!(@arms ($Ev) ($cur, ev) [ ] [ ] $crate::gen_statem!(@arms ($Ev) ($cur, ev, $cx) [ ] [ ]
$( on $st => { $($rows)* } )+); $( on $st => { $($rows)* } )+);
match next { match next {
$crate::gen_statem::Resolution::To(s) if s == $cur => { $crate::gen_statem::Resolution::To(s) if s == $cur => {
@@ -896,7 +1192,7 @@ macro_rules! gen_statem {
// is the `Info` silent-drop — last and broadest, so per-state `info` rows // is the `Info` silent-drop — last and broadest, so per-state `info` rows
// stay reachable; cast/call/timeouts get no fallback, so a forgotten pair is // stay reachable; cast/call/timeouts get no fallback, so a forgotten pair is
// still E0004). // still E0004).
(@arms ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ]) => { (@arms ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ]) => {
{ {
// Phase 1 — postpone routing (borrow-only). A guard on a postpone // Phase 1 — postpone routing (borrow-only). A guard on a postpone
// row runs here, against by-ref bindings, so it must not depend on // row runs here, against by-ref bindings, so it must not depend on
@@ -913,140 +1209,247 @@ macro_rules! gen_statem {
// returned for it. // returned for it.
match ($ss, $se) { match ($ss, $se) {
$($arms)* $($arms)*
// Macro-injected defaults, last and broadest so per-state rows
// stay reachable; a user's own catch-all row may shadow them
// entirely, hence the allow.
#[allow(unreachable_patterns)]
(_, $Ev::Info(_)) => $crate::gen_statem::Resolution::Unhandled, (_, $Ev::Info(_)) => $crate::gen_statem::Resolution::Unhandled,
#[allow(unreachable_patterns)]
(_, $Ev::Exit(_)) => $crate::gen_statem::Resolution::Unhandled,
#[allow(unreachable_patterns)]
(_, $Ev::Shutdown) => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) },
} }
} }
}; };
// Open an on-block: remember its state pat, drain its rows, then continue. // Open an on-block: remember its state pat, drain its rows, then continue.
(@arms ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] (@arms ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ]
on $st:pat => { $($rows:tt)* } $($more:tt)* on $st:pat => { $($rows:tt)* } $($more:tt)*
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) [ $($arms)* ] [ $($post)* ] ($st) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx) [ $($arms)* ] [ $($post)* ] ($st)
{ $($rows)* } { $($more)* }) { $($rows)* } { $($more)* })
}; };
// ===== @rows: drain one on-block's rows, threading both accs ============= // ===== @rows: drain one on-block's rows, threading both accs =============
// cast, explicit refusal // cast, explicit refusal
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ cast $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* } { cast $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Cast($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ] [ $($arms)* ($st, $Ev::Cast($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ]
[ $($post)* ] [ $($post)* ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// cast, stop (end the machine normally after this event)
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ cast $ev:pat $(if $g:expr)? => stop , $($rows:tt)* } { $($more:tt)* }
) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Cast($ev)) $(if $g)? => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) }, ]
[ $($post)* ]
($st) { $($rows)* } { $($more)* })
};
// cast, postpone (defer the event; the replay in a later state handles it) // cast, postpone (defer the event; the replay in a later state handles it)
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ cast $ev:pat $(if $g:expr)? => postpone , $($rows:tt)* } { $($more:tt)* } { cast $ev:pat $(if $g:expr)? => postpone , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Cast($ev)) $(if $g)? => unreachable!("postponed event is replayed, not dispatched here"), ] [ $($arms)* ($st, $Ev::Cast($ev)) $(if $g)? => unreachable!("postponed event is replayed, not dispatched here"), ]
[ $($post)* ($st, $Ev::Cast($ev)) $(if $g)? => true, ] [ $($post)* ($st, $Ev::Cast($ev)) $(if $g)? => true, ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// cast, transition / stay / branch // cast, transition / stay / branch
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ cast $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* } { cast $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Cast($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ] [ $($arms)* ($st, $Ev::Cast($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ]
[ $($post)* ] [ $($post)* ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// call, explicit refusal // call, explicit refusal
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ call $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* } { call $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Call($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ] [ $($arms)* ($st, $Ev::Call($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ]
[ $($post)* ] [ $($post)* ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// call, stop (end the machine normally after this event)
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ call $ev:pat $(if $g:expr)? => stop , $($rows:tt)* } { $($more:tt)* }
) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Call($ev)) $(if $g)? => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) }, ]
[ $($post)* ]
($st) { $($rows)* } { $($more)* })
};
// call, postpone (the Reply rides inside the event onto the queue) // call, postpone (the Reply rides inside the event onto the queue)
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ call $ev:pat $(if $g:expr)? => postpone , $($rows:tt)* } { $($more:tt)* } { call $ev:pat $(if $g:expr)? => postpone , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Call($ev)) $(if $g)? => unreachable!("postponed event is replayed, not dispatched here"), ] [ $($arms)* ($st, $Ev::Call($ev)) $(if $g)? => unreachable!("postponed event is replayed, not dispatched here"), ]
[ $($post)* ($st, $Ev::Call($ev)) $(if $g)? => true, ] [ $($post)* ($st, $Ev::Call($ev)) $(if $g)? => true, ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// call, transition / stay / branch // call, transition / stay / branch
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ call $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* } { call $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Call($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ] [ $($arms)* ($st, $Ev::Call($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ]
[ $($post)* ] [ $($post)* ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// info, explicit refusal // info, explicit refusal
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ info $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* } { info $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Info($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ] [ $($arms)* ($st, $Ev::Info($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ]
[ $($post)* ] [ $($post)* ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// info, stop (end the machine normally after this event)
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ info $ev:pat $(if $g:expr)? => stop , $($rows:tt)* } { $($more:tt)* }
) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Info($ev)) $(if $g)? => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) }, ]
[ $($post)* ]
($st) { $($rows)* } { $($more)* })
};
// info, postpone // info, postpone
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ info $ev:pat $(if $g:expr)? => postpone , $($rows:tt)* } { $($more:tt)* } { info $ev:pat $(if $g:expr)? => postpone , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Info($ev)) $(if $g)? => unreachable!("postponed event is replayed, not dispatched here"), ] [ $($arms)* ($st, $Ev::Info($ev)) $(if $g)? => unreachable!("postponed event is replayed, not dispatched here"), ]
[ $($post)* ($st, $Ev::Info($ev)) $(if $g)? => true, ] [ $($post)* ($st, $Ev::Info($ev)) $(if $g)? => true, ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// info, transition / stay / branch // info, transition / stay / branch
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ info $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* } { info $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Info($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ] [ $($arms)* ($st, $Ev::Info($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ]
[ $($post)* ] [ $($post)* ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// state_timeout, explicit refusal (unit event — no pattern; not postponable) // state_timeout, explicit refusal (unit event — no pattern; not postponable)
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ state_timeout $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* } { state_timeout $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::StateTimeout) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ] [ $($arms)* ($st, $Ev::StateTimeout) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ]
[ $($post)* ] [ $($post)* ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// state_timeout, stop (end the machine normally after this event)
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ state_timeout $(if $g:expr)? => stop , $($rows:tt)* } { $($more:tt)* }
) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::StateTimeout) $(if $g)? => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) }, ]
[ $($post)* ]
($st) { $($rows)* } { $($more)* })
};
// state_timeout, transition / stay / branch // state_timeout, transition / stay / branch
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ state_timeout $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* } { state_timeout $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::StateTimeout) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ] [ $($arms)* ($st, $Ev::StateTimeout) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ]
[ $($post)* ] [ $($post)* ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// timeout, explicit refusal (pattern matches the name; not postponable) // timeout, explicit refusal (pattern matches the name; not postponable)
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ timeout $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* } { timeout $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Timeout($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ] [ $($arms)* ($st, $Ev::Timeout($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ]
[ $($post)* ] [ $($post)* ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// timeout, stop (end the machine normally after this event)
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ timeout $ev:pat $(if $g:expr)? => stop , $($rows:tt)* } { $($more:tt)* }
) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Timeout($ev)) $(if $g)? => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) }, ]
[ $($post)* ]
($st) { $($rows)* } { $($more)* })
};
// timeout, transition / stay / branch // timeout, transition / stay / branch
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ timeout $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* } { timeout $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Timeout($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ] [ $($arms)* ($st, $Ev::Timeout($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ]
[ $($post)* ] [ $($post)* ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// shutdown, explicit refusal (unit event — no pattern; not postponable; a state with no shutdown row defaults to `stop`)
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ shutdown $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* }
) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Shutdown) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ]
[ $($post)* ]
($st) { $($rows)* } { $($more)* })
};
// shutdown, stop (end the machine normally after this event)
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ shutdown $(if $g:expr)? => stop , $($rows:tt)* } { $($more:tt)* }
) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Shutdown) $(if $g)? => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) }, ]
[ $($post)* ]
($st) { $($rows)* } { $($more)* })
};
// shutdown, transition / stay / branch
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ shutdown $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* }
) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Shutdown) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ]
[ $($post)* ]
($st) { $($rows)* } { $($more)* })
};
// exit, explicit refusal (pattern matches the ExitSignal; not postponable)
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ exit $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* }
) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Exit($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ]
[ $($post)* ]
($st) { $($rows)* } { $($more)* })
};
// exit, stop (end the machine normally after this event)
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ exit $ev:pat $(if $g:expr)? => stop , $($rows:tt)* } { $($more:tt)* }
) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Exit($ev)) $(if $g)? => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) }, ]
[ $($post)* ]
($st) { $($rows)* } { $($more)* })
};
// exit, transition / stay / branch
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ exit $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* }
) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Exit($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ]
[ $($post)* ]
($st) { $($rows)* } { $($more)* })
};
// this block is drained: hand the remaining on-blocks back to @arms // this block is drained: hand the remaining on-blocks back to @arms
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ } { $($more:tt)* } { } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@arms ($Ev) ($ss, $se) [ $($arms)* ] [ $($post)* ] $($more)*) $crate::gen_statem!(@arms ($Ev) ($ss, $se, $cx) [ $($arms)* ] [ $($post)* ] $($more)*)
}; };
} }
+2 -2
View File
@@ -191,8 +191,8 @@ pub struct ActorInfo {
/// [`Stack::new`](crate::stack::Stack::new) rounds them. /// [`Stack::new`](crate::stack::Stack::new) rounds them.
#[derive(Debug, Clone, Copy, PartialEq, Eq)] #[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub struct StackInfo { pub struct StackInfo {
/// Usable stack size ([`SpawnOpts::stack_reserve`] /// Usable stack size ([`SpawnOpts::stack_reserve`](crate::SpawnOpts::stack_reserve)
/// (crate::SpawnOpts::stack_reserve) or the Config/default). /// or the Config/default).
pub reserve: usize, pub reserve: usize,
/// PROT_NONE guard below the usable region. /// PROT_NONE guard below the usable region.
pub guard: usize, pub guard: usize,
+20 -9
View File
@@ -11,9 +11,20 @@
//! //!
//! See `LOOM.md` for the design intent and the deferred-for-later list. //! See `LOOM.md` for the design intent and the deferred-for-later list.
// Docs are part of the contract: broken/private intra-doc links fail `cargo doc`,
// and every doctest is compiled with `deny(warnings)` under `cargo test --doc`.
#![deny(rustdoc::broken_intra_doc_links)]
#![deny(rustdoc::private_intra_doc_links)]
#![deny(rustdoc::redundant_explicit_links)]
#![deny(rustdoc::invalid_codeblock_attributes)]
#![deny(rustdoc::invalid_rust_codeblocks)]
#![doc(test(attr(deny(warnings))))]
pub mod actor; pub mod actor;
pub mod causal; pub mod causal;
pub mod channel; pub mod channel;
#[cfg(feature = "cluster")]
pub mod cluster;
pub mod context; pub mod context;
pub mod gen_server; pub mod gen_server;
pub mod gen_statem; pub mod gen_statem;
@@ -60,10 +71,10 @@ pub use channel::{
pub use gen_server::{ pub use gen_server::{
call, cast, shutdown, whereis_server, CallError, CallTimeoutError, CastError, GenServer, call, cast, shutdown, whereis_server, CallError, CallTimeoutError, CastError, GenServer,
GenServerBuilder, GenServerCtx, GenServerName, GenServerRef, NamedGenServerBuilder, GenServerBuilder, GenServerCtx, GenServerName, GenServerRef, NamedGenServerBuilder,
TimerHandle, Watcher, ShutdownAction, StopHandle, TimerHandle, Watcher,
}; };
pub use gen_statem::{ pub use gen_statem::{
CallError as GenStatemCallError, Cx, GenStatemRef, Machine, Reply, Resolution, CallError as GenStatemCallError, Cx, GenStatemName, GenStatemRef, Machine, Reply, Resolution,
SendError as GenStatemSendError, SendError as GenStatemSendError,
}; };
pub use introspect::{ pub use introspect::{
@@ -85,15 +96,15 @@ pub use registry::{
install, lookup_as, register, resolve_name, send, send_dyn, send_to, unregister, whereis, install, lookup_as, register, resolve_name, send, send_dyn, send_to, unregister, whereis,
NameResolution, RegisterError, SendError, NameResolution, RegisterError, SendError,
}; };
pub use runtime::{init, Config, Runtime}; pub use runtime::{init, Config, Runtime, RuntimeHandle};
pub use scheduler::{ pub use scheduler::{
block_on_io, cancel_timer, request_stop, run, self_pid, send_after, send_after_named, block_on_io, cancel_timer, request_shutdown, request_stop, run, self_pid, send_after,
send_after_named_wall, send_after_wall, sleep, sleep_wall, spawn, spawn_addr, spawn_addr_with, send_after_named, send_after_named_wall, send_after_wall, sleep, sleep_wall, spawn, spawn_addr,
spawn_under, spawn_under_with, spawn_with, try_spawn, try_spawn_under_with, wait_readable, spawn_addr_with, spawn_monitor, spawn_monitor_with, spawn_under, spawn_under_with, spawn_with,
wait_readable_timeout, wait_writable, wait_writable_timeout, yield_now, FdArm, JoinError, try_spawn, try_spawn_under_with, wait_readable, wait_readable_timeout, wait_writable,
JoinHandle, SpawnError, SpawnOpts, wait_writable_timeout, yield_now, FdArm, JoinError, JoinHandle, SpawnError, SpawnOpts,
}; };
pub use supervisor::{ChildSpec, OneForOne, Restart, Signal, Strategy}; pub use supervisor::{ChildSpec, OneForOne, Restart, Shutdown, Signal, Strategy};
pub use timer::TimerId; pub use timer::TimerId;
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
+82 -42
View File
@@ -98,6 +98,11 @@ pub enum DownReason {
Panic, Panic,
/// The target was cooperatively cancelled via `request_stop`. /// The target was cooperatively cancelled via `request_stop`.
Stopped, Stopped,
/// A graceful shutdown was requested via `request_shutdown`. Only ever
/// appears in an [`ExitSignal`](crate::link::ExitSignal) delivered to a
/// trapping actor — never in a [`Down`]: a target that honours the request
/// exits *normally*, one that does not trap is `Stopped`.
Shutdown,
/// The target was already gone (finished and reclaimed, or never alive) /// The target was already gone (finished and reclaimed, or never alive)
/// at the moment `monitor()` was called. /// at the moment `monitor()` was called.
NoProc, NoProc,
@@ -140,46 +145,79 @@ pub struct Monitor {
/// Monitor `target`. Returns a [`Monitor`] whose `rx` receives exactly one /// Monitor `target`. Returns a [`Monitor`] whose `rx` receives exactly one
/// [`Down`]. /// [`Down`].
/// ///
/// To monitor a child you are spawning yourself, prefer
/// [`spawn_monitor`](crate::spawn_monitor): `spawn` followed by `monitor` on
/// the returned pid can race the child's death and observe `NoProc` instead
/// of its real reason.
///
/// If `target` is still live, the `Down` arrives when it terminates. If /// If `target` is still live, the `Down` arrives when it terminates. If
/// `target` is already gone, a [`DownReason::NoProc`] `Down` is queued /// `target` is already gone, a [`DownReason::NoProc`] `Down` is queued
/// immediately so the caller's `rx.recv()` returns without parking. /// immediately so the caller's `rx.recv()` returns without parking.
pub fn monitor<A>(target: Pid<A>) -> Monitor { pub fn monitor<A>(target: Pid<A>) -> Monitor {
let target = target.erase(); let target = target.erase();
let (tx, rx) = channel::<Down>(); let (tx, rx) = channel::<Down>();
let id = with_runtime(|inner| inner.alloc_monitor_id());
// Implementation note: registration happens under the target's cold if !register_monitor(target, id, &tx) {
// lock. `tx.clone()` takes the channel's own lock, a Channel-class
// RawMutex, which is explicitly permitted under a Leaf (cold) lock by
// the lock order documented in raw_mutex.rs. We must still not *send*
// under the lock, since `Sender::send` can unpark a parked receiver,
// and there's no reason to nest that.
let (id, registered) = with_runtime(|inner| {
let id = inner.alloc_monitor_id();
let registered = match inner.slot_at(target) {
Some(slot) => {
let mut cold = slot.cold.lock();
if slot.is_live_for(target) {
cold.monitors.push((id, tx.clone()));
true
} else {
false
}
}
None => false,
};
(id, registered)
});
if !registered {
let _ = tx.send(Down { let _ = tx.send(Down {
pid: target, pid: target,
reason: DownReason::NoProc, reason: DownReason::NoProc,
}); });
} }
Monitor { id, target, rx } Monitor { id, target, rx }
} }
/// Register a monitor `id` on `target` that delivers its `Down` to `tx` — the
/// primitive under [`monitor`], split out so a caller can fan many monitors
/// into ONE channel (process groups: every membership's death lands on the
/// reaper's single inbox). Returns `false` if `target` is already gone, in
/// which case nothing is registered and the caller decides what to queue
/// (`monitor` sends `NoProc`). The caller allocates `id` up front so it can
/// record the registration *before* arming it.
///
/// Implementation note: registration happens under the target's cold lock.
/// `tx.clone()` takes the channel's own lock, a Channel-class RawMutex, which
/// is explicitly permitted under a Leaf (cold) lock by the lock order
/// documented in raw_mutex.rs. We must still not *send* under the lock, since
/// `Sender::send` can unpark a parked receiver, and there's no reason to nest
/// that.
pub(crate) fn register_monitor(target: Pid, id: MonitorId, tx: &Sender<Down>) -> bool {
with_runtime(|inner| match inner.slot_at(target) {
Some(slot) => {
let mut cold = slot.cold.lock();
if slot.is_live_for(target) {
cold.monitors.push((id, tx.clone()));
true
} else {
false
}
}
None => false,
})
}
/// Remove registration `id` from `target` — the primitive under
/// [`demonitor`], for callers that hold only the id (see
/// [`register_monitor`]). `None` if the registration is not there: already
/// fired, already removed, or the slot has moved on to a new tenant.
///
/// The registration is removed under the target's cold lock, but the
/// `Sender` is moved *out* and dropped only after the lock is released.
/// Dropping the last sender runs `Sender::drop`, which may unpark a parked
/// receiver; legal under a cold lock, but pointless to nest.
pub(crate) fn unregister_monitor(target: Pid, id: MonitorId) -> Option<MonitorId> {
let removed: Option<(MonitorId, Sender<Down>)> = with_runtime(|inner| {
let slot = inner.slot_at(target)?;
let mut cold = slot.cold.lock();
if slot.generation() != target.generation() {
return None; // slot reused; the Down already fired
}
let pos = cold.monitors.iter().position(|(mid, _)| *mid == id)?;
Some(cold.monitors.remove(pos))
});
// `removed`'s sender drops here, outside the lock.
removed.map(|(id, _sender)| id)
}
/// Flag `target`'s tenancy as watchable: its death will stamp the slot's /// Flag `target`'s tenancy as watchable: its death will stamp the slot's
/// terminal record (see [`terminal_reason`]), exactly as registering a name /// terminal record (see [`terminal_reason`]), exactly as registering a name
/// does. The bridge calls this wherever a smarm pid is *encoded across the /// does. The bridge calls this wherever a smarm pid is *encoded across the
@@ -210,6 +248,23 @@ pub fn mark_watchable<A>(target: Pid<A>) {
}); });
} }
/// Whether `target` is live *and* its tenancy is watchable. The cluster's
/// remote-monitor admission check (RFC 010 c12): a peer may monitor a pid only
/// if that pid was exposed or crossed the wire (the D12 set-sites), and a
/// live-but-unwatchable pid answers exactly like a dead one — no liveness leak
/// beyond what `watchable` already grants. Same context contract as
/// [`monitor`].
#[cfg(feature = "cluster")]
pub(crate) fn is_watchable<A>(target: Pid<A>) -> bool {
let target = target.erase();
with_runtime(|inner| {
inner.slot_at(target).is_some_and(|slot| {
let cold = slot.cold.lock();
slot.is_live_for(target) && cold.watchable
})
})
}
/// The terminal [`DownReason`] of the tenancy `target` names, if that tenancy /// The terminal [`DownReason`] of the tenancy `target` names, if that tenancy
/// ever registered a name and is the *most recent named* death of its slot: /// ever registered a name and is the *most recent named* death of its slot:
/// finalize stamps the slot with `(generation, reason)` for once-registered /// finalize stamps the slot with `(generation, reason)` for once-registered
@@ -249,20 +304,5 @@ pub fn terminal_reason<A>(target: Pid<A>) -> Option<DownReason> {
/// instead of, or in addition to, calling this: dropping the [`Monitor`] /// instead of, or in addition to, calling this: dropping the [`Monitor`]
/// closes its receiver and any queued notice is discarded with it. /// closes its receiver and any queued notice is discarded with it.
pub fn demonitor(m: &Monitor) -> Option<MonitorId> { pub fn demonitor(m: &Monitor) -> Option<MonitorId> {
// Implementation note: the registration is removed under the target's unregister_monitor(m.target, m.id)
// cold lock, but the `Sender` is moved *out* and dropped only after the
// lock is released. Dropping the last sender runs `Sender::drop`, which
// may unpark a parked receiver; legal under a cold lock, but pointless
// to nest.
let removed: Option<(MonitorId, Sender<Down>)> = with_runtime(|inner| {
let slot = inner.slot_at(m.target)?;
let mut cold = slot.cold.lock();
if slot.generation() != m.target.generation() {
return None; // slot reused; the Down already fired
}
let pos = cold.monitors.iter().position(|(mid, _)| *mid == m.id)?;
Some(cold.monitors.remove(pos))
});
// `removed`'s sender drops here, outside the lock.
removed.map(|(id, _sender)| id)
} }
+463 -248
View File
@@ -99,10 +99,13 @@
//! ## Identity and clustering //! ## Identity and clustering
//! //!
//! A group member is described by a [`Member`] — a [`Pid`] plus a [`NodeId`] and //! A group member is described by a [`Member`] — a [`Pid`] plus a [`NodeId`] and
//! an [`Incarnation`]. Today everything is single-node, those two fields are //! an [`Incarnation`]. Everything on this page is **local**: you pass and
//! fixed defaults, and you only ever pass and receive a plain [`Pid`]: the extra //! receive plain [`Pid`]s, and [`members`] / [`pick`] / [`dispatch`] only ever
//! identity is carried so this API will not have to change when groups learn to //! name actors on this node (Erlang's `get_local_members`). With the `cluster`
//! span a cluster. //! feature a group also holds the members other nodes have announced, carried
//! under their [`NodeId`]; those never surface here — the cluster-wide reads
//! live in [`cluster::pg`](crate::cluster::pg) (`members_all` and friends) and
//! return a `Local | Remote` member type, since a [`Pid`] cannot hold a remote.
//! //!
//! ## Running context //! ## Running context
//! //!
@@ -110,10 +113,11 @@
//! from inside [`run`](crate::run) (that is, on an actor thread). Calling one //! from inside [`run`](crate::run) (that is, on an actor thread). Calling one
//! from outside a running runtime panics. //! from outside a running runtime panics.
use crate::monitor::{demonitor, monitor, Monitor}; use crate::channel::{channel, Sender};
use crate::monitor::{register_monitor, unregister_monitor, Down, DownReason, MonitorId};
use crate::pid::{assert_type, Addressable, Pid}; use crate::pid::{assert_type, Addressable, Pid};
use crate::registry::{send_to, SendError}; use crate::registry::{send_to, SendError};
use crate::scheduler::with_runtime; use crate::scheduler::{spawn_under, with_runtime};
use std::collections::HashMap; use std::collections::HashMap;
/// A cluster node handle. A `u32` integer handle, *not* an interned atom — the /// A cluster node handle. A `u32` integer handle, *not* an interned atom — the
@@ -186,13 +190,15 @@ pub struct Member {
pub pid: Pid, pub pid: Pid,
} }
/// One membership: a [`Member`] and the [`Monitor`] that watches its liveness. /// One membership: a [`Member`] and the id of the monitor that watches its
/// The monitor lives *alongside* the group entry so a group is /// liveness. The monitor's `Down` is delivered to the group reaper's single
/// self-contained: draining the membership tells us whether the member is /// inbox (see [`ProcessGroups::deaths`]), so the membership carries only what
/// still alive, and dropping the membership drops its monitor. /// [`leave`] needs to tear the registration down: the id.
struct Membership { pub(crate) struct Membership {
member: Member, pub(crate) member: Member,
monitor: Monitor, /// `None` for a remote member (cluster): the origin node is its liveness
/// authority; nothing here watches it.
pub(crate) monitor: Option<MonitorId>,
} }
/// The store: `name → multiset<Member>`. Within a single group a `Member` /// The store: `name → multiset<Member>`. Within a single group a `Member`
@@ -202,46 +208,60 @@ struct Membership {
/// ///
/// Locking discipline. Held under one Leaf-class `RawMutex` on `RuntimeInner`, /// Locking discipline. Held under one Leaf-class `RawMutex` on `RuntimeInner`,
/// mirroring the registry, and never held together with another Leaf lock (it /// mirroring the registry, and never held together with another Leaf lock (it
/// never touches the registry or a slot's cold lock). The two operations that /// never touches the registry or a slot's cold lock). Monitor registration and
/// do need another lock are kept off the group-lock path: /// removal take the target's cold lock (also Leaf), so they run *before* /
/// *after* the group lock, never under it — see [`join`] for the ordering that
/// makes that safe. Nothing under this lock ever touches a channel.
/// ///
/// - `monitor()` / `demonitor()` take the target's cold lock (also Leaf), so /// Eviction is *eager*: every membership's monitor delivers to the one
/// they run *before* / *after* the group lock, never under it. /// `deaths` channel, drained by a per-run reaper actor that sweeps the dead
/// - draining a monitor with `try_recv` takes the channel's Channel-class /// pid out of every group the moment its `Down` is scheduled. The read path
/// lock, which the lock order permits *under* a Leaf; a channel critical /// keeps a slot-liveness backstop for the window between a death and the
/// section only does the lock-free unpark protocol, so no Leaf ever nests /// reaper's turn.
/// under it.
///
/// Evicted and rejected [`Monitor`]s are therefore dropped only *after* the
/// group lock is released, so a receiver-drop never runs a wakeup under the
/// lock — the same discipline as `demonitor`.
pub(crate) struct ProcessGroups { pub(crate) struct ProcessGroups {
groups: HashMap<String, Vec<Membership>>, groups: HashMap<String, Vec<Membership>>,
/// The reaper's inboxes: every membership monitor is registered against
/// a clone of `deaths`. `None` until the first `join` of a run spawns
/// the reaper; a stale one (receiver gone with the previous run's
/// teardown) is detected via `receiver_alive` and replaced.
reaper: Option<ReaperInboxes>,
/// `NodeId → node name` for every peer with members in the store, kept
/// by the pg actor under this lock, so a stored remote member can be
/// rendered back to its wire identity without asking anyone.
#[cfg(feature = "cluster")]
node_names: HashMap<NodeId, String>,
} }
impl ProcessGroups { impl ProcessGroups {
pub(crate) fn new() -> Self { pub(crate) fn new() -> Self {
Self { Self {
groups: HashMap::new(), groups: HashMap::new(),
reaper: None,
#[cfg(feature = "cluster")]
node_names: HashMap::new(),
} }
} }
/// Insert `ms` into `group`. Idempotent on the *member*: if the member is /// Forget the reaper. Called at the start of every `run()` so a stopped
/// already present the new membership is handed back (`Some`) so the caller /// reaper from a previous run is never sent to; `join` respawns.
/// can tear its now-redundant monitor down outside the lock; `None` means pub(crate) fn reset_reaper(&mut self) {
/// it was inserted. self.reaper = None;
fn join(&mut self, group: &str, ms: Membership) -> Option<Membership> { }
/// Insert `ms` into `group`. Idempotent on the *member*: `false` means the
/// member was already present and nothing changed; `true` means inserted.
pub(crate) fn join(&mut self, group: &str, ms: Membership) -> bool {
let v = self.groups.entry(group.to_owned()).or_default(); let v = self.groups.entry(group.to_owned()).or_default();
if v.iter().any(|e| e.member == ms.member) { if v.iter().any(|e| e.member == ms.member) {
return Some(ms); return false;
} }
v.push(ms); v.push(ms);
None true
} }
/// Remove `member`'s membership from `group`, returning it (so the caller /// Remove `member`'s membership from `group`, returning it (so the caller
/// can `demonitor` it outside the lock). An emptied group is pruned. /// can unregister its monitor outside the lock). An emptied group is pruned.
fn leave(&mut self, group: &str, member: Member) -> Option<Membership> { pub(crate) fn leave(&mut self, group: &str, member: Member) -> Option<Membership> {
let v = self.groups.get_mut(group)?; let v = self.groups.get_mut(group)?;
let pos = v.iter().position(|e| e.member == member)?; let pos = v.iter().position(|e| e.member == member)?;
let removed = v.remove(pos); let removed = v.remove(pos);
@@ -252,20 +272,23 @@ impl ProcessGroups {
} }
/// The one dumb eviction primitive: drop every member matching `pred` from /// The one dumb eviction primitive: drop every member matching `pred` from
/// every group, pruning emptied groups, and return the evicted memberships' /// every group, pruning emptied groups, and return the evicted
/// monitors for the caller to drop outside the lock. The primitive does not /// memberships with the group each was in. The primitive does not know
/// know *why* a member leaves; that is the caller's concern. Its callers are /// *why* a member leaves; that is the caller's concern. Its callers are
/// the death hook (`reap_group`) and, once clustering lands, an /// the reaper (a local death) and the cluster's node-down / re-sync
/// incarnation-eviction sweep — both over this same predicate path, which is /// sweeps — all over this same predicate path, which is the whole reason
/// the whole reason to shape eviction as a predicate. Insertion order within /// to shape eviction as a predicate. Insertion order within a group is
/// a group is preserved (`members` / `pick` are order-stable). /// preserved (`members` / `pick` are order-stable).
fn remove_where(&mut self, mut pred: impl FnMut(&Member) -> bool) -> Vec<Monitor> { pub(crate) fn remove_where(
&mut self,
mut pred: impl FnMut(&Member) -> bool,
) -> Vec<(String, Membership)> {
let mut evicted = Vec::new(); let mut evicted = Vec::new();
self.groups.retain(|_, v| { self.groups.retain(|g, v| {
let mut i = 0; let mut i = 0;
while i < v.len() { while i < v.len() {
if pred(&v[i].member) { if pred(&v[i].member) {
evicted.push(v.remove(i).monitor); evicted.push((g.clone(), v.remove(i)));
} else { } else {
i += 1; i += 1;
} }
@@ -275,38 +298,6 @@ impl ProcessGroups {
evicted evicted
} }
/// Drain-on-contact death hook. The registry can prune a stale binding
/// lazily, on contact, because it only ever resolves one binding at a time;
/// a group is *iterated* — `members` fans out to everyone — so it must not
/// carry a dead member across a broadcast. Every group operation reaps the
/// group it touches first.
///
/// Drains every membership monitor in `group` with a non-blocking
/// `try_recv`: a delivered `Down` (any reason) or a closed channel means
/// that member is dead. On the first death detected, sweep *all* of the
/// dead pids out of *every* group via [`remove_where`] — a death is removed
/// from each group it joined, not just the one being touched. Returns the
/// evicted monitors to drop outside the lock.
fn reap_group(&mut self, group: &str) -> Vec<Monitor> {
let dead: Vec<Pid> = {
let Some(v) = self.groups.get(group) else {
return Vec::new();
};
v.iter()
.filter_map(|e| match e.monitor.rx.try_recv() {
// A Down arrived, or the channel closed and drained: dead.
Ok(Some(_)) | Err(_) => Some(e.member.pid),
// Empty but open — the sender still lives in the slot: alive.
Ok(None) => None,
})
.collect()
};
if dead.is_empty() {
return Vec::new();
}
self.remove_where(|m| dead.contains(&m.pid))
}
/// Raw enumeration of a group's members — no liveness filtering. Used by /// Raw enumeration of a group's members — no liveness filtering. Used by
/// tests to assert storage state independently of the read-path backstop. /// tests to assert storage state independently of the read-path backstop.
#[cfg(test)] #[cfg(test)]
@@ -317,16 +308,24 @@ impl ProcessGroups {
.unwrap_or_default() .unwrap_or_default()
} }
/// Live members of `group`, in insertion order. The `is_live` oracle is the /// Live members of `group` **on `node`**, in insertion order. The
/// read-path backstop: a member whose slot is already dead is /// `is_live` oracle is the read-path backstop: a member whose slot is
/// dropped from the *result* even if its `Down` has not been drained yet. /// already dead is dropped from the *result* even if the reaper has not
/// Backstop only — the entry stays in storage; eviction is the monitor's /// swept it yet. Backstop only — the entry stays in storage; eviction is
/// job (`reap_group`). /// the reaper's job. The node filter is what keeps the local API local:
fn members_where(&self, group: &str, mut is_live: impl FnMut(Pid) -> bool) -> Vec<Pid> { /// a remote member's `pid` is another node's slot bits, meaningless to
/// `is_live` and to any local send.
fn members_where(
&self,
group: &str,
node: NodeId,
mut is_live: impl FnMut(Pid) -> bool,
) -> Vec<Pid> {
self.groups self.groups
.get(group) .get(group)
.map(|v| { .map(|v| {
v.iter() v.iter()
.filter(|e| e.member.node == node)
.map(|e| e.member.pid) .map(|e| e.member.pid)
.filter(|&p| is_live(p)) .filter(|&p| is_live(p))
.collect() .collect()
@@ -334,19 +333,210 @@ impl ProcessGroups {
.unwrap_or_default() .unwrap_or_default()
} }
/// The first live member of `group` in insertion order — stateless /// The first live member of `group` on `node` in insertion order —
/// first-live `pick`, with the same read-path backstop as `members_where`. /// stateless first-live `pick`, with the same read-path backstop and node
fn first_member_where(&self, group: &str, mut is_live: impl FnMut(Pid) -> bool) -> Option<Pid> { /// filter as `members_where`.
fn first_member_where(
&self,
group: &str,
node: NodeId,
mut is_live: impl FnMut(Pid) -> bool,
) -> Option<Pid> {
self.groups self.groups
.get(group)? .get(group)?
.iter() .iter()
.filter(|e| e.member.node == node)
.map(|e| e.member.pid) .map(|e| e.member.pid)
.find(|&p| is_live(p)) .find(|&p| is_live(p))
} }
} }
/// The store's cluster-side surface: raw reads the pg actor needs to speak
/// for this node (`Sync`, membership checks) and the peer-name memo. One
/// `cfg` block: everything here exists only when there is a mesh.
#[cfg(feature = "cluster")]
impl ProcessGroups {
/// Does `group` hold `member` right now? (Raw storage, no liveness.)
pub(crate) fn contains(&self, group: &str, member: &Member) -> bool {
self.groups
.get(group)
.is_some_and(|v| v.iter().any(|e| e.member == *member))
}
/// Every stored member of `group`, any node, insertion order. Raw storage.
pub(crate) fn all_of(&self, group: &str) -> Vec<Member> {
self.groups
.get(group)
.map(|v| v.iter().map(|e| e.member).collect())
.unwrap_or_default()
}
/// `(group, [pid])` for every group with a member on `node` — the
/// `Sync` payload. Raw storage; groups with no such member are omitted.
pub(crate) fn groups_on(&self, node: NodeId) -> Vec<(String, Vec<Pid>)> {
let mut out: Vec<(String, Vec<Pid>)> = self
.groups
.iter()
.filter_map(|(g, v)| {
let pids: Vec<Pid> = v
.iter()
.filter(|e| e.member.node == node)
.map(|e| e.member.pid)
.collect();
(!pids.is_empty()).then(|| (g.clone(), pids))
})
.collect();
out.sort_by(|a, b| a.0.cmp(&b.0));
out
}
/// Record / forget the name behind a peer's `NodeId`.
pub(crate) fn set_node_name(&mut self, node: NodeId, name: String) {
self.node_names.insert(node, name);
}
pub(crate) fn forget_node_name(&mut self, node: NodeId) {
self.node_names.remove(&node);
}
pub(crate) fn node_name(&self, node: NodeId) -> Option<&str> {
self.node_names.get(&node).map(String::as_str)
}
}
/// The group reaper: one detached actor per run, spawned by the first `join`,
/// parked on the shared `deaths` inbox. Every local membership's monitor
/// delivers here, so a death is swept out of *every* group it joined as soon
/// as the reaper is scheduled — no group operation has to happen first.
/// Sweeps by `(node, pid)`: only local members, since a remote member's pid
/// bits are meaningless here. Exits when the last sender is gone, i.e. never
/// during a run (the store holds one); the run's teardown stops it like any
/// other parked actor. Spawned under `ROOT_PID` so its exit signal is absorbed
/// rather than delivered to whichever supervisor's child happened to join
/// first.
///
/// Under `cluster` the same actor is the node's **pg actor** (RFC 010 Phase
/// 5, c15): it also drains a control inbox of local join/leave announcements,
/// the membership stream and the exposed `"pg"` inbox — see
/// [`crate::cluster::pg`]. Its store-side sweep is unchanged.
#[cfg(not(feature = "cluster"))]
fn reaper(rx: crate::channel::Receiver<Down>, ctl: crate::channel::Receiver<PgEvent>) {
// No mesh: nothing to tell about joins/leaves. Drop the control inbox
// so announcements are refused at the sender rather than queued.
drop(ctl);
while let Ok(down) = rx.recv() {
sweep_local_death(down.pid);
// Evicted memberships hold only ids; their monitors have fired.
}
}
/// What the local API tells the reaper besides deaths (which arrive as
/// [`Down`] on their own inbox — that channel's type is fixed by the monitor
/// primitive, so the two cannot be one enum). The default reaper has no use
/// for these; the cluster's pg actor broadcasts them (RFC 010 Phase 5).
// The default reaper never looks inside — that is the point, not a bug.
#[cfg_attr(not(feature = "cluster"), allow(dead_code))]
pub(crate) enum PgEvent {
/// `join` inserted `pid` into `group`. The consumer re-checks the store
/// before acting on it.
Joined { group: String, pid: Pid },
/// `leave` removed `pid` from `group`.
Left { group: String, pid: Pid },
/// `cluster::start` has the manager up and the local identity set: take
/// a membership subscription, register + expose the `"pg"` name, and
/// start speaking to peers.
#[cfg(feature = "cluster")]
Attach,
}
/// Evict the local member `pid` from every group. The reaper's one store
/// operation; returns what was evicted with its group (the cluster's
/// `Leave` broadcast wants both).
pub(crate) fn sweep_local_death(pid: Pid) -> Vec<(String, Membership)> {
with_runtime(|inner| {
let node = inner.node_id;
inner
.process_groups
.lock()
.remove_where(|m| m.node == node && m.pid == pid)
})
}
/// The reaper's inboxes. `deaths` is the liveness authority for the set
/// (`ctl` is created and dropped with it, on the same actor).
#[derive(Clone)]
pub(crate) struct ReaperInboxes {
pub(crate) deaths: Sender<Down>,
/// The control inbox: local `join`/`leave` announce here (see
/// [`PgEvent`]). The default reaper closes it on entry.
pub(crate) ctl: Sender<PgEvent>,
}
impl ReaperInboxes {
fn alive(&self) -> bool {
self.deaths.receiver_alive()
}
}
/// Live senders for the reaper's inboxes, spawning the reaper if this run has
/// none yet. Two racing first-spawns may both spawn; the loser's senders drop
/// on return, its spare reaper sees a closed inbox and exits.
pub(crate) fn reaper_inboxes() -> ReaperInboxes {
let existing = with_runtime(|inner| {
let pg = inner.process_groups.lock();
pg.reaper.clone().filter(ReaperInboxes::alive)
});
if let Some(r) = existing {
return r;
}
let (tx, rx) = channel::<Down>();
let (ctl_tx, ctl_rx) = channel::<PgEvent>();
// Detached: the handle drops here. The reaper's lifetime is the run's.
// The ONE seam between the local store and the cluster: same inboxes,
// different body.
#[cfg(not(feature = "cluster"))]
let _ = spawn_under(crate::runtime::ROOT_PID, move || reaper(rx, ctl_rx));
#[cfg(feature = "cluster")]
let _ = spawn_under(crate::runtime::ROOT_PID, move || {
crate::cluster::pg::actor(rx, ctl_rx)
});
let fresh = ReaperInboxes {
deaths: tx,
ctl: ctl_tx,
};
with_runtime(|inner| {
let mut pg = inner.process_groups.lock();
match &pg.reaper {
Some(r) if r.alive() => r.clone(),
_ => {
pg.reaper = Some(fresh.clone());
fresh
}
}
})
}
/// A live sender for the reaper's `deaths` inbox (spawning it if needed).
fn deaths_sender() -> Sender<Down> {
reaper_inboxes().deaths
}
/// Announce a local group change to the reaper, if this run has one. A
/// closed inbox is the default reaper (uninterested) or a run tearing down.
fn announce(msg: PgEvent) {
let ctl = with_runtime(|inner| {
inner
.process_groups
.lock()
.reaper
.as_ref()
.map(|r| r.ctl.clone())
});
if let Some(ctl) = ctl {
let _ = ctl.send(msg);
}
}
/// Build the full member identity for `pid` from runtime identity. /// Build the full member identity for `pid` from runtime identity.
fn member_for(inner: &crate::runtime::RuntimeInner, pid: Pid) -> Member { pub(crate) fn member_for(inner: &crate::runtime::RuntimeInner, pid: Pid) -> Member {
Member { Member {
node: inner.node_id, node: inner.node_id,
incarnation: inner.incarnation, incarnation: inner.incarnation,
@@ -358,7 +548,7 @@ fn member_for(inner: &crate::runtime::RuntimeInner, pid: Pid) -> Member {
/// no lock — identical to the registry's guard. The read-path backstop: a /// no lock — identical to the registry's guard. The read-path backstop: a
/// generation is never reused, so a dead member is detectable independently of /// generation is never reused, so a dead member is detectable independently of
/// whether its monitor `Down` has been drained yet. /// whether its monitor `Down` has been drained yet.
fn live(inner: &crate::runtime::RuntimeInner, pid: Pid) -> bool { pub(crate) fn live(inner: &crate::runtime::RuntimeInner, pid: Pid) -> bool {
inner.slot_at(pid).is_some_and(|s| s.is_live_for(pid)) inner.slot_at(pid).is_some_and(|s| s.is_live_for(pid))
} }
@@ -367,61 +557,76 @@ fn live(inner: &crate::runtime::RuntimeInner, pid: Pid) -> bool {
/// added the membership, `false` if it was already a member. /// added the membership, `false` if it was already a member.
/// ///
/// Installs a monitor on `pid` so the actor's death evicts it from the group /// Installs a monitor on `pid` so the actor's death evicts it from the group
/// automatically — you never have to remove a dead member yourself. A redundant /// automatically — you never have to remove a dead member yourself. Joining a
/// (idempotent) join tears its extra monitor back down. /// pid that is already dead is accepted and evicted the same way (via a
/// `NoProc` notice), so it never shows up in a read.
/// ///
/// Panics if called outside `Runtime::run()`. /// Panics if called outside `Runtime::run()`.
pub fn join<A>(group: impl Into<String>, pid: Pid<A>) -> bool { pub fn join<A>(group: impl Into<String>, pid: Pid<A>) -> bool {
let group = group.into(); let group = group.into();
let pid = pid.erase(); let pid = pid.erase();
// Install the monitor BEFORE taking the group lock: monitor() acquires the let deaths = deaths_sender();
// target's cold lock (Leaf), and two Leaf locks are never held at once. The // Record the membership BEFORE arming its monitor: the reaper sweeps by
// registration races `finalize_actor` under that cold lock exactly as every // pid on the first `Down`, so a `Down` that could precede the entry would
// other monitor does, so no death can slip between the join and the monitor // leave a corpse in storage forever (visible to no read — the backstop
// being in place. // hides it — but a leak, and once groups are clustered a member that
let mon = monitor(pid); // would be announced). Arming after insertion means every `Down` finds
// its entry. The monitor id is allocated up front so `leave` can tear the
let (rejected, reaped) = with_runtime(|inner| { // registration down even if it lands in the tiny window before arming (an
// orphaned registration is harmless: its `Down` names a pid whose
// membership is gone, and the sweep finds nothing).
let id = with_runtime(|inner| inner.alloc_monitor_id());
let inserted = with_runtime(|inner| {
let ms = Membership { let ms = Membership {
member: member_for(inner, pid), member: member_for(inner, pid),
monitor: mon, monitor: Some(id),
}; };
let mut pg = inner.process_groups.lock(); inner.process_groups.lock().join(&group, ms)
let reaped = pg.reap_group(&group); });
let rejected = pg.join(&group, ms); if !inserted {
(rejected, reaped) return false;
}
// Tell the reaper (the cluster's pg actor re-checks the store before it
// broadcasts, so a `leave`/death that overtakes this announcement is
// never advertised as a join).
announce(PgEvent::Joined {
group: group.clone(),
pid,
}); });
// Outside the group lock: drop the reaped (dead) monitors, and if this join // Outside the group lock: registration takes the target's cold lock (Leaf).
// was redundant, demonitor + drop the extra monitor we just installed. // The registration races `finalize_actor` under that cold lock exactly as
drop(reaped); // every other monitor does, so no death can slip between the join and the
match rejected { // monitor being in place.
Some(dup) => { if !register_monitor(pid, id, &deaths) {
demonitor(&dup.monitor); // Already gone: queue the notice ourselves, exactly as `monitor` does.
false let _ = deaths.send(Down {
} pid,
None => true, reason: DownReason::NoProc,
});
} }
true
} }
/// Drop `pid`'s membership of `group`. Returns whether a membership was /// Drop `pid`'s membership of `group`. Returns whether a membership was
/// removed. The membership's monitor is demonitored and dropped. /// removed. The membership's monitor registration is torn down.
/// ///
/// Panics if called outside `Runtime::run()`. /// Panics if called outside `Runtime::run()`.
pub fn leave<A>(group: &str, pid: Pid<A>) -> bool { pub fn leave<A>(group: &str, pid: Pid<A>) -> bool {
let pid = pid.erase(); let pid = pid.erase();
let (removed, reaped) = with_runtime(|inner| { let removed = with_runtime(|inner| {
let member = member_for(inner, pid); let member = member_for(inner, pid);
let mut pg = inner.process_groups.lock(); inner.process_groups.lock().leave(group, member)
let reaped = pg.reap_group(group);
let removed = pg.leave(group, member);
(removed, reaped)
}); });
drop(reaped);
match removed { match removed {
Some(ms) => { Some(ms) => {
demonitor(&ms.monitor); if let Some(id) = ms.monitor {
unregister_monitor(pid, id);
}
announce(PgEvent::Left {
group: group.to_owned(),
pid,
});
true true
} }
None => false, None => false,
@@ -431,21 +636,19 @@ pub fn leave<A>(group: &str, pid: Pid<A>) -> bool {
/// Every live member of `group`, in the order they joined. Returns an empty /// Every live member of `group`, in the order they joined. Returns an empty
/// vector if the group does not exist or has no live members. /// vector if the group does not exist or has no live members.
/// ///
/// Dead members are never returned: the group is pruned of anything that has /// Dead members are never returned: the reaper evicts a member as soon as its
/// died before the read, and as a backstop a member whose slot is already dead /// death is processed, and as a backstop a member whose slot is already dead
/// is dropped from the result even in the brief window before its death has /// is dropped from the result even in the brief window before the reaper's
/// been fully processed. /// turn.
/// ///
/// Panics if called outside `Runtime::run()`. /// Panics if called outside `Runtime::run()`.
pub fn members(group: &str) -> Vec<Pid> { pub fn members(group: &str) -> Vec<Pid> {
let (pids, reaped) = with_runtime(|inner| { with_runtime(|inner| {
let mut pg = inner.process_groups.lock(); inner
let reaped = pg.reap_group(group); .process_groups
let pids = pg.members_where(group, |pid| live(inner, pid)); .lock()
(pids, reaped) .members_where(group, inner.node_id, |pid| live(inner, pid))
}); })
drop(reaped);
pids
} }
/// One live member of `group`, or `None` if the group is empty (or every /// One live member of `group`, or `None` if the group is empty (or every
@@ -455,14 +658,12 @@ pub fn members(group: &str) -> Vec<Pid> {
/// ///
/// Panics if called outside `Runtime::run()`. /// Panics if called outside `Runtime::run()`.
pub fn pick(group: &str) -> Option<Pid> { pub fn pick(group: &str) -> Option<Pid> {
let (picked, reaped) = with_runtime(|inner| { with_runtime(|inner| {
let mut pg = inner.process_groups.lock(); inner
let reaped = pg.reap_group(group); .process_groups
let picked = pg.first_member_where(group, |pid| live(inner, pid)); .lock()
(picked, reaped) .first_member_where(group, inner.node_id, |pid| live(inner, pid))
}); })
drop(reaped);
picked
} }
/// Typed [`pick`]: one live member of `group` as a [`Pid<A>`](Pid). /// Typed [`pick`]: one live member of `group` as a [`Pid<A>`](Pid).
@@ -505,8 +706,8 @@ pub fn dispatch<A: Addressable>(group: &str, msg: A::Msg) -> Result<Pid<A>, Send
#[cfg(test)] #[cfg(test)]
mod tests { mod tests {
use super::*; use super::*;
use crate::channel::{channel, Sender}; use crate::scheduler::spawn;
use crate::monitor::{Down, DownReason, MonitorId}; use std::time::{Duration, Instant};
fn member(index: u32, generation: u32) -> Member { fn member(index: u32, generation: u32) -> Member {
Member { Member {
@@ -516,33 +717,21 @@ mod tests {
} }
} }
/// A synthetic membership with a real (but slot-less) monitor channel. The /// A synthetic membership: the store never looks at the id.
/// returned `Sender` stands in for the slot's `Down` sender: hold it to fn synth(index: u32, generation: u32) -> Membership {
/// keep the member "alive" (`try_recv` → `Ok(None)`), `send` a `Down` to Membership {
/// simulate death, or `drop` it to simulate a drained/closed channel.
fn synth(index: u32, generation: u32) -> (Membership, Sender<Down>) {
let pid = Pid::new(index, generation);
let (tx, rx) = channel::<Down>();
let ms = Membership {
member: member(index, generation), member: member(index, generation),
monitor: Monitor { monitor: Some(MonitorId(0)),
id: MonitorId(0), }
target: pid,
rx,
},
};
(ms, tx)
} }
#[test] #[test]
fn join_is_idempotent_within_a_group() { fn join_is_idempotent_within_a_group() {
let mut pg = ProcessGroups::new(); let mut pg = ProcessGroups::new();
let (a, _ta) = synth(1, 0); assert!(pg.join("workers", synth(1, 0)), "first join inserts");
let (b, _tb) = synth(1, 0);
assert!(pg.join("workers", a).is_none(), "first join inserts");
assert!( assert!(
pg.join("workers", b).is_some(), !pg.join("workers", synth(1, 0)),
"second identical join is handed back" "second identical join is refused"
); );
assert_eq!(pg.members_of("workers"), vec![member(1, 0)]); assert_eq!(pg.members_of("workers"), vec![member(1, 0)]);
} }
@@ -550,12 +739,9 @@ mod tests {
#[test] #[test]
fn same_pid_in_many_groups_is_independent() { fn same_pid_in_many_groups_is_independent() {
let mut pg = ProcessGroups::new(); let mut pg = ProcessGroups::new();
let (a, _ta) = synth(1, 0); pg.join("a", synth(1, 0));
let (b, _tb) = synth(1, 0); pg.join("b", synth(1, 0));
let (c, _tc) = synth(2, 0); pg.join("b", synth(2, 0));
pg.join("a", a);
pg.join("b", b);
pg.join("b", c);
assert_eq!(pg.members_of("a"), vec![member(1, 0)]); assert_eq!(pg.members_of("a"), vec![member(1, 0)]);
assert_eq!(pg.members_of("b"), vec![member(1, 0), member(2, 0)]); assert_eq!(pg.members_of("b"), vec![member(1, 0), member(2, 0)]);
} }
@@ -564,11 +750,9 @@ mod tests {
fn distinct_generations_are_distinct_members() { fn distinct_generations_are_distinct_members() {
// ABA guard: same slot index, different generation = different actor. // ABA guard: same slot index, different generation = different actor.
let mut pg = ProcessGroups::new(); let mut pg = ProcessGroups::new();
let (a, _ta) = synth(1, 0); assert!(pg.join("g", synth(1, 0)));
let (b, _tb) = synth(1, 1);
assert!(pg.join("g", a).is_none());
assert!( assert!(
pg.join("g", b).is_none(), pg.join("g", synth(1, 1)),
"different generation is a distinct member" "different generation is a distinct member"
); );
assert_eq!(pg.members_of("g"), vec![member(1, 0), member(1, 1)]); assert_eq!(pg.members_of("g"), vec![member(1, 0), member(1, 1)]);
@@ -577,10 +761,8 @@ mod tests {
#[test] #[test]
fn leave_removes_one_membership_and_prunes_empty_groups() { fn leave_removes_one_membership_and_prunes_empty_groups() {
let mut pg = ProcessGroups::new(); let mut pg = ProcessGroups::new();
let (a, _ta) = synth(1, 0); pg.join("g", synth(1, 0));
let (b, _tb) = synth(2, 0); pg.join("g", synth(2, 0));
pg.join("g", a);
pg.join("g", b);
assert!(pg.leave("g", member(1, 0)).is_some()); assert!(pg.leave("g", member(1, 0)).is_some());
assert_eq!(pg.members_of("g"), vec![member(2, 0)]); assert_eq!(pg.members_of("g"), vec![member(2, 0)]);
assert!( assert!(
@@ -598,7 +780,7 @@ mod tests {
#[test] #[test]
fn remove_where_sweeps_every_group() { fn remove_where_sweeps_every_group() {
let mut pg = ProcessGroups::new(); let mut pg = ProcessGroups::new();
for (g, (m, _t)) in [ for (g, m) in [
("a", synth(1, 0)), ("a", synth(1, 0)),
("a", synth(2, 0)), ("a", synth(2, 0)),
("b", synth(1, 0)), ("b", synth(1, 0)),
@@ -616,104 +798,137 @@ mod tests {
#[test] #[test]
fn remove_where_can_match_an_incarnation_sweep() { fn remove_where_can_match_an_incarnation_sweep() {
// Shape check for the later evict_incarnation(node, inc) caller. // Shape check for the node-down / incarnation sweep caller.
let mut pg = ProcessGroups::new(); let mut pg = ProcessGroups::new();
let pid = Pid::new(1, 0); let stale = Membership {
let (tx, rx) = channel::<Down>();
let dead = Membership {
member: Member { member: Member {
node: DEFAULT_NODE_ID, node: DEFAULT_NODE_ID,
incarnation: Incarnation::new(7), incarnation: Incarnation::new(7),
pid, pid: Pid::new(1, 0),
},
monitor: Monitor {
id: MonitorId(0),
target: pid,
rx,
}, },
monitor: Some(MonitorId(0)),
}; };
let _keep = tx; pg.join("g", stale);
let (live, _tl) = synth(2, 0); pg.join("g", synth(2, 0));
pg.join("g", dead);
pg.join("g", live);
let evicted = pg.remove_where(|mem| mem.incarnation == Incarnation::new(7)); let evicted = pg.remove_where(|mem| mem.incarnation == Incarnation::new(7));
assert_eq!(evicted.len(), 1); assert_eq!(evicted.len(), 1);
assert_eq!(pg.members_of("g"), vec![member(2, 0)]); assert_eq!(pg.members_of("g"), vec![member(2, 0)]);
} }
#[test] #[test]
fn reap_keeps_live_members() { fn read_backstop_hides_a_member_the_reaper_has_not_yet_swept() {
let mut pg = ProcessGroups::new(); let mut pg = ProcessGroups::new();
let (a, _ta) = synth(1, 0); // sender held: member stays alive pg.join("g", synth(1, 0));
pg.join("a", a); pg.join("g", synth(2, 0));
assert!(pg.reap_group("a").is_empty(), "no deaths");
assert_eq!(pg.members_of("a"), vec![member(1, 0)]);
}
#[test]
fn reap_evicts_a_dead_member_and_sweeps_all_its_groups() {
let mut pg = ProcessGroups::new();
let (a1, ta1) = synth(1, 0); // pid 1 in group a
let (a2, _ta2) = synth(2, 0); // pid 2 in group a (stays alive)
let (b1, _tb1) = synth(1, 0); // pid 1 in group b
pg.join("a", a1);
pg.join("a", a2);
pg.join("b", b1);
// pid 1 dies: its group-a monitor receives a Down. Its group-b monitor
// has not — reap must still sweep pid 1 out of b by the pid predicate.
ta1.send(Down {
pid: Pid::new(1, 0),
reason: DownReason::Exit,
})
.unwrap();
let evicted = pg.reap_group("a");
assert_eq!(
evicted.len(),
2,
"pid 1's memberships in both a and b are evicted"
);
assert_eq!(pg.members_of("a"), vec![member(2, 0)]);
assert!(pg.members_of("b").is_empty(), "swept from b too; pruned");
}
#[test]
fn reap_treats_a_closed_channel_as_dead() {
let mut pg = ProcessGroups::new();
let (a, ta) = synth(1, 0);
pg.join("a", a);
drop(ta); // sender gone, queue empty → try_recv = Err(RecvError) = dead
let evicted = pg.reap_group("a");
assert_eq!(evicted.len(), 1);
assert!(pg.members_of("a").is_empty());
}
#[test]
fn read_backstop_hides_a_member_the_monitor_has_not_yet_reaped() {
let mut pg = ProcessGroups::new();
// Both senders held: reap_group would see Ok(None) and evict neither.
let (a, _ta) = synth(1, 0);
let (b, _tb) = synth(2, 0);
pg.join("g", a);
pg.join("g", b);
// The slot-word oracle already reports pid 1 dead (finalize window), // The slot-word oracle already reports pid 1 dead (finalize window),
// ahead of any Down delivery. // ahead of the reaper's turn.
let dead = Pid::new(1, 0); let dead = Pid::new(1, 0);
let oracle = |pid: Pid| pid != dead; let oracle = |pid: Pid| pid != dead;
assert_eq!( assert_eq!(
pg.members_where("g", oracle), pg.members_where("g", DEFAULT_NODE_ID, oracle),
vec![Pid::new(2, 0)], vec![Pid::new(2, 0)],
"dead pid filtered from read" "dead pid filtered from read"
); );
assert_eq!( assert_eq!(
pg.first_member_where("g", oracle), pg.first_member_where("g", DEFAULT_NODE_ID, oracle),
Some(Pid::new(2, 0)), Some(Pid::new(2, 0)),
"pick skips the dead first member" "pick skips the dead first member"
); );
// Backstop does not evict — that stays the monitor's job; raw storage // Backstop does not evict — that stays the reaper's job; raw storage
// still holds both until reap runs. // still holds both until it runs.
assert_eq!(pg.members_of("g"), vec![member(1, 0), member(2, 0)]); assert_eq!(pg.members_of("g"), vec![member(1, 0), member(2, 0)]);
} }
// ---- reaper: eager eviction against a live runtime ----
/// Raw storage view for a group, bypassing the read-path backstop.
fn stored(group: &str) -> Vec<Member> {
with_runtime(|inner| inner.process_groups.lock().members_of(group))
}
/// Cooperative wait (`smarm::sleep`, never an OS block) until `pred`.
fn wait_until(what: &str, mut pred: impl FnMut() -> bool) {
let deadline = Instant::now() + Duration::from_secs(2);
while !pred() {
assert!(Instant::now() < deadline, "timed out waiting for: {what}");
crate::sleep(Duration::from_millis(1));
}
}
#[test]
fn a_death_is_swept_from_storage_without_any_group_operation() {
crate::run(|| {
let (tx, rx) = channel::<()>();
let w = spawn(move || {
rx.recv().unwrap();
});
let pid = w.pid();
join("a", pid);
join("b", pid);
assert_eq!(stored("a"), vec![member_for_test(pid)]);
tx.send(()).unwrap();
w.join().unwrap();
// No members()/pick()/join() on a or b from here on: the reaper
// alone must clear both.
wait_until("reaper sweeps a and b", || {
stored("a").is_empty() && stored("b").is_empty()
});
});
}
#[test]
fn a_dead_at_join_pid_is_swept_from_storage() {
crate::run(|| {
let h = spawn(|| {});
let pid = h.pid();
h.join().unwrap();
assert!(join("late", pid), "join is accepted; eviction is uniform");
wait_until("reaper sweeps the NoProc member", || {
stored("late").is_empty()
});
});
}
#[test]
fn leave_then_death_does_not_disturb_a_rejoined_group() {
// A monitor unregistered by `leave` must not fire later; the pid's
// fresh membership after re-join is swept exactly once, by its own
// monitor, on death.
crate::run(|| {
let (tx, rx) = channel::<()>();
let w = spawn(move || {
rx.recv().unwrap();
});
let pid = w.pid();
join("g", pid);
assert!(leave("g", pid));
assert!(join("g", pid));
assert_eq!(members("g"), vec![pid]);
tx.send(()).unwrap();
w.join().unwrap();
wait_until("reaper sweeps g", || stored("g").is_empty());
});
}
#[test]
fn reaper_is_respawned_for_a_second_run_of_the_same_runtime() {
let rt = crate::runtime::init(crate::runtime::Config::exact(1));
let body = || {
let h = spawn(|| {});
let pid = h.pid();
h.join().unwrap();
join("g", pid);
wait_until("reaper sweeps g", || stored("g").is_empty());
};
rt.run(body);
rt.run(body);
}
fn member_for_test(pid: Pid) -> Member {
with_runtime(|inner| member_for(inner, pid))
}
} }
+42 -1
View File
@@ -76,7 +76,7 @@ pub struct Pid<A = Erased> {
impl Pid<Erased> { impl Pid<Erased> {
/// Build an untyped pid from raw numbers. The runtime mints identities /// Build an untyped pid from raw numbers. The runtime mints identities
/// here; typing happens at typed-actor boundaries via [`Pid::from_raw`]. /// here; typing happens at typed-actor boundaries via `Pid::from_raw`.
#[inline] #[inline]
pub const fn new(index: u32, generation: u32) -> Self { pub const fn new(index: u32, generation: u32) -> Self {
Self { Self {
@@ -292,3 +292,44 @@ mod typed_pid_tests {
assert_send_sync::<Name<CounterMsg>>(); assert_send_sync::<Name<CounterMsg>>();
} }
} }
// ---- RFC 010 c10: pids auto-serialize (cluster feature) ---------------------
/// A local `Pid<A>` serializes as a
/// [`RemotePid<A>`](crate::cluster::remote::RemotePid): the wire form stamps
/// this node's name and incarnation from the ambient runtime, so a pid can
/// sit inside any message field and reply-to needs no ceremony (RFC 010 §3,
/// "sugar not a bear trap"). Serializing a pid also marks it **watchable**
/// — the wire crossing is the cluster's `mark_watchable` set-site (D12), the
/// exact analog of the membrane crossing.
///
/// Must run inside `run()` (the ambient identity lives on the runtime); a
/// runtime without a cluster identity cannot serialize a pid at all — it is
/// a serialize error, surfacing as the send's `Encode` failure — rather than
/// a `("", 0)` stamp that every peer would silently drop.
#[cfg(feature = "cluster")]
impl<A: 'static> serde::Serialize for Pid<A> {
fn serialize<S: serde::Serializer>(&self, s: S) -> Result<S::Ok, S::Error> {
crate::cluster::remote::RemotePid::<A>::from_local(*self)
.ok_or_else(|| serde::ser::Error::custom("pid serialized with no local node identity"))?
.serialize(s)
}
}
/// Deserializing into a `Pid<A>` is the **collapse**: it succeeds only when
/// the wire pid names this very node (name and incarnation both), and is a
/// decode error otherwise — a foreign pid cannot become a local `Pid`.
/// Fields that may hold a pid from anywhere are `RemotePid<A>`.
#[cfg(feature = "cluster")]
impl<'de, A: 'static> serde::Deserialize<'de> for Pid<A> {
fn deserialize<D: serde::Deserializer<'de>>(d: D) -> Result<Self, D::Error> {
let rp = crate::cluster::remote::RemotePid::<A>::deserialize(d)?;
rp.local().ok_or_else(|| {
serde::de::Error::custom(format!(
"pid {}@{} is not local to this node",
rp.index(),
rp.node()
))
})
}
}
+53 -6
View File
@@ -105,18 +105,36 @@ pub(crate) fn clear_current_slot() {
/// (RFC 019 §7), which additionally relies on this being a plain load of a /// (RFC 019 §7), which additionally relies on this being a plain load of a
/// const-initialized TLS Cell (no lazy init, no allocation, no dtor): safe /// const-initialized TLS Cell (no lazy init, no allocation, no dtor): safe
/// from a signal handler. /// from a signal handler.
#[inline] #[inline(never)]
pub(crate) fn current_slot_ptr() -> *const crate::runtime::Slot { pub(crate) fn current_slot_ptr() -> *const crate::runtime::Slot {
crate::context::tls_fence();
CURRENT_SLOT.with(|c| c.get()) CURRENT_SLOT.with(|c| c.get())
} }
/// Swap the preemption gate, returning the previous value. The one accessor
/// for `PREEMPTION_ENABLED` from actor context (`NoPreempt`, `RawMutex`,
/// `with_runtime`, trace): `#[inline(never)]` + fence, see `context` docs.
#[inline(never)]
pub(crate) fn preemption_swap(enabled: bool) -> bool {
crate::context::tls_fence();
PREEMPTION_ENABLED.with(|c| c.replace(enabled))
}
/// Read the preemption gate (debug assertions on the queue paths).
#[inline(never)]
pub(crate) fn preemption_enabled() -> bool {
crate::context::tls_fence();
PREEMPTION_ENABLED.with(|c| c.get())
}
/// RFC 007 (`smarm-causal`) — push the slice start forward by `cycles`, so /// RFC 007 (`smarm-causal`) — push the slice start forward by `cycles`, so
/// virtually-injected delay spun inside `maybe_preempt` does not count against /// virtually-injected delay spun inside `maybe_preempt` does not count against
/// the actor's timeslice (the clock-correction half of the RFC: the runtime /// the actor's timeslice (the clock-correction half of the RFC: the runtime
/// owns this clock, so it can subtract its own perturbation). /// owns this clock, so it can subtract its own perturbation).
#[cfg(feature = "smarm-causal")] #[cfg(feature = "smarm-causal")]
#[inline] #[inline(never)]
pub(crate) fn extend_timeslice(cycles: u64) { pub(crate) fn extend_timeslice(cycles: u64) {
crate::context::tls_fence();
TIMESLICE_START.with(|c| c.set(c.get().wrapping_add(cycles))); TIMESLICE_START.with(|c| c.set(c.get().wrapping_add(cycles)));
} }
@@ -124,8 +142,9 @@ pub(crate) fn extend_timeslice(cycles: u64) {
/// no-op if no actor is bound (the scheduler's own stack). Reached only from /// no-op if no actor is bound (the scheduler's own stack). Reached only from
/// the slice-expiry branch, which is already the yield path, so its cost is /// the slice-expiry branch, which is already the yield path, so its cost is
/// irrelevant. /// irrelevant.
#[inline] #[inline(never)]
fn note_overrun() { fn note_overrun() {
crate::context::tls_fence();
let p = CURRENT_SLOT.with(|c| c.get()); let p = CURRENT_SLOT.with(|c| c.get());
// SAFETY: `p` is null (no actor on-CPU) or a pointer to the on-CPU actor's // SAFETY: `p` is null (no actor on-CPU) or a pointer to the on-CPU actor's
// slot in the fixed slab. The slot is not reclaimed while the actor runs // slot in the fixed slab. The slot is not reclaimed while the actor runs
@@ -141,8 +160,9 @@ fn note_overrun() {
/// outside an actor (null slot). One TLS load + one Relaxed load/store on a /// outside an actor (null slot). One TLS load + one Relaxed load/store on a
/// cache line the receiving thread already owns — no atomic RMW, no lock. Same /// cache line the receiving thread already owns — no atomic RMW, no lock. Same
/// slot-lifetime safety argument as `note_overrun`. /// slot-lifetime safety argument as `note_overrun`.
#[inline] #[inline(never)]
pub(crate) fn note_message_received() { pub(crate) fn note_message_received() {
crate::context::tls_fence();
let p = CURRENT_SLOT.with(|c| c.get()); let p = CURRENT_SLOT.with(|c| c.get());
if !p.is_null() { if !p.is_null() {
unsafe { (*p).record_message() }; unsafe { (*p).record_message() };
@@ -158,8 +178,9 @@ pub(crate) fn note_message_received() {
/// Called from `maybe_preempt` (amortised, the `check!()`/alloc path) and from /// Called from `maybe_preempt` (amortised, the `check!()`/alloc path) and from
/// the wakeup side of every blocking park (`park_current`/`yield_now`), which /// the wakeup side of every blocking park (`park_current`/`yield_now`), which
/// is past the prep-to-park window — so it can never lose a wakeup. /// is past the prep-to-park window — so it can never lose a wakeup.
#[inline] #[inline(never)]
pub fn check_cancelled() { pub fn check_cancelled() {
crate::context::tls_fence();
let p = CURRENT_STOP.with(|c| c.get()); let p = CURRENT_STOP.with(|c| c.get());
// SAFETY: `p` is either null (no actor on-CPU — the scheduler clears it on // SAFETY: `p` is either null (no actor on-CPU — the scheduler clears it on
// every return) or a pointer into the on-CPU actor's `Arc<AtomicBool>` // every return) or a pointer into the on-CPU actor's `Arc<AtomicBool>`
@@ -198,8 +219,27 @@ pub(crate) fn elapsed_slice_cycles() -> u64 {
rdtsc().saturating_sub(TIMESLICE_START.with(|c| c.get())) rdtsc().saturating_sub(TIMESLICE_START.with(|c| c.get()))
} }
/// Read the TSC, unserialised. The core may execute this before earlier
/// instructions retire (or after later ones start), so a single stamp can
/// land tens of cycles — worst case a stalled load's worth — early or late.
/// That is negligible against every consumer in this module: the timeslice
/// arm/expiry compare against a ~10^5-cycle slice, and an early stamp only
/// makes a slice look *more* used (expires marginally sooner, never later).
/// The `lfence` this used to carry was a pipeline drain paid on every
/// resume; the one place that needs it is causal-site attribution, which
/// opts in via [`rdtsc_serialising`].
#[inline(always)] #[inline(always)]
pub fn rdtsc() -> u64 { pub fn rdtsc() -> u64 {
// SAFETY: x86-64 only (this crate is x86-64 Linux only).
unsafe { core::arch::x86_64::_rdtsc() }
}
/// Read the TSC after all prior instructions have completed locally.
/// Use where a stamp bounds an interval attributed to *code* — a speculative
/// early read would credit the tail of that code to whatever comes next.
/// Costs a pipeline drain; keep it off the per-resume path.
#[inline(always)]
pub fn rdtsc_serialising() -> u64 {
unsafe { unsafe {
// SAFETY: x86-64 only. `lfence` serialises the instruction stream so // SAFETY: x86-64 only. `lfence` serialises the instruction stream so
// we don't measure time before prior instructions retire. // we don't measure time before prior instructions retire.
@@ -247,8 +287,13 @@ unsafe impl GlobalAlloc for PreemptingAllocator {
/// the actor would then park, and the wakeup would be lost. Library /// the actor would then park, and the wakeup would be lost. Library
/// code that touches the parking primitives must keep its prep-to-park /// code that touches the parking primitives must keep its prep-to-park
/// regions allocation-free and check!()-free. /// regions allocation-free and check!()-free.
#[inline(always)] ///
/// `#[inline(never)]`: this touches thread-locals and can switch threads in
/// the middle; inlined into a caller's loop the TLS base would be hoisted
/// across the switch (see `context` module docs). The call is the price.
#[inline(never)]
pub fn maybe_preempt() { pub fn maybe_preempt() {
crate::context::tls_fence();
ALLOC_COUNT.with(|c| { ALLOC_COUNT.with(|c| {
let n = c.get(); let n = c.get();
if n == 0 { if n == 0 {
@@ -297,7 +342,9 @@ pub fn maybe_preempt() {
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
/// Force-expire the timeslice so the next RDTSC check preempts. /// Force-expire the timeslice so the next RDTSC check preempts.
#[inline(never)]
pub fn expire_timeslice_for_test() { pub fn expire_timeslice_for_test() {
crate::context::tls_fence();
TIMESLICE_START.with(|c| c.set(0)); TIMESLICE_START.with(|c| c.set(0));
ALLOC_COUNT.with(|c| c.set(0)); ALLOC_COUNT.with(|c| c.set(0));
} }
+12 -4
View File
@@ -75,8 +75,13 @@ thread_local! {
static CHANNELS_HELD: std::cell::Cell<u32> = const { std::cell::Cell::new(0) }; static CHANNELS_HELD: std::cell::Cell<u32> = const { std::cell::Cell::new(0) };
} }
#[inline] // Debug-only TLS bookkeeping: out of line only when it has a body (release
// builds must not pay a call for an empty function).
#[cfg_attr(debug_assertions, inline(never))]
#[cfg_attr(not(debug_assertions), inline(always))]
fn order_check_acquire(class: LockClass) { fn order_check_acquire(class: LockClass) {
#[cfg(debug_assertions)]
crate::context::tls_fence();
#[cfg(debug_assertions)] #[cfg(debug_assertions)]
match class { match class {
LockClass::Leaf => LEAVES_HELD.with(|l| { LockClass::Leaf => LEAVES_HELD.with(|l| {
@@ -111,8 +116,11 @@ fn order_check_acquire(class: LockClass) {
let _ = class; let _ = class;
} }
#[inline] #[cfg_attr(debug_assertions, inline(never))]
#[cfg_attr(not(debug_assertions), inline(always))]
fn order_check_release(class: LockClass) { fn order_check_release(class: LockClass) {
#[cfg(debug_assertions)]
crate::context::tls_fence();
#[cfg(debug_assertions)] #[cfg(debug_assertions)]
match class { match class {
LockClass::Leaf => LEAVES_HELD.with(|c| c.set(c.get() - 1)), LockClass::Leaf => LEAVES_HELD.with(|c| c.set(c.get() - 1)),
@@ -157,7 +165,7 @@ impl<T> RawMutex<T> {
pub(crate) fn lock(&self) -> RawMutexGuard<'_, T> { pub(crate) fn lock(&self) -> RawMutexGuard<'_, T> {
// Enter NoPreempt *before* acquiring, so a preemption can't fire // Enter NoPreempt *before* acquiring, so a preemption can't fire
// between acquisition and guard construction. // between acquisition and guard construction.
let prev_preempt = crate::preempt::PREEMPTION_ENABLED.with(|c| c.replace(false)); let prev_preempt = crate::preempt::preemption_swap(false);
order_check_acquire(self.class); order_check_acquire(self.class);
if self if self
.state .state
@@ -237,7 +245,7 @@ impl<T> Drop for RawMutexGuard<'_, T> {
self.m.unlock(); self.m.unlock();
order_check_release(self.m.class); order_check_release(self.m.class);
// Restore preemption only after the lock is released. // Restore preemption only after the lock is released.
crate::preempt::PREEMPTION_ENABLED.with(|c| c.set(self.prev_preempt)); crate::preempt::preemption_swap(self.prev_preempt);
} }
} }
+38 -1
View File
@@ -10,7 +10,7 @@
//! directly, without ever having been handed a `Pid`. //! directly, without ever having been handed a `Pid`.
//! //!
//! ``` //! ```
//! use smarm::{channel, register, run, send, spawn, unregister, whereis, Name}; //! use smarm::{channel, register, run, send, spawn, whereis, Name};
//! //!
//! const COUNTER: Name<u64> = Name::new("counter"); //! const COUNTER: Name<u64> = Name::new("counter");
//! //!
@@ -215,6 +215,7 @@ impl<M> std::error::Error for SendError<M> {}
trait ErasedSender: Send { trait ErasedSender: Send {
fn as_any(&self) -> &dyn Any; fn as_any(&self) -> &dyn Any;
fn queued_len(&self) -> usize; fn queued_len(&self) -> usize;
fn receiver_alive(&self) -> bool;
} }
impl<M: Send + 'static> ErasedSender for Sender<M> { impl<M: Send + 'static> ErasedSender for Sender<M> {
@@ -224,6 +225,9 @@ impl<M: Send + 'static> ErasedSender for Sender<M> {
fn queued_len(&self) -> usize { fn queued_len(&self) -> usize {
Sender::queued_len(self) Sender::queued_len(self)
} }
fn receiver_alive(&self) -> bool {
Sender::receiver_alive(self)
}
} }
/// One typed channel of an actor, type-erased. Concretely a `Sender<M>` filed /// One typed channel of an actor, type-erased. Concretely a `Sender<M>` filed
@@ -442,6 +446,18 @@ pub(crate) fn register_with<M: Send + 'static>(
/// name) and [`install`] (which does not). A leftover mailbox at this slot /// name) and [`install`] (which does not). A leftover mailbox at this slot
/// index from a dead prior incarnation (pid mismatch) is replaced wholesale. /// index from a dead prior incarnation (pid mismatch) is replaced wholesale.
/// Caller holds the registry lock and has established that `me` is live. /// Caller holds the registry lock and has established that `me` is live.
///
/// **One channel per message type per actor.** Publishing a second `M`
/// channel on the same live actor replaces the first — and if the first's
/// receiver is still alive, that replacement drops its last sender, closing
/// it, and any `recv`/`select` on it then returns "closed" immediately and
/// forever: a silent hot loop that starves the scheduler. That is never
/// intended, so it panics here (found the hard way in RFC 010 c9, where two
/// `Name<String>`s registered on one actor did exactly this). Replacing a
/// channel whose receiver is already gone is fine (an actor re-registering
/// after dropping its old inbox) and stays silent. To hold two names of the
/// same type, register them from two actors, or bind both names to one
/// cloned sender.
fn publish_channel<M: Send + 'static>(reg: &mut Registry, me: Pid, tx: Sender<M>) { fn publish_channel<M: Send + 'static>(reg: &mut Registry, me: Pid, tx: Sender<M>) {
let mb = reg let mb = reg
.by_index .by_index
@@ -450,6 +466,16 @@ fn publish_channel<M: Send + 'static>(reg: &mut Registry, me: Pid, tx: Sender<M>
if mb.pid != me { if mb.pid != me {
*mb = Mailbox::new(me); *mb = Mailbox::new(me);
} }
if let Some(existing) = mb.channels.get(&TypeId::of::<M>()) {
assert!(
!existing.sender.receiver_alive() || same_channel::<M>(existing, &tx),
"smarm: actor {me:?} already publishes a live channel for message type `{}`; \
a second one would replace and CLOSE the first (its receiver would then \
read as closed forever). Register the second name from another actor, or \
bind both names to a clone of the same sender.",
type_name::<M>()
);
}
mb.channels.insert( mb.channels.insert(
TypeId::of::<M>(), TypeId::of::<M>(),
Channel { Channel {
@@ -459,6 +485,17 @@ fn publish_channel<M: Send + 'static>(reg: &mut Registry, me: Pid, tx: Sender<M>
); );
} }
/// True if `existing` and `tx` are senders of the very same channel (a
/// cloned sender bound under a second name is the sanctioned way to hold two
/// names of one type on one actor).
fn same_channel<M: Send + 'static>(existing: &Channel, tx: &Sender<M>) -> bool {
existing
.sender
.as_any()
.downcast_ref::<Sender<M>>()
.is_some_and(|old| old.same_channel(tx))
}
/// Publish the current actor's `Sender<A::Msg>` into its mailbox **without** /// Publish the current actor's `Sender<A::Msg>` into its mailbox **without**
/// binding a name, and hand back the typed [`Pid<A>`] that addresses this /// binding a name, and hand back the typed [`Pid<A>`] that addresses this
/// actor directly. /// actor directly.
+162 -23
View File
@@ -2,9 +2,9 @@
//! features (no runtime dispatch — the scheduler's pop loop is the hottest //! features (no runtime dispatch — the scheduler's pop loop is the hottest
//! code in the runtime): //! code in the runtime):
//! //!
//! - `rq-mutex` (default) — `Mutex<VecDeque>`. The control/baseline: //! - `rq-mutex` — `Mutex<VecDeque>`. The control/baseline:
//! strictly FIFO, trivially correct, one global lock. //! strictly FIFO, trivially correct, one global lock.
//! - `rq-mpmc` — a single hand-rolled Vyukov bounded MPMC ring (per-cell //! - `rq-mpmc` (default) — a single hand-rolled Vyukov bounded MPMC ring (per-cell
//! sequence numbers). Strict FIFO, lock-free, one hot //! sequence numbers). Strict FIFO, lock-free, one hot
//! enqueue/dequeue cache-line pair. //! enqueue/dequeue cache-line pair.
//! - `rq-striped` — M Vyukov rings with fetch-add ticket distribution. //! - `rq-striped` — M Vyukov rings with fetch-add ticket distribution.
@@ -18,7 +18,7 @@
//! //!
//! All variants are compiled unconditionally (so every build runs every //! All variants are compiled unconditionally (so every build runs every
//! variant's unit tests); the feature only picks which one the runtime uses //! variant's unit tests); the feature only picks which one the runtime uses
//! via the [`RunQueue`] alias. //! via the `RunQueue` alias.
//! //!
//! # Contract (shared by all variants) //! # Contract (shared by all variants)
//! //!
@@ -51,7 +51,7 @@
//! - `len()` is approximate (stats only). //! - `len()` is approximate (stats only).
use crate::pid::Pid; use crate::pid::Pid;
use crate::sync_shim::{AtomicUsize, Ordering, UnsafeCell}; use crate::sync_shim::{fence, AtomicUsize, Ordering, UnsafeCell};
use std::mem::MaybeUninit; use std::mem::MaybeUninit;
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -61,20 +61,20 @@ use std::mem::MaybeUninit;
#[cfg(not(any(feature = "rq-mutex", feature = "rq-mpmc", feature = "rq-striped")))] #[cfg(not(any(feature = "rq-mutex", feature = "rq-mpmc", feature = "rq-striped")))]
compile_error!( compile_error!(
"smarm: no run queue selected. Enable exactly one of the features \ "smarm: no run queue selected. Enable exactly one of the features \
`rq-mutex` (default), `rq-mpmc`, `rq-striped`." `rq-mpmc` (default), `rq-mutex`, `rq-striped`."
); );
#[cfg(all(feature = "rq-mutex", feature = "rq-mpmc"))] #[cfg(all(feature = "rq-mutex", feature = "rq-mpmc"))]
compile_error!( compile_error!(
"smarm: features `rq-mutex` and `rq-mpmc` are mutually exclusive \ "smarm: features `rq-mutex` and `rq-mpmc` are mutually exclusive \
(use --no-default-features to drop the default `rq-mutex`)." (use --no-default-features to drop the default `rq-mpmc`)."
); );
#[cfg(all(feature = "rq-mutex", feature = "rq-striped"))] #[cfg(all(feature = "rq-mutex", feature = "rq-striped"))]
compile_error!( compile_error!("smarm: features `rq-mutex` and `rq-striped` are mutually exclusive.");
"smarm: features `rq-mutex` and `rq-striped` are mutually exclusive \
(use --no-default-features to drop the default `rq-mutex`)."
);
#[cfg(all(feature = "rq-mpmc", feature = "rq-striped"))] #[cfg(all(feature = "rq-mpmc", feature = "rq-striped"))]
compile_error!("smarm: features `rq-mpmc` and `rq-striped` are mutually exclusive."); compile_error!(
"smarm: features `rq-mpmc` and `rq-striped` are mutually exclusive \
(use --no-default-features to drop the default `rq-mpmc`)."
);
#[cfg(feature = "rq-mutex")] #[cfg(feature = "rq-mutex")]
pub(crate) type RunQueue = MutexQueue; pub(crate) type RunQueue = MutexQueue;
@@ -86,12 +86,52 @@ pub(crate) type RunQueue = StripedRing;
#[inline] #[inline]
fn assert_no_preempt() { fn assert_no_preempt() {
debug_assert!( debug_assert!(
!crate::preempt::PREEMPTION_ENABLED.with(|c| c.get()), !crate::preempt::preemption_enabled(),
"run-queue op with preemption enabled — a switch mid-op stalls or \ "run-queue op with preemption enabled — a switch mid-op stalls or \
corrupts the queue; route through with_runtime or scheduler context" corrupts the queue; route through with_runtime or scheduler context"
); );
} }
/// Escalating wait for transient ring stalls: a peer preempted by the OS
/// inside its ~100ns claim→publish window (finding 13). Spin first (the
/// common stall is a peer that is merely slow, gone within a few hundred
/// cycles), then donate the timeslice — under oversubscription pure
/// spinning STARVES the descheduled peer of the CPU it needs to publish
/// (measured in the finding-13 soak: 10⁶ pure spins can outlast the very
/// stall they prolong). OS-level yielding is orthogonal to
/// `assert_no_preempt`, which guards smarm signal preemption only.
struct Backoff(u32);
impl Backoff {
const SPIN_LIMIT: u32 = 6;
fn new() -> Self {
Self(0)
}
/// Total waits so far — lets bounded callers cap the yield phase.
fn steps(&self) -> u32 {
self.0
}
#[inline]
fn wait(&mut self) {
#[cfg(loom)]
// Loom has no notion of spinning time; every wait is a scheduling
// point so the model explores the stalled peer's progress.
loom::thread::yield_now();
#[cfg(not(loom))]
if self.0 <= Self::SPIN_LIMIT {
for _ in 0..1u32 << self.0 {
std::hint::spin_loop();
}
} else {
std::thread::yield_now();
}
self.0 += 1;
}
}
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// rq-mutex — the baseline // rq-mutex — the baseline
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -204,15 +244,57 @@ impl MpmcRing {
pub fn push(&self, pid: Pid) { pub fn push(&self, pid: Pid) {
assert_no_preempt(); assert_no_preempt();
assert!( if self.try_push(pid) {
self.try_push(pid), return;
"smarm: run queue overflow — occupancy exceeded the slab bound, \ }
which the at-most-once-enqueued invariant forbids. This is a \ self.push_slow(pid);
runtime bug (double enqueue), not a capacity tuning problem."
);
} }
/// One full claim attempt; `false` only if the ring is full. /// Cold path: the cell at `enqueue_pos` is still a lap behind. Two
/// worlds are indistinguishable at the cell (finding 13): a consumer
/// preempted between its `dequeue_pos` claim and its seq release while
/// the ring lapped onto that cell (transient), or a genuine occupancy
/// overflow (a runtime bug). The COUNTERS discriminate — the same shape
/// as crossbeam `ArrayQueue::push`'s fence + opposite-counter check:
/// occupancy < capacity ⇒ transient ⇒ wait for the stalled peer;
/// occupancy ≥ capacity ⇒ the at-most-once-enqueued invariant really is
/// broken ⇒ panic (a legal push starts from occupancy ≤ max_actors − 1).
#[cold]
fn push_slow(&self, pid: Pid) {
let mut backoff = Backoff::new();
loop {
// Order the counter reads after the failed cell read.
// `enqueue_pos` is loaded BEFORE `dequeue_pos`, so a pop racing
// us can only make the computed occupancy an UNDERestimate —
// conservative in the safe direction (never a spurious panic;
// a genuine violation is persistent and caught next lap).
fence(Ordering::SeqCst);
let enq = self.enqueue_pos.0.load(Ordering::SeqCst);
let deq = self.dequeue_pos.0.load(Ordering::SeqCst);
let occ = enq.wrapping_sub(deq);
assert!(
occ <= self.mask,
"smarm: run queue occupancy {} reached capacity {} — a pid \
was enqueued more than once (or more pids exist than \
max_actors); the at-most-once-enqueued invariant is broken \
(enq={} deq={})",
occ,
self.mask + 1,
enq,
deq
);
backoff.wait();
if self.try_push(pid) {
return;
}
}
}
/// One full claim attempt; `false` when the cell at `enqueue_pos` is
/// still a lap behind — which means EITHER genuinely full OR a
/// lap-stalled consumer (finding 13). The cell cannot tell the two
/// apart; callers must disambiguate via the counters (`push_slow`) or
/// tolerate refusal (`StripedRing`'s probe).
fn try_push(&self, pid: Pid) -> bool { fn try_push(&self, pid: Pid) -> bool {
let mut pos = self.enqueue_pos.0.load(Ordering::Relaxed); let mut pos = self.enqueue_pos.0.load(Ordering::Relaxed);
loop { loop {
@@ -246,6 +328,7 @@ impl MpmcRing {
pub fn pop(&self) -> Option<Pid> { pub fn pop(&self) -> Option<Pid> {
assert_no_preempt(); assert_no_preempt();
let mut backoff = Backoff::new();
let mut pos = self.dequeue_pos.0.load(Ordering::Relaxed); let mut pos = self.dequeue_pos.0.load(Ordering::Relaxed);
loop { loop {
let cell = &self.buf[pos & self.mask]; let cell = &self.buf[pos & self.mask];
@@ -270,9 +353,31 @@ impl MpmcRing {
Err(actual) => pos = actual, Err(actual) => pos = actual,
} }
} else if diff < 0 { } else if diff < 0 {
// Empty (or the producer at this cell hasn't published yet — // The cell at `pos` is unpublished: either the queue is
// a snapshot miss the caller's idle-retry loop absorbs). // empty, or the producer that claimed it was preempted
return None; // inside its claim→publish window (finding 13). The cell
// cannot tell the two apart; the counters can.
fence(Ordering::SeqCst);
let deq = self.dequeue_pos.0.load(Ordering::SeqCst);
if deq != pos {
// Stale head — no verdict; re-probe at the real head.
pos = deq;
continue;
}
if self.enqueue_pos.0.load(Ordering::SeqCst) == pos {
return None; // counters agree: genuinely empty
}
// Producer mid-publish. Wait briefly, then report None
// anyway — a DELIBERATE bounded deviation from crossbeam's
// unbounded retry: a spurious None is correctness-benign
// here (the stalled push completes and its RFC 018
// enqueue-wake re-wakes a parked scheduler), a parked
// scheduler beats a yielding one, and StripedRing's pop
// probe must not hang on one stripe.
if backoff.steps() > Backoff::SPIN_LIMIT + 8 {
return None;
}
backoff.wait();
} else { } else {
pos = self.dequeue_pos.0.load(Ordering::Relaxed); pos = self.dequeue_pos.0.load(Ordering::Relaxed);
} }
@@ -342,6 +447,7 @@ impl StripedRing {
// guarantees a free stripe exists, so the outer loop terminates. // guarantees a free stripe exists, so the outer loop terminates.
// The retry-from-home lap handles the racy case where every stripe // The retry-from-home lap handles the racy case where every stripe
// momentarily refused us. // momentarily refused us.
let mut backoff = Backoff::new();
loop { loop {
for i in 0..=self.stripe_mask { for i in 0..=self.stripe_mask {
let s = &self.stripes[(home + i) & self.stripe_mask]; let s = &self.stripes[(home + i) & self.stripe_mask];
@@ -349,7 +455,13 @@ impl StripedRing {
return; return;
} }
} }
std::hint::spin_loop(); // Every stripe refused this lap: either transiently full (the
// headroom argument above guarantees a genuinely free stripe
// exists) or the probes landed on lap-stalled cells (finding
// 13 — try_push cannot tell the two apart). Waiting is correct
// either way; the backoff escalates to an OS yield so stalled
// consumers get the CPU they need to release their cells.
backoff.wait();
} }
} }
@@ -560,6 +672,33 @@ mod loom_tests {
}); });
} }
/// Finding 13 regression: a consumer stalled between its dequeue_pos
/// claim and its seq release must NOT make a lapping producer conclude
/// "full" (the old code asserted here — occupancy never exceeded the
/// bound; the cell just hadn't been recycled). Capacity 2: fill, pop
/// once on each thread, then push a third element — in the
/// interleavings where a pop's release is still pending, that push
/// laps onto the stalled cell and must wait, not die.
#[test]
fn mpmc_lap_onto_stalled_consumer_completes() {
loom::model(|| {
let q = Arc::new(MpmcRing::with_capacity(2));
q.push(pid(0));
q.push(pid(1));
let q2 = q.clone();
let h = thread::spawn(move || q2.pop().expect("ring has two elements"));
let a = q.pop().expect("ring has two elements");
// enqueue_pos = 2 → cell 0: laps onto the other thread's cell
// whenever its release is delayed.
q.push(pid(2));
let b = h.join().unwrap();
let c = q.pop().expect("the lapping push must have landed");
let mut got = vec![a.index(), b.index(), c.index()];
got.sort_unstable();
assert_eq!(got, vec![0, 1, 2]);
});
}
/// Producer races a consumer on a single element: the consumer either /// Producer races a consumer on a single element: the consumer either
/// gets it or sees a clean None — never a torn/duplicated element. /// gets it or sees a clean None — never a torn/duplicated element.
#[test] #[test]
+310 -78
View File
@@ -76,7 +76,7 @@
//! time (`rq-mutex` / `rq-mpmc` / `rq-striped`). Queue ops require //! time (`rq-mutex` / `rq-mpmc` / `rq-striped`). Queue ops require
//! preemption disabled (debug-asserted there); when the mutex variant is in //! preemption disabled (debug-asserted there); when the mutex variant is in
//! play it is the innermost lock — nothing else is acquired under it. //! play it is the innermost lock — nothing else is acquired under it.
//! - Per-slot `cold` locks ([`RawMutex`], non-poisoning, guard enters //! - Per-slot `cold` locks (`RawMutex`, non-poisoning, guard enters
//! `NoPreempt`) guard the lifecycle collections. **Leaf rule: never hold //! `NoPreempt`) guard the lifecycle collections. **Leaf rule: never hold
//! two cold locks at once** — `finalize_actor`'s link cascade and `link()` //! two cold locks at once** — `finalize_actor`'s link cascade and `link()`
//! lock peers one at a time (correctness arguments at the call sites). //! lock peers one at a time (correctness arguments at the call sites).
@@ -116,7 +116,7 @@ use crate::actor::{
take_last_outcome, Actor, Outcome, take_last_outcome, Actor, Outcome,
}; };
use crate::channel::Sender; use crate::channel::Sender;
use crate::context::{get_actor_sp, set_actor_sp, switch_to_actor}; use crate::context::switch_to_actor;
use crate::io::IoThread; use crate::io::IoThread;
use crate::monitor::{Down, DownReason, MonitorId}; use crate::monitor::{Down, DownReason, MonitorId};
use crate::pid::Pid; use crate::pid::Pid;
@@ -127,7 +127,7 @@ use crate::supervisor::Signal;
use crate::timer::Timers; use crate::timer::Timers;
use std::sync::atomic::{AtomicBool, AtomicPtr, AtomicU32, AtomicU64, AtomicUsize, Ordering}; use std::sync::atomic::{AtomicBool, AtomicPtr, AtomicU32, AtomicU64, AtomicUsize, Ordering};
use std::sync::{Arc, Mutex}; use std::sync::{Arc, Mutex, Weak};
use std::thread; use std::thread;
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -144,13 +144,13 @@ pub const DEFAULT_MAX_ACTORS: usize = 16_384;
/// use smarm::runtime::Config; /// use smarm::runtime::Config;
/// ///
/// // Use all available CPUs (default): /// // Use all available CPUs (default):
/// let c = Config::default(); /// let _all = Config::default();
/// ///
/// // Exactly 4 scheduler threads: /// // Exactly 4 scheduler threads:
/// let c = Config::exact(4); /// let _four = Config::exact(4);
/// ///
/// // Between 2 and 8, clamped to available parallelism: /// // Between 2 and 8, clamped to available parallelism:
/// let c = Config::new(2, 8, None); /// let _clamped = Config::new(2, 8, None);
/// ``` /// ```
#[derive(Clone, Debug)] #[derive(Clone, Debug)]
pub struct Config { pub struct Config {
@@ -182,7 +182,7 @@ impl Config {
stack_reserve: DEFAULT_STACK_RESERVE, stack_reserve: DEFAULT_STACK_RESERVE,
stack_guard: DEFAULT_STACK_GUARD, stack_guard: DEFAULT_STACK_GUARD,
max_actors: DEFAULT_MAX_ACTORS, max_actors: DEFAULT_MAX_ACTORS,
wake_slot: false, wake_slot: true,
node_id: crate::pg::DEFAULT_NODE_ID, node_id: crate::pg::DEFAULT_NODE_ID,
incarnation: crate::pg::DEFAULT_INCARNATION, incarnation: crate::pg::DEFAULT_INCARNATION,
} }
@@ -205,7 +205,7 @@ impl Config {
stack_reserve: DEFAULT_STACK_RESERVE, stack_reserve: DEFAULT_STACK_RESERVE,
stack_guard: DEFAULT_STACK_GUARD, stack_guard: DEFAULT_STACK_GUARD,
max_actors: DEFAULT_MAX_ACTORS, max_actors: DEFAULT_MAX_ACTORS,
wake_slot: false, wake_slot: true,
node_id: crate::pg::DEFAULT_NODE_ID, node_id: crate::pg::DEFAULT_NODE_ID,
incarnation: crate::pg::DEFAULT_INCARNATION, incarnation: crate::pg::DEFAULT_INCARNATION,
} }
@@ -283,7 +283,7 @@ impl Config {
/// thread's slot; it is resumed next on that core and inherits the /// thread's slot; it is resumed next on that core and inherits the
/// remainder of the waker's timeslice. Scheduler-context wakes /// remainder of the waker's timeslice. Scheduler-context wakes
/// (timer/IO drain) and spawns always go to the shared queue. /// (timer/IO drain) and spawns always go to the shared queue.
/// Default: `false` (off until the slot shootout accepts it). /// Default: `true` (accepted 2026-08-18, history.md finding 17/18).
pub fn wake_slot(mut self, on: bool) -> Self { pub fn wake_slot(mut self, on: bool) -> Self {
self.wake_slot = on; self.wake_slot = on;
self self
@@ -333,7 +333,7 @@ impl Default for Config {
stack_reserve: DEFAULT_STACK_RESERVE, stack_reserve: DEFAULT_STACK_RESERVE,
stack_guard: DEFAULT_STACK_GUARD, stack_guard: DEFAULT_STACK_GUARD,
max_actors: DEFAULT_MAX_ACTORS, max_actors: DEFAULT_MAX_ACTORS,
wake_slot: false, wake_slot: true,
node_id: crate::pg::DEFAULT_NODE_ID, node_id: crate::pg::DEFAULT_NODE_ID,
incarnation: crate::pg::DEFAULT_INCARNATION, incarnation: crate::pg::DEFAULT_INCARNATION,
} }
@@ -356,6 +356,25 @@ pub struct SchedulerStats {
pub slot_hits: AtomicU64, pub slot_hits: AtomicU64,
/// RFC 005: slot occupants displaced to the shared queue by a newer wake. /// RFC 005: slot occupants displaced to the shared queue by a newer wake.
pub slot_displacements: AtomicU64, pub slot_displacements: AtomicU64,
// --- wake-path diagnostics (target 5). Relaxed counters, cheap. ---
/// unpark hit Parked from actor context → slot_push.
pub unpark_slot: AtomicU64,
/// unpark hit Parked from scheduler/foreign context → shared enqueue.
pub unpark_queue: AtomicU64,
/// unpark hit Running → RunningNotified (peer had not parked yet).
pub unpark_notified: AtomicU64,
/// park_return found the flag consumed → shared re-enqueue.
pub park_flag_consumed: AtomicU64,
/// Yield-intent re-enqueues.
pub yield_requeues: AtomicU64,
/// `enqueue` tail wake actually delivered a futex permit.
pub enqueue_wakes: AtomicU64,
/// Chain-rule wake actually delivered a permit.
pub chain_wakes: AtomicU64,
/// Times this scheduler entered Pop::Idle (about to futex-park).
pub idle_parks: AtomicU64,
/// Idle parks whose recheck aborted (WorkFound) — no futex.
pub idle_recheck_hits: AtomicU64,
} }
impl SchedulerStats { impl SchedulerStats {
@@ -365,8 +384,44 @@ impl SchedulerStats {
run_queue_len: AtomicU64::new(0), run_queue_len: AtomicU64::new(0),
slot_hits: AtomicU64::new(0), slot_hits: AtomicU64::new(0),
slot_displacements: AtomicU64::new(0), slot_displacements: AtomicU64::new(0),
unpark_slot: AtomicU64::new(0),
unpark_queue: AtomicU64::new(0),
unpark_notified: AtomicU64::new(0),
park_flag_consumed: AtomicU64::new(0),
yield_requeues: AtomicU64::new(0),
enqueue_wakes: AtomicU64::new(0),
chain_wakes: AtomicU64::new(0),
idle_parks: AtomicU64::new(0),
idle_recheck_hits: AtomicU64::new(0),
} }
} }
fn reset(&self) {
for c in [
&self.slot_hits,
&self.slot_displacements,
&self.unpark_slot,
&self.unpark_queue,
&self.unpark_notified,
&self.park_flag_consumed,
&self.yield_requeues,
&self.enqueue_wakes,
&self.chain_wakes,
&self.idle_parks,
&self.idle_recheck_hits,
] {
c.store(0, Ordering::Relaxed);
}
}
}
/// Bump a per-thread diagnostic counter on the calling scheduler thread.
macro_rules! diag {
($inner:expr, $field:ident) => {
$inner.stats[sched_slot()]
.$field
.fetch_add(1, Ordering::Relaxed)
};
} }
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -422,6 +477,34 @@ impl RuntimeStats {
.map(|s| s.slot_displacements.load(Ordering::Relaxed)) .map(|s| s.slot_displacements.load(Ordering::Relaxed))
.sum() .sum()
} }
/// Wake-path diagnostics summed across threads, as `name=value` pairs
/// (target 5 instrumentation). Reset at the start of each `run()`.
pub fn wake_diag(&self) -> String {
let sum = |f: fn(&SchedulerStats) -> &AtomicU64| -> u64 {
self.inner
.stats
.iter()
.map(|s| f(s).load(Ordering::Relaxed))
.sum()
};
format!(
"slot_hits={} displaced={} unpark_slot={} unpark_queue={} unpark_notified={} \
flag_consumed={} yield_requeues={} enqueue_wakes={} chain_wakes={} \
idle_parks={} idle_recheck_hits={}",
sum(|s| &s.slot_hits),
sum(|s| &s.slot_displacements),
sum(|s| &s.unpark_slot),
sum(|s| &s.unpark_queue),
sum(|s| &s.unpark_notified),
sum(|s| &s.park_flag_consumed),
sum(|s| &s.yield_requeues),
sum(|s| &s.enqueue_wakes),
sum(|s| &s.chain_wakes),
sum(|s| &s.idle_parks),
sum(|s| &s.idle_recheck_hits),
)
}
} }
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -873,6 +956,15 @@ impl Slot {
} }
fn take_closure(&self) -> Option<Closure> { fn take_closure(&self) -> Option<Closure> {
// Fast path: every resume after the first (the overwhelming case)
// finds null. A plain load suffices to prove it — `store_closure`
// runs only before `publish_queued`, whose Release/Acquire pairing
// with the claimer's `try_claim` orders it before this call, so no
// writer can race the load within an occupancy. This keeps the
// locked RMW (full barrier, ~20+ cycles) off the per-resume path.
if self.closure.load(Ordering::Relaxed).is_null() {
return None;
}
let raw = self.closure.swap(std::ptr::null_mut(), Ordering::Acquire); let raw = self.closure.swap(std::ptr::null_mut(), Ordering::Acquire);
if raw.is_null() { if raw.is_null() {
None None
@@ -904,17 +996,10 @@ pub(crate) struct RuntimeInner {
pub(crate) live_actors: AtomicU32, pub(crate) live_actors: AtomicU32,
/// Packed `(index << 32 | generation)` of the run's root (initial) actor, /// Packed `(index << 32 | generation)` of the run's root (initial) actor,
/// or `u64::MAX` (the ROOT_PID sentinel) before one is set. When this actor /// or `u64::MAX` (the ROOT_PID sentinel) before one is set. When this actor
/// finalizes it flags `root_exited`; the scheduler's idle verdict then /// finalizes, `finalize_actor` runs the root-exit shutdown (see
/// stops the remaining (parked-forever) actors. Set once per `run()`, right /// `shutdown_forest_roots`). Set once per `run()`, right after the initial
/// after the initial spawn. /// spawn.
pub(crate) root_bits: AtomicU64, pub(crate) root_bits: AtomicU64,
/// Set when the root actor finalizes; read by the scheduler's idle verdict
/// to trigger the one-shot teardown sweep. Reset per `run()`.
pub(crate) root_exited: AtomicBool,
/// Guards the teardown sweep to fire at most once per run (a parked-forever
/// remainder that survives the sweep falls through to the normal idle wait
/// rather than busy-spinning). Reset per `run()`.
pub(crate) root_swept: AtomicBool,
/// Timer heap. Independent lock: never nested with any other. /// Timer heap. Independent lock: never nested with any other.
pub(crate) timers: Mutex<Timers>, pub(crate) timers: Mutex<Timers>,
/// IO subsystem. `None` between runs. Lock order: io before everything. /// IO subsystem. `None` between runs. Lock order: io before everything.
@@ -962,6 +1047,17 @@ pub(crate) struct RuntimeInner {
/// checks under it read only the atomic slot word, and the eviction path /// checks under it read only the atomic slot word, and the eviction path
/// keeps it off the send path. /// keeps it off the send path.
pub(crate) process_groups: RawMutex<crate::pg::ProcessGroups>, pub(crate) process_groups: RawMutex<crate::pg::ProcessGroups>,
/// RFC 010 c8: the exposure registry (exposed names + type-hash decoders).
/// RawMutex Leaf, same discipline as `process_groups`; decoders run under
/// it and are leaf-only by contract (they decode and send — `send_dyn`
/// takes `registry`, never this). cfg-gated: zero-cost-when-off (c1).
#[cfg(feature = "cluster")]
pub(crate) exposure: RawMutex<crate::cluster::expose::ExposureState>,
/// RFC 010 c9: the outbound table (node name → the connection's
/// dedicated `Sender<Frame>`), manager-maintained. Leaf; the send happens
/// outside the lock. cfg-gated like `exposure`.
#[cfg(feature = "cluster")]
pub(crate) outbound: RawMutex<crate::cluster::remote::Outbound>,
/// Recycled stacks waiting to be reused by the next spawn. /// Recycled stacks waiting to be reused by the next spawn.
pub(crate) stack_pool: RawMutex<Vec<crate::stack::Stack>>, pub(crate) stack_pool: RawMutex<Vec<crate::stack::Stack>>,
/// Maximum number of stacks to retain in the pool. /// Maximum number of stacks to retain in the pool.
@@ -1004,8 +1100,6 @@ impl RuntimeInner {
free: RawMutex::new(free), free: RawMutex::new(free),
live_actors: AtomicU32::new(0), live_actors: AtomicU32::new(0),
root_bits: AtomicU64::new(u64::MAX), root_bits: AtomicU64::new(u64::MAX),
root_exited: AtomicBool::new(false),
root_swept: AtomicBool::new(false),
timers: Mutex::new(timers), timers: Mutex::new(timers),
io: Mutex::new(None), io: Mutex::new(None),
next_monitor_id: AtomicU64::new(0), next_monitor_id: AtomicU64::new(0),
@@ -1022,6 +1116,10 @@ impl RuntimeInner {
node_id, node_id,
incarnation, incarnation,
process_groups: RawMutex::new(crate::pg::ProcessGroups::new()), process_groups: RawMutex::new(crate::pg::ProcessGroups::new()),
#[cfg(feature = "cluster")]
exposure: RawMutex::new(crate::cluster::expose::ExposureState::new()),
#[cfg(feature = "cluster")]
outbound: RawMutex::new(crate::cluster::remote::Outbound::new()),
stack_pool: RawMutex::new(Vec::new()), stack_pool: RawMutex::new(Vec::new()),
stack_pool_cap, stack_pool_cap,
stack_reserve: crate::stack::round_to_pages(stack_reserve), stack_reserve: crate::stack::round_to_pages(stack_reserve),
@@ -1080,7 +1178,9 @@ impl RuntimeInner {
// pure-compute hot path pays (almost) nothing. Bias is over-wake: // pure-compute hot path pays (almost) nothing. Bias is over-wake:
// a spurious wake costs one futex round-trip and a failed pop; a // a spurious wake costs one futex round-trip and a failed pop; a
// missed wake would cost a stranded actor. // missed wake would cost a stranded actor.
self.coord.wake_one_if_idle(); if self.coord.wake_one_if_idle() {
diag!(self, enqueue_wakes);
}
} }
/// Make `pid` runnable if it is parked; coalesce or defer otherwise. /// Make `pid` runnable if it is parked; coalesce or defer otherwise.
@@ -1137,12 +1237,15 @@ impl RuntimeInner {
// (install_actor enqueues directly), so they bypass the // (install_actor enqueues directly), so they bypass the
// slot by construction. // slot by construction.
if self.wake_slot && crate::actor::current_pid().is_some() { if self.wake_slot && crate::actor::current_pid().is_some() {
diag!(self, unpark_slot);
self.slot_push(pid); self.slot_push(pid);
} else { } else {
diag!(self, unpark_queue);
self.enqueue(pid); self.enqueue(pid);
} }
} }
Unpark::Notified => { Unpark::Notified => {
diag!(self, unpark_notified);
crate::te!(crate::trace::Event::UnparkDeferred(pid)); crate::te!(crate::trace::Event::UnparkDeferred(pid));
} }
Unpark::Noop => {} Unpark::Noop => {}
@@ -1157,9 +1260,11 @@ impl RuntimeInner {
/// Displacement: the NEW wake takes the slot (newest is hottest; the old /// Displacement: the NEW wake takes the slot (newest is hottest; the old
/// occupant was about to lose its locality window anyway) and the old /// occupant was about to lose its locality window anyway) and the old
/// occupant is pushed to the shared queue. /// occupant is pushed to the shared queue.
#[inline(never)]
fn slot_push(&self, pid: Pid) { fn slot_push(&self, pid: Pid) {
crate::context::tls_fence();
debug_assert!( debug_assert!(
!crate::preempt::PREEMPTION_ENABLED.with(|c| c.get()), !crate::preempt::preemption_enabled(),
"slot_push with preemption enabled — a switch mid-op could \ "slot_push with preemption enabled — a switch mid-op could \
migrate the actor and split the slot access across threads" migrate the actor and split the slot access across threads"
); );
@@ -1173,11 +1278,9 @@ impl RuntimeInner {
let displaced = WAKE_SLOT.with(|s| s.replace(Some(pid))); let displaced = WAKE_SLOT.with(|s| s.replace(Some(pid)));
crate::te!(crate::trace::Event::SlotPush(pid)); crate::te!(crate::trace::Event::SlotPush(pid));
if let Some(old) = displaced { if let Some(old) = displaced {
SCHED_SLOT.with(|s| { self.stats[sched_slot()]
self.stats[s.get()] .slot_displacements
.slot_displacements .fetch_add(1, Ordering::Relaxed);
.fetch_add(1, Ordering::Relaxed)
});
self.enqueue(old); self.enqueue(old);
} }
} }
@@ -1330,19 +1433,20 @@ impl Runtime {
// RFC 005: slot counters reset at the START of a run (not the end), // RFC 005: slot counters reset at the START of a run (not the end),
// so `stats()` read after `run()` returns reports that run's totals. // so `stats()` read after `run()` returns reports that run's totals.
for stat in &self.inner.stats { for stat in &self.inner.stats {
stat.slot_hits.store(0, Ordering::Relaxed); stat.reset();
stat.slot_displacements.store(0, Ordering::Relaxed);
} }
// Spawn the initial actor through the public spawn path (which // Spawn the initial actor through the public spawn path (which
// requires a running runtime in the thread-local). // requires a running runtime in the thread-local).
RUNTIME.with(|r| *r.borrow_mut() = Some(self.inner.clone())); RUNTIME.with(|r| *r.borrow_mut() = Some(self.inner.clone()));
let initial_handle = crate::scheduler::spawn(f); let initial_handle = crate::scheduler::spawn(f);
// The initial actor is the run's root: when it exits, remaining actors // The initial actor is the run's root: its exit means "the program is
// are stopped so the run winds down (see finalize_actor / schedule_loop). // done" — every remaining top-level actor is asked to shut down (see
self.inner.root_exited.store(false, Ordering::Relaxed); // `finalize_actor` / `shutdown_forest_roots`).
self.inner.root_swept.store(false, Ordering::Relaxed);
self.inner.set_root(initial_handle.pid()); self.inner.set_root(initial_handle.pid());
// A previous run's group reaper was stopped with that run; forget it
// so the first `join` of this run spawns a fresh one.
self.inner.process_groups.lock().reset_reaper();
// Launch N-1 extra scheduler threads, named `smarm-sched-{slot}` so // Launch N-1 extra scheduler threads, named `smarm-sched-{slot}` so
// they are identifiable in `/proc/<pid>/task/*/comm`, stack dumps and // they are identifiable in `/proc/<pid>/task/*/comm`, stack dumps and
@@ -1457,6 +1561,71 @@ impl Runtime {
inner: self.inner.clone(), inner: self.inner.clone(),
} }
} }
/// A `Send + Sync` handle to this runtime, usable from any thread —
/// including threads that are not smarm schedulers (an OS-signal handler
/// thread, an external event source). Grab it *before* [`run`](Self::run)
/// and hand it to e.g. a signal thread; that thread can then
/// [`request_stop`](RuntimeHandle::request_stop) the runtime's top
/// supervisor to drive an ordered shutdown from outside the runtime.
///
/// The in-runtime primitives ([`scheduler::request_stop`](crate::request_stop)
/// and friends) reach the runtime through a thread-local that is unset on
/// any non-scheduler thread, so they are silent no-ops off-runtime; this
/// handle carries its own reference and closes that gap.
pub fn handle(&self) -> RuntimeHandle {
RuntimeHandle {
inner: Arc::downgrade(&self.inner),
}
}
}
// ---------------------------------------------------------------------------
// RuntimeHandle — off-runtime wake/stop
// ---------------------------------------------------------------------------
/// A `Send + Sync` handle to a [`Runtime`], obtained from
/// [`Runtime::handle`]. Lets a thread that is *not* a smarm scheduler thread
/// drive a cooperative stop into the runtime — the off-runtime counterpart to
/// [`scheduler::request_stop`](crate::request_stop).
///
/// Holds a [`Weak`] to the runtime, for the same reason the IO backend does
/// (RFC 018): a lingering handle can never keep the runtime's slot table alive
/// and can never block [`Runtime::run`] from finishing. Once the `Runtime` is
/// dropped every method is a harmless no-op — the same end state as calling
/// `request_stop` on an actor that has already exited.
#[derive(Clone)]
pub struct RuntimeHandle {
inner: Weak<RuntimeInner>,
}
impl RuntimeHandle {
/// Ask `pid` to stop cooperatively, from any thread. The off-runtime
/// equivalent of [`scheduler::request_stop`](crate::request_stop): it sets
/// the target's stop flag and wakes it, so a parked actor unwinds at its
/// next checkpoint exactly as it would for an in-runtime stop. A no-op if
/// the runtime has been dropped, or if the actor has already exited.
pub fn request_stop<A>(&self, pid: Pid<A>) {
let pid = pid.erase();
// Upgrade the Weak per call, like the IO backend does (io.rs): a live
// runtime yields the inner and we drive the same stop the in-runtime
// path would; a dropped runtime makes this a no-op.
if let Some(inner) = self.inner.upgrade() {
crate::scheduler::request_stop_inner(&inner, pid);
}
}
/// Ask `pid` to shut down gracefully, from any thread. The off-runtime
/// equivalent of [`scheduler::request_shutdown`](crate::request_shutdown);
/// the delivered [`ExitSignal`](crate::ExitSignal) carries `from ==
/// ROOT_PID`, since no actor made the request. A no-op if the runtime has
/// been dropped, or if the actor has already exited.
pub fn request_shutdown<A>(&self, pid: Pid<A>) {
let pid = pid.erase();
if let Some(inner) = self.inner.upgrade() {
crate::scheduler::request_shutdown_inner(&inner, pid, ROOT_PID);
}
}
} }
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -1495,7 +1664,17 @@ pub(crate) enum YieldIntent {
Park, Park,
} }
/// This scheduler thread's stats index. Read from actor context by every
/// `stat!()` on the enqueue/wake paths, hence out of line (`context` docs).
#[inline(never)]
fn sched_slot() -> usize {
crate::context::tls_fence();
SCHED_SLOT.with(|s| s.get())
}
#[inline(never)]
pub(crate) fn set_yield_intent(i: YieldIntent) { pub(crate) fn set_yield_intent(i: YieldIntent) {
crate::context::tls_fence();
YIELD_INTENT.with(|c| c.set(i)); YIELD_INTENT.with(|c| c.set(i));
} }
@@ -1619,6 +1798,10 @@ pub(crate) fn install_actor(
stack: crate::stack::Stack, stack: crate::stack::Stack,
supervisor: Pid, supervisor: Pid,
closure: Closure, closure: Closure,
monitor: Option<(
crate::monitor::MonitorId,
crate::channel::Sender<crate::monitor::Down>,
)>,
) -> Pid { ) -> Pid {
let slot = &inner.slots[idx as usize]; let slot = &inner.slots[idx as usize];
let gen = slot.generation(); // stable: we own the vacant slot via the free list let gen = slot.generation(); // stable: we own the vacant slot via the free list
@@ -1646,6 +1829,12 @@ pub(crate) fn install_actor(
cold.outstanding_handles = 1; cold.outstanding_handles = 1;
cold.outcome = None; cold.outcome = None;
cold.pending_io_result = None; cold.pending_io_result = None;
// `spawn_monitor`: register before publish, so no scheduler can run
// (and finalize) the child before its monitor exists. Same slot as a
// `monitor()` registration; the send-from-finalize path is unchanged.
if let Some(m) = monitor {
cold.monitors.push(m);
}
} }
slot.sp.store(sp, Ordering::Relaxed); slot.sp.store(sp, Ordering::Relaxed);
// RFC 019: a fresh incarnation starts with its high-water at the fresh // RFC 019: a fresh incarnation starts with its high-water at the fresh
@@ -1851,14 +2040,12 @@ fn finalize_actor(inner: &Arc<RuntimeInner>, pid: Pid, outcome: Outcome) {
// Reclaim if no outstanding handles (re-verified inside). // Reclaim if no outstanding handles (re-verified inside).
reclaim_slot(inner, pid); reclaim_slot(inner, pid);
// Root-exit teardown is DEFERRED to the scheduler's idle verdict, not done // Root exit = the program is done. Ask every top-level survivor to shut
// here: stopping eagerly would cut off actors that still have queued work // down, right here, before the live-count decrement below: any wake this
// (they'd unwind on the stop before draining their mailbox). Flagging it // produces is then ordered before `live_actors` can be observed at its
// instead lets the run queue drain naturally first; only the parked-forever // decremented value, same as every other wakeup finalize issues.
// remainder (e.g. a server pinned alive by a registered name) is then
// stopped, once nothing runnable is left. See `schedule_loop`.
if inner.is_root(pid) { if inner.is_root(pid) {
inner.root_exited.store(true, Ordering::Release); shutdown_forest_roots(inner, pid);
} }
// The decrement is LAST: every wakeup this finalize produced (joiners, // The decrement is LAST: every wakeup this finalize produced (joiners,
@@ -1869,16 +2056,78 @@ fn finalize_actor(inner: &Arc<RuntimeInner>, pid: Pid, outcome: Outcome) {
debug_assert!(prev >= 1, "live_actors underflow — double finalize"); debug_assert!(prev >= 1, "live_actors underflow — double finalize");
} }
/// Cooperatively stop every live actor — the root-exit teardown sweep, run from /// The root-exit shutdown. Delivers [`request_shutdown`](crate::request_shutdown)
/// `schedule_loop` once the run queue is empty after the root has exited. Each /// to every **forest root**: each live actor whose recorded parent
/// [`request_stop_inner`](crate::scheduler::request_stop_inner) re-verifies the /// (`Actor::supervisor` — the spawner for a plain `spawn`, the supervisor for
/// target under its cold lock, so the racy per-slot generation read is safe: a /// `spawn_under`) is the run itself (`ROOT_PID`) or is no longer live. Actors
/// vacant, dead, or reused slot no-ops. The swept actors unpark, unwind at their /// under a live parent are not addressed — that parent is responsible for
/// next observation point, and finalize, dropping `live_actors` to zero. /// them: a supervisor traps and runs its ordered, policy-driven shutdown; a
fn stop_live_actors(inner: &Arc<RuntimeInner>) { /// bare parent that dies takes non-trapping children with it via the next
/// pass of this same rule only if it dies *now*, so a parent that outlives
/// this scan and later dies leaves its subtree to itself (Erlang semantics: an
/// unlinked spawn is nobody's child).
///
/// Semantics per target follow `request_shutdown`: a trapping actor receives
/// `ExitSignal { from: root, reason: Shutdown }` and may finish work — drain,
/// keep its timers ticking, then stop itself; a non-trapping one is stopped
/// outright. There is no second, forcing sweep: an actor that traps and never
/// stops keeps the run alive by design (put it under a supervisor with a
/// `Shutdown::Timeout` policy if that is not wanted). Runs once, on the root's
/// finalize path, so it races only against actors that are still running —
/// each `request_shutdown_inner` re-verifies its target under the cold lock,
/// so a slot that dies or is reused mid-scan is a no-op.
fn shutdown_forest_roots(inner: &Arc<RuntimeInner>, root: Pid) {
for idx in 0..inner.slots.len() as u32 { for idx in 0..inner.slots.len() as u32 {
let pid = Pid::new(idx, inner.slots[idx as usize].generation()); let slot = &inner.slots[idx as usize];
crate::scheduler::request_stop_inner(inner, pid); let pid = Pid::new(idx, slot.generation());
if pid == root {
continue;
}
// Lock-free reject first. The slab is `max_actors` entries (16_384 by
// default) and is almost entirely vacant at root exit, so locking every
// slot's `cold` to discover `actor == None` made this scan cost one
// uncontended mutex round-trip per slot — a fixed ~150 µs per run on a
// 5900X, and `general.rs` times `init` + `run` together, so it landed
// in every smarm bench number. A non-Live slot has no actor to shut
// down, and the lock is taken again below for the ones that do.
//
// This does not weaken the sweep. `is_live_for` is a snapshot, so a
// slot can go Live just after we pass it — but that was already true
// of a spawn landing after the scan finished, and there is no second
// sweep either way (an actor that outlives this scan is its spawner's
// business, per the rule below).
if !slot.is_live_for(pid) {
continue;
}
// Read the parent under the cold lock (generation-verified); act
// outside it — `request_shutdown_inner` sends and may unpark.
let parent = {
let cold = slot.cold.lock();
if slot.generation() != pid.generation() {
continue;
}
match cold.actor.as_ref() {
Some(a) => a.supervisor,
None => continue,
}
};
let parent_live = inner
.slot_at(parent)
.is_some_and(|ps| ps.is_live_for(parent));
if !parent_live {
// `_probe` so the trace can say what each leftover was: under
// `smarm-trace` every swept actor is a `root_sweep` line — the
// visibility that makes "a forgotten actor costs only a slot, never
// a hung run" a checkable claim rather than a hope.
let found = crate::scheduler::request_shutdown_inner_probe(inner, pid, root);
#[cfg_attr(not(feature = "smarm-trace"), allow(unused_variables))]
if let Some(trapping) = found {
crate::te!(crate::trace::Event::RootSweep {
target: pid,
trapping
});
}
}
} }
} }
@@ -1962,9 +2211,6 @@ fn schedule_loop(inner: &Arc<RuntimeInner>, slot_idx: usize) {
Got(Pid), Got(Pid),
Idle, Idle,
AllDone, AllDone,
/// Root has exited and nothing is runnable: stop the parked-forever
/// remainder, then re-pop. Fires at most once per run.
RootDrain,
} }
// 2a. RFC 005: drain this thread's wake slot before touching the // 2a. RFC 005: drain this thread's wake slot before touching the
@@ -2014,15 +2260,6 @@ fn schedule_loop(inner: &Arc<RuntimeInner>, slot_idx: usize) {
let live = inner.live_actors.load(Ordering::Acquire); let live = inner.live_actors.load(Ordering::Acquire);
if live == 0 && io_out == 0 { if live == 0 && io_out == 0 {
Pop::AllDone Pop::AllDone
} else if inner.root_exited.load(Ordering::Acquire)
&& !inner.root_swept.swap(true, Ordering::AcqRel)
{
// Root gone and nothing runnable — the live remainder
// are parked-forever daemons (Queued actors with pending
// work drained before the queue emptied). Stop them so
// the run can end. One-shot: a survivor falls through to
// the idle wait below on the next pass.
Pop::RootDrain
} else { } else {
Pop::Idle Pop::Idle
} }
@@ -2050,13 +2287,6 @@ fn schedule_loop(inner: &Arc<RuntimeInner>, slot_idx: usize) {
inner.coord.wake_all(); inner.coord.wake_all();
return; return;
} }
Pop::RootDrain => {
// Root has exited and nothing is runnable: stop the
// parked-forever remainder, then loop back to re-pop the
// now-runnable (stopping) actors.
stop_live_actors(inner);
continue;
}
Pop::Idle => { Pop::Idle => {
// Something is still in flight. Park on our own futex // Something is still in flight. Park on our own futex
// until a producer wakes us (enqueue tail), a deadline // until a producer wakes us (enqueue tail), a deadline
@@ -2082,15 +2312,17 @@ fn schedule_loop(inner: &Arc<RuntimeInner>, slot_idx: usize) {
// The mandatory post-publish re-check: a producer that // The mandatory post-publish re-check: a producer that
// enqueued (or a verdict input that flipped) before it // enqueued (or a verdict input that flipped) before it
// could see our idle bit has left us the evidence. // could see our idle bit has left us the evidence.
let _ = inner.coord.park(slot_idx, tk_deadline, || { stats.idle_parks.fetch_add(1, Ordering::Relaxed);
let pr = inner.coord.park(slot_idx, tk_deadline, || {
!inner.run_queue.is_empty() !inner.run_queue.is_empty()
|| (inner.live_actors.load(Ordering::Acquire) == 0 || (inner.live_actors.load(Ordering::Acquire) == 0
&& inner.io_outstanding.load(Ordering::Acquire) == 0 && inner.io_outstanding.load(Ordering::Acquire) == 0
&& inner.io_fd_waiters.load(Ordering::Acquire) == 0) && inner.io_fd_waiters.load(Ordering::Acquire) == 0)
|| (inner.root_exited.load(Ordering::Acquire)
&& !inner.root_swept.load(Ordering::Acquire))
|| inner.coord.deadline_due() || inner.coord.deadline_due()
}); });
if matches!(pr, crate::park::ParkResult::WorkFound) {
stats.idle_recheck_hits.fetch_add(1, Ordering::Relaxed);
}
if tk_deadline.is_some() { if tk_deadline.is_some() {
// Hand the role back BEFORE firing: pop_due can run // Hand the role back BEFORE firing: pop_due can run
// `Send` thunks that insert new timers, and the // `Send` thunks that insert new timers, and the
@@ -2125,8 +2357,8 @@ fn schedule_loop(inner: &Arc<RuntimeInner>, slot_idx: usize) {
// miss here is safe — the enqueue that created the surplus already // miss here is safe — the enqueue that created the surplus already
// issued its own wake (RFC 018 no-lost-wake); this only sharpens // issued its own wake (RFC 018 no-lost-wake); this only sharpens
// parallelism latency. // parallelism latency.
if !inner.run_queue.is_empty() { if !inner.run_queue.is_empty() && inner.coord.wake_one_if_idle() {
inner.coord.wake_one_if_idle(); stats.chain_wakes.fetch_add(1, Ordering::Relaxed);
} }
let slot = match inner.slot_at(pid) { let slot = match inner.slot_at(pid) {
@@ -2150,7 +2382,6 @@ fn schedule_loop(inner: &Arc<RuntimeInner>, slot_idx: usize) {
.current_pid_index .current_pid_index
.store(pid.index(), Ordering::Relaxed); .store(pid.index(), Ordering::Relaxed);
set_actor_sp(sp);
set_current_pid(pid); set_current_pid(pid);
crate::preempt::set_current_stop(stop_flag); crate::preempt::set_current_stop(stop_flag);
crate::preempt::set_current_slot(slot as *const Slot); crate::preempt::set_current_slot(slot as *const Slot);
@@ -2176,7 +2407,7 @@ fn schedule_loop(inner: &Arc<RuntimeInner>, slot_idx: usize) {
crate::causal::on_resume(slot); crate::causal::on_resume(slot);
crate::te!(crate::trace::Event::Resume(pid)); crate::te!(crate::trace::Event::Resume(pid));
unsafe { switch_to_actor() }; let saved_sp = unsafe { switch_to_actor(sp) };
PREEMPTION_ENABLED.with(|c| c.set(false)); PREEMPTION_ENABLED.with(|c| c.set(false));
// RFC 016 Chunk 2: charge the cycles this resume consumed to the actor // RFC 016 Chunk 2: charge the cycles this resume consumed to the actor
@@ -2190,7 +2421,6 @@ fn schedule_loop(inner: &Arc<RuntimeInner>, slot_idx: usize) {
crate::preempt::clear_current_slot(); crate::preempt::clear_current_slot();
let intent = YIELD_INTENT.with(|c| c.get()); let intent = YIELD_INTENT.with(|c| c.get());
let saved_sp = get_actor_sp();
slot.sp.store(saved_sp, Ordering::Relaxed); slot.sp.store(saved_sp, Ordering::Relaxed);
// RFC 019 §2: sampled high-water — one branch + at most one store // RFC 019 §2: sampled high-water — one branch + at most one store
// into the line the store above just dirtied. Relaxed and advisory; // into the line the store above just dirtied. Relaxed and advisory;
@@ -2217,6 +2447,7 @@ fn schedule_loop(inner: &Arc<RuntimeInner>, slot_idx: usize) {
// arriving mid-run coalesces into the re-queue. // arriving mid-run coalesces into the re-queue.
crate::te!(crate::trace::Event::Yield(pid)); crate::te!(crate::trace::Event::Yield(pid));
slot.word.yield_return(gen); slot.word.yield_return(gen);
diag!(inner, yield_requeues);
inner.enqueue(pid); inner.enqueue(pid);
} }
YieldIntent::Park => { YieldIntent::Park => {
@@ -2257,6 +2488,7 @@ fn schedule_loop(inner: &Arc<RuntimeInner>, slot_idx: usize) {
#[cfg(feature = "smarm-causal")] #[cfg(feature = "smarm-causal")]
crate::causal::on_deschedule(slot, false); crate::causal::on_deschedule(slot, false);
crate::te!(crate::trace::Event::UnparkFlagConsumed(pid)); crate::te!(crate::trace::Event::UnparkFlagConsumed(pid));
diag!(inner, park_flag_consumed);
inner.enqueue(pid); inner.enqueue(pid);
} }
} }
+162 -17
View File
@@ -71,7 +71,7 @@ use crate::pid::{Name, Pid};
use crate::runtime::{self, RuntimeInner, YieldIntent, RUNTIME}; use crate::runtime::{self, RuntimeInner, YieldIntent, RUNTIME};
use crate::supervisor::Signal; use crate::supervisor::Signal;
use std::sync::atomic::Ordering; use std::sync::atomic::Ordering;
use std::sync::Arc; use std::sync::{Arc, Weak};
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// with_runtime / try_with_runtime // with_runtime / try_with_runtime
@@ -85,8 +85,10 @@ use std::sync::Arc;
// released on the wrong thread's copy of the thread-local, corrupting its // released on the wrong thread's copy of the thread-local, corrupting its
// borrow count. `f` is also always runtime bookkeeping that should run to // borrow count. `f` is also always runtime bookkeeping that should run to
// completion without the actor being suspended or unwound partway through. // completion without the actor being suspended or unwound partway through.
#[inline(never)]
pub(crate) fn with_runtime<R>(f: impl FnOnce(&Arc<RuntimeInner>) -> R) -> R { pub(crate) fn with_runtime<R>(f: impl FnOnce(&Arc<RuntimeInner>) -> R) -> R {
let prev = crate::preempt::PREEMPTION_ENABLED.with(|c| c.replace(false)); crate::context::tls_fence();
let prev = crate::preempt::preemption_swap(false);
let result = RUNTIME.with(|r| { let result = RUNTIME.with(|r| {
let b = r.borrow(); let b = r.borrow();
let inner = match b.as_ref() { let inner = match b.as_ref() {
@@ -95,17 +97,19 @@ pub(crate) fn with_runtime<R>(f: impl FnOnce(&Arc<RuntimeInner>) -> R) -> R {
}; };
f(inner) f(inner)
}); });
crate::preempt::PREEMPTION_ENABLED.with(|c| c.set(prev)); crate::preempt::preemption_swap(prev);
result result
} }
// Borrow the runtime if present, otherwise `None`. Used on cleanup paths // Borrow the runtime if present, otherwise `None`. Used on cleanup paths
// (e.g. a channel's Drop impl during teardown) that may run after the // (e.g. a channel's Drop impl during teardown) that may run after the
// runtime has already gone away. Same preemption gate as `with_runtime`. // runtime has already gone away. Same preemption gate as `with_runtime`.
#[inline(never)]
pub(crate) fn try_with_runtime<R>(f: impl FnOnce(&Arc<RuntimeInner>) -> R) -> Option<R> { pub(crate) fn try_with_runtime<R>(f: impl FnOnce(&Arc<RuntimeInner>) -> R) -> Option<R> {
let prev = crate::preempt::PREEMPTION_ENABLED.with(|c| c.replace(false)); crate::context::tls_fence();
let prev = crate::preempt::preemption_swap(false);
let result = RUNTIME.with(|r| r.borrow().as_ref().map(f)); let result = RUNTIME.with(|r| r.borrow().as_ref().map(f));
crate::preempt::PREEMPTION_ENABLED.with(|c| c.set(prev)); crate::preempt::preemption_swap(prev);
result result
} }
@@ -263,6 +267,7 @@ impl Drop for JoinHandle {
/// ``` /// ```
/// use smarm::SpawnOpts; /// use smarm::SpawnOpts;
/// let opts = SpawnOpts { stack_reserve: Some(8 * 1024 * 1024), ..SpawnOpts::default() }; /// let opts = SpawnOpts { stack_reserve: Some(8 * 1024 * 1024), ..SpawnOpts::default() };
/// assert_eq!(opts.stack_reserve, Some(8 * 1024 * 1024));
/// ``` /// ```
/// ///
/// Both sizes are page-rounded. The reserve is *virtual* (demand-paged): /// Both sizes are page-rounded. The reserve is *virtual* (demand-paged):
@@ -286,8 +291,8 @@ pub struct SpawnOpts {
#[non_exhaustive] #[non_exhaustive]
#[derive(Clone, Copy, Debug, PartialEq, Eq)] #[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub enum SpawnError { pub enum SpawnError {
/// The fixed actor slab ([`Config::max_actors`] /// The fixed actor slab ([`Config::max_actors`](crate::runtime::Config::max_actors))
/// (crate::runtime::Config::max_actors)) is full: every slot is claimed /// is full: every slot is claimed
/// by a live actor. This is a routine overload condition, not an /// by a live actor. This is a routine overload condition, not an
/// invariant violation — shed the unit of work (close the socket, /// invariant violation — shed the unit of work (close the socket,
/// return a 503) and try again once actors have died. /// return a 503) and try again once actors have died.
@@ -362,7 +367,7 @@ pub fn spawn_under_with<A>(
let pid = with_runtime(|inner| { let pid = with_runtime(|inner| {
let idx = inner.allocate_slot(); // panics loudly on slab exhaustion let idx = inner.allocate_slot(); // panics loudly on slab exhaustion
crate::runtime::install_actor(inner, idx, sp, stack, supervisor, closure) crate::runtime::install_actor(inner, idx, sp, stack, supervisor, closure, None)
}); });
JoinHandle { JoinHandle {
@@ -371,6 +376,52 @@ pub fn spawn_under_with<A>(
} }
} }
/// [`spawn`] and [`monitor`](crate::monitor()) the child in one step, with no
/// window in which the child can die unobserved.
///
/// `spawn` followed by `monitor(h.pid())` races: on a multi-scheduler
/// runtime the child can run to completion before the monitor registers, and
/// a monitor on a dead pid delivers [`DownReason::NoProc`](crate::DownReason::NoProc)
/// — the real reason (Exit vs Panic) is lost. Here the monitor is registered
/// on the child's slot *before* the child is published to any run queue, so
/// the `Down` always carries the child's actual termination reason. This is
/// Erlang's `spawn_monitor/1`.
pub fn spawn_monitor(f: impl FnOnce() + Send + 'static) -> (JoinHandle, crate::monitor::Monitor) {
spawn_monitor_with(SpawnOpts::default(), f)
}
/// [`spawn_monitor`] with per-actor stack shape overrides (RFC 019).
pub fn spawn_monitor_with(
opts: SpawnOpts,
f: impl FnOnce() + Send + 'static,
) -> (JoinHandle, crate::monitor::Monitor) {
let parent = current_pid().unwrap_or_else(|| with_runtime(|_| crate::runtime::ROOT_PID));
let (tx, rx) = crate::channel::channel::<crate::monitor::Down>();
let stack = with_runtime(|inner| crate::runtime::acquire_stack(inner, opts));
let sp = init_actor_stack(stack.top(), crate::actor::trampoline);
let closure: crate::runtime::Closure = Box::new(f);
let (pid, id) = with_runtime(|inner| {
let idx = inner.allocate_slot(); // panics loudly on slab exhaustion
let id = inner.alloc_monitor_id();
let pid =
crate::runtime::install_actor(inner, idx, sp, stack, parent, closure, Some((id, tx)));
(pid, id)
});
(
JoinHandle {
pid,
consumed: false,
},
crate::monitor::Monitor {
id,
target: pid,
rx,
},
)
}
/// [`spawn`] that reports a full actor slab instead of panicking. /// [`spawn`] that reports a full actor slab instead of panicking.
/// ///
/// Behaviour parity with [`spawn`] in every case except one: when the fixed /// Behaviour parity with [`spawn`] in every case except one: when the fixed
@@ -429,7 +480,7 @@ pub fn try_spawn_under_with<A>(
claimed.0 = None; // install_actor takes ownership of the slot from here claimed.0 = None; // install_actor takes ownership of the slot from here
let pid = with_runtime(|inner| { let pid = with_runtime(|inner| {
crate::runtime::install_actor(inner, idx, sp, stack, supervisor, closure) crate::runtime::install_actor(inner, idx, sp, stack, supervisor, closure, None)
}); });
Ok(JoinHandle { Ok(JoinHandle {
@@ -544,6 +595,36 @@ pub(crate) fn unpark_at(pid: Pid, epoch: u32) {
let _ = try_with_runtime(|inner| inner.unpark_at(pid, epoch)); let _ = try_with_runtime(|inner| inner.unpark_at(pid, epoch));
} }
// The current actor's runtime as a `Weak`, for a waker that must reach the
// runtime from a foreign thread later. A channel captures this ONCE, the first
// time its receiver parks, so a cross-thread `send` can wake without the
// `RUNTIME` thread-local (unset off a scheduler thread). `None` off a
// scheduler thread, where there is nothing to capture.
//
// Deliberately not called per park: `Arc::downgrade` plus the matching drop is
// a locked RMW pair on one globally shared counter, and the park/unpark
// round-trip is the hot path of every channel workload.
pub(crate) fn runtime_weak() -> Option<Weak<RuntimeInner>> {
try_with_runtime(Arc::downgrade)
}
// Epoch-matched wake of `pid` from a waker that may or may not be on a
// scheduler thread. On a scheduler thread we take the thread-local path
// (preemption-gated, slot-eligible); off one that path is a silent no-op, so
// we reach the runtime through `rt` — the `Weak` the waker captured while it
// was in-runtime. Mirrors the IO backend's cross-context wake (io.rs, RFC 018).
pub(crate) fn unpark_at_via(pid: Pid, epoch: u32, rt: impl FnOnce() -> Option<Weak<RuntimeInner>>) {
if try_with_runtime(|inner| inner.unpark_at(pid, epoch)).is_some() {
return;
}
// Off a scheduler thread only: `rt` re-takes the waker's lock to read the
// captured `Weak`, which is why it is a closure and not a value — the
// in-runtime path above must not pay for it.
if let Some(inner) = rt().as_ref().and_then(Weak::upgrade) {
inner.unpark_at(pid, epoch);
}
}
// Open a new wait for the current actor and return its wait identity // Open a new wait for the current actor and return its wait identity
// ("epoch"). Call once per wait, before registering with any waker. Lock-free, // ("epoch"). Call once per wait, before registering with any waker. Lock-free,
// so it's legal to call while already holding another internal lock. // so it's legal to call while already holding another internal lock.
@@ -577,10 +658,12 @@ pub(crate) fn retire_wait() {
/// [`JoinHandle::join`] reports it as a normal, non-error exit: cooperative /// [`JoinHandle::join`] reports it as a normal, non-error exit: cooperative
/// stop is a controlled shutdown, not a failure. /// stop is a controlled shutdown, not a failure.
/// ///
/// This is exactly the mechanism `gen_server` shutdown, supervisor restarts, /// This is the *hard* stop — OTP's `exit(Pid, kill)`. It is what a supervisor
/// and structured teardown are built from: reach for [`GenServerRef::shutdown`](crate::GenServerRef::shutdown) /// falls back to when a child overstays its [`Shutdown`](crate::supervisor::Shutdown)
/// or a [`supervisor`](crate::supervisor) instead of calling this directly /// grace period. For a stop the target gets to prepare for, use
/// where those apply. /// [`request_shutdown`]; for structured teardown, reach for
/// [`GenServerRef::shutdown`](crate::GenServerRef::shutdown) or a
/// [`supervisor`](crate::supervisor) instead of calling this directly.
/// ///
/// Because it's cooperative, an actor stuck in a tight loop with no /// Because it's cooperative, an actor stuck in a tight loop with no
/// blocking call, no [`check!`](crate::check), and no allocation cannot be /// blocking call, no [`check!`](crate::check), and no allocation cannot be
@@ -592,7 +675,8 @@ pub fn request_stop<A>(pid: Pid<A>) {
} }
// The core of `request_stop`, taking the runtime directly so it can also be // The core of `request_stop`, taking the runtime directly so it can also be
// driven from inside the runtime itself (the root-exit sweep) without // driven from inside the runtime itself (the RuntimeHandle path, supervisor
// sweeps) without
// re-borrowing the thread-local. Sets the stop flag under the target's lock // re-borrowing the thread-local. Sets the stop flag under the target's lock
// (a generation mismatch, or no live actor there, makes it a no-op) and // (a generation mismatch, or no live actor there, makes it a no-op) and
// wakes the target. // wakes the target.
@@ -612,6 +696,68 @@ pub(crate) fn request_stop_inner(inner: &RuntimeInner, pid: Pid) {
} }
} }
/// Ask an actor to shut down gracefully — OTP's `exit(Pid, shutdown)`, where
/// [`request_stop`] is `exit(Pid, kill)`.
///
/// If the target has called [`trap_exit`](crate::trap_exit), it receives an
/// [`ExitSignal`](crate::ExitSignal) with reason
/// [`DownReason::Shutdown`](crate::DownReason::Shutdown) on its trap inbox and
/// keeps running: the request is advisory, and the target is expected to wind
/// down and exit normally in its own time (a supervisor bounds that time with
/// its child's [`Shutdown`](crate::supervisor::Shutdown) policy and falls back
/// to `request_stop`). A target that is not trapping is stopped exactly as by
/// `request_stop`. A dead pid is a no-op.
///
/// The signal's `from` is the calling actor, or `ROOT_PID` when driven from
/// outside the runtime (see [`RuntimeHandle::request_shutdown`](crate::RuntimeHandle::request_shutdown)).
pub fn request_shutdown<A>(pid: Pid<A>) {
let pid = pid.erase();
let from = current_pid().unwrap_or(crate::runtime::ROOT_PID);
let _ = try_with_runtime(|inner| request_shutdown_inner(inner, pid, from));
}
// The core of `request_shutdown`. Reads the target's trap sender under its
// cold lock (generation-verified), then acts outside the lock: a trap send
// may unpark the receiver, and `request_stop_inner` re-takes the lock.
pub(crate) fn request_shutdown_inner(inner: &RuntimeInner, pid: Pid, from: Pid) {
request_shutdown_inner_probe(inner, pid, from);
}
/// [`request_shutdown_inner`], reporting what it found: `Some(true)` if the
/// target was trapping (got the signal), `Some(false)` if it was stopped
/// outright, `None` if there was nothing live at `pid`.
pub(crate) fn request_shutdown_inner_probe(
inner: &RuntimeInner,
pid: Pid,
from: Pid,
) -> Option<bool> {
let trap = match inner.slot_at(pid) {
Some(slot) => {
let cold = slot.cold.lock();
if slot.generation() == pid.generation() {
cold.actor.as_ref().map(|a| a.trap.clone())
} else {
None // stale pid: nothing there to shut down
}
}
None => None,
};
match trap {
Some(Some(tx)) => {
let _ = tx.send(crate::link::ExitSignal {
from,
reason: crate::monitor::DownReason::Shutdown,
});
Some(true)
}
Some(None) => {
request_stop_inner(inner, pid);
Some(false)
}
None => None,
}
}
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// NoPreempt // NoPreempt
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -624,14 +770,13 @@ pub struct NoPreempt(bool);
impl NoPreempt { impl NoPreempt {
pub fn enter() -> Self { pub fn enter() -> Self {
let prev = crate::preempt::PREEMPTION_ENABLED.with(|c| c.replace(false)); NoPreempt(crate::preempt::preemption_swap(false))
NoPreempt(prev)
} }
} }
impl Drop for NoPreempt { impl Drop for NoPreempt {
fn drop(&mut self) { fn drop(&mut self) {
crate::preempt::PREEMPTION_ENABLED.with(|c| c.set(self.0)); crate::preempt::preemption_swap(self.0);
} }
} }
+189 -68
View File
@@ -140,7 +140,8 @@ impl Signal {
} }
} }
use crate::channel::channel; use crate::channel::{channel, RecvTimeoutError};
use crate::monitor::DownReason;
use std::collections::{HashMap, VecDeque}; use std::collections::{HashMap, VecDeque};
use std::sync::Arc; use std::sync::Arc;
use std::time::{Duration, Instant}; use std::time::{Duration, Instant};
@@ -165,15 +166,53 @@ pub enum Restart {
pub struct ChildSpec { pub struct ChildSpec {
start: Arc<dyn Fn() + Send + Sync + 'static>, start: Arc<dyn Fn() + Send + Sync + 'static>,
restart: Restart, restart: Restart,
shutdown: Shutdown,
} }
impl ChildSpec { impl ChildSpec {
/// A child with the given restart policy and the default
/// [`Shutdown::Timeout`] of 5 seconds.
pub fn new(restart: Restart, start: impl Fn() + Send + Sync + 'static) -> Self { pub fn new(restart: Restart, start: impl Fn() + Send + Sync + 'static) -> Self {
Self { Self {
start: Arc::new(start), start: Arc::new(start),
restart, restart,
shutdown: Shutdown::default(),
} }
} }
/// Set how the supervisor stops this child (see [`Shutdown`]). A child
/// that is itself a supervisor should use [`Shutdown::Infinity`] so its
/// own subtree gets its full grace periods.
pub fn shutdown(mut self, shutdown: Shutdown) -> Self {
self.shutdown = shutdown;
self
}
}
/// How a supervisor stops a child it is taking down — the OTP child-spec
/// `shutdown` value. Applies to every supervisor-initiated stop: the ordered
/// shutdown of the whole set and the sibling cycling of
/// [`Strategy::OneForAll`] / [`Strategy::RestForOne`].
///
/// A graceful stop is a [`request_shutdown`](crate::request_shutdown): a child
/// that traps exits receives the request as a message and winds down in its
/// own time; one that does not is stopped outright.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum Shutdown {
/// `request_stop` immediately; no request, no grace period.
BrutalKill,
/// `request_shutdown`, wait up to the duration for the child to exit, then
/// `request_stop` it. The default, at 5 seconds.
Timeout(Duration),
/// `request_shutdown` and wait however long the child takes. Use for a
/// child supervisor, whose subtree has its own timeouts.
Infinity,
}
impl Default for Shutdown {
fn default() -> Self {
Shutdown::Timeout(Duration::from_secs(5))
}
} }
/// How a supervisor reacts when one child terminates and a restart is due. /// How a supervisor reacts when one child terminates and a restart is due.
@@ -244,23 +283,35 @@ impl OneForOne {
} }
/// Run the supervision loop on the current actor. Returns when every child /// Run the supervision loop on the current actor. Returns when every child
/// has reached a terminal, non-restartable state, or when the restart /// has reached a terminal, non-restartable state, when the restart
/// intensity cap is tripped. /// intensity cap is tripped, or when the supervisor is asked to shut down
/// (a [`request_shutdown`](crate::request_shutdown) — from its own
/// supervisor, or from the app). On every one of those exits the survivors
/// are stopped in reverse start order, each per its
/// [`Shutdown`] policy, before this returns.
///
/// The supervisor traps exits for the length of the loop (that is how the
/// shutdown request reaches it as a message). Should the supervisor itself
/// be hard-stopped with [`request_stop`](crate::request_stop), it unwinds
/// without waiting for anything — but a drop guard hard-stops its live
/// children on the way out, so the subtree is not orphaned (a child
/// supervisor unwinds the same way, recursively).
pub fn run(self) { pub fn run(self) {
let me = crate::scheduler::self_pid(); let me = crate::scheduler::self_pid();
let (tx, rx) = channel::<Signal>(); let (tx, rx) = channel::<Signal>();
crate::scheduler::register_supervisor_channel(me, tx); crate::scheduler::register_supervisor_channel(me, tx);
let exits = crate::link::trap_exit();
// pid -> index into `self.children`, for the children currently alive. // pid -> index into `self.children`, for the children currently alive.
let mut by_pid: HashMap<Pid, usize> = HashMap::new(); let mut live = Live::default();
let mut active: usize = 0; let mut active: usize = 0;
// Sliding window of recent restart instants, for the intensity cap. // Sliding window of recent restart instants, for the intensity cap.
let mut restarts: Vec<Instant> = Vec::new(); let mut restarts: Vec<Instant> = Vec::new();
let start_child = |idx: usize, by_pid: &mut HashMap<Pid, usize>| { let start_child = |idx: usize, live: &mut Live| {
let start = self.children[idx].start.clone(); let start = self.children[idx].start.clone();
let h = crate::scheduler::spawn_under(me, move || (start)()); let h = crate::scheduler::spawn_under(me, move || (start)());
by_pid.insert(h.pid(), idx); live.insert(h.pid(), idx);
// We supervise via the signal funnel, not by joining; drop the // We supervise via the signal funnel, not by joining; drop the
// handle so the child's slot is reclaimed promptly on death (the // handle so the child's slot is reclaimed promptly on death (the
// termination Signal is delivered before reclamation regardless). // termination Signal is delivered before reclamation regardless).
@@ -268,28 +319,105 @@ impl OneForOne {
}; };
for idx in 0..self.children.len() { for idx in 0..self.children.len() {
start_child(idx, &mut by_pid); start_child(idx, &mut live);
active += 1; active += 1;
} }
// A signal that arrives while we are awaiting stop-confirmations (for a // A signal that arrives while we are awaiting stop-confirmations (for a
// child we are *not* currently stopping) is stashed here and processed // child we are *not* currently stopping) is stashed here and processed
// by the main loop before it blocks on `recv` again. // by the main loop before it blocks again.
let mut pending: VecDeque<Signal> = VecDeque::new(); let mut pending: VecDeque<Signal> = VecDeque::new();
let next_signal = |pending: &mut VecDeque<Signal>| -> Option<Signal> {
if let Some(s) = pending.pop_front() { // Stop one child per its policy and wait for its termination signal.
Some(s) // Signals for other pids that arrive meanwhile are stashed. Bounded by
} else { // construction: `request_stop` (used directly, or as the fallback once
rx.recv().ok() // the grace period lapses) always produces a signal.
let stop_child = |pid: Pid, idx: usize, pending: &mut VecDeque<Signal>| {
let await_one = |deadline: Option<Instant>, pending: &mut VecDeque<Signal>| -> bool {
loop {
let sig = match pending.iter().position(|s| s.pid() == pid) {
Some(i) => pending.remove(i),
None => match deadline {
None => rx.recv().ok(),
Some(dl) => {
match rx.recv_timeout(dl.saturating_duration_since(Instant::now()))
{
Ok(s) => Some(s),
Err(RecvTimeoutError::Timeout) => return false,
Err(RecvTimeoutError::Disconnected) => None,
}
}
},
};
match sig {
Some(s) if s.pid() == pid => return true,
Some(s) => pending.push_back(s),
None => return true, // funnel closed: nothing more can arrive
}
}
};
match self.children[idx].shutdown {
Shutdown::BrutalKill => {
crate::scheduler::request_stop(pid);
await_one(None, pending);
}
Shutdown::Timeout(grace) => {
crate::scheduler::request_shutdown(pid);
if !await_one(Some(Instant::now() + grace), pending) {
crate::scheduler::request_stop(pid);
await_one(None, pending);
}
}
Shutdown::Infinity => {
crate::scheduler::request_shutdown(pid);
await_one(None, pending);
}
}
};
// Stop a set of children in reverse start order, one at a time.
let stop_set =
|set: &mut Vec<(Pid, usize)>, live: &mut Live, pending: &mut VecDeque<Signal>| {
set.sort_unstable_by_key(|x| std::cmp::Reverse(x.1));
for (pid, idx) in set.iter() {
live.remove(pid);
stop_child(*pid, *idx, pending);
}
};
// Wait for the next event: a stashed signal, a child signal, or a
// shutdown request. `Ok(sig)`, or `Err(())` when we must wind down.
let next_event = |pending: &mut VecDeque<Signal>| -> Result<Signal, ()> {
loop {
if let Some(s) = pending.pop_front() {
return Ok(s);
}
// The trap inbox is arm 0: a shutdown request is noticed even
// under a flood of child signals.
match crate::channel::select(&[&exits, &rx]) {
0 => match exits.try_recv() {
Ok(Some(sig)) if sig.reason == DownReason::Shutdown => return Err(()),
// Any other exit signal (a linked peer's death — a
// supervisor links nothing itself, but may be linked
// to) is not ours to act on; a closed trap inbox is
// impossible while `exits` is held here.
_ => {}
},
_ => match rx.try_recv() {
Ok(Some(s)) => return Ok(s),
Ok(None) => {}
Err(_) => return Err(()), // funnel closed: nothing left to supervise
},
}
} }
}; };
while active > 0 { while active > 0 {
let sig = match next_signal(&mut pending) { let sig = match next_event(&mut pending) {
Some(s) => s, Ok(s) => s,
None => break, // mailbox closed: nothing left to supervise Err(()) => break,
}; };
let idx = match by_pid.remove(&sig.pid()) { let idx = match live.remove(&sig.pid()) {
Some(i) => i, Some(i) => i,
None => continue, // stray/duplicate signal None => continue, // stray/duplicate signal
}; };
@@ -321,76 +449,69 @@ impl OneForOne {
restarts.push(now); restarts.push(now);
// Which *live* siblings get cycled along with the failed child. // Which *live* siblings get cycled along with the failed child.
// (The failed child is already gone — removed from `by_pid` above.) // (The failed child is already gone — removed from `live` above.)
let mut to_stop: Vec<(Pid, usize)> = match self.strategy { let mut to_stop: Vec<(Pid, usize)> = match self.strategy {
Strategy::OneForOne => Vec::new(), Strategy::OneForOne => Vec::new(),
Strategy::OneForAll => by_pid.iter().map(|(p, i)| (*p, *i)).collect(), Strategy::OneForAll => live.iter().map(|(p, i)| (*p, *i)).collect(),
Strategy::RestForOne => by_pid Strategy::RestForOne => live
.iter() .iter()
.filter(|(_, i)| **i > idx) .filter(|(_, i)| **i > idx)
.map(|(p, i)| (*p, *i)) .map(|(p, i)| (*p, *i))
.collect(), .collect(),
}; };
// Stop survivors in reverse start order (highest child index first).
to_stop.sort_unstable_by_key(|x| std::cmp::Reverse(x.1));
// The set we will restart: the failed child plus every sibling we // The set we will restart: the failed child plus every sibling we
// are about to stop, restarted in start (ascending index) order. // are about to stop, restarted in start (ascending index) order.
let mut restart_set: Vec<usize> = Vec::with_capacity(to_stop.len() + 1); let mut restart_set: Vec<usize> = Vec::with_capacity(to_stop.len() + 1);
restart_set.push(idx); restart_set.push(idx);
restart_set.extend(to_stop.iter().map(|(_, i)| *i));
// Request stops, then await each survivor's termination signal // Stop the survivors (each per its policy, reverse start order),
// before restarting. `request_stop` on an already-dead pid is a // then restart the whole set in start order. Net effect on
// no-op; in that case its (already-sent) Exit signal serves as the // `active`: one child died (idx), `to_stop.len()` were stopped,
// confirmation. Any signal for a pid we are *not* awaiting is // and `restart_set.len() == 1 + to_stop.len()` are started — so
// stashed for the main loop.
let mut awaiting: Vec<Pid> = Vec::with_capacity(to_stop.len());
for (pid, cidx) in &to_stop {
by_pid.remove(pid);
restart_set.push(*cidx);
crate::scheduler::request_stop(*pid);
awaiting.push(*pid);
}
while !awaiting.is_empty() {
let s = match next_signal(&mut pending) {
Some(s) => s,
None => break, // mailbox closed mid-await; stop waiting
};
if let Some(pos) = awaiting.iter().position(|p| *p == s.pid()) {
awaiting.swap_remove(pos);
} else {
pending.push_back(s);
}
}
// Restart the whole set in start order. Net effect on `active`:
// one child died (idx), `to_stop.len()` were stopped, and
// `restart_set.len() == 1 + to_stop.len()` are started — so
// `active` is unchanged and needs no adjustment here. // `active` is unchanged and needs no adjustment here.
stop_set(&mut to_stop, &mut live, &mut pending);
restart_set.sort_unstable(); restart_set.sort_unstable();
for cidx in restart_set { for cidx in restart_set {
start_child(cidx, &mut by_pid); start_child(cidx, &mut live);
} }
} }
// Ordered shutdown: stop any survivors in reverse start order and await // Ordered shutdown: stop any survivors in reverse start order, each per
// their termination. On the normal `active == 0` exit `by_pid` is empty // its policy. On the normal `active == 0` exit `live` is empty and this
// and this is a no-op; on a cap-trip or mailbox-closed break it tears // is a no-op; on a shutdown request, a cap-trip, or a closed funnel it
// the remaining children down deterministically instead of leaking them. // tears the remaining children down deterministically.
let mut survivors: Vec<(Pid, usize)> = by_pid.iter().map(|(p, i)| (*p, *i)).collect(); let mut survivors: Vec<(Pid, usize)> = live.iter().map(|(p, i)| (*p, *i)).collect();
survivors.sort_unstable_by_key(|x| std::cmp::Reverse(x.1)); stop_set(&mut survivors, &mut live, &mut pending);
let mut awaiting: Vec<Pid> = Vec::with_capacity(survivors.len()); }
for (pid, _) in &survivors { }
crate::scheduler::request_stop(*pid);
awaiting.push(*pid); /// The live children of a supervisor, with a drop guard: if the supervisor is
} /// unwound (a hard `request_stop`, or a panic in the loop) its children are
while !awaiting.is_empty() { /// hard-stopped rather than orphaned. Fire-and-forget by necessity — a guard
let s = match next_signal(&mut pending) { /// running mid-unwind cannot park to await anything.
Some(s) => s, #[derive(Default)]
None => break, struct Live(HashMap<Pid, usize>);
};
if let Some(pos) = awaiting.iter().position(|p| *p == s.pid()) { impl std::ops::Deref for Live {
awaiting.swap_remove(pos); type Target = HashMap<Pid, usize>;
fn deref(&self) -> &Self::Target {
&self.0
}
}
impl std::ops::DerefMut for Live {
fn deref_mut(&mut self) -> &mut Self::Target {
&mut self.0
}
}
impl Drop for Live {
fn drop(&mut self) {
if std::thread::panicking() {
for pid in self.0.keys() {
crate::scheduler::request_stop(*pid);
} }
} }
} }
+37 -8
View File
@@ -46,17 +46,32 @@ mod inner {
#[derive(Clone, Debug)] #[derive(Clone, Debug)]
pub enum Event { pub enum Event {
// Actor lifecycle // Actor lifecycle
Spawn { parent: Pid, child: Pid }, Spawn {
parent: Pid,
child: Pid,
},
Resume(Pid), Resume(Pid),
Yield(Pid), Yield(Pid),
Park(Pid), Park(Pid),
Done(Pid), Done(Pid),
/// Root exit found a live forest root (an actor nobody supervises)
/// and delivered `request_shutdown` to it. `trapping` says whether it
/// got the chance to drain (true) or was stopped outright (false).
/// Every such line is an actor whose lifetime was nobody's business
/// but the runtime's — the way to *see* unsupervised leftovers.
RootSweep {
target: Pid,
trapping: bool,
},
// Wakeup paths // Wakeup paths
UnparkDirect(Pid), // unpark() saw Parked -> re-queued immediately UnparkDirect(Pid), // unpark() saw Parked -> re-queued immediately
UnparkDeferred(Pid), // unpark() saw Runnable -> set pending_unpark flag UnparkDeferred(Pid), // unpark() saw Runnable -> set pending_unpark flag
UnparkFlagConsumed(Pid), // scheduler saw flag on Park -> re-queued instead UnparkFlagConsumed(Pid), // scheduler saw flag on Park -> re-queued instead
// Channel // Channel
Send { sender: Pid, receiver: Option<Pid> }, Send {
sender: Pid,
receiver: Option<Pid>,
},
RecvPark(Pid), RecvPark(Pid),
RecvWake(Pid), RecvWake(Pid),
// Queue // Queue
@@ -65,6 +80,13 @@ mod inner {
// RFC 005 wake slot // RFC 005 wake slot
SlotPush(Pid), // actor-context wake parked in the waking thread's slot SlotPush(Pid), // actor-context wake parked in the waking thread's slot
SlotPop(Pid), // scheduler resumed a pid from its own slot SlotPop(Pid), // scheduler resumed a pid from its own slot
// Cluster (RFC 010): the conn actor's verdict on one inbound frame —
// local knowledge only, never on the wire; the label is
// `InboundVerdict::label()`. No pid: a refused frame has none.
ClusterInbound(&'static str),
// Cluster (RFC 010): the connector's verdict on one dial attempt —
// `"ok"` or `DialError::label()`. No pid.
ClusterDial(&'static str),
} }
// ----------------------------------------------------------------------- // -----------------------------------------------------------------------
@@ -161,18 +183,16 @@ mod inner {
// Hot path // Hot path
// ----------------------------------------------------------------------- // -----------------------------------------------------------------------
#[inline(never)]
pub fn record(event: Event) { pub fn record(event: Event) {
crate::context::tls_fence();
// Disable preemption for the entire duration of record(). Any // Disable preemption for the entire duration of record(). Any
// allocation here (mutex internals, channel send, lazy init) would // allocation here (mutex internals, channel send, lazy init) would
// trigger PreemptingAllocator -> maybe_preempt -> switch_to_scheduler, // trigger PreemptingAllocator -> maybe_preempt -> switch_to_scheduler,
// which would try to re-acquire inner.shared (already held at many // which would try to re-acquire inner.shared (already held at many
// te!() call sites) -> deadlock. Guard at the very top, before any // te!() call sites) -> deadlock. Guard at the very top, before any
// allocation-capable call. // allocation-capable call.
let was_enabled = crate::preempt::PREEMPTION_ENABLED.with(|e| { let was_enabled = crate::preempt::preemption_swap(false);
let v = e.get();
e.set(false);
v
});
LOCAL_STATE.with(|cell| { LOCAL_STATE.with(|cell| {
let mut opt = cell.borrow_mut(); let mut opt = cell.borrow_mut();
@@ -194,7 +214,7 @@ mod inner {
} }
}); });
crate::preempt::PREEMPTION_ENABLED.with(|e| e.set(was_enabled)); crate::preempt::preemption_swap(was_enabled);
} }
// ----------------------------------------------------------------------- // -----------------------------------------------------------------------
@@ -253,6 +273,13 @@ mod inner {
Event::Yield(p) => ("yield".into(), p.index()), Event::Yield(p) => ("yield".into(), p.index()),
Event::Park(p) => ("park".into(), p.index()), Event::Park(p) => ("park".into(), p.index()),
Event::Done(p) => ("done".into(), p.index()), Event::Done(p) => ("done".into(), p.index()),
Event::RootSweep { target, trapping } => (
format!(
"root_sweep {}",
if *trapping { "shutdown" } else { "stopped" }
),
target.index(),
),
Event::UnparkDirect(p) => ("unpark_direct".into(), p.index()), Event::UnparkDirect(p) => ("unpark_direct".into(), p.index()),
Event::UnparkDeferred(p) => ("unpark_deferred".into(), p.index()), Event::UnparkDeferred(p) => ("unpark_deferred".into(), p.index()),
Event::UnparkFlagConsumed(p) => ("unpark_flag_consumed".into(), p.index()), Event::UnparkFlagConsumed(p) => ("unpark_flag_consumed".into(), p.index()),
@@ -271,6 +298,8 @@ mod inner {
Event::Dequeue(p) => ("dequeue".into(), p.index()), Event::Dequeue(p) => ("dequeue".into(), p.index()),
Event::SlotPush(p) => ("slot_push".into(), p.index()), Event::SlotPush(p) => ("slot_push".into(), p.index()),
Event::SlotPop(p) => ("slot_pop".into(), p.index()), Event::SlotPop(p) => ("slot_pop".into(), p.index()),
Event::ClusterInbound(v) => (format!("cluster_inbound {v}"), 0),
Event::ClusterDial(v) => (format!("cluster_dial {v}"), 0),
} }
} }
+4 -2
View File
@@ -138,10 +138,12 @@ fn channel_ops_interleaved_with_monitor_churn_multi_thread() {
let tx = tx.clone(); let tx = tx.clone();
handles.push(spawn(move || { handles.push(spawn(move || {
// Short-lived target whose death fires the monitor below. // Short-lived target whose death fires the monitor below.
let t = spawn(move || { // spawn_monitor: registered before publish, so the Down is
// the finalize-sent Exit this test is about, never NoProc
// (spawn-then-monitor raced ~8% at 4 threads).
let (t, m) = smarm::spawn_monitor(move || {
tx.send(i).unwrap(); tx.send(i).unwrap();
}); });
let m = smarm::monitor(t.pid());
t.join().unwrap(); t.join().unwrap();
// Down delivery exercises send-from-finalize. // Down delivery exercises send-from-finalize.
let d = m.rx.recv().unwrap(); let d = m.rx.recv().unwrap();
+117
View File
@@ -0,0 +1,117 @@
//! RFC 010 c6a — connection-actor lifecycle against the manager table.
//!
//! The handshake is bypassed here (c6b wires it): each connection is
//! constructed already-established over a real localhost TCP pair, handed a
//! fabricated `Peer`, and spawned. `spawn_established` registers it with the
//! manager, which takes its handle and monitors it, so the table reflects the
//! connection while it lives and reaps it on any exit path. This proves three
//! things at once: a live connection shows up, a commanded `Disconnect`
//! removes exactly that one, and a peer close (EOF, no command) removes the
//! other.
//!
//! TCP parks the calling actor, so everything runs inside `smarm::run`; the
//! single-threaded runtime is fine because every wait is a cooperative fd park.
#![cfg(feature = "cluster")]
use std::time::Duration;
use smarm::cluster::envelope::NodeMeta;
use smarm::cluster::handshake::Peer;
use smarm::cluster::manager::{Call, Manager, Reply, MANAGER};
use smarm::cluster::spawn_established;
use smarm::cluster::transport::tcp::TcpTransport;
use smarm::cluster::transport::{Conn, FramedConn, Transport};
use smarm::cluster::Timing;
use smarm::gen_server::{self, GenServerBuilder};
use smarm::pg::Incarnation;
use smarm::{run, sleep};
/// A fabricated post-handshake peer identity. Only `node_name` matters to the
/// manager table; the rest is filler until c7 consumes it.
fn peer(name: &str) -> Peer {
Peer {
node_name: name.to_string(),
incarnation: Incarnation::new(1),
meta: NodeMeta {
role: "test".to_string(),
region: "test".to_string(),
},
}
}
/// One established transport pair over localhost. Relies on TCP backlog so the
/// sequential dial-then-accept needs no concurrent acceptor (same assumption as
/// the c3 conformance suite).
fn pair(t: &dyn Transport) -> (Box<dyn Conn>, Box<dyn Conn>) {
let mut l = t.listen("127.0.0.1:0").unwrap();
let a = t.dial(&l.local_addr()).unwrap();
let b = l.accept().unwrap();
(a, b)
}
/// Poll the manager until its peer set matches `expected` (sorted), or fail.
/// The bound is generous against a sub-millisecond real cost.
fn wait_peers(expected: &[&str]) {
let want: Vec<String> = expected.iter().map(|s| s.to_string()).collect();
for _ in 0..2000 {
if let Ok(Reply::Peers(got)) = gen_server::call(MANAGER, Call::Peers) {
if got == want {
return;
}
}
sleep(Duration::from_millis(1));
}
let got = gen_server::call(MANAGER, Call::Peers);
panic!("timed out waiting for peers == {want:?}; last = {got:?}");
}
#[test]
fn connection_up_commanded_shutdown_and_eof_all_reflected_in_table() {
run(|| {
// The manager, started plainly and reachable at its well-known name.
// (The supervised subtree in `cluster::start` is permanent by design;
// a plainly-started manager lets this test terminate cleanly.)
let mgr = GenServerBuilder::new(Manager::new())
.named(MANAGER)
.start()
.expect("manager name is free");
let t = TcpTransport;
let (a1, b1) = pair(&t);
let (a2, b2) = pair(&t);
// Manage the `a` ends as peers node-b and node-c; keep the `b` far ends
// open so neither socket is closed from the far side yet.
spawn_established(FramedConn::new(a1), peer("node-b"), Timing::default())
.expect("node-b registers");
spawn_established(FramedConn::new(a2), peer("node-c"), Timing::default())
.expect("node-c registers");
// Up: both connections register and the table shows them.
wait_peers(&["node-b", "node-c"]);
// A commanded disconnect reaps exactly its own connection: the
// manager drops that entry's handle and the actor stops.
assert!(matches!(
gen_server::call(
MANAGER,
Call::Disconnect {
name: "node-b".to_string()
}
),
Ok(Reply::Disconnected)
));
wait_peers(&["node-c"]);
// A peer close (EOF) reaps the other with no command at all.
drop(b2);
wait_peers(&[]);
// node-b's far end stayed open until here, so its removal above was the
// disconnect command and not an EOF.
drop(b1);
// All connection actors have exited; stop the manager so `run` returns.
mgr.shutdown();
});
}
+180
View File
@@ -0,0 +1,180 @@
//! RFC 010 c6c — heartbeat send + fixed-timeout liveness + teardown.
//!
//! Each case runs one real connection actor over an in-process localhost TCP
//! pair, with the far end held as a raw `FramedConn` (no actor) so the test
//! controls exactly what — if anything — the peer says. That gives the three
//! protocol-visible facts direct handles: heartbeats appear on the wire
//! unprompted; a mute peer is torn down (and reaped from the manager table)
//! once `LIVENESS_TIMEOUT` empties; and a peer that does nothing but send
//! heartbeats keeps the connection alive past that same window.
//!
//! Loopback has no fd and cannot drive liveness (documented on the actor),
//! so everything here is TCP. TCP parks the calling actor, so everything
//! runs inside `smarm::run`.
#![cfg(feature = "cluster")]
use std::time::{Duration, Instant};
use smarm::cluster::conn::{HEARTBEAT_INTERVAL, LIVENESS_TIMEOUT};
use smarm::cluster::envelope::{Frame, NodeMeta};
use smarm::cluster::handshake::Peer;
use smarm::cluster::manager::{Call, Manager, Reply, MANAGER};
use smarm::cluster::spawn_established;
use smarm::cluster::transport::tcp::TcpTransport;
use smarm::cluster::transport::{Conn, FramedConn, Transport};
use smarm::cluster::Timing;
use smarm::gen_server::{self, GenServerBuilder};
use smarm::pg::Incarnation;
use smarm::{run, sleep, spawn};
/// A fabricated post-handshake peer identity (same shape as the c6a suite).
fn peer(name: &str) -> Peer {
Peer {
node_name: name.to_string(),
incarnation: Incarnation::new(1),
meta: NodeMeta {
role: "test".to_string(),
region: "test".to_string(),
},
}
}
/// One established transport pair over localhost (TCP backlog covers the
/// sequential dial-then-accept, as in the c3 conformance suite).
fn pair(t: &dyn Transport) -> (Box<dyn Conn>, Box<dyn Conn>) {
let mut l = t.listen("127.0.0.1:0").unwrap();
let a = t.dial(&l.local_addr()).unwrap();
let b = l.accept().unwrap();
(a, b)
}
fn peers() -> Vec<String> {
match gen_server::call(MANAGER, Call::Peers) {
Ok(Reply::Peers(p)) => p,
other => panic!("manager unreachable: {other:?}"),
}
}
/// Poll until the manager's peer set matches `expected` (sorted) or `budget`
/// runs out.
fn wait_peers(expected: &[&str], budget: Duration) {
let want: Vec<String> = expected.iter().map(|s| s.to_string()).collect();
let deadline = Instant::now() + budget;
while Instant::now() < deadline {
if peers() == want {
return;
}
sleep(Duration::from_millis(10));
}
panic!(
"timed out waiting for peers == {want:?}; last = {:?}",
peers()
);
}
/// The actor emits heartbeats unprompted: the raw far end, saying nothing,
/// sees a `Frame::Heartbeat` well within one interval (the first goes out at
/// spawn).
#[test]
fn heartbeats_are_sent_unprompted() {
run(|| {
let mgr = GenServerBuilder::new(Manager::new())
.named(MANAGER)
.start()
.expect("manager name is free");
let (a, b) = pair(&TcpTransport);
spawn_established(FramedConn::new(a), peer("hb-send"), Timing::default())
.expect("register");
let mut far = FramedConn::new(b);
let frame = far
.recv_deadline(Instant::now() + HEARTBEAT_INTERVAL)
.expect("a heartbeat before one interval elapses");
assert_eq!(frame, Some(Frame::Heartbeat));
// Teardown: closing the far end is an EOF at the actor.
far.close();
wait_peers(&[], Duration::from_secs(2));
mgr.shutdown();
});
}
/// A mute peer is dead: no inbound frame for `LIVENESS_TIMEOUT` tears the
/// connection down and the manager's monitor reaps the table entry. The
/// entry is still present well inside the window — the teardown is the
/// timer, not an accident of setup.
#[test]
fn mute_peer_is_torn_down_after_liveness_timeout() {
run(|| {
let mgr = GenServerBuilder::new(Manager::new())
.named(MANAGER)
.start()
.expect("manager name is free");
let (a, b) = pair(&TcpTransport);
spawn_established(FramedConn::new(a), peer("mute"), Timing::default()).expect("register");
// Held open and silent: no frames, no EOF. (Unread inbound
// heartbeats sit in kernel buffers; they are 5 bytes each.)
let _far = FramedConn::new(b);
// Well inside the window the connection is still up.
sleep(LIVENESS_TIMEOUT / 2);
assert_eq!(peers(), vec!["mute".to_string()], "torn down too early");
// ...and once the window empties it is gone. Generous budget over
// the remaining half-window.
wait_peers(&[], LIVENESS_TIMEOUT);
mgr.shutdown();
});
}
/// Heartbeats alone keep a connection alive past `LIVENESS_TIMEOUT`: a far
/// end that sends `Frame::Heartbeat` at the interval (and nothing else)
/// holds the entry; when it goes quiet, liveness finally fires.
#[test]
fn heartbeats_keep_the_connection_alive() {
run(|| {
let mgr = GenServerBuilder::new(Manager::new())
.named(MANAGER)
.start()
.expect("manager name is free");
let (a, b) = pair(&TcpTransport);
spawn_established(FramedConn::new(a), peer("kept"), Timing::default()).expect("register");
// The far heartbeat pump: interval-paced sends until told to stop,
// then holds the socket open, silent, so the eventual teardown is
// liveness — not EOF.
let (ctl_tx, ctl_rx) = smarm::channel::channel::<()>();
spawn(move || {
let mut far = FramedConn::new(b);
// Phase 1: heartbeat at the interval until the first signal.
while matches!(ctl_rx.try_recv(), Ok(None)) {
far.send(&Frame::Heartbeat).expect("far send");
sleep(HEARTBEAT_INTERVAL);
}
// Phase 2: silent but with the socket held open — dropping
// `far` here would EOF the actor and mask the liveness path.
// Exits when the test's closure ends and drops `ctl_tx` (an
// eternal park would stop `run` from ever returning).
while matches!(ctl_rx.try_recv(), Ok(None)) {
sleep(Duration::from_millis(20));
}
});
// Past the liveness window with margin: still up.
sleep(LIVENESS_TIMEOUT + LIVENESS_TIMEOUT / 2);
assert_eq!(
peers(),
vec!["kept".to_string()],
"liveness fired despite heartbeats"
);
// Silence the pump; liveness now empties and the entry goes.
ctl_tx.send(()).expect("pump alive");
wait_peers(&[], LIVENESS_TIMEOUT * 2);
mgr.shutdown();
// `ctl_tx` drops here, releasing the pump's phase-2 wait.
});
}
+482
View File
@@ -0,0 +1,482 @@
//! RFC 010 c6b — the handshake on the accept/connect path.
//!
//! Path-level tests drive [`dial_handshake`]/[`accept_handshake`] over the
//! loopback transport on plain threads (its intended use — synchronous, no
//! runtime). Integration tests run the manager-backed [`dial`] and
//! [`spawn_acceptor`] over real localhost TCP inside `smarm::run`, and the
//! two-node case as subprocesses via the c4 harness. Flake budget: see
//! tests/common/mod.rs.
#![cfg(feature = "cluster")]
mod common;
use std::sync::mpsc;
use std::time::{Duration, Instant};
use common::{maybe_child, spawn_node, WAIT};
use smarm::cluster::connect::{
accept_handshake, dial, dial_handshake, spawn_acceptor, DialError, HandshakeError,
HANDSHAKE_TIMEOUT,
};
use smarm::cluster::envelope::{Frame, NodeMeta, RejectReason};
use smarm::cluster::handshake::{Local, PeerStanding};
use smarm::cluster::manager::{Call, Manager, Reply, MANAGER};
use smarm::cluster::transport::loopback::LoopbackTransport;
use smarm::cluster::transport::tcp::TcpTransport;
use smarm::cluster::transport::{FramedConn, Transport};
use smarm::cluster::Timing;
use smarm::gen_server::{self, GenServerBuilder};
use smarm::pg::Incarnation;
use smarm::{run, sleep};
const ROLES: &[(&str, fn())] = &[
("hs_listener", role_hs_listener),
("hs_dialer", role_hs_dialer),
];
const HASH: u64 = 0xC6B0_C6B0_C6B0_C6B0;
fn local(name: &str) -> Local {
Local {
node_name: name.into(),
incarnation: Incarnation::new(3),
build_hash: HASH,
meta: NodeMeta {
role: "test".into(),
region: "test".into(),
},
}
}
/// A loopback conn pair as `FramedConn`s, ready for a threaded handshake.
fn loopback_pair() -> (FramedConn, FramedConn) {
let t = LoopbackTransport::default();
let mut l = t.listen("hs").unwrap();
let dialer = FramedConn::new(t.dial("hs").unwrap());
let accepted = FramedConn::new(l.accept().unwrap());
(dialer, accepted)
}
/// Far-future deadline for loopback paths, where it cannot fire anyway.
fn no_deadline() -> Instant {
Instant::now() + Duration::from_secs(3600)
}
/// Park the node forever: it has announced everything the parent asserts on,
/// and must now hold its connection open until SIGKILLed.
fn park() -> ! {
loop {
sleep(Duration::from_secs(1));
}
}
/// Cooperative bounded receive across the closure/actor boundary. A blocking
/// `std::mpsc` wait would park the OS thread and starve the single-threaded
/// scheduler, so every wait inside `run` polls with [`sleep`] instead.
fn poll_recv<T>(rx: &mpsc::Receiver<T>, what: &str) -> T {
let deadline = Instant::now() + WAIT;
loop {
match rx.try_recv() {
Ok(v) => return v,
Err(mpsc::TryRecvError::Empty) => {
assert!(Instant::now() < deadline, "timed out waiting for {what}");
sleep(Duration::from_millis(1));
}
Err(mpsc::TryRecvError::Disconnected) => panic!("channel closed waiting for {what}"),
}
}
}
// ---------------------------------------------------------------------------
// Path level, over loopback on plain threads
// ---------------------------------------------------------------------------
#[test]
fn loopback_happy_path_establishes_both_ends() {
maybe_child(ROLES);
let (mut dialer, mut accepted) = loopback_pair();
let responder = std::thread::spawn(move || {
accept_handshake(
&mut accepted,
local("node-b"),
|name| {
assert_eq!(name, "node-a");
PeerStanding::Free
},
no_deadline(),
)
});
let peer_of_dialer = dial_handshake(&mut dialer, &local("node-a"), no_deadline()).unwrap();
let peer_of_acceptor = responder.join().unwrap().unwrap();
assert_eq!(peer_of_dialer.node_name, "node-b");
assert_eq!(peer_of_acceptor.node_name, "node-a");
}
#[test]
fn loopback_hash_mismatch_rejected_with_frame_then_eof() {
maybe_child(ROLES);
let (mut dialer, mut accepted) = loopback_pair();
let mut wrong = local("node-b");
wrong.build_hash ^= 1;
let responder = std::thread::spawn(move || {
accept_handshake(&mut accepted, wrong, |_| PeerStanding::Free, no_deadline())
});
// The dial side receives the reject frame — the compatibility anchor.
match dial_handshake(&mut dialer, &local("node-a"), no_deadline()) {
Err(HandshakeError::Rejected(RejectReason::HashMismatch)) => {}
other => panic!("expected HashMismatch reject, got {other:?}"),
}
match responder.join().unwrap() {
Err(HandshakeError::Rejected(RejectReason::HashMismatch)) => {}
other => panic!("expected accept side to report the reject, got {other:?}"),
}
}
#[test]
fn loopback_tie_break_loser_closed_silently() {
maybe_child(ROLES);
// The inbound dial is from "node-z"; we are "node-a" with our own dial to
// node-z in flight. dial_wins("node-z", "node-a") is false, so the
// inbound loses: closed with no frame at all.
let (mut dialer, mut accepted) = loopback_pair();
let responder = std::thread::spawn(move || {
accept_handshake(
&mut accepted,
local("node-a"),
|_| PeerStanding::Dialing,
no_deadline(),
)
});
// Silent close: the dial side sees EOF, never a frame.
match dial_handshake(&mut dialer, &local("node-z"), no_deadline()) {
Err(HandshakeError::Closed) => {}
other => panic!("expected silent close (Closed), got {other:?}"),
}
match responder.join().unwrap() {
Err(HandshakeError::TieBreakLoss) => {}
other => panic!("expected TieBreakLoss on the accept side, got {other:?}"),
}
}
#[test]
fn loopback_read_ahead_past_hello_survives_into_established_conn() {
maybe_child(ROLES);
// The buffer trap, proven: the dialer coalesces Hello + Heartbeat before
// the responder's first read, so the Heartbeat lands in the shared
// FramedConn's decode buffer during the handshake. The dialer sends
// nothing afterwards — the post-handshake recv can only succeed if the
// read-ahead travelled with the FramedConn.
let (mut dialer, mut accepted) = loopback_pair();
let (_init, hello) = smarm::cluster::handshake::Initiator::new(&local("node-a"));
dialer.send(&hello).unwrap();
dialer.send(&Frame::Heartbeat).unwrap();
// Both frames are buffered before the responder reads at all.
let (tx, rx) = mpsc::channel();
std::thread::spawn(move || {
let peer = accept_handshake(
&mut accepted,
local("node-b"),
|_| PeerStanding::Free,
no_deadline(),
)
.unwrap();
let next = accepted.recv();
let _ = tx.send((peer, next));
});
// A bounded wait: if the Heartbeat were NOT carried in the buffer, the
// recv above would block forever (the dialer stays open and silent).
let (peer, next) = rx
.recv_timeout(Duration::from_secs(5))
.expect("read-ahead lost: post-handshake recv blocked");
assert_eq!(peer.node_name, "node-a");
match next {
Ok(Some(Frame::Heartbeat)) => {}
other => panic!("expected the read-ahead Heartbeat, got {other:?}"),
}
drop(dialer);
}
// ---------------------------------------------------------------------------
// Deadline + manager integration, over TCP inside the runtime
// ---------------------------------------------------------------------------
#[test]
fn tcp_silent_peer_times_out_on_the_accept_path() {
maybe_child(ROLES);
run(|| {
let t = TcpTransport;
let mut l = t.listen("127.0.0.1:0").unwrap();
// Connect and then say nothing at all.
let silent = t.dial(&l.local_addr()).unwrap();
let mut accepted = FramedConn::new(l.accept().unwrap());
let (tx, rx) = mpsc::channel();
smarm::spawn(move || {
let r = accept_handshake(
&mut accepted,
local("node-b"),
|_| PeerStanding::Free,
Instant::now() + Duration::from_millis(200),
);
let _ = tx.send(r);
});
match poll_recv(&rx, "accept-path outcome") {
Err(HandshakeError::TimedOut) => {}
other => panic!("expected TimedOut, got {other:?}"),
}
drop(silent);
});
}
/// Poll the manager until its peer set matches `expected` (sorted), or fail.
fn wait_peers(expected: &[&str]) {
let want: Vec<String> = expected.iter().map(|s| s.to_string()).collect();
for _ in 0..5000 {
if let Ok(Reply::Peers(got)) = gen_server::call(MANAGER, Call::Peers) {
if got == want {
return;
}
}
sleep(Duration::from_millis(1));
}
let got = gen_server::call(MANAGER, Call::Peers);
panic!("timed out waiting for peers == {want:?}; last = {got:?}");
}
#[test]
fn tcp_duplicate_name_rejected_by_acceptor() {
maybe_child(ROLES);
run(|| {
let mgr = GenServerBuilder::new(Manager::new())
.named(MANAGER)
.start()
.expect("manager name is free");
let listener = TcpTransport.listen("127.0.0.1:0").unwrap();
let acceptor = spawn_acceptor(listener, local("node-b"), Timing::default());
let addr = acceptor.local_addr().to_string();
// First dial offering "dup-node": establishes and registers.
let mut first = FramedConn::new(TcpTransport.dial(&addr).unwrap());
let peer = dial_handshake(
&mut first,
&local("dup-node"),
Instant::now() + HANDSHAKE_TIMEOUT,
)
.unwrap();
assert_eq!(peer.node_name, "node-b");
wait_peers(&["dup-node"]);
// Second dial offering the same name: deterministic NameTaken.
let mut second = FramedConn::new(TcpTransport.dial(&addr).unwrap());
match dial_handshake(
&mut second,
&local("dup-node"),
Instant::now() + HANDSHAKE_TIMEOUT,
) {
Err(HandshakeError::Rejected(RejectReason::NameTaken)) => {}
other => panic!("expected NameTaken, got {other:?}"),
}
// The established connection was untouched by the rejected one.
wait_peers(&["dup-node"]);
// Teardown: the acceptor owns no connections, so the established one
// is torn down through the table.
acceptor.shutdown();
assert!(matches!(
gen_server::call(
MANAGER,
Call::Disconnect {
name: "dup-node".to_string()
}
),
Ok(Reply::Disconnected)
));
wait_peers(&[]);
first.close();
mgr.shutdown();
});
}
#[test]
fn dial_intent_cleared_when_dialer_dies() {
maybe_child(ROLES);
run(|| {
let mgr = GenServerBuilder::new(Manager::new())
.named(MANAGER)
.start()
.expect("manager name is free");
let (begun_tx, begun_rx) = mpsc::channel();
let (go_tx, go_rx) = mpsc::channel::<()>();
smarm::spawn(move || {
let me = smarm::self_pid();
match gen_server::call(
MANAGER,
Call::DialBegin {
name: "ghost".into(),
pid: me,
},
) {
Ok(Reply::DialBegan(true)) => {}
other => panic!("DialBegin failed: {other:?}"),
}
let _ = begun_tx.send(());
let () = poll_recv(&go_rx, "go signal");
panic!("dialer dies mid-dial");
});
poll_recv(&begun_rx, "DialBegin done");
// While the dialer lives, the intent is visible.
match gen_server::call(
MANAGER,
Call::Standing {
peer_name: "ghost".into(),
},
) {
Ok(Reply::Standing(s)) => assert_eq!(s, PeerStanding::Dialing),
other => panic!("PeerStanding failed: {other:?}"),
}
// Kill it; the monitor must clear the intent without cooperation.
go_tx.send(()).unwrap();
let deadline = Instant::now() + WAIT;
loop {
match gen_server::call(
MANAGER,
Call::Standing {
peer_name: "ghost".into(),
},
) {
Ok(Reply::Standing(s)) if s != PeerStanding::Dialing => break,
_ if Instant::now() > deadline => {
panic!("dial intent not cleared after dialer death")
}
_ => sleep(Duration::from_millis(1)),
}
}
mgr.shutdown();
});
}
// ---------------------------------------------------------------------------
// Two nodes, two processes: the integrated dial against a real acceptor
// ---------------------------------------------------------------------------
/// Announce, then park forever. Neither role ever tears its connection
/// down: a table entry only exists while the *peer* holds its side open, so
/// any teardown here would retract the other node's observation before it
/// had made it. The parent reaps both with SIGKILL once it has both
/// announcements (see [`common::Node`]'s `Drop`).
fn role_hs_listener() {
run(|| {
let _mgr = GenServerBuilder::new(Manager::new())
.named(MANAGER)
.start()
.expect("manager name is free");
let listener = TcpTransport.listen("127.0.0.1:0").unwrap();
let acceptor = spawn_acceptor(listener, local("node-b"), Timing::default());
println!("LISTENING {}", acceptor.local_addr());
wait_peers(&["node-a"]);
println!("PEERS node-a");
park();
});
}
fn role_hs_dialer() {
let addr = std::env::var("SMARM_PEER_ADDR").expect("SMARM_PEER_ADDR not set");
run(move || {
let _mgr = GenServerBuilder::new(Manager::new())
.named(MANAGER)
.start()
.expect("manager name is free");
let (tx, rx) = mpsc::channel();
smarm::spawn(move || {
let r = dial(
&TcpTransport,
&addr,
"node-b",
&local("node-a"),
Timing::default(),
);
let _ = tx.send(r);
});
if let Err(e) = poll_recv(&rx, "dial outcome") {
println!("DIAL failed: {e:?}");
std::process::exit(3);
}
wait_peers(&["node-b"]);
println!("PEERS node-b");
park();
});
}
#[test]
fn two_node_integrated_handshake_over_tcp() {
maybe_child(ROLES);
let mut listener = spawn_node("hs_listener", &[]);
let addr = listener.wait_listening();
let mut dialer = spawn_node("hs_dialer", &[("SMARM_PEER_ADDR", &addr)]);
// Each node reports its own table naming the other: a real dial against a
// real acceptor established in both directions. Both nodes then park —
// clean-exit behaviour is the c4 harness's own smoke test, and demanding
// it here would mean a teardown, which is exactly what cannot be ordered
// safely across two processes. Dropping the nodes SIGKILLs them.
dialer.wait_line("PEERS node-b", |l| l == "PEERS node-b");
listener.wait_line("PEERS node-a", |l| l == "PEERS node-a");
}
// ---------------------------------------------------------------------------
// Integrated-dial guardrails (no acceptor involved)
// ---------------------------------------------------------------------------
#[test]
fn concurrent_dial_to_same_name_refused() {
maybe_child(ROLES);
run(|| {
let mgr = GenServerBuilder::new(Manager::new())
.named(MANAGER)
.start()
.expect("manager name is free");
let (begun_tx, begun_rx) = mpsc::channel();
let (go_tx, go_rx) = mpsc::channel::<()>();
// First dialer parks with the intent held (it never connects —
// 'holding the intent' is all this test needs from it).
smarm::spawn(move || {
let me = smarm::self_pid();
assert!(matches!(
gen_server::call(
MANAGER,
Call::DialBegin {
name: "node-x".into(),
pid: me,
}
),
Ok(Reply::DialBegan(true))
));
let _ = begun_tx.send(());
let () = poll_recv(&go_rx, "go signal");
let _ = gen_server::call(
MANAGER,
Call::DialEnd {
name: "node-x".into(),
},
);
});
poll_recv(&begun_rx, "DialBegin done");
// Second integrated dial to the same name: refused before connecting
// (the addr is unroutable on purpose — it must never be dialed).
let (tx, rx) = mpsc::channel();
smarm::spawn(move || {
let r = dial(
&TcpTransport,
"127.0.0.1:1",
"node-x",
&local("node-a"),
Timing::default(),
);
let _ = tx.send(r);
});
match poll_recv(&rx, "second dial outcome") {
Err(DialError::AlreadyDialing) => {}
other => panic!("expected AlreadyDialing, got {other:?}"),
}
go_tx.send(()).unwrap();
mgr.shutdown();
});
}
+115
View File
@@ -0,0 +1,115 @@
//! RFC 010 — a seed whose address answers as a *different* name
//! (`DialError::PeerNameMismatch`) is dialed once and then parked: the
//! connector must not redial it on backoff forever.
//!
//! Observed from the misdialed peer: each such dial establishes at the
//! responder (it registers, `node_up`), then the dialer closes on the name
//! check (`node_down`) — one membership blip per attempt. Cross-process: a
//! *server* named `server` subscribes and reports; a *client* on fast
//! timing (50–500ms backoff) seeds `("wrongname", server_addr)`. After the
//! first blip the server counts further `NodeUp`s across 2s — several
//! backoff periods. Parked ⇒ zero. Negative-control-verified: with the park
//! stubbed out the count is ≥ 1 in the same window.
#![cfg(feature = "cluster")]
mod common;
use common::{maybe_child, spawn_node};
use smarm::cluster::envelope::NodeMeta;
use smarm::cluster::membership::{subscribe, NodeEvent};
use smarm::cluster::{start, Config, StaticSeeds, Timing};
use std::time::{Duration, Instant};
const ROLES: &[(&str, fn())] = &[("server", role_server), ("client", role_client)];
fn meta() -> NodeMeta {
NodeMeta {
role: "mismatch".into(),
region: "local".into(),
}
}
fn timing() -> Timing {
Timing {
initial_backoff: Duration::from_millis(50),
max_backoff: Duration::from_millis(500),
..Timing::default()
}
}
fn role_server() {
smarm::run(|| {
let cluster = start(Config {
node_name: "server".into(),
meta: meta(),
listen_addr: std::env::var("SMARM_LISTEN_ADDR")
.unwrap_or_else(|_| "127.0.0.1:0".into()),
strategy: Box::new(StaticSeeds::new(Vec::<(String, String)>::new())),
timing: timing(),
})
.expect("binds");
let ev = subscribe().unwrap();
println!("LISTENING {}", cluster.local_addr());
// First blip: the misdialed client establishes, then closes on us.
loop {
match ev.rx.recv() {
Ok(NodeEvent::NodeDown(i)) if i.name == "client" => break,
Ok(_) => continue,
Err(_) => panic!("manager gone"),
}
}
println!("BLIP");
// Now count further NodeUps across several backoff periods.
let mut more = 0usize;
let t0 = Instant::now();
while t0.elapsed() < Duration::from_millis(2000) {
match ev.rx.try_recv() {
Ok(Some(NodeEvent::NodeUp(i))) if i.name == "client" => more += 1,
Ok(_) => {}
Err(_) => panic!("manager gone"),
}
smarm::sleep(Duration::from_millis(50));
}
println!("MORE {more}");
loop {
smarm::sleep(Duration::from_secs(3600));
}
});
}
fn role_client() {
let server_addr = std::env::var("SMARM_SERVER_ADDR").expect("SMARM_SERVER_ADDR");
smarm::run(move || {
let _cluster = start(Config {
node_name: "client".into(),
meta: meta(),
listen_addr: "127.0.0.1:0".into(),
strategy: Box::new(StaticSeeds::new(vec![(
"wrongname".to_string(),
server_addr,
)])),
timing: timing(),
})
.expect("binds");
println!("CLIENT UP");
loop {
smarm::sleep(Duration::from_secs(3600));
}
});
}
#[test]
fn mismatched_seed_is_dialed_once_then_parked() {
maybe_child(ROLES);
let mut server = spawn_node("server", &[]);
let saddr = server.wait_listening();
let mut client = spawn_node("client", &[("SMARM_SERVER_ADDR", &saddr)]);
client.wait_line("CLIENT UP", |l| l == "CLIENT UP");
server.wait_line("BLIP", |l| l == "BLIP");
let line = server.wait_line("MORE", |l| l.starts_with("MORE "));
let more: usize = line.split_whitespace().nth(1).unwrap().parse().unwrap();
assert_eq!(
more, 0,
"mismatched seed was redialed {more}× after being parked"
);
}
+379
View File
@@ -0,0 +1,379 @@
//! RFC 010 c13 — connection-loss synthesis.
//!
//! Local suite (`run()`, no network): the read-side backstop. A
//! `RemoteMonitor` whose channel closes without a notice reads as
//! `Disconnected` exactly once (a `Monitor` command that reached the conn
//! actor's inbox but was never processed — the drain gap); after
//! `demonitor_remote` a closed channel stays a plain `Err`, never a notice.
//!
//! Cross-process: the headline contrast — an actor's own death gives its
//! TRUE reason, loss of the LINK gives `Disconnected` (both a commanded
//! `Disconnect` and a SIGKILLed peer process are `Disconnected` from the
//! monitor's view: nobody is left to say otherwise). Reconnect does not
//! resurrect: the old monitor yields nothing more, proven by stream ORDER
//! (a fresh monitor over the new link delivers first). The ignored test
//! trips liveness by SIGSTOP and then drops the link too, asserting exactly
//! one notice for one monitor.
#![cfg(feature = "cluster")]
mod common;
use common::{maybe_child, spawn_node};
use smarm::cluster::envelope::NodeMeta;
use smarm::cluster::expose::{expose, expose_type};
use smarm::cluster::manager::{Call, Reply, MANAGER};
use smarm::cluster::membership::{subscribe, NodeEvent};
use smarm::cluster::remote::{
self, demonitor_remote, monitor_remote, send_to_remote, RemoteName, RemotePid,
};
use smarm::cluster::{start, Config, RemoteDownReason, StaticSeeds, Timing};
use smarm::pg::Incarnation;
use smarm::{
channel, gen_server, install, register, run, spawn, Addressable, DownReason, Erased, Name, Pid,
};
use std::collections::HashMap;
use std::time::Duration;
// ---- message types (hand-rolled serde; the crate is derive-less) ---------
#[derive(Debug)]
struct Ctl {
cmd: String,
reply_to: RemotePid<Client>,
}
#[derive(Debug)]
struct Answer {
text: String,
pid: Option<RemotePid<Erased>>,
}
struct Client;
impl Addressable for Client {
type Msg = Answer;
}
impl serde::Serialize for Ctl {
fn serialize<S: serde::Serializer>(&self, s: S) -> Result<S::Ok, S::Error> {
use serde::ser::SerializeTuple;
let mut t = s.serialize_tuple(2)?;
t.serialize_element(&self.cmd)?;
t.serialize_element(&self.reply_to)?;
t.end()
}
}
impl<'de> serde::Deserialize<'de> for Ctl {
fn deserialize<D: serde::Deserializer<'de>>(d: D) -> Result<Self, D::Error> {
let (cmd, reply_to) = <(String, RemotePid<Client>)>::deserialize(d)?;
Ok(Ctl { cmd, reply_to })
}
}
impl serde::Serialize for Answer {
fn serialize<S: serde::Serializer>(&self, s: S) -> Result<S::Ok, S::Error> {
use serde::ser::SerializeTuple;
let mut t = s.serialize_tuple(2)?;
t.serialize_element(&self.text)?;
t.serialize_element(&self.pid)?;
t.end()
}
}
impl<'de> serde::Deserialize<'de> for Answer {
fn deserialize<D: serde::Deserializer<'de>>(d: D) -> Result<Self, D::Error> {
let (text, pid) = <(String, Option<RemotePid<Erased>>)>::deserialize(d)?;
Ok(Answer { text, pid })
}
}
// ================= local suite =========================================
/// A `Monitor` command handed to the connection but never processed (its
/// receiver dropped unread) reads as `Disconnected` — once. A second read
/// is the ordinary closed-channel `Err`, so "exactly one notice" holds.
#[test]
fn unread_command_reads_as_disconnected_once() {
maybe_child(ROLES);
run(|| {
remote::set_local_identity("me", Incarnation::new(7));
let (probe_tx, _probe_rx) = channel();
let inbox =
remote::bind_outbound_probe_with_monitors("peer", Incarnation::new(5), probe_tx);
let target = RemotePid::<Erased>::from_parts("peer", Incarnation::new(5), 9, 1);
let m = monitor_remote(target.clone());
assert!(
matches!(m.try_recv(), Ok(None)),
"command is in flight, no notice yet"
);
drop(inbox); // the conn actor died with the command unread
let d = m.recv().unwrap();
assert_eq!(d.pid, target);
assert_eq!(d.reason, RemoteDownReason::Disconnected);
assert!(
m.recv().is_err(),
"second read is closed, not a second notice"
);
assert!(m.try_recv().is_err());
});
}
/// After `demonitor_remote`, a closed channel is a closed channel: no
/// notice is synthesized for a monitor the caller cancelled.
#[test]
fn cancelled_monitor_never_synthesizes() {
maybe_child(ROLES);
run(|| {
remote::set_local_identity("me", Incarnation::new(7));
let (probe_tx, _probe_rx) = channel();
let inbox =
remote::bind_outbound_probe_with_monitors("peer", Incarnation::new(5), probe_tx);
let target = RemotePid::<Erased>::from_parts("peer", Incarnation::new(5), 9, 1);
let m = monitor_remote(target);
demonitor_remote(&m);
drop(inbox);
assert!(m.recv().is_err());
assert!(m.try_recv().is_err());
});
}
// ================= cross-process ======================================
const ROLES: &[(&str, fn())] = &[
("server", role_server),
("client", role_client),
("client_stop", role_client_stop),
];
const CTL: Name<Ctl> = Name::new("c13.ctl");
fn cfg(name: &str, seeds: Vec<(String, String)>) -> Config {
Config {
node_name: name.to_string(),
meta: NodeMeta {
role: "c13".into(),
region: "local".into(),
},
listen_addr: std::env::var("SMARM_LISTEN_ADDR").unwrap_or_else(|_| "127.0.0.1:0".into()),
strategy: Box::new(StaticSeeds::new(seeds)),
timing: timing(),
}
}
/// The p11 knobs make the liveness test fast: both roles of that test are
/// spawned with `SMARM_FAST_TIMING=1` and agree on a 100ms heartbeat /
/// 500ms liveness window. Everything else runs the shipping defaults.
fn timing() -> Timing {
if std::env::var_os("SMARM_FAST_TIMING").is_some() {
Timing {
heartbeat_interval: Duration::from_millis(100),
liveness_timeout: Duration::from_millis(500),
initial_backoff: Duration::from_millis(50),
max_backoff: Duration::from_millis(500),
..Timing::default()
}
} else {
Timing::default()
}
}
fn wait_up(events: &smarm::cluster::membership::MembershipEvents, who: &str) {
loop {
match events.rx.recv() {
Ok(NodeEvent::NodeUp(i)) if i.name == who => return,
Ok(_) => continue,
Err(_) => panic!("manager gone"),
}
}
}
fn disconnect(name: &str) {
assert!(matches!(
gen_server::call(
MANAGER,
Call::Disconnect {
name: name.to_string()
}
),
Ok(Reply::Disconnected)
));
}
/// Server: `spawn` ⇒ a parked worker (answer carries its pid);
/// `kill:<index>` releases it, whereupon it returns (Exit).
fn role_server() {
smarm::run(move || {
let cluster = start(cfg("server", vec![])).expect("binds");
println!("LISTENING {}", cluster.local_addr());
let (tx, rx) = channel::<Ctl>();
register(CTL, tx).unwrap();
expose(CTL);
println!("READY");
let mut workers: HashMap<u32, smarm::channel::Sender<()>> = HashMap::new();
loop {
let ctl = rx.recv().unwrap();
println!("CTL {}", ctl.cmd);
let (text, pid): (String, Option<RemotePid<Erased>>) = match ctl.cmd.as_str() {
"spawn" => {
let (go_tx, go_rx) = channel::<()>();
let p: Pid = spawn(move || {
let _ = go_rx.recv();
})
.pid();
workers.insert(p.index(), go_tx);
(
"ok".into(),
Some(RemotePid::from_local(p).expect("identity set")),
)
}
other => {
let idx: u32 = other.strip_prefix("kill:").unwrap().parse().unwrap();
if let Some(go) = workers.remove(&idx) {
let _ = go.send(());
}
("killed".into(), None)
}
};
send_to_remote(ctl.reply_to, Answer { text, pid }).unwrap();
}
});
}
/// Client-side setup shared by both client roles: join, expose the reply
/// path, hand back an `ask` closure and the membership stream.
fn client_setup() -> (
smarm::cluster::Cluster,
smarm::cluster::membership::MembershipEvents,
impl Fn(&str) -> Answer,
) {
let server_addr = std::env::var("SMARM_SERVER_ADDR").expect("SMARM_SERVER_ADDR");
let cluster = start(cfg("client", vec![("server".into(), server_addr)])).expect("binds");
let ev = subscribe().unwrap();
wait_up(&ev, "server");
let (tx, rx) = channel::<Answer>();
let me: Pid<Client> = install::<Client>(tx);
expose_type::<Answer>();
let ask = move |cmd: &str| -> Answer {
remote::send(
RemoteName::new("server", CTL),
Ctl {
cmd: cmd.into(),
reply_to: RemotePid::from_local(me).expect("identity set"),
},
)
.unwrap();
rx.recv().unwrap()
};
(cluster, ev, ask)
}
fn role_client() {
smarm::run(move || {
let (_cluster, ev, ask) = client_setup();
// 1. Headline: actor death ⇒ TRUE reason; link cut ⇒ Disconnected.
let a = ask("spawn").pid.unwrap();
let b = ask("spawn").pid.unwrap();
let ma = monitor_remote(a.clone());
let mb = monitor_remote(b.clone());
ask(&format!("kill:{}", a.index()));
let d = ma.recv().unwrap();
assert_eq!(d.pid, a);
println!("DOWN actor {:?}", d.reason);
disconnect("server");
let d = mb.recv().unwrap();
assert_eq!(d.pid, b);
println!("DOWN link {:?}", d.reason);
// 2. Reconnect does not resurrect. The connector redials on
// node_down; over the NEW link a fresh monitor delivers, while
// the old one (already answered) yields nothing further — order
// proves it, and `b` is even still alive on the server.
wait_up(&ev, "server");
println!("RECONNECTED");
let c = ask("spawn").pid.unwrap();
let mc = monitor_remote(c.clone());
ask(&format!("kill:{}", b.index()));
ask(&format!("kill:{}", c.index()));
assert_eq!(mc.recv().unwrap().reason, DownReason::Exit.into());
let stray = matches!(mb.try_recv(), Ok(Some(_)));
println!("RESURRECT stray={stray}");
// 3. Peer PROCESS killed ⇒ Disconnected too (nobody is left to send
// Down): the parent SIGKILLs the server once it sees the marker.
let e = ask("spawn").pid.unwrap();
let me_ = monitor_remote(e.clone());
println!("KILL SERVER NOW");
let d = me_.recv().unwrap();
assert_eq!(d.pid, e);
println!("DOWN procdeath {:?}", d.reason);
println!("CLIENT DONE");
loop {
smarm::sleep(Duration::from_secs(3600));
}
});
}
/// The slow role: liveness expiry (peer SIGSTOPped) followed by the link
/// dropping for real (peer SIGKILLed) — one monitor, exactly one notice.
fn role_client_stop() {
smarm::run(move || {
let (_cluster, _ev, ask) = client_setup();
let a = ask("spawn").pid.unwrap();
let ma = monitor_remote(a.clone());
println!("STOP SERVER NOW");
let d = ma.recv().unwrap(); // liveness expiry, ~liveness_timeout
assert_eq!(d.pid, a);
println!("DOWN stopped {:?}", d.reason);
println!("KILL SERVER NOW");
// Give the drop every chance to produce a second notice, then look.
smarm::sleep(Duration::from_secs(1));
let dup = matches!(ma.try_recv(), Ok(Some(_)));
println!("DUPLICATE dup={dup}");
println!("CLIENT DONE");
loop {
smarm::sleep(Duration::from_secs(3600));
}
});
}
/// The Phase 4 c13 gate: partition vs. death distinguishable; nothing
/// survives reconnect; a dead peer process is a Disconnected too.
#[test]
fn link_loss_is_disconnected_and_does_not_survive_reconnect() {
maybe_child(ROLES);
let mut server = spawn_node("server", &[]);
let saddr = server.wait_listening();
server.wait_line("READY", |l| l == "READY");
let mut client = spawn_node("client", &[("SMARM_SERVER_ADDR", &saddr)]);
client.wait_line("DOWN actor Local(Exit)", |l| l == "DOWN actor Local(Exit)");
client.wait_line("DOWN link Disconnected", |l| l == "DOWN link Disconnected");
client.wait_line("RECONNECTED", |l| l == "RECONNECTED");
client.wait_line("RESURRECT stray=false", |l| l == "RESURRECT stray=false");
client.wait_line("KILL SERVER NOW", |l| l == "KILL SERVER NOW");
server.kill();
client.wait_line("DOWN procdeath Disconnected", |l| {
l == "DOWN procdeath Disconnected"
});
client.wait_line("CLIENT DONE", |l| l == "CLIENT DONE");
}
/// Covers the invariant the headline test cannot: liveness expiry and the
/// transport drop both firing for the same connection yield ONE notice.
/// Runs on the fast [`timing`] (both roles) — was `#[ignore]`d at the 4s
/// default until the p11 knobs landed.
#[test]
fn timeout_then_drop_yields_one_notice() {
maybe_child(ROLES);
let fast = ("SMARM_FAST_TIMING", "1");
let mut server = spawn_node("server", &[fast]);
let saddr = server.wait_listening();
server.wait_line("READY", |l| l == "READY");
let mut client = spawn_node("client_stop", &[("SMARM_SERVER_ADDR", &saddr), fast]);
client.wait_line("STOP SERVER NOW", |l| l == "STOP SERVER NOW");
let spid = server.pid().expect("server alive") as libc::pid_t;
assert_eq!(unsafe { libc::kill(spid, libc::SIGSTOP) }, 0);
client.wait_line("DOWN stopped Disconnected", |l| {
l == "DOWN stopped Disconnected"
});
client.wait_line("KILL SERVER NOW", |l| l == "KILL SERVER NOW");
server.kill(); // SIGKILL works on a stopped process; Drop would too
client.wait_line("DUPLICATE dup=false", |l| l == "DUPLICATE dup=false");
client.wait_line("CLIENT DONE", |l| l == "CLIENT DONE");
}
+161
View File
@@ -0,0 +1,161 @@
//! RFC 010 — `Discovery::Withdrawn`: a strategy retracts a candidate and the
//! connector stops dialing it.
//!
//! Cross-process: a plain *server* node, and a *client* whose strategy is a
//! script: announce a decoy `(ghost, addr)` where `addr` is a raw
//! `TcpListener` the client itself holds (an OS thread accepts and
//! immediately closes, so every dial fails at handshake and the connector
//! keeps retrying on backoff — the accept count is the dial count); after a
//! beat, withdraw the decoy and announce the real server. The client waits
//! for the server's `node_up` — which is *after* the withdrawal in the
//! strategy's own stream — then watches the decoy's accept count stay flat
//! across a window longer than the pending backoff. Before withdrawal it
//! must have been climbing (≥ 1), or the negative proves nothing.
#![cfg(feature = "cluster")]
mod common;
use common::{maybe_child, spawn_node};
use smarm::channel::Sender;
use smarm::cluster::discovery::{Discovery, Strategy};
use smarm::cluster::envelope::NodeMeta;
use smarm::cluster::membership::{subscribe, NodeEvent};
use smarm::cluster::{start, Config, StaticSeeds, Timing};
use std::net::TcpListener;
use std::sync::atomic::{AtomicUsize, Ordering};
use std::sync::Arc;
use std::time::{Duration, Instant};
const ROLES: &[(&str, fn())] = &[("server", role_server), ("client", role_client)];
fn meta() -> NodeMeta {
NodeMeta {
role: "withdraw".into(),
region: "local".into(),
}
}
fn role_server() {
smarm::run(|| {
let cluster = start(Config {
node_name: "server".into(),
meta: meta(),
listen_addr: std::env::var("SMARM_LISTEN_ADDR")
.unwrap_or_else(|_| "127.0.0.1:0".into()),
strategy: Box::new(StaticSeeds::new(Vec::<(String, String)>::new())),
timing: Timing::default(),
})
.expect("binds");
println!("LISTENING {}", cluster.local_addr());
loop {
smarm::sleep(Duration::from_secs(3600));
}
});
}
/// Scripted strategy: decoy, pause, withdraw decoy, real server, done.
struct Script {
decoy: String,
server: String,
}
impl Strategy for Script {
fn run(self: Box<Self>, out: Sender<Discovery>) {
let _ = out.send(Discovery::Candidate {
name: "ghost".into(),
addr: self.decoy.clone(),
});
// Long enough for the 250ms/500ms retries to land: ≥ 3 dials.
smarm::sleep(Duration::from_millis(1100));
let _ = out.send(Discovery::Withdrawn {
name: "ghost".into(),
addr: self.decoy,
});
let _ = out.send(Discovery::Candidate {
name: "server".into(),
addr: self.server,
});
}
}
fn role_client() {
let server_addr = std::env::var("SMARM_SERVER_ADDR").expect("SMARM_SERVER_ADDR");
// The decoy: accept-and-close on an OS thread; count every accept.
let decoy = TcpListener::bind("127.0.0.1:0").unwrap();
let decoy_addr = decoy.local_addr().unwrap().to_string();
let dials = Arc::new(AtomicUsize::new(0));
let counter = dials.clone();
std::thread::spawn(move || {
for conn in decoy.incoming() {
counter.fetch_add(1, Ordering::SeqCst);
drop(conn);
}
});
smarm::run(move || {
let _cluster = start(Config {
node_name: "client".into(),
meta: meta(),
listen_addr: "127.0.0.1:0".into(),
strategy: Box::new(Script {
decoy: decoy_addr,
server: server_addr,
}),
timing: Timing::default(),
})
.expect("binds");
let ev = subscribe().unwrap();
loop {
match ev.rx.recv() {
Ok(NodeEvent::NodeUp(i)) if i.name == "server" => break,
Ok(_) => continue,
Err(_) => panic!("manager gone"),
}
}
// The withdrawal preceded the server candidate in the strategy's
// stream, so it has been applied. Any dial that started before it
// is bounded by the connect+handshake deadlines; let it drain, then
// hold the count flat across a window longer than the pending
// backoff would be (1s at this point, 2s next).
let before = dials.load(Ordering::SeqCst);
smarm::sleep(Duration::from_millis(500));
let settled = dials.load(Ordering::SeqCst);
let t0 = Instant::now();
while t0.elapsed() < Duration::from_millis(3000) {
smarm::sleep(Duration::from_millis(100));
}
let after = dials.load(Ordering::SeqCst);
println!("WITHDRAWN before={before} settled={settled} after={after}");
println!("CLIENT DONE");
loop {
smarm::sleep(Duration::from_secs(3600));
}
});
}
#[test]
fn withdrawn_candidate_is_no_longer_dialed() {
maybe_child(ROLES);
let mut server = spawn_node("server", &[]);
let saddr = server.wait_listening();
let mut client = spawn_node("client", &[("SMARM_SERVER_ADDR", &saddr)]);
let line = client.wait_line("WITHDRAWN", |l| l.starts_with("WITHDRAWN "));
let mut nums = line
.split_whitespace()
.skip(1)
.map(|kv| kv.split_once('=').unwrap().1.parse::<usize>().unwrap());
let (before, settled, after) = (
nums.next().unwrap(),
nums.next().unwrap(),
nums.next().unwrap(),
);
assert!(
before >= 1,
"decoy was never dialed; the negative proves nothing: {line}"
);
assert_eq!(
settled, after,
"connector kept dialing a withdrawn candidate: {line}"
);
client.wait_line("CLIENT DONE", |l| l == "CLIENT DONE");
}
+276
View File
@@ -0,0 +1,276 @@
//! RFC 010 c2 — owned envelope tests (roadmap: per-frame roundtrip,
//! truncation mid-field, unknown tag, length prefix lying long and short,
//! zero-length payload, adversarial lengths).
#![cfg(feature = "cluster")]
use serde::{Deserialize, Serialize};
use smarm::cluster::envelope::{
decode_payload, encode_payload, DecodeError, Frame, NodeMeta, RejectReason, MAX_FRAME_LEN,
PROTO_VERSION,
};
use smarm::cluster::RemoteDownReason;
use smarm::monitor::DownReason;
use smarm::pg::Incarnation;
fn meta() -> NodeMeta {
NodeMeta {
role: "worker".into(),
region: "eu-west".into(),
}
}
fn all_frames() -> Vec<Frame> {
vec![
Frame::Hello {
proto_version: PROTO_VERSION,
build_hash: 0xDEAD_BEEF_CAFE_F00D,
node_name: "alpha".into(),
incarnation: Incarnation::new(7),
meta: meta(),
},
Frame::HelloAck {
node_name: "beta".into(),
incarnation: Incarnation::new(9),
meta: meta(),
},
Frame::HelloReject {
reason: RejectReason::NameTaken,
},
Frame::Heartbeat,
Frame::Send {
index: 42,
generation: 3,
type_hash: 0x1234_5678_9ABC_DEF0,
payload: vec![1, 2, 3, 4, 5],
},
Frame::SendNamed {
name: "the_counter".into(),
type_hash: 0xFFFF_0000_FFFF_0000,
payload: vec![],
},
Frame::Monitor {
monitor_id: 77,
index: 42,
generation: 3,
},
Frame::Demonitor { monitor_id: 77 },
Frame::Down {
monitor_id: 77,
reason: RemoteDownReason::Local(DownReason::Panic),
},
Frame::Down {
monitor_id: 78,
reason: RemoteDownReason::Disconnected,
},
]
}
fn encode_one(f: &Frame) -> Vec<u8> {
let mut buf = Vec::new();
f.encode(&mut buf).unwrap();
buf
}
#[test]
fn per_frame_roundtrip() {
for f in all_frames() {
let buf = encode_one(&f);
let (decoded, consumed) = Frame::decode(&buf).unwrap().unwrap();
assert_eq!(decoded, f, "roundtrip mismatch");
assert_eq!(consumed, buf.len(), "consumed != buffer length for {f:?}");
}
}
#[test]
fn back_to_back_frames_decode_sequentially() {
let mut buf = Vec::new();
for f in all_frames() {
f.encode(&mut buf).unwrap();
}
let mut off = 0;
let mut decoded = Vec::new();
while off < buf.len() {
let (f, n) = Frame::decode(&buf[off..]).unwrap().unwrap();
decoded.push(f);
off += n;
}
assert_eq!(decoded, all_frames());
assert_eq!(off, buf.len());
}
#[test]
fn heartbeat_golden_bytes() {
// Locks the layout: u32 LE length prefix, then the tag byte.
let buf = encode_one(&Frame::Heartbeat);
assert_eq!(buf, vec![1, 0, 0, 0, 4]);
}
#[test]
fn zero_length_payload_roundtrips() {
let f = Frame::Send {
index: 0,
generation: 0,
type_hash: 0,
payload: vec![],
};
let buf = encode_one(&f);
let (decoded, consumed) = Frame::decode(&buf).unwrap().unwrap();
assert_eq!(decoded, f);
assert_eq!(consumed, buf.len());
}
#[test]
fn incomplete_is_none_not_error() {
let buf = encode_one(&all_frames()[0]);
// Every strict prefix short of the full frame must report "need more".
for cut in 0..buf.len() {
assert_eq!(
Frame::decode(&buf[..cut]).unwrap(),
None,
"cut at {cut} should be incomplete"
);
}
}
#[test]
fn unknown_frame_tag() {
let buf = vec![1, 0, 0, 0, 250];
assert_eq!(Frame::decode(&buf), Err(DecodeError::UnknownTag(250)));
}
#[test]
fn unknown_enum_tags() {
// HelloReject with a bogus reason tag.
let buf = vec![2, 0, 0, 0, 3, 99];
assert_eq!(
Frame::decode(&buf),
Err(DecodeError::UnknownEnumTag {
what: "RejectReason",
tag: 99
})
);
// Down with a bogus reason tag (id = 0u64).
let mut buf = vec![10, 0, 0, 0, 9];
buf.extend_from_slice(&0u64.to_le_bytes());
buf.push(200);
assert_eq!(
Frame::decode(&buf),
Err(DecodeError::UnknownEnumTag {
what: "RemoteDownReason",
tag: 200
})
);
}
#[test]
fn length_prefix_lying_long_with_bytes_present_is_trailing() {
let mut buf = encode_one(&Frame::Heartbeat);
// Declare 3 extra body bytes and actually supply them.
let declared = u32::from_le_bytes([buf[0], buf[1], buf[2], buf[3]]) + 3;
buf[0..4].copy_from_slice(&declared.to_le_bytes());
buf.extend_from_slice(&[0xAA, 0xBB, 0xCC]);
assert_eq!(Frame::decode(&buf), Err(DecodeError::Trailing { extra: 3 }));
}
#[test]
fn length_prefix_lying_long_without_bytes_is_incomplete() {
// Indistinguishable from a partial read — must be None, not an error.
let mut buf = encode_one(&Frame::Heartbeat);
let declared = u32::from_le_bytes([buf[0], buf[1], buf[2], buf[3]]) + 3;
buf[0..4].copy_from_slice(&declared.to_le_bytes());
assert_eq!(Frame::decode(&buf).unwrap(), None);
}
#[test]
fn length_prefix_lying_short_truncates_a_field() {
let f = &all_frames()[0]; // Hello: plenty of fields to cut into
let mut buf = encode_one(f);
let declared = u32::from_le_bytes([buf[0], buf[1], buf[2], buf[3]]);
let lie = declared - 4; // cut mid-field
buf[0..4].copy_from_slice(&lie.to_le_bytes());
assert_eq!(Frame::decode(&buf), Err(DecodeError::Truncated));
}
#[test]
fn truncation_mid_string_field() {
// A frame whose declared length is intact but whose inner string length
// runs past the body: SendNamed claiming a 1000-byte name in a tiny body.
let mut body = vec![6u8]; // TAG_SEND_NAMED
body.extend_from_slice(&1000u16.to_le_bytes());
body.extend_from_slice(b"short");
let mut buf = (body.len() as u32).to_le_bytes().to_vec();
buf.extend_from_slice(&body);
assert_eq!(Frame::decode(&buf), Err(DecodeError::Truncated));
}
#[test]
fn adversarial_lengths() {
// Length prefix of u32::MAX: reject as oversized, do not wait for 4 GiB.
let buf = [0xFF, 0xFF, 0xFF, 0xFF, 0];
assert_eq!(
Frame::decode(&buf),
Err(DecodeError::FrameTooLarge {
declared: u32::MAX as usize
})
);
// Just over the cap: also rejected.
let over = (MAX_FRAME_LEN as u32 + 1).to_le_bytes();
assert!(matches!(
Frame::decode(&over),
Err(DecodeError::FrameTooLarge { .. })
));
// Zero-length frame: there is no tag byte; corrupt, not incomplete.
let buf = [0, 0, 0, 0];
assert_eq!(Frame::decode(&buf), Err(DecodeError::EmptyFrame));
}
#[test]
fn invalid_utf8_in_string_field() {
let mut buf = encode_one(&Frame::SendNamed {
name: "abcd".into(),
type_hash: 0,
payload: vec![],
});
// name bytes start after: 4 (len) + 1 (tag) + 2 (str len) = offset 7
buf[7] = 0xFF;
assert_eq!(Frame::decode(&buf), Err(DecodeError::Utf8));
}
#[derive(Debug, PartialEq, Serialize, Deserialize)]
struct Ping {
seq: u64,
label: String,
}
#[test]
fn payload_seam_roundtrip() {
let ping = Ping {
seq: 31337,
label: "hello".into(),
};
let blob = encode_payload(&ping).unwrap();
// Carry it through a real frame, as it will travel in c9.
let f = Frame::Send {
index: 1,
generation: 1,
type_hash: 0xABCD,
payload: blob,
};
let buf = encode_one(&f);
let (decoded, _) = Frame::decode(&buf).unwrap().unwrap();
let Frame::Send { payload, .. } = decoded else {
panic!("wrong frame");
};
let back: Ping = decode_payload(&payload).unwrap();
assert_eq!(back, ping);
}
#[test]
fn payload_seam_rejects_truncated_blob() {
let blob = encode_payload(&Ping {
seq: 1,
label: "x".into(),
})
.unwrap();
assert!(decode_payload::<Ping>(&blob[..blob.len() - 1]).is_err());
}
+161
View File
@@ -0,0 +1,161 @@
//! RFC 010 c8 — exposure registry + type hashing. Purely local, no network.
//!
//! Payload types are std types (`String`, `u64`) because the crate's serde is
//! deliberately derive-less (`default-features = false`) — user crates bring
//! their own derive; the contract here is `DeserializeOwned`.
//!
//! The hash-stability test re-execs the current binary (the c4 harness): the
//! guarantee under test is "stable across runs in the SAME binary" — exactly
//! what the build-hash handshake reduces the mesh to — not stability across
//! builds, which the scope guard explicitly rejects.
#![cfg(feature = "cluster")]
mod common;
use common::{maybe_child, spawn_node};
use smarm::cluster::envelope::encode_payload;
use smarm::cluster::expose::{
decode_deliver, decoder_registered, expose, expose_type, exposed_hash, exposed_names,
type_hash, DeliverError,
};
use smarm::monitor::{monitor, terminal_reason, DownReason};
use smarm::{channel, register, run, spawn, Name};
const ROLES: &[(&str, fn())] = &[("hasher", role_hasher)];
/// Print the hashes this process computes; the parent (a different run of
/// the same binary) compares against its own.
fn role_hasher() {
println!("HASH-STRING {}", type_hash::<String>());
println!("HASH-U64 {}", type_hash::<u64>());
}
const GREETER: Name<String> = Name::new("expose-test.greeter");
/// Exposed and unexposed lookup, the returned hash, and the audit listing.
#[test]
fn exposed_and_unexposed_lookup() {
maybe_child(ROLES);
run(|| {
let h = expose(GREETER);
assert_eq!(h, type_hash::<String>());
assert_eq!(exposed_hash("expose-test.greeter"), Some(h));
assert_eq!(exposed_hash("never-exposed"), None);
assert!(exposed_names().contains(&("expose-test.greeter", h)));
});
}
/// Distinct types land on distinct hashes (FNV over distinct TypeIds — a
/// smoke assertion; a collision would degrade to a decode error, never a
/// misroute, per RFC §3).
#[test]
fn distinct_types_distinct_hashes() {
maybe_child(ROLES);
run(|| {
assert_ne!(type_hash::<String>(), type_hash::<u64>());
assert_ne!(type_hash::<String>(), type_hash::<Vec<u8>>());
});
}
/// The decode-and-deliver contract: a registered hash decodes into the
/// target's typed channel; an unknown hash, corrupt bytes, and a missing
/// channel each fail without delivering — `WrongChannel`, never a misroute.
#[test]
fn decoder_registration_and_delivery() {
maybe_child(ROLES);
run(|| {
let h_string = expose_type::<String>();
let h_u64 = expose_type::<u64>();
assert!(decoder_registered(h_string));
assert!(!decoder_registered(h_string.wrapping_add(1)));
// A live actor with a String channel (registered from its own body,
// announced via a ready signal — the tests/registry.rs idiom).
let (ready_tx, ready_rx) = channel::<()>();
let (stop_tx, stop_rx) = channel::<()>();
let (msg_tx, msg_rx) = channel::<String>();
let pid = spawn(move || {
register(Name::<String>::new("expose-test.sink"), msg_tx).unwrap();
ready_tx.send(()).unwrap();
let _ = stop_rx.recv();
})
.pid();
ready_rx.recv().unwrap();
// Happy path: decode + deliver through the published channel.
let bytes = encode_payload("hello across the seam").unwrap();
decode_deliver(h_string, pid, &bytes).unwrap();
assert_eq!(msg_rx.recv().unwrap(), "hello across the seam");
// Unknown hash: nothing was registered under it.
assert!(matches!(
decode_deliver(h_string.wrapping_add(1), pid, &bytes),
Err(DeliverError::UnknownType)
));
// Corrupt bytes: the decoder fails before any send.
assert!(matches!(
decode_deliver(h_string, pid, &[0xff; 3]),
Err(DeliverError::Decode(_))
));
// Right decoder, wrong channel: the actor has no u64 channel, so the
// decoded value is refused — the NoChannel guarantee.
let u64_bytes = encode_payload(&7u64).unwrap();
assert!(matches!(
decode_deliver(h_u64, pid, &u64_bytes),
Err(DeliverError::WrongChannel)
));
stop_tx.send(()).unwrap();
});
}
/// `expose` and the bridge crossing agree on the resulting set: both funnel
/// the pid-boundary mark through the watchable machinery, so an exposed
/// name's holder dies with a terminal record — the exact observable
/// `mark_watchable` guarantees the membrane. (For named holders the mark is
/// already stamped by `register` itself; this pins the shared contract.)
#[test]
fn expose_and_bridge_crossing_agree_on_the_set() {
maybe_child(ROLES);
run(|| {
let (ready_tx, ready_rx) = channel::<()>();
let (stop_tx, stop_rx) = channel::<()>();
let (msg_tx, _msg_rx) = channel::<String>();
let pid = spawn(move || {
register(GREETER, msg_tx).unwrap();
ready_tx.send(()).unwrap();
let _ = stop_rx.recv();
})
.pid();
ready_rx.recv().unwrap();
expose(GREETER);
let m = monitor(pid);
stop_tx.send(()).unwrap();
assert_eq!(m.rx.recv().unwrap().reason, DownReason::Exit);
assert_eq!(terminal_reason(pid), Some(DownReason::Exit));
});
}
/// Hash stability across runs in the same binary: a re-exec of this binary
/// computes the same hashes this process does.
#[test]
fn hash_stable_across_runs_in_same_binary() {
maybe_child(ROLES);
let (mine_string, mine_u64) = {
// Computing a TypeId hash needs no runtime, but keep the contract
// uniform with real call sites.
(type_hash::<String>(), type_hash::<u64>())
};
let mut child = spawn_node("hasher", &[]);
let line = child.wait_line("HASH-STRING", |l| l.starts_with("HASH-STRING "));
assert_eq!(
line["HASH-STRING ".len()..].parse::<u64>().unwrap(),
mine_string
);
let line = child.wait_line("HASH-U64", |l| l.starts_with("HASH-U64 "));
assert_eq!(line["HASH-U64 ".len()..].parse::<u64>().unwrap(), mine_u64);
child.wait_exit();
}
+240
View File
@@ -0,0 +1,240 @@
//! RFC 010 c5 — handshake state-machine tests (roadmap: happy path; hash
//! mismatch; proto-version mismatch; name already claimed; simultaneous-connect
//! tie-break; garbage before Hello). Pure — no IO, no actors, no runtime.
#![cfg(feature = "cluster")]
use smarm::cluster::envelope::{Frame, NodeMeta, RejectReason, PROTO_VERSION};
use smarm::cluster::handshake::{
dial_wins, Initiator, InitiatorOutcome, Local, PeerStanding, Responder, ResponderOutcome,
};
use smarm::pg::Incarnation;
const HASH: u64 = 0xDEAD_BEEF_CAFE_F00D;
fn local(name: &str) -> Local {
Local {
node_name: name.into(),
incarnation: Incarnation::new(7),
build_hash: HASH,
meta: NodeMeta {
role: "worker".into(),
region: "eu-west".into(),
},
}
}
/// The Hello that `Initiator::new(&local(name))` emits, built by hand.
fn hello_from(name: &str) -> Frame {
let l = local(name);
Frame::Hello {
proto_version: PROTO_VERSION,
build_hash: l.build_hash,
node_name: l.node_name,
incarnation: l.incarnation,
meta: l.meta,
}
}
#[test]
fn happy_path_establishes_both_ends() {
// alpha dials beta.
let (initiator, hello) = Initiator::new(&local("alpha"));
assert_eq!(hello, hello_from("alpha"), "initiator emits its identity");
let responder = Responder::new(local("beta"));
let (reply, peer) = match responder.on_frame(hello, PeerStanding::Free) {
ResponderOutcome::Accepted { reply, peer } => (reply, peer),
other => panic!("expected Accepted, got {other:?}"),
};
assert_eq!(peer.node_name, "alpha");
assert_eq!(peer.incarnation, Incarnation::new(7));
assert_eq!(peer.meta.role, "worker");
// The ack carries the responder's identity, no hash/version (one-sided
// check — sound because equality is symmetric).
let l = local("beta");
assert_eq!(
reply,
Frame::HelloAck {
node_name: l.node_name,
incarnation: l.incarnation,
meta: l.meta,
}
);
match initiator.on_frame(reply) {
InitiatorOutcome::Established(peer) => {
assert_eq!(peer.node_name, "beta");
assert_eq!(peer.incarnation, Incarnation::new(7));
assert_eq!(peer.meta.region, "eu-west");
}
other => panic!("expected Established, got {other:?}"),
}
}
#[test]
fn hash_mismatch_rejected() {
let responder = Responder::new(local("beta"));
let hello = Frame::Hello {
proto_version: PROTO_VERSION,
build_hash: HASH ^ 1,
node_name: "alpha".into(),
incarnation: Incarnation::new(7),
meta: local("alpha").meta,
};
match responder.on_frame(hello, PeerStanding::Free) {
ResponderOutcome::Rejected { reply, reason } => {
assert_eq!(reason, RejectReason::HashMismatch);
assert_eq!(reply, Frame::HelloReject { reason });
}
other => panic!("expected Rejected, got {other:?}"),
}
// The dialer side of the same story: a reject frame comes back.
let (initiator, _hello) = Initiator::new(&local("alpha"));
match initiator.on_frame(Frame::HelloReject {
reason: RejectReason::HashMismatch,
}) {
InitiatorOutcome::Rejected(RejectReason::HashMismatch) => {}
other => panic!("expected Rejected(HashMismatch), got {other:?}"),
}
}
#[test]
fn proto_version_mismatch_rejected_and_checked_first() {
// Both proto and hash wrong: proto wins — nothing after the version can
// be trusted, and HelloReject is the cross-version compatibility anchor.
let responder = Responder::new(local("beta"));
let hello = Frame::Hello {
proto_version: PROTO_VERSION + 1,
build_hash: HASH ^ 1,
node_name: "alpha".into(),
incarnation: Incarnation::new(7),
meta: local("alpha").meta,
};
match responder.on_frame(hello, PeerStanding::Free) {
ResponderOutcome::Rejected { reason, .. } => {
assert_eq!(reason, RejectReason::ProtoVersion);
}
other => panic!("expected Rejected, got {other:?}"),
}
}
#[test]
fn claimed_name_rejected() {
let responder = Responder::new(local("beta"));
let ctx = PeerStanding::Claimed;
match responder.on_frame(hello_from("alpha"), ctx) {
ResponderOutcome::Rejected { reason, .. } => {
assert_eq!(reason, RejectReason::NameTaken);
}
other => panic!("expected Rejected, got {other:?}"),
}
}
#[test]
fn own_name_offered_rejected_as_name_taken() {
// Self-connect or genuine collision: the responder's own name arrives.
let responder = Responder::new(local("beta"));
match responder.on_frame(hello_from("beta"), PeerStanding::Free) {
ResponderOutcome::Rejected { reason, .. } => {
assert_eq!(reason, RejectReason::NameTaken);
}
other => panic!("expected Rejected, got {other:?}"),
}
}
#[test]
fn hash_checked_before_name() {
// Wrong hash AND claimed name: hash wins (validity before identity).
let responder = Responder::new(local("beta"));
let hello = Frame::Hello {
proto_version: PROTO_VERSION,
build_hash: HASH ^ 1,
node_name: "alpha".into(),
incarnation: Incarnation::new(7),
meta: local("alpha").meta,
};
let ctx = PeerStanding::Claimed;
match responder.on_frame(hello, ctx) {
ResponderOutcome::Rejected { reason, .. } => {
assert_eq!(reason, RejectReason::HashMismatch);
}
other => panic!("expected Rejected, got {other:?}"),
}
}
#[test]
fn dial_wins_is_deterministic_and_antisymmetric() {
// The smaller name's dial survives; both ends compute the same verdict.
assert!(dial_wins("alpha", "beta"));
assert!(!dial_wins("beta", "alpha"));
for (a, b) in [("a", "b"), ("node-1", "node-2"), ("x", "xx")] {
assert_ne!(dial_wins(a, b), dial_wins(b, a), "({a}, {b})");
}
}
#[test]
fn simultaneous_connect_exactly_one_side_accepts() {
// alpha and beta dial each other at once. Each responder sees the peer's
// Hello while its own dial is in flight.
let ctx = PeerStanding::Dialing;
// On beta: inbound is alpha's dial; alpha < beta, so the inbound wins.
let on_beta = Responder::new(local("beta")).on_frame(hello_from("alpha"), ctx);
assert!(
matches!(on_beta, ResponderOutcome::Accepted { .. }),
"beta must accept alpha's dial, got {on_beta:?}"
);
// On alpha: inbound is beta's dial; it loses — close silently, no frame
// (ratified: both ends can compute the outcome, a reject adds nothing).
let on_alpha = Responder::new(local("alpha")).on_frame(hello_from("beta"), ctx);
assert!(
matches!(on_alpha, ResponderOutcome::TieBreakLoss),
"alpha must silently drop beta's dial, got {on_alpha:?}"
);
}
#[test]
fn tiebreak_loss_only_applies_when_dialing() {
// Same inbound Hello, no dial in flight: plain accept.
let on_alpha = Responder::new(local("alpha")).on_frame(hello_from("beta"), PeerStanding::Free);
assert!(matches!(on_alpha, ResponderOutcome::Accepted { .. }));
}
#[test]
fn garbage_before_hello_fails_without_reply() {
// Any valid-but-wrong frame before Hello is a protocol violation: close,
// no reject frame. (Undecodable bytes are the codec's Err, not ours.)
for frame in [
Frame::Heartbeat,
Frame::HelloAck {
node_name: "alpha".into(),
incarnation: Incarnation::new(7),
meta: local("alpha").meta,
},
Frame::Demonitor { monitor_id: 3 },
] {
let out = Responder::new(local("beta")).on_frame(frame.clone(), PeerStanding::Free);
match out {
ResponderOutcome::Failed(f) => assert_eq!(f, frame),
other => panic!("expected Failed({frame:?}), got {other:?}"),
}
}
}
#[test]
fn garbage_before_ack_fails_the_initiator() {
for frame in [
Frame::Heartbeat,
hello_from("beta"),
Frame::Demonitor { monitor_id: 3 },
] {
let (initiator, _hello) = Initiator::new(&local("alpha"));
match initiator.on_frame(frame.clone()) {
InitiatorOutcome::Failed(f) => assert_eq!(f, frame),
other => panic!("expected Failed({frame:?}), got {other:?}"),
}
}
}
+270
View File
@@ -0,0 +1,270 @@
//! RFC 010 c7a — membership events and the view, at the manager.
//!
//! Same construction as the c6a lifecycle suite: the handshake is bypassed,
//! connections are built already-established over localhost TCP pairs with
//! fabricated `Peer`s, and the manager is started plainly so the test can
//! terminate. What is under test is the membership layer that c7 adds to the
//! manager: `node_up`/`node_down` events to subscribers (snapshot-then-stream),
//! the view, and NodeId identity — memoized per `(name, incarnation)`, so a
//! reconnect blip keeps its id and a restart (new incarnation) gets a fresh one.
#![cfg(feature = "cluster")]
use std::time::Duration;
use smarm::cluster::envelope::NodeMeta;
use smarm::cluster::handshake::Peer;
use smarm::cluster::manager::{Call, Manager, Reply, MANAGER};
use smarm::cluster::membership::{subscribe, view, MembershipEvents, NodeEvent};
use smarm::cluster::spawn_established;
use smarm::cluster::transport::tcp::TcpTransport;
use smarm::cluster::transport::{Conn, FramedConn, Transport};
use smarm::cluster::Timing;
use smarm::gen_server::{self, GenServerBuilder};
use smarm::pg::{Incarnation, NodeId};
use smarm::run;
/// A fabricated post-handshake peer identity, with the incarnation under the
/// test's control (it is identity-bearing here, unlike in the c6a suite).
fn peer(name: &str, inc: u32) -> Peer {
Peer {
node_name: name.to_string(),
incarnation: Incarnation::new(inc),
meta: NodeMeta {
role: "test".to_string(),
region: "test".to_string(),
},
}
}
/// One established transport pair over localhost (TCP backlog covers the
/// sequential dial-then-accept, as in the c3 conformance suite).
fn pair(t: &dyn Transport) -> (Box<dyn Conn>, Box<dyn Conn>) {
let mut l = t.listen("127.0.0.1:0").unwrap();
let a = t.dial(&l.local_addr()).unwrap();
let b = l.accept().unwrap();
(a, b)
}
/// The next event, or a panic naming the wait. The bound is generous against
/// a sub-millisecond real cost.
fn next_event(ev: &MembershipEvents, waiting_for: &str) -> NodeEvent {
ev.rx
.recv_timeout(Duration::from_secs(5))
.unwrap_or_else(|e| panic!("timed out waiting for {waiting_for}: {e:?}"))
}
/// Assert the subscription is drained: no event is pending.
fn assert_quiet(ev: &MembershipEvents) {
assert!(matches!(ev.rx.try_recv(), Ok(None)));
}
fn disconnect(name: &str) {
assert!(matches!(
gen_server::call(
MANAGER,
Call::Disconnect {
name: name.to_string()
}
),
Ok(Reply::Disconnected)
));
}
/// Live subscription: an empty snapshot, then `NodeUp` on registration and
/// `NodeDown` (same id) on commanded disconnect and on peer EOF alike.
#[test]
fn subscriber_sees_up_and_down() {
run(|| {
let mgr = GenServerBuilder::new(Manager::new())
.named(MANAGER)
.start()
.expect("manager name is free");
let ev = subscribe().expect("manager is up");
assert_quiet(&ev); // nothing live: the snapshot is empty
let t = TcpTransport;
let (a1, b1) = pair(&t);
let (a2, b2) = pair(&t);
spawn_established(FramedConn::new(a1), peer("node-b", 1), Timing::default())
.expect("node-b registers");
spawn_established(FramedConn::new(a2), peer("node-c", 1), Timing::default())
.expect("node-c registers");
let up_b = match next_event(&ev, "node_up(node-b)") {
NodeEvent::NodeUp(info) => {
assert_eq!(info.name, "node-b");
assert_eq!(info.incarnation, Incarnation::new(1));
assert_eq!(info.meta.role, "test");
info
}
other => panic!("expected node_up(node-b), got {other:?}"),
};
let up_c = match next_event(&ev, "node_up(node-c)") {
NodeEvent::NodeUp(info) => {
assert_eq!(info.name, "node-c");
info
}
other => panic!("expected node_up(node-c), got {other:?}"),
};
assert_ne!(up_b.node, up_c.node, "distinct peers get distinct ids");
// Commanded disconnect: down with node-b's id.
disconnect("node-b");
assert_eq!(
next_event(&ev, "node_down(node-b)"),
NodeEvent::NodeDown(up_b.clone())
);
// Peer EOF, no command: down with node-c's id.
drop(b2);
assert_eq!(
next_event(&ev, "node_down(node-c)"),
NodeEvent::NodeDown(up_c.clone())
);
assert_quiet(&ev);
drop(b1);
mgr.shutdown();
});
}
/// Snapshot-then-stream: a subscriber arriving after connections established
/// receives one `NodeUp` per live peer before anything else, and the view
/// call agrees with it.
#[test]
fn late_subscriber_gets_snapshot_and_view_agrees() {
run(|| {
let mgr = GenServerBuilder::new(Manager::new())
.named(MANAGER)
.start()
.expect("manager name is free");
let t = TcpTransport;
let (a1, b1) = pair(&t);
let (a2, b2) = pair(&t);
spawn_established(FramedConn::new(a1), peer("node-b", 1), Timing::default())
.expect("node-b registers");
spawn_established(FramedConn::new(a2), peer("node-c", 1), Timing::default())
.expect("node-c registers");
let ev = subscribe().expect("manager is up");
let mut names = Vec::new();
for _ in 0..2 {
match next_event(&ev, "a snapshot node_up") {
NodeEvent::NodeUp(info) => names.push(info.name),
other => panic!("expected a snapshot node_up, got {other:?}"),
}
}
names.sort();
assert_eq!(names, ["node-b", "node-c"]);
assert_quiet(&ev); // the snapshot is exactly the live set
let mut v = view().expect("manager is up");
v.sort_by(|a, b| a.name.cmp(&b.name));
assert_eq!(v.len(), 2);
assert_eq!(v[0].name, "node-b");
assert_eq!(v[1].name, "node-c");
disconnect("node-b");
disconnect("node-c");
drop((b1, b2));
// Drain the two downs so the subscription ends quiet.
let _ = next_event(&ev, "node_down");
let _ = next_event(&ev, "node_down");
mgr.shutdown();
});
}
/// NodeId identity: a restart (same name, new incarnation) is a NEW id — the
/// ghost and its successor are distinguishable — while a reconnect blip (same
/// name, same incarnation) keeps its id.
#[test]
fn restart_gets_new_id_blip_keeps_id() {
run(|| {
let mgr = GenServerBuilder::new(Manager::new())
.named(MANAGER)
.start()
.expect("manager name is free");
let ev = subscribe().expect("manager is up");
let t = TcpTransport;
let id = |e: NodeEvent, what: &str| -> NodeId {
match e {
NodeEvent::NodeUp(info) => info.node,
other => panic!("expected node_up ({what}), got {other:?}"),
}
};
let down_id = |e: NodeEvent, what: &str| -> NodeId {
match e {
NodeEvent::NodeDown(info) => info.node,
other => panic!("expected node_down ({what}), got {other:?}"),
}
};
// Up at incarnation 1, then the peer dies (EOF).
let (a1, b1) = pair(&t);
spawn_established(FramedConn::new(a1), peer("node-b", 1), Timing::default())
.expect("registers");
let id1 = id(next_event(&ev, "node_up inc 1"), "inc 1");
drop(b1);
assert_eq!(down_id(next_event(&ev, "node_down inc 1"), "inc 1"), id1);
// Restart: new incarnation, new id — the ghost's id is not reused.
let (a2, b2) = pair(&t);
spawn_established(FramedConn::new(a2), peer("node-b", 2), Timing::default())
.expect("registers");
let id2 = id(next_event(&ev, "node_up inc 2"), "inc 2");
assert_ne!(
id1, id2,
"a restarted node must be distinguishable from its ghost"
);
// Blip: the same incarnation reconnects and keeps its id.
disconnect("node-b");
assert_eq!(down_id(next_event(&ev, "node_down inc 2"), "inc 2"), id2);
let (a3, b3) = pair(&t);
spawn_established(FramedConn::new(a3), peer("node-b", 2), Timing::default())
.expect("registers");
let id3 = id(next_event(&ev, "node_up after blip"), "blip");
assert_eq!(
id2, id3,
"a reconnect at the same incarnation is the same node"
);
disconnect("node-b");
let _ = next_event(&ev, "final node_down");
drop((b2, b3));
mgr.shutdown();
});
}
/// A dropped subscriber is pruned on the next emit and never disturbs the
/// manager or a live subscriber.
#[test]
fn dead_subscriber_is_pruned() {
run(|| {
let mgr = GenServerBuilder::new(Manager::new())
.named(MANAGER)
.start()
.expect("manager name is free");
let dead = subscribe().expect("manager is up");
drop(dead);
let live = subscribe().expect("manager is up");
let t = TcpTransport;
let (a1, b1) = pair(&t);
spawn_established(FramedConn::new(a1), peer("node-b", 1), Timing::default())
.expect("registers");
match next_event(&live, "node_up despite a dead co-subscriber") {
NodeEvent::NodeUp(info) => assert_eq!(info.name, "node-b"),
other => panic!("expected node_up, got {other:?}"),
}
disconnect("node-b");
let _ = next_event(&live, "node_down");
drop(b1);
mgr.shutdown();
});
}
+175
View File
@@ -0,0 +1,175 @@
//! RFC 010 c7 — the Phase 2 gate: a 3-node mesh under the subprocess
//! harness, repeatable.
//!
//! Each node process runs the integrated `cluster::start` (manager +
//! acceptor + connector + static seeds), subscribes to membership like any
//! consumer, and announces protocol-visible facts as lines:
//! `LISTENING <addr>`, `MEMBER-UP <name> inc=<n>`, `MEMBER-DOWN <name>`.
//! Then it **parks forever** — cross-process teardown is retractable state
//! (binding trap), so the parent SIGKILLs via `Node`'s `Drop` and clean exit
//! stays the c4 harness's own smoke test.
//!
//! Ports: nodes bind `:0` and report, so the mesh is built by seeding each
//! node with the previously-reported addresses (n1: no seeds; n2: n1;
//! n3: n1+n2 — inbound covers the reverse edges). The late-seed test is the
//! one exception: the parent pre-reserves a port by binding-and-closing it,
//! seeds one node with it, then starts the second node on that exact
//! address. In principle another process could steal the port in the gap;
//! in practice the window is microseconds on a local runner — accepted, and
//! confined to that one test.
#![cfg(feature = "cluster")]
mod common;
use common::{maybe_child, spawn_node, Node};
use smarm::cluster::envelope::NodeMeta;
use smarm::cluster::membership::{subscribe, NodeEvent};
use smarm::cluster::{start, Config, StaticSeeds, Timing};
use std::time::Duration;
const ROLES: &[(&str, fn())] = &[("node", role_node)];
/// A mesh node: identity and seeds from env, membership events to stdout,
/// park forever (the parent reaps).
fn role_node() {
let name = std::env::var("SMARM_NODE_NAME").expect("SMARM_NODE_NAME not set");
let listen = std::env::var("SMARM_LISTEN_ADDR").unwrap_or_else(|_| "127.0.0.1:0".to_string());
// Seeds: comma-separated `name=addr` pairs; empty or unset means none.
let seeds: Vec<(String, String)> = std::env::var("SMARM_SEEDS")
.unwrap_or_default()
.split(',')
.filter(|s| !s.is_empty())
.map(|s| {
let (n, a) = s.split_once('=').expect("seed must be name=addr");
(n.to_string(), a.to_string())
})
.collect();
smarm::run(move || {
let cluster = start(Config {
node_name: name,
meta: NodeMeta {
role: "mesh-test".to_string(),
region: "local".to_string(),
},
listen_addr: listen,
strategy: Box::new(StaticSeeds::new(seeds)),
timing: Timing::default(),
})
.expect("listener binds");
println!("LISTENING {}", cluster.local_addr());
let events = subscribe().expect("manager is up");
loop {
match events.rx.recv() {
Ok(NodeEvent::NodeUp(info)) => {
println!("MEMBER-UP {} inc={}", info.name, info.incarnation.get());
}
Ok(NodeEvent::NodeDown(info)) => {
println!("MEMBER-DOWN {}", info.name);
}
Err(_) => break, // manager gone; park below regardless
}
}
loop {
smarm::sleep(Duration::from_secs(3600));
}
});
}
fn spawn_mesh_node(name: &str, seeds: &str, listen: Option<&str>) -> Node {
let mut env: Vec<(&str, &str)> = vec![("SMARM_NODE_NAME", name), ("SMARM_SEEDS", seeds)];
if let Some(addr) = listen {
env.push(("SMARM_LISTEN_ADDR", addr));
}
spawn_node("node", &env)
}
/// Wait for `MEMBER-UP <peer> inc=<n>` and return the incarnation.
fn wait_member_up(node: &mut Node, peer: &str) -> u32 {
let prefix = format!("MEMBER-UP {peer} inc=");
let line = node.wait_line(&format!("MEMBER-UP {peer}"), |l| l.starts_with(&prefix));
line[prefix.len()..].parse().expect("incarnation parses")
}
fn wait_member_down(node: &mut Node, peer: &str) {
let want = format!("MEMBER-DOWN {peer}");
node.wait_line(&want, |l| l == want);
}
/// The gate, plus the kill and restart facts, as one mesh's life: three
/// nodes form a full mesh (every node sees both others up); killing one
/// yields `node_down` at both survivors; its restart under the same name
/// arrives as a NEW incarnation — the ghost and its successor are
/// distinguishable at every observer.
#[test]
fn three_node_mesh_forms_then_kill_then_restart_distinguishable() {
maybe_child(ROLES);
let mut n1 = spawn_mesh_node("node-1", "", None);
let a1 = n1.wait_listening();
let mut n2 = spawn_mesh_node("node-2", &format!("node-1={a1}"), None);
let a2 = n2.wait_listening();
let mut n3 = spawn_mesh_node("node-3", &format!("node-1={a1},node-2={a2}"), None);
let _a3 = n3.wait_listening();
// Full mesh: each node reports both peers up (dialed or inbound alike).
wait_member_up(&mut n1, "node-2");
let inc3_at_n1 = wait_member_up(&mut n1, "node-3");
wait_member_up(&mut n2, "node-1");
let inc3_at_n2 = wait_member_up(&mut n2, "node-3");
wait_member_up(&mut n3, "node-1");
wait_member_up(&mut n3, "node-2");
assert_eq!(
inc3_at_n1, inc3_at_n2,
"one node, one incarnation, all observers"
);
// Kill node-3 (SIGKILL via Drop): node_down at both survivors.
drop(n3);
wait_member_down(&mut n1, "node-3");
wait_member_down(&mut n2, "node-3");
// Restart node-3 under the same name: it re-dials its seeds and comes
// up everywhere as a new incarnation — never the ghost's.
let mut n3b = spawn_mesh_node("node-3", &format!("node-1={a1},node-2={a2}"), None);
let _ = n3b.wait_listening();
let inc3b_at_n1 = wait_member_up(&mut n1, "node-3");
let inc3b_at_n2 = wait_member_up(&mut n2, "node-3");
assert_eq!(inc3b_at_n1, inc3b_at_n2);
assert_ne!(
inc3_at_n1, inc3b_at_n1,
"a restarted node must be distinguishable from its ghost"
);
wait_member_up(&mut n3b, "node-1");
wait_member_up(&mut n3b, "node-2");
}
/// A seed that is unreachable at start is not fatal: the connector retries
/// on backoff, and when a node finally appears at that address, the mesh
/// edge forms.
#[test]
fn seed_unreachable_at_start_then_arriving_later() {
maybe_child(ROLES);
// Pre-reserve an address by binding and immediately closing it (see the
// module docs for the accepted steal window). Dials to it are refused
// until node-b starts there.
let reserved = {
let l = std::net::TcpListener::bind("127.0.0.1:0").expect("bind");
l.local_addr().expect("addr").to_string()
};
let mut a = spawn_mesh_node("node-a", &format!("node-b={reserved}"), None);
let _ = a.wait_listening();
// Let a few refused attempts happen before the seed comes up, so the
// retry path is what forms the edge (backoff cap 5s < harness WAIT 10s).
std::thread::sleep(Duration::from_millis(600));
let mut b = spawn_mesh_node("node-b", "", Some(&reserved));
let _ = b.wait_listening();
wait_member_up(&mut a, "node-b");
wait_member_up(&mut b, "node-a");
}
+359
View File
@@ -0,0 +1,359 @@
//! RFC 010 c12 — remote monitors.
//!
//! Local suite (`run()`, no network): the immediate answers — no connection
//! ⇒ `Disconnected`, dead incarnation ⇒ `NoProc` — and the self-node
//! collapse (a plain local monitor underneath, incl. `demonitor_remote`).
//!
//! Cross-process: a *server* exposes a control name and spawns workers on
//! request, replying with each worker's pid (via `RemotePid::from_local`,
//! the D12 set-site) or, for the deliberately unshipped one, only its raw
//! slot numbers. The *client* monitors them and asserts: kill ⇒ the true
//! reason (Exit / Panic); a corpse ⇒ its recorded terminal reason, not
//! NoProc; a live pid that never crossed the wire ⇒ NoProc (no liveness
//! leak); a demonitor racing the kill ⇒ no notice, proven by stream ORDER
//! (a later notice on the same connection arrives while the earlier slot
//! is still empty), not by sleeping.
#![cfg(feature = "cluster")]
mod common;
use common::{maybe_child, spawn_node};
use smarm::cluster::envelope::NodeMeta;
use smarm::cluster::expose::{expose, expose_type};
use smarm::cluster::membership::{subscribe, NodeEvent};
use smarm::cluster::remote::{
self, demonitor_remote, monitor_remote, send_to_remote, RemoteName, RemotePid,
};
use smarm::cluster::{start, Config, RemoteDownReason, StaticSeeds, Timing};
use smarm::pg::Incarnation;
use smarm::{channel, install, register, run, spawn, Addressable, DownReason, Erased, Name, Pid};
use std::collections::HashMap;
use std::time::Duration;
// ---- message types (hand-rolled serde; the crate is derive-less) ---------
#[derive(Debug)]
struct Ctl {
cmd: String,
reply_to: RemotePid<Client>,
}
#[derive(Debug)]
struct Answer {
text: String,
pid: Option<RemotePid<Erased>>,
}
struct Client;
impl Addressable for Client {
type Msg = Answer;
}
impl serde::Serialize for Ctl {
fn serialize<S: serde::Serializer>(&self, s: S) -> Result<S::Ok, S::Error> {
use serde::ser::SerializeTuple;
let mut t = s.serialize_tuple(2)?;
t.serialize_element(&self.cmd)?;
t.serialize_element(&self.reply_to)?;
t.end()
}
}
impl<'de> serde::Deserialize<'de> for Ctl {
fn deserialize<D: serde::Deserializer<'de>>(d: D) -> Result<Self, D::Error> {
let (cmd, reply_to) = <(String, RemotePid<Client>)>::deserialize(d)?;
Ok(Ctl { cmd, reply_to })
}
}
impl serde::Serialize for Answer {
fn serialize<S: serde::Serializer>(&self, s: S) -> Result<S::Ok, S::Error> {
use serde::ser::SerializeTuple;
let mut t = s.serialize_tuple(2)?;
t.serialize_element(&self.text)?;
t.serialize_element(&self.pid)?;
t.end()
}
}
impl<'de> serde::Deserialize<'de> for Answer {
fn deserialize<D: serde::Deserializer<'de>>(d: D) -> Result<Self, D::Error> {
let (text, pid) = <(String, Option<RemotePid<Erased>>)>::deserialize(d)?;
Ok(Answer { text, pid })
}
}
// ================= local suite =========================================
/// No connection to the pid's node: `Disconnected` at once — the remote
/// analog of NoProc, and the first thing c11's variant is for.
#[test]
fn unconnected_node_is_disconnected_immediately() {
maybe_child(ROLES);
run(|| {
remote::set_local_identity("me", Incarnation::new(7));
let ghost = RemotePid::<Erased>::from_parts("nowhere", Incarnation::new(1), 3, 1);
let m = monitor_remote(ghost.clone());
let d = m.recv().unwrap();
assert_eq!(d.pid, ghost);
assert_eq!(d.reason, RemoteDownReason::Disconnected);
});
}
/// The node is connected but the pid names an earlier incarnation: the
/// actor is a known corpse (RFC v2 §3), so `NoProc` at once — never
/// `Disconnected`, nothing on the wire.
#[test]
fn dead_incarnation_is_noproc_immediately() {
maybe_child(ROLES);
run(|| {
remote::set_local_identity("me", Incarnation::new(7));
let (probe_tx, probe_rx) = channel();
remote::bind_outbound_probe("peer", Incarnation::new(5), probe_tx);
let stale = RemotePid::<Erased>::from_parts("peer", Incarnation::new(4), 9, 1);
let m = monitor_remote(stale);
assert_eq!(m.recv().unwrap().reason, DownReason::NoProc.into());
assert!(probe_rx.try_recv().unwrap().is_none(), "no frame emitted");
});
}
/// A self-node pid collapses to an ordinary local monitor: the true reason
/// on exit, and `demonitor_remote` cancels it.
#[test]
fn self_node_pid_collapses_to_local_monitor() {
maybe_child(ROLES);
run(|| {
remote::set_local_identity("me", Incarnation::new(7));
let (go_tx, go_rx) = channel::<()>();
let (go2_tx, go2_rx) = channel::<()>();
let a = spawn(move || {
let _ = go_rx.recv();
})
.pid();
let b = spawn(move || {
let _ = go2_rx.recv();
})
.pid();
let ma = monitor_remote(RemotePid::from_local(a).expect("identity set"));
let mb = monitor_remote(RemotePid::from_local(b).expect("identity set"));
assert_ne!(ma.id, mb.id);
assert!(ma.target.local() == Some(a));
demonitor_remote(&mb);
go2_tx.send(()).unwrap();
go_tx.send(()).unwrap();
let d = ma.recv().unwrap();
assert_eq!(d.reason, DownReason::Exit.into());
assert_eq!(d.pid.local(), Some(a));
// `a` is down (its notice arrived), and `b` was killed first on the
// same scheduler — a notice for `b` would be here by now. After a
// demonitor the channel is closed-empty (`Err`), like the local one.
assert!(matches!(mb.try_recv(), Ok(None) | Err(_)));
});
}
// ================= cross-process ======================================
const ROLES: &[(&str, fn())] = &[("server", role_server), ("client", role_client)];
const CTL: Name<Ctl> = Name::new("c12.ctl");
fn cfg(name: &str, seeds: Vec<(String, String)>) -> Config {
Config {
node_name: name.to_string(),
meta: NodeMeta {
role: "c12".into(),
region: "local".into(),
},
listen_addr: std::env::var("SMARM_LISTEN_ADDR").unwrap_or_else(|_| "127.0.0.1:0".into()),
strategy: Box::new(StaticSeeds::new(seeds)),
timing: Timing::default(),
}
}
fn wait_up(events: &smarm::cluster::membership::MembershipEvents, who: &str) {
loop {
match events.rx.recv() {
Ok(NodeEvent::NodeUp(i)) if i.name == who => return,
Ok(_) => continue,
Err(_) => panic!("manager gone"),
}
}
}
/// Server commands (all answered to `reply_to`):
/// - `spawn:exit` / `spawn:panic` — a parked worker; `kill:<index>` releases
/// it, whereupon it returns / panics. Answer carries its pid.
/// - `spawn:corpse` — a worker that has already exited when the answer is
/// sent; the pid was shipped (watchable) before it died.
/// - `spawn:unwatched` — a parked worker whose pid is NEVER shipped; the
/// answer carries only `text = "slot:<index>:<generation>"`.
fn role_server() {
smarm::run(move || {
let cluster = start(cfg("server", vec![])).expect("binds");
println!("LISTENING {}", cluster.local_addr());
let (tx, rx) = channel::<Ctl>();
register(CTL, tx).unwrap();
expose(CTL);
println!("READY");
let mut workers: HashMap<u32, smarm::channel::Sender<()>> = HashMap::new();
loop {
let ctl = rx.recv().unwrap();
println!("CTL {}", ctl.cmd);
let (text, pid): (String, Option<RemotePid<Erased>>) = match ctl.cmd.as_str() {
"spawn:exit" | "spawn:panic" => {
let panic = ctl.cmd == "spawn:panic";
let (go_tx, go_rx) = channel::<()>();
let p: Pid = spawn(move || {
let _ = go_rx.recv();
if panic {
panic!("worker asked to panic");
}
})
.pid();
workers.insert(p.index(), go_tx);
(
"ok".into(),
Some(RemotePid::from_local(p).expect("identity set")),
)
}
"spawn:corpse" => {
let p: Pid = spawn(|| {}).pid();
let rp = RemotePid::from_local(p).expect("identity set"); // shipped ⇒ watchable
let m = smarm::monitor(p);
let _ = m.rx.recv(); // dead before the answer goes out
("ok".into(), Some(rp))
}
"spawn:unwatched" => {
let (go_tx, go_rx) = channel::<()>();
let p: Pid = spawn(move || {
let _ = go_rx.recv();
})
.pid();
workers.insert(p.index(), go_tx);
(format!("slot:{}:{}", p.index(), p.generation()), None)
}
other => {
let idx: u32 = other.strip_prefix("kill:").unwrap().parse().unwrap();
if let Some(go) = workers.remove(&idx) {
let _ = go.send(());
}
("killed".into(), None)
}
};
send_to_remote(ctl.reply_to, Answer { text, pid }).unwrap();
}
});
}
fn role_client() {
let server_addr = std::env::var("SMARM_SERVER_ADDR").expect("SMARM_SERVER_ADDR");
smarm::run(move || {
let _cluster = start(cfg("client", vec![("server".into(), server_addr)])).expect("binds");
let ev = subscribe().unwrap();
wait_up(&ev, "server");
let (tx, rx) = channel::<Answer>();
let me: Pid<Client> = install::<Client>(tx);
expose_type::<Answer>();
let ask = |cmd: &str| -> Answer {
remote::send(
RemoteName::new("server", CTL),
Ctl {
cmd: cmd.into(),
reply_to: RemotePid::from_local(me).expect("identity set"),
},
)
.unwrap();
rx.recv().unwrap()
};
let server_inc = ev_incarnation();
// 1. kill ⇒ true reason (Exit).
let a = ask("spawn:exit").pid.unwrap();
let ma = monitor_remote(a.clone());
ask(&format!("kill:{}", a.index()));
let d = ma.recv().unwrap();
assert_eq!(d.pid, a);
println!("DOWN exit {:?}", d.reason);
// 2. kill ⇒ true reason (Panic).
let b = ask("spawn:panic").pid.unwrap();
let mb = monitor_remote(b.clone());
ask(&format!("kill:{}", b.index()));
println!("DOWN panic {:?}", mb.recv().unwrap().reason);
// 3. corpse ⇒ recorded terminal reason, not NoProc.
let c = ask("spawn:corpse").pid.unwrap();
println!("DOWN corpse {:?}", monitor_remote(c).recv().unwrap().reason);
// 4. live but never shipped/exposed ⇒ NoProc (no leak); a made-up
// slot on the same node ⇒ NoProc too, indistinguishably.
let ans = ask("spawn:unwatched");
let mut it = ans.text.strip_prefix("slot:").unwrap().split(':');
let (idx, gen): (u32, u32) = (
it.next().unwrap().parse().unwrap(),
it.next().unwrap().parse().unwrap(),
);
let hidden = RemotePid::<Erased>::from_parts("server", server_inc, idx, gen);
println!(
"DOWN hidden {:?}",
monitor_remote(hidden).recv().unwrap().reason
);
let bogus = RemotePid::<Erased>::from_parts("server", server_inc, 100_000, 1);
println!(
"DOWN bogus {:?}",
monitor_remote(bogus).recv().unwrap().reason
);
// 5. demonitor races the kill: no notice for `d1`, proven by order —
// `d2`'s notice (same connection, later) arrives while `d1`'s
// slot is still empty.
let d1 = ask("spawn:exit").pid.unwrap();
let m1 = monitor_remote(d1.clone());
demonitor_remote(&m1);
ask(&format!("kill:{}", d1.index()));
let d2 = ask("spawn:exit").pid.unwrap();
let m2 = monitor_remote(d2.clone());
ask(&format!("kill:{}", d2.index()));
assert_eq!(m2.recv().unwrap().reason, DownReason::Exit.into());
// Closed-empty (`Err`) or open-empty (`Ok(None)`) both mean no notice.
let stray = matches!(m1.try_recv(), Ok(Some(_)));
println!("DEMONITOR stray={stray}");
println!("CLIENT DONE");
loop {
smarm::sleep(Duration::from_secs(3600));
}
});
}
/// The server's incarnation as this node sees it — for building pids by hand.
fn ev_incarnation() -> Incarnation {
smarm::cluster::membership::view()
.expect("manager up")
.into_iter()
.find(|i| i.name == "server")
.map(|i| i.incarnation)
.expect("server in view")
}
/// The Phase 4 c12 gate: remote monitors report the true reason, honour
/// corpses, leak nothing for unshipped pids, and cancel cleanly.
#[test]
fn remote_monitors_report_true_reasons() {
maybe_child(ROLES);
let mut server = spawn_node("server", &[]);
let saddr = server.wait_listening();
server.wait_line("READY", |l| l == "READY");
let mut client = spawn_node("client", &[("SMARM_SERVER_ADDR", &saddr)]);
client.wait_line("DOWN exit Local(Exit)", |l| l == "DOWN exit Local(Exit)");
client.wait_line("DOWN panic Local(Panic)", |l| {
l == "DOWN panic Local(Panic)"
});
client.wait_line("DOWN corpse Local(Exit)", |l| {
l == "DOWN corpse Local(Exit)"
});
client.wait_line("DOWN hidden Local(NoProc)", |l| {
l == "DOWN hidden Local(NoProc)"
});
client.wait_line("DOWN bogus Local(NoProc)", |l| {
l == "DOWN bogus Local(NoProc)"
});
client.wait_line("DEMONITOR stray=false", |l| l == "DEMONITOR stray=false");
client.wait_line("CLIENT DONE", |l| l == "CLIENT DONE");
}
+254
View File
@@ -0,0 +1,254 @@
//! RFC 010 c15 — distributed pg: sync on `NodeUp`, incremental
//! `Join`/`Leave`, eager eviction announced, `NodeDown` sweep.
//!
//! Two nodes. The *origin* joins two local workers to `"pool"` before the
//! *observer* connects (so the observer's view comes from `Sync`), exposes a
//! `"go"` command inbox and then does exactly what the observer tells it:
//! kill one worker, join a third, leave with the second. The observer drives
//! that script through the cluster itself and asserts every step from
//! `members_all` — never touching the group on its own side, except once to
//! prove a mixed local+remote group reads correctly and that `members` stays
//! local. `dispatch_any` is exercised both ways: into the origin's worker
//! (remote pick, `send_to_remote`) and, once the origin is gone, into the
//! observer's own (local pick, `send_to`). Finally the parent SIGKILLs the
//! origin: the observer must sweep every remote member on `NodeDown`.
#![cfg(feature = "cluster")]
mod common;
use common::{maybe_child, spawn_node};
use smarm::cluster::envelope::NodeMeta;
use smarm::cluster::expose::{expose, expose_type};
use smarm::cluster::membership::{subscribe, NodeEvent};
use smarm::cluster::remote::{self, RemoteName};
use smarm::cluster::{
dispatch_any, members_all, pick_any, start, Config, DispatchAnyError, GroupMember, StaticSeeds,
Timing,
};
use smarm::{channel, join, leave, members, register, send_to, spawn_addr, Addressable, Name, Pid};
use std::time::{Duration, Instant};
const GO: Name<u8> = Name::new("go");
const POOL: &str = "pool";
/// A pool worker's message: `"die"` stops it, anything else is printed.
#[derive(Debug, PartialEq)]
struct Job(String);
struct Worker;
impl Addressable for Worker {
type Msg = Job;
}
impl serde::Serialize for Job {
fn serialize<S: serde::Serializer>(&self, s: S) -> Result<S::Ok, S::Error> {
self.0.serialize(s)
}
}
impl<'de> serde::Deserialize<'de> for Job {
fn deserialize<D: serde::Deserializer<'de>>(d: D) -> Result<Self, D::Error> {
String::deserialize(d).map(Job)
}
}
const ROLES: &[(&str, fn())] = &[("origin", role_origin), ("observer", role_observer)];
fn cfg(name: &str, seeds: Vec<(String, String)>) -> Config {
Config {
node_name: name.into(),
meta: NodeMeta {
role: "c15".into(),
region: "local".into(),
},
listen_addr: "127.0.0.1:0".into(),
strategy: Box::new(StaticSeeds::new(seeds)),
timing: Timing::default(),
}
}
/// A pool worker: prints every job it is handed, exits on `"die"`.
fn worker() -> Pid<Worker> {
spawn_addr::<Worker>(|rx| {
while let Ok(Job(s)) = rx.recv() {
if s == "die" {
return;
}
println!("JOB {s}");
}
})
}
fn role_origin() {
smarm::run(|| {
let cluster = start(cfg("origin", vec![])).expect("binds");
// Remote dispatch lands here only for a type this node accepts.
expose_type::<Job>();
let w1 = worker();
let w2 = worker();
assert!(join(POOL, w1));
assert!(join(POOL, w2));
let (go_tx, go_rx) = channel::<u8>();
register(GO, go_tx).unwrap();
expose(GO);
println!("LISTENING {}", cluster.local_addr());
println!("JOINED 2");
loop {
match go_rx.recv().unwrap() {
1 => {
send_to(w1, Job("die".into())).unwrap();
println!("KILLED w1");
}
2 => {
assert!(leave(POOL, w2));
println!("LEFT w2");
}
3 => {
let w3 = worker();
assert!(join(POOL, w3));
println!("JOINED w3");
}
n => panic!("unknown command {n}"),
}
}
});
}
fn remote_count(group: &str) -> usize {
members_all(group)
.iter()
.filter(|m| matches!(m, GroupMember::Remote(_)))
.count()
}
/// Cooperative poll until `pred`; panics (with the last view) on timeout.
fn wait_view(what: &str, group: &str, pred: impl Fn(&[GroupMember]) -> bool) {
let deadline = Instant::now() + Duration::from_secs(5);
loop {
let v = members_all(group);
if pred(&v) {
return;
}
assert!(
Instant::now() < deadline,
"timed out waiting for {what}; view = {v:?}"
);
smarm::sleep(Duration::from_millis(5));
}
}
fn role_observer() {
let origin_addr = std::env::var("SMARM_ORIGIN_ADDR").expect("SMARM_ORIGIN_ADDR");
smarm::run(move || {
let _cluster = start(cfg("observer", vec![("origin".into(), origin_addr)])).expect("binds");
let ev = subscribe().unwrap();
loop {
match ev.rx.recv().unwrap() {
NodeEvent::NodeUp(i) if i.name == "origin" => break,
_ => {}
}
}
let go = |n: u8| remote::send(RemoteName::new("origin", GO), n).unwrap();
// Sync: both pre-existing members arrive with no join on this side.
wait_view("sync of 2 remote members", POOL, |v| {
v.len() == 2 && v.iter().all(|m| matches!(m, GroupMember::Remote(_)))
});
let synced = members_all(POOL);
assert!(synced.iter().all(|m| match m {
GroupMember::Remote(p) => p.node() == "origin",
GroupMember::Local(_) => false,
}));
println!("SEES 2");
// Origin-side death: the origin's reaper announces the leave.
go(1);
wait_view("death evicted on observer", POOL, |v| v.len() == 1);
println!("SEES 1 after death");
// Incremental Join.
go(3);
wait_view("incremental join", POOL, |v| v.len() == 2);
println!("SEES 2 after join");
// Voluntary Leave.
go(2);
wait_view("incremental leave", POOL, |v| v.len() == 1);
println!("SEES 1 after leave");
// Mixed group: our own member sits beside the remote one in
// `members_all`; `members` stays local-only.
let me = worker();
assert!(join(POOL, me));
wait_view("mixed local+remote", POOL, |v| {
v.len() == 2 && v.contains(&GroupMember::Local(me.erase()))
});
assert_eq!(
members(POOL),
vec![me.erase()],
"local API never shows remotes"
);
assert_eq!(remote_count(POOL), 1);
println!("MIXED ok");
// dispatch_any: the store's first entry is the origin's w3 (it was
// announced before we joined), so the pick is remote and the job
// crosses the wire — the origin's worker prints it.
let picked = pick_any(POOL).expect("pool has members");
assert!(
matches!(picked, GroupMember::Remote(_)),
"first entry is remote: {picked:?}"
);
let reached = dispatch_any::<Worker>(POOL, Job("from-observer".into())).unwrap();
assert_eq!(reached, picked);
println!("DISPATCHED remote");
println!("PARK");
// Parent SIGKILLs the origin now: NodeDown must sweep its member,
// ours must survive.
wait_view("node_down sweep", POOL, |v| {
v == [GroupMember::Local(me.erase())]
});
assert_eq!(members(POOL), vec![me.erase()]);
println!("SWEPT");
// Now the only member is ours: a local pick, a local send.
let reached = dispatch_any::<Worker>(POOL, Job("local".into())).unwrap();
assert_eq!(reached, GroupMember::Local(me.erase()));
// And an empty group hands the message back.
match dispatch_any::<Worker>("nobody", Job("lost".into())) {
Err(DispatchAnyError::NoMember(Job(s))) => assert_eq!(s, "lost"),
other => panic!("expected NoMember, got {other:?}"),
}
println!("DISPATCHED local");
loop {
smarm::sleep(Duration::from_secs(3600));
}
});
}
/// The Phase 5 gate: sync, join, leave, death, node_down — all observed from
/// the peer, none of them a group operation on the peer — plus dispatch_any
/// reaching a remote member and a local one.
#[test]
fn groups_span_two_nodes() {
maybe_child(ROLES);
let mut origin = spawn_node("origin", &[]);
let addr = origin.wait_listening();
origin.wait_line("JOINED 2", |l| l == "JOINED 2");
let mut observer = spawn_node("observer", &[("SMARM_ORIGIN_ADDR", &addr)]);
observer.wait_line("SEES 2", |l| l == "SEES 2");
origin.wait_line("KILLED w1", |l| l == "KILLED w1");
observer.wait_line("SEES 1 after death", |l| l == "SEES 1 after death");
origin.wait_line("JOINED w3", |l| l == "JOINED w3");
observer.wait_line("SEES 2 after join", |l| l == "SEES 2 after join");
origin.wait_line("LEFT w2", |l| l == "LEFT w2");
observer.wait_line("SEES 1 after leave", |l| l == "SEES 1 after leave");
observer.wait_line("MIXED ok", |l| l == "MIXED ok");
observer.wait_line("DISPATCHED remote", |l| l == "DISPATCHED remote");
origin.wait_line("JOB from-observer", |l| l == "JOB from-observer");
observer.wait_line("PARK", |l| l == "PARK");
origin.kill();
observer.wait_line("SWEPT", |l| l == "SWEPT");
// Order between the root's line and the worker's is scheduling; wait
// for the later one to be certain both happened.
observer.wait_line("DISPATCHED local", |l| l == "DISPATCHED local");
observer.wait_line("JOB local", |l| l == "JOB local");
}
+356
View File
@@ -0,0 +1,356 @@
//! RFC 010 c10 — pid targeting + auto-serialization. The Phase 3 gate:
//! cross-node call/reply with no ceremony, under the subprocess harness.
//!
//! Local suite (`run()`, no network): serialize/deserialize shapes,
//! self-collapse, the outside-runtime contract, the local send-site
//! incarnation check with a probe proving **no frame is emitted**.
//!
//! Cross-process: two nodes. The *server* exposes a `Name<Req>`; the
//! *client* sends a `Req` carrying its own `Pid<Reply>` (auto-serialized to
//! a `RemotePid` on the wire); the server replies via `send_to_remote`
//! straight back to that pid — no name at the client end, no ceremony. A
//! third-node roundtrip: the client's pid travels client→server→relay→
//! server→client, and still delivers.
#![cfg(feature = "cluster")]
mod common;
use common::{maybe_child, spawn_node};
use smarm::cluster::envelope::{encode_payload, Frame, NodeMeta};
use smarm::cluster::expose::{expose, type_hash};
use smarm::cluster::membership::{subscribe, NodeEvent};
use smarm::cluster::remote::{self, send_to_remote, RemoteName, RemotePid, ToRemoteError};
use smarm::cluster::{start, Config, StaticSeeds, Timing};
use smarm::pg::Incarnation;
use smarm::{channel, install, register, run, Addressable, Name, Pid};
use std::time::Duration;
// ---- message types (std-only payloads; the crate's serde is derive-less,
// so wire types are hand-rolled with serde's tuple/seq API via `serde::ser`
// impls below — the same thing a user's derive would generate) ------------
/// A request carrying a reply-to. Serialize/Deserialize are written by hand
/// here for exactly one reason: this crate deliberately does not pull in
/// serde-derive. Field 1 is the auto-serializing pid.
#[derive(Debug, PartialEq)]
struct Req {
text: String,
reply_to: RemotePid<Replier>,
}
#[derive(Debug, PartialEq)]
struct Reply(String);
struct Replier;
impl Addressable for Replier {
type Msg = Reply;
}
impl serde::Serialize for Req {
fn serialize<S: serde::Serializer>(&self, s: S) -> Result<S::Ok, S::Error> {
use serde::ser::SerializeTuple;
let mut t = s.serialize_tuple(2)?;
t.serialize_element(&self.text)?;
t.serialize_element(&self.reply_to)?;
t.end()
}
}
impl<'de> serde::Deserialize<'de> for Req {
fn deserialize<D: serde::Deserializer<'de>>(d: D) -> Result<Self, D::Error> {
let (text, reply_to) = <(String, RemotePid<Replier>)>::deserialize(d)?;
Ok(Req { text, reply_to })
}
}
impl serde::Serialize for Reply {
fn serialize<S: serde::Serializer>(&self, s: S) -> Result<S::Ok, S::Error> {
self.0.serialize(s)
}
}
impl<'de> serde::Deserialize<'de> for Reply {
fn deserialize<D: serde::Deserializer<'de>>(d: D) -> Result<Self, D::Error> {
String::deserialize(d).map(Reply)
}
}
// ================= local suite =========================================
/// A local `Pid<A>` serializes as a `RemotePid<A>` stamped with this node's
/// identity; deserializing it back on the same node collapses to the same
/// local pid (`local()` is `Some`, `Pid` round-trips).
#[test]
fn local_pid_serializes_and_collapses_on_self() {
maybe_child(ROLES);
run(|| {
// The local identity is set by cluster::start; the local suite sets
// it directly.
remote::set_local_identity("me", Incarnation::new(7));
let (tx, _rx) = channel::<Reply>();
let me: Pid<Replier> = install::<Replier>(tx);
let bytes = encode_payload(&me).unwrap();
let rp: RemotePid<Replier> = smarm::cluster::envelope::decode_payload(&bytes).unwrap();
assert_eq!(rp.node(), "me");
assert_eq!(rp.incarnation(), Incarnation::new(7));
assert_eq!(
rp.local(),
Some(me),
"self-node pid collapses to the local pid"
);
// Deserializing straight into Pid<A> works for a self-node pid...
let back: Pid<Replier> = smarm::cluster::envelope::decode_payload(&bytes).unwrap();
assert_eq!(back, me);
// ...and FAILS for a foreign one (collapse is literal: node == self).
let foreign = RemotePid::<Replier>::from_parts("elsewhere", Incarnation::new(1), 3, 1);
let fbytes = encode_payload(&foreign).unwrap();
assert!(smarm::cluster::envelope::decode_payload::<Pid<Replier>>(&fbytes).is_err());
assert_eq!(foreign.local(), None);
});
}
/// `send_to_remote` short-circuits locally for a self-node pid — the
/// zero-copy-equivalent collapse: the message object itself lands in the
/// local channel, no encode, no frame.
#[test]
fn send_to_remote_collapses_locally_for_self() {
maybe_child(ROLES);
run(|| {
remote::set_local_identity("me", Incarnation::new(7));
let (tx, rx) = channel::<Reply>();
let me: Pid<Replier> = install::<Replier>(tx);
let rp = RemotePid::from_local(me).expect("identity set");
// Probe the outbound path: nothing must be handed to any connection.
let (probe_tx, probe_rx) = channel::<Frame>();
remote::bind_outbound_probe("me", Incarnation::new(7), probe_tx);
send_to_remote(rp, Reply("hi".into())).unwrap();
assert_eq!(rx.recv().unwrap(), Reply("hi".into()));
assert!(
matches!(probe_rx.try_recv(), Ok(None)),
"no frame for a local collapse"
);
});
}
/// RFC v2 §3: a `RemotePid` whose incarnation is not the current one for its
/// node fails at the local send site with `DeadIncarnation`, and NO frame
/// is emitted — asserted on a probe sender bound as that node's outbound.
#[test]
fn stale_incarnation_rejected_locally_no_frame() {
maybe_child(ROLES);
run(|| {
remote::set_local_identity("me", Incarnation::new(7));
let (probe_tx, probe_rx) = channel::<Frame>();
remote::bind_outbound_probe("peer", Incarnation::new(5), probe_tx);
let stale = RemotePid::<Replier>::from_parts("peer", Incarnation::new(4), 9, 1);
match send_to_remote(stale, Reply("late".into())) {
Err(ToRemoteError::DeadIncarnation(Reply(s))) => assert_eq!(s, "late"),
other => panic!("expected DeadIncarnation, got {other:?}"),
}
assert!(
matches!(probe_rx.try_recv(), Ok(None)),
"stale pid must emit no frame"
);
// The current incarnation goes through: a Send frame with the pid's
// (index, generation) and Reply's hash lands on the probe.
let live = RemotePid::<Replier>::from_parts("peer", Incarnation::new(5), 9, 1);
send_to_remote(live, Reply("now".into())).unwrap();
match probe_rx.recv().unwrap() {
Frame::Send {
index,
generation,
type_hash: h,
payload,
} => {
assert_eq!((index, generation), (9, 1));
assert_eq!(h, type_hash::<Reply>());
let r: Reply = smarm::cluster::envelope::decode_payload(&payload).unwrap();
assert_eq!(r, Reply("now".into()));
}
f => panic!("expected Send, got {f:?}"),
}
// Unknown node: NotConnected, no frame anywhere.
let nowhere = RemotePid::<Replier>::from_parts("nowhere", Incarnation::new(1), 1, 1);
assert!(matches!(
send_to_remote(nowhere, Reply("x".into())),
Err(ToRemoteError::NotConnected(_))
));
});
}
// ================= cross-process gate ==================================
const ROLES: &[(&str, fn())] = &[
("server", role_server),
("client", role_client),
("relay", role_relay),
];
const ECHO: Name<Req> = Name::new("c10.echo");
const RELAY: Name<Req> = Name::new("c10.relay");
fn cfg(name: &str, seeds: Vec<(String, String)>) -> Config {
Config {
node_name: name.to_string(),
meta: NodeMeta {
role: "c10".into(),
region: "local".into(),
},
listen_addr: std::env::var("SMARM_LISTEN_ADDR").unwrap_or_else(|_| "127.0.0.1:0".into()),
strategy: Box::new(StaticSeeds::new(seeds)),
timing: Timing::default(),
}
}
fn wait_up(events: &smarm::cluster::membership::MembershipEvents, who: &str) {
loop {
match events.rx.recv() {
Ok(NodeEvent::NodeUp(i)) if i.name == who => return,
Ok(_) => continue,
Err(_) => panic!("manager gone"),
}
}
}
/// Server: exposes ECHO; each Req is answered by `send_to_remote` to its
/// reply_to — the server never learns a name for the client. If the Req text
/// starts with "via-relay:", it forwards the whole Req (reply_to and all) to
/// the relay node instead, which sends it back here; the second arrival is
/// answered normally. That is the pid's third-node roundtrip.
fn role_server() {
let relay_addr = std::env::var("SMARM_RELAY_ADDR").ok();
smarm::run(move || {
let seeds = relay_addr
.map(|a| vec![("relay".to_string(), a)])
.unwrap_or_default();
let cluster = start(cfg("server", seeds)).expect("binds");
println!("LISTENING {}", cluster.local_addr());
let (tx, rx) = channel::<Req>();
register(ECHO, tx).unwrap();
expose(ECHO);
println!("READY");
loop {
let req = rx.recv().unwrap();
if let Some(rest) = req.text.strip_prefix("via-relay:") {
let fwd = Req {
text: format!("relayed:{rest}"),
reply_to: req.reply_to,
};
remote::send(RemoteName::new("relay", RELAY), fwd).unwrap();
println!("FORWARDED");
continue;
}
println!("REQ {}", req.text);
send_to_remote(req.reply_to, Reply(format!("echo:{}", req.text))).unwrap();
}
});
}
/// Relay: exposes RELAY; bounces every Req straight back to the server's
/// ECHO, untouched. The client's pid inside it now crosses relay→server.
fn role_relay() {
let server_addr = std::env::var("SMARM_SERVER_ADDR").expect("SMARM_SERVER_ADDR");
smarm::run(move || {
let cluster = start(cfg("relay", vec![("server".into(), server_addr)])).expect("binds");
println!("LISTENING {}", cluster.local_addr());
let (tx, rx) = channel::<Req>();
register(RELAY, tx).unwrap();
expose(RELAY);
let ev = subscribe().unwrap();
wait_up(&ev, "server");
println!("READY");
loop {
let req = rx.recv().unwrap();
println!("RELAYING {}", req.text);
remote::send(RemoteName::new("server", ECHO), req).unwrap();
}
});
}
/// Client: connects to server, installs a Reply inbox on its own pid,
/// declares it accepts `Reply` (`expose_type` — the RFC's one kept piece of
/// ceremony: nothing is remotely deliverable by default), sends a Req with
/// `reply_to = my pid` (auto-serialized), awaits the reply.
fn role_client() {
let server_addr = std::env::var("SMARM_SERVER_ADDR").expect("SMARM_SERVER_ADDR");
let via_relay = std::env::var("SMARM_VIA_RELAY").is_ok();
smarm::run(move || {
let _cluster = start(cfg("client", vec![("server".into(), server_addr)])).expect("binds");
let ev = subscribe().unwrap();
wait_up(&ev, "server");
println!("MEMBER-UP server");
let (tx, rx) = channel::<Reply>();
let me: Pid<Replier> = install::<Replier>(tx);
// The one deliberate line: a pid-targeted inbound is deliverable only
// for types this node has said it accepts (RFC §4, the safety).
smarm::cluster::expose::expose_type::<Reply>();
let text = if via_relay { "via-relay:ping" } else { "ping" };
remote::send(
RemoteName::new("server", ECHO),
Req {
text: text.into(),
reply_to: RemotePid::from_local(me).expect("identity set"),
},
)
.unwrap();
println!("SENT");
let Reply(s) = rx.recv().unwrap();
println!("REPLY {s}");
loop {
smarm::sleep(Duration::from_secs(3600));
}
});
}
/// The gate: cross-node call/reply with no ceremony.
#[test]
fn cross_node_call_reply_no_ceremony() {
maybe_child(ROLES);
let mut server = spawn_node("server", &[]);
let saddr = server.wait_listening();
server.wait_line("READY", |l| l == "READY");
let mut client = spawn_node("client", &[("SMARM_SERVER_ADDR", &saddr)]);
client.wait_line("SENT", |l| l == "SENT");
server.wait_line("REQ ping", |l| l == "REQ ping");
client.wait_line("REPLY echo:ping", |l| l == "REPLY echo:ping");
}
/// The client's pid, round-tripped through a third node, still delivers.
#[test]
fn pid_roundtrips_through_third_node() {
maybe_child(ROLES);
// Relay needs the server address; server needs the relay address —
// pre-reserve the relay port (same accepted micro-window as cluster_mesh).
let relay_addr = {
let l = std::net::TcpListener::bind("127.0.0.1:0").unwrap();
l.local_addr().unwrap().to_string()
};
let mut server = spawn_node("server", &[("SMARM_RELAY_ADDR", &relay_addr)]);
let saddr = server.wait_listening();
server.wait_line("READY", |l| l == "READY");
let mut relay = spawn_node(
"relay",
&[
("SMARM_SERVER_ADDR", &saddr),
("SMARM_LISTEN_ADDR", &relay_addr),
],
);
let _ = relay.wait_listening();
relay.wait_line("READY", |l| l == "READY");
let mut client = spawn_node(
"client",
&[("SMARM_SERVER_ADDR", &saddr), ("SMARM_VIA_RELAY", "1")],
);
client.wait_line("SENT", |l| l == "SENT");
server.wait_line("FORWARDED", |l| l == "FORWARDED");
relay.wait_line("RELAYING", |l| l.starts_with("RELAYING"));
server.wait_line("REQ relayed:ping", |l| l == "REQ relayed:ping");
client.wait_line("REPLY echo:relayed:ping", |l| {
l == "REPLY echo:relayed:ping"
});
}
+221
View File
@@ -0,0 +1,221 @@
//! RFC 010 c9 — remote `Name` sends: the outbound seam and the single
//! inbound name-resolution seam, cross-process.
//!
//! Two node processes each run the integrated `cluster::start`. The
//! *receiver* registers a `String` inbox under a name and exposes it (and
//! registers a second name it does NOT expose); the *sender* waits for
//! `node_up`, then sends. Facts cross as stdout lines: `LISTENING <addr>`,
//! `MEMBER-UP <name>`, `GOT <payload>`, `SEND-RESULT <case> <verdict>`.
//! Roles park forever afterwards (retractable-state trap); the parent
//! SIGKILLs via `Drop`.
//!
//! What is asserted at each end (roadmap-binding):
//! - cross-node name-send delivers the payload;
//! - an unexposed name is unreachable — the receiver's inbox stays empty
//! even though the name IS registered locally;
//! - a wrong type hash is a decode failure at the receiver, never a
//! misroute — the `String` inbox does not see a `u64` delivered under a
//! made-up hash, nor a `u64` under `u64`'s hash;
//! - a send to a disconnected (never-connected) node fails locally with
//! `NotConnected`, and `Ok(())` means only "handed to the transport".
//!
//! Timing note for the "stays empty" assertions: they are proven by
//! ORDERING, not by waiting — the sender emits the negative-case frames
//! BEFORE the positive one on the same connection (in-order stream), so when
//! the receiver has seen the positive payload, the negatives have already
//! been processed and refused. No sleep-and-hope.
#![cfg(feature = "cluster")]
mod common;
use common::{maybe_child, spawn_node, Node};
use smarm::cluster::envelope::NodeMeta;
use smarm::cluster::expose::expose;
use smarm::cluster::membership::{subscribe, NodeEvent};
use smarm::cluster::remote::{send_remote_raw, RemoteName, RemoteSendError};
use smarm::cluster::{start, Config, StaticSeeds, Timing};
use smarm::{channel, register, Name};
use std::time::Duration;
const ROLES: &[(&str, fn())] = &[("receiver", role_receiver), ("sender", role_sender)];
const INBOX: Name<String> = Name::new("c9.inbox");
const HIDDEN: Name<String> = Name::new("c9.hidden");
fn base_config(name: &str, seeds: Vec<(String, String)>) -> Config {
Config {
node_name: name.to_string(),
meta: NodeMeta {
role: "c9".to_string(),
region: "local".to_string(),
},
listen_addr: "127.0.0.1:0".to_string(),
strategy: Box::new(StaticSeeds::new(seeds)),
timing: Timing::default(),
}
}
/// Receiver: register + expose INBOX; register HIDDEN unexposed **in a
/// separate actor** (one actor holds one channel per message type — a
/// second `register` of the same `M` on one actor silently replaces the
/// first, closing it); print every payload that lands in either.
fn role_receiver() {
smarm::run(|| {
let cluster = start(base_config("recv", vec![])).expect("listener binds");
println!("LISTENING {}", cluster.local_addr());
// HIDDEN's holder: its own actor, so its String channel does not
// displace INBOX's on the root actor.
let (hidden_ready_tx, hidden_ready_rx) = channel::<()>();
smarm::spawn(move || {
let (hid_tx, hid_rx) = channel::<String>();
register(HIDDEN, hid_tx).unwrap();
hidden_ready_tx.send(()).unwrap();
loop {
match hid_rx.recv() {
Ok(s) => println!("GOT-HIDDEN {s}"),
Err(_) => break,
}
}
});
hidden_ready_rx.recv().unwrap();
let (in_tx, in_rx) = channel::<String>();
register(INBOX, in_tx).unwrap();
expose(INBOX);
println!("READY");
loop {
match in_rx.recv() {
Ok(s) => println!("GOT {s}"),
Err(_) => break,
}
}
loop {
smarm::sleep(Duration::from_secs(3600));
}
});
}
/// Sender: connect to recv, wait for node_up, then in this ORDER on the one
/// connection: hidden-name send, wrong-hash sends (two flavours), then the
/// positive send. Plus a send to a node that is not connected at all.
fn role_sender() {
let recv_addr = std::env::var("SMARM_RECV_ADDR").expect("SMARM_RECV_ADDR");
smarm::run(move || {
let _cluster = start(base_config("send", vec![("recv".to_string(), recv_addr)]))
.expect("listener binds");
let events = subscribe().expect("manager is up");
loop {
match events.rx.recv() {
Ok(NodeEvent::NodeUp(info)) if info.name == "recv" => break,
Ok(_) => continue,
Err(_) => panic!("manager gone"),
}
}
println!("MEMBER-UP recv");
// Not connected: purely local knowledge, no frame leaves.
let ghost: RemoteName<String> = RemoteName::new("nowhere", INBOX);
let r = smarm::cluster::remote::send(ghost, "lost".to_string());
println!(
"SEND-RESULT not-connected {}",
match r {
Err(RemoteSendError::NotConnected(_)) => "NotConnected",
Ok(()) => "Ok",
Err(_) => "OtherErr",
}
);
// Unexposed name at the peer: the frame goes (local knowledge can't
// know the peer's exposed set) and the peer refuses it.
let hidden: RemoteName<String> = RemoteName::new("recv", HIDDEN);
let r = smarm::cluster::remote::send(hidden, "should not land".to_string());
println!(
"SEND-RESULT hidden {}",
if r.is_ok() { "Ok" } else { "Err" }
);
// Wrong hash, two flavours: (a) a u64 payload under a made-up hash
// (unknown type at the peer); (b) a u64 payload under u64's real
// hash against a String-typed name (decoder known, wrong channel).
// Both are raw sends — the typed API cannot express them, by design.
let bogus = 0xdead_beef_u64;
let r = send_remote_raw(
"recv",
"c9.inbox",
bogus,
&smarm::cluster::envelope::encode_payload(&7u64).unwrap(),
);
println!(
"SEND-RESULT wrong-hash-unknown {}",
if r.is_ok() { "Ok" } else { "Err" }
);
let r = send_remote_raw(
"recv",
"c9.inbox",
smarm::cluster::expose::type_hash::<u64>(),
&smarm::cluster::envelope::encode_payload(&7u64).unwrap(),
);
println!(
"SEND-RESULT wrong-hash-known {}",
if r.is_ok() { "Ok" } else { "Err" }
);
// Positive: last on the stream, so its arrival proves the negatives
// were already processed.
let inbox: RemoteName<String> = RemoteName::new("recv", INBOX);
let r = smarm::cluster::remote::send(inbox, "hello from send".to_string());
println!(
"SEND-RESULT positive {}",
if r.is_ok() { "Ok" } else { "Err" }
);
loop {
smarm::sleep(Duration::from_secs(3600));
}
});
}
fn wait_send_result(node: &mut Node, case: &str) -> String {
let prefix = format!("SEND-RESULT {case} ");
let line = node.wait_line(&prefix, |l| l.starts_with(&prefix));
line[prefix.len()..].to_string()
}
#[test]
fn remote_name_send_delivers_and_refusals_never_misroute() {
maybe_child(ROLES);
let mut recv = spawn_node("receiver", &[]);
let addr = recv.wait_listening();
recv.wait_line("READY", |l| l == "READY");
let mut send = spawn_node("sender", &[("SMARM_RECV_ADDR", &addr)]);
send.wait_line("MEMBER-UP recv", |l| l == "MEMBER-UP recv");
// Local-knowledge-only failure for an unknown node.
assert_eq!(wait_send_result(&mut send, "not-connected"), "NotConnected");
// Every frame-bearing send is Ok — Ok means "handed to the transport",
// nothing about what the peer does with it (RFC §3, documented here).
assert_eq!(wait_send_result(&mut send, "hidden"), "Ok");
assert_eq!(wait_send_result(&mut send, "wrong-hash-unknown"), "Ok");
assert_eq!(wait_send_result(&mut send, "wrong-hash-known"), "Ok");
assert_eq!(wait_send_result(&mut send, "positive"), "Ok");
// The positive payload lands...
recv.wait_line("GOT hello from send", |l| l == "GOT hello from send");
// ...and, by stream ordering, every negative before it was refused: no
// GOT for the wrong-hash frames, no GOT-HIDDEN at all. The transcript
// up to this point is the proof.
let transcript = recv.transcript();
let gots: Vec<&str> = transcript
.iter()
.map(|s| s.as_str())
.filter(|l| l.starts_with("GOT"))
.collect();
assert_eq!(
gots,
["GOT hello from send"],
"exactly one delivery, the exposed one"
);
}
+272
View File
@@ -0,0 +1,272 @@
//! RFC 010 c3 — transport conformance suite, run against both shipped impls
//! (TCP and in-memory loopback), plus impl-specific cases.
//!
//! Shared suite (roadmap): frame roundtrips through the framed codec, framing
//! across a split write, coalesced frames in one write, peer-close mid-frame
//! (must error, not EOF), clean close at a frame boundary (EOF as `Ok(None)`).
//!
//! The TCP impl parks the calling actor, so its runs live inside `smarm::run`;
//! loopback blocks the OS thread and runs as plain tests.
#![cfg(feature = "cluster")]
use smarm::cluster::envelope::Frame;
use smarm::cluster::transport::loopback::LoopbackTransport;
use smarm::cluster::transport::tcp::TcpTransport;
use smarm::cluster::transport::{Conn, FramedConn, RecvError, Transport};
// ---------------------------------------------------------------------------
// Helpers
// ---------------------------------------------------------------------------
/// Listener + dial + accept against one transport, both conns returned.
/// Relies on dial not requiring a concurrent accept (TCP backlog / loopback
/// queue), so a single thread or actor can hold both ends.
fn pair(t: &dyn Transport, addr: &str) -> (Box<dyn Conn>, Box<dyn Conn>) {
let mut l = t.listen(addr).unwrap();
let a = t.dial(&l.local_addr()).unwrap();
let b = l.accept().unwrap();
(a, b)
}
fn frames() -> Vec<Frame> {
vec![
Frame::Heartbeat,
Frame::Send {
index: 42,
generation: 3,
type_hash: 0x1234_5678_9ABC_DEF0,
payload: vec![1, 2, 3, 4, 5],
},
Frame::SendNamed {
name: "the_counter".into(),
type_hash: 0xFFFF_0000_FFFF_0000,
payload: vec![],
},
Frame::Demonitor { monitor_id: 77 },
]
}
fn encode(f: &Frame) -> Vec<u8> {
let mut out = Vec::new();
f.encode(&mut out).unwrap();
out
}
// ---------------------------------------------------------------------------
// Shared conformance suite — generic over an established pair
// ---------------------------------------------------------------------------
fn suite_roundtrip(a: Box<dyn Conn>, b: Box<dyn Conn>) {
let mut fa = FramedConn::new(a);
let mut fb = FramedConn::new(b);
// a -> b, then b -> a: both directions carry every frame shape.
for f in frames() {
fa.send(&f).unwrap();
assert_eq!(fb.recv().unwrap().unwrap(), f);
}
for f in frames() {
fb.send(&f).unwrap();
assert_eq!(fa.recv().unwrap().unwrap(), f);
}
}
fn suite_split_write(mut a: Box<dyn Conn>, b: Box<dyn Conn>) {
let f = Frame::Send {
index: 7,
generation: 1,
type_hash: 0xAB,
payload: vec![9; 64],
};
let bytes = encode(&f);
// Split inside the length prefix, then inside the body: the reader must
// reassemble regardless of where the boundary falls.
a.write_all(&bytes[..2]).unwrap();
a.write_all(&bytes[2..10]).unwrap();
a.write_all(&bytes[10..]).unwrap();
let mut fb = FramedConn::new(b);
assert_eq!(fb.recv().unwrap().unwrap(), f);
}
fn suite_coalesced(mut a: Box<dyn Conn>, b: Box<dyn Conn>) {
let f1 = Frame::Heartbeat;
let f2 = Frame::Demonitor { monitor_id: 5 };
let mut bytes = encode(&f1);
bytes.extend_from_slice(&encode(&f2));
a.write_all(&bytes).unwrap();
let mut fb = FramedConn::new(b);
assert_eq!(fb.recv().unwrap().unwrap(), f1);
assert_eq!(fb.recv().unwrap().unwrap(), f2);
}
fn suite_close_mid_frame(mut a: Box<dyn Conn>, b: Box<dyn Conn>) {
let bytes = encode(&Frame::Send {
index: 1,
generation: 1,
type_hash: 1,
payload: vec![0; 128],
});
a.write_all(&bytes[..bytes.len() / 2]).unwrap();
a.close();
let mut fb = FramedConn::new(b);
match fb.recv() {
Err(RecvError::TruncatedByPeer) => {}
other => panic!("expected TruncatedByPeer, got {other:?}"),
}
}
fn suite_clean_close(mut a: Box<dyn Conn>, b: Box<dyn Conn>) {
let f = Frame::Heartbeat;
a.write_all(&encode(&f)).unwrap();
a.close();
let mut fb = FramedConn::new(b);
// The buffered frame is still delivered, then EOF at the boundary.
assert_eq!(fb.recv().unwrap().unwrap(), f);
assert!(fb.recv().unwrap().is_none());
}
fn run_suite(t: &dyn Transport, addr: &str) {
let (a, b) = pair(t, addr);
suite_roundtrip(a, b);
let (a, b) = pair(t, addr);
suite_split_write(a, b);
let (a, b) = pair(t, addr);
suite_coalesced(a, b);
let (a, b) = pair(t, addr);
suite_close_mid_frame(a, b);
let (a, b) = pair(t, addr);
suite_clean_close(a, b);
}
// ---------------------------------------------------------------------------
// Loopback — plain tests, no runtime
// ---------------------------------------------------------------------------
#[test]
fn loopback_conformance() {
// Fresh transport per pair() call is fine, but one instance must also
// support sequential re-listen on distinct addresses.
let t = LoopbackTransport::default();
run_suite(&t, "alpha");
}
#[test]
fn loopback_dial_unknown_addr_refused() {
let t = LoopbackTransport::default();
let err = t.dial("nobody-home").unwrap_err();
assert_eq!(err.kind(), std::io::ErrorKind::ConnectionRefused);
}
#[test]
fn loopback_addr_in_use() {
let t = LoopbackTransport::default();
let _l = t.listen("alpha").unwrap();
let err = t.listen("alpha").unwrap_err();
assert_eq!(err.kind(), std::io::ErrorKind::AddrInUse);
}
#[test]
fn loopback_listener_drop_frees_addr_and_refuses_dial() {
let t = LoopbackTransport::default();
let l = t.listen("alpha").unwrap();
drop(l);
let err = t.dial("alpha").unwrap_err();
assert_eq!(err.kind(), std::io::ErrorKind::ConnectionRefused);
// Address is reusable after the listener is gone.
let _l2 = t.listen("alpha").unwrap();
}
#[test]
fn loopback_write_after_peer_close_broken_pipe() {
let t = LoopbackTransport::default();
let (mut a, mut b) = pair(&t, "alpha");
b.close();
let err = a.write_all(&[1, 2, 3]).unwrap_err();
assert_eq!(err.kind(), std::io::ErrorKind::BrokenPipe);
}
#[test]
fn loopback_cross_thread_blocking_read() {
// Reader blocks on an empty pipe until the writer thread delivers.
let t = LoopbackTransport::default();
let (a, b) = pair(&t, "alpha");
let mut fb = FramedConn::new(b);
let writer = std::thread::spawn(move || {
let mut a = a;
std::thread::sleep(std::time::Duration::from_millis(30));
a.write_all(&encode(&Frame::Heartbeat)).unwrap();
});
assert_eq!(fb.recv().unwrap().unwrap(), Frame::Heartbeat);
writer.join().unwrap();
}
// ---------------------------------------------------------------------------
// TCP — inside the runtime (read/write park the calling actor)
// ---------------------------------------------------------------------------
#[test]
fn tcp_conformance() {
smarm::run(|| {
run_suite(&TcpTransport, "127.0.0.1:0");
});
}
#[test]
fn tcp_dial_refused() {
smarm::run(|| {
// Bind to an OS-assigned port, learn it, close the listener, dial it.
let addr = {
let l = TcpTransport.listen("127.0.0.1:0").unwrap();
l.local_addr()
};
let err = TcpTransport.dial(&addr).unwrap_err();
assert_eq!(err.kind(), std::io::ErrorKind::ConnectionRefused);
});
}
#[test]
fn tcp_bad_addr_rejected_without_resolution() {
// Addresses are opaque pre-resolved strings; the c9 seam resolves names.
// A hostname is therefore invalid input here, not something to resolve.
let err = TcpTransport.dial("localhost:1234").unwrap_err();
assert_eq!(err.kind(), std::io::ErrorKind::InvalidInput);
}
#[test]
fn tcp_local_addr_reports_real_port() {
let l = TcpTransport.listen("127.0.0.1:0").unwrap();
let addr = l.local_addr();
let port: u16 = addr.rsplit(':').next().unwrap().parse().unwrap();
assert_ne!(port, 0);
}
#[test]
fn tcp_big_frame_across_socket_buffers() {
// A payload far beyond socket buffer sizes forces genuine fragmentation
// and write backpressure: writer and reader must run concurrently.
smarm::run(|| {
let (tx, rx) = smarm::channel::<Frame>();
let payload = vec![0xA5u8; 4 * 1024 * 1024];
let f = Frame::Send {
index: 9,
generation: 2,
type_hash: 0xC0FFEE,
payload,
};
let mut l = TcpTransport.listen("127.0.0.1:0").unwrap();
let addr = l.local_addr();
let fw = f.clone();
let writer = smarm::spawn(move || {
let mut fa = FramedConn::new(TcpTransport.dial(&addr).unwrap());
fa.send(&fw).unwrap();
});
let reader = smarm::spawn(move || {
let mut fb = FramedConn::new(l.accept().unwrap());
let got = fb.recv().unwrap().unwrap();
tx.send(got).unwrap();
});
let got = rx.recv().unwrap();
assert_eq!(got, f);
writer.join().unwrap();
reader.join().unwrap();
});
}
+108
View File
@@ -0,0 +1,108 @@
//! RFC 010 c4 — two-node harness smoke tests.
//!
//! Roadmap: "spawn two, handshake-less connect, both exit clean." The
//! listener node binds port 0 and announces its concrete address; the
//! dialer connects raw (no Hello — c5 doesn't exist yet), pushes one
//! Heartbeat through the real framed codec, and closes. Assertions are on
//! protocol-visible lines only. Flake budget: see tests/common/mod.rs.
#![cfg(feature = "cluster")]
mod common;
use common::{maybe_child, spawn_node};
use smarm::cluster::envelope::Frame;
use smarm::cluster::transport::tcp::TcpTransport;
use smarm::cluster::transport::{FramedConn, Transport};
const ROLES: &[(&str, fn())] = &[
("listener", role_listener),
("dialer", role_dialer),
("hang", role_hang),
("fail", role_fail),
];
fn role_listener() {
smarm::run(|| {
let mut l = TcpTransport.listen("127.0.0.1:0").unwrap();
println!("LISTENING {}", l.local_addr());
let mut fc = FramedConn::new(l.accept().unwrap());
match fc.recv() {
Ok(Some(Frame::Heartbeat)) => println!("RECV heartbeat"),
other => {
println!("RECV unexpected: {other:?}");
std::process::exit(3);
}
}
match fc.recv() {
Ok(None) => println!("PEER-CLOSED clean"),
other => {
println!("PEER-CLOSED unexpected: {other:?}");
std::process::exit(3);
}
}
});
println!("EXIT ok");
}
fn role_dialer() {
let addr = std::env::var("SMARM_PEER_ADDR").expect("SMARM_PEER_ADDR not set");
smarm::run(move || {
let mut fc = FramedConn::new(TcpTransport.dial(&addr).unwrap());
fc.send(&Frame::Heartbeat).unwrap();
fc.close();
println!("SENT heartbeat");
});
println!("EXIT ok");
}
fn role_hang() {
println!("HANGING");
loop {
std::thread::sleep(std::time::Duration::from_secs(3600));
}
}
fn role_fail() {
std::process::exit(7);
}
/// The roadmap smoke test: two real processes, raw transport connect, one
/// frame across, clean close observed on both sides, both exit 0.
#[test]
fn two_nodes_connect_and_exit_clean() {
maybe_child(ROLES);
let mut listener = spawn_node("listener", &[]);
let addr = listener.wait_listening();
let mut dialer = spawn_node("dialer", &[("SMARM_PEER_ADDR", &addr)]);
dialer.wait_line("SENT heartbeat", |l| l == "SENT heartbeat");
listener.wait_line("RECV heartbeat", |l| l == "RECV heartbeat");
listener.wait_line("clean peer close", |l| l == "PEER-CLOSED clean");
dialer.wait_exit_ok();
listener.wait_exit_ok();
}
/// Reap guarantee: dropping a Node kills a hung child — no orphan survives
/// a panicking test.
#[test]
fn drop_reaps_hung_node() {
maybe_child(ROLES);
let mut node = spawn_node("hang", &[]);
node.wait_line("HANGING", |l| l == "HANGING");
let pid = node.pid().expect("live child has a pid") as libc::pid_t;
drop(node);
// After Drop's kill+wait the pid is fully reaped: signalling it fails
// with ESRCH (pid-reuse in this instant is not a realistic race).
let rc = unsafe { libc::kill(pid, 0) };
assert_eq!(rc, -1, "process still signallable after Drop");
let errno = std::io::Error::last_os_error().raw_os_error();
assert_eq!(errno, Some(libc::ESRCH), "expected ESRCH, got {errno:?}");
}
/// Nonzero child exits surface as statuses, not hangs or panics.
#[test]
fn nonzero_exit_is_reported() {
maybe_child(ROLES);
let mut node = spawn_node("fail", &[]);
let status = node.wait_exit();
assert_eq!(status.code(), Some(7));
}
+250
View File
@@ -0,0 +1,250 @@
//! RFC 010 c4 — subprocess multi-node test harness.
//!
//! The runtime is a process singleton, so two real nodes means two
//! processes. This harness re-execs the *current test binary* as node
//! processes (precedent: tests/stack_diag.rs), tails their output live,
//! waits on protocol-visible lines, and reaps reliably no matter how the
//! test dies.
//!
//! Usage, per test file:
//!
//! - Declare roles as plain `fn()`s. A role prints protocol-visible facts
//! as single lines (Rust's piped stdout is line-buffered, so `println!`
//! is enough) and exits.
//! - **Every** `#[test]` in the file starts with
//! [`maybe_child`]`(ROLES)` — in the child re-exec, whichever test
//! libtest runs first performs the role and exits before the rest of the
//! suite runs (children are spawned with `--test-threads=1 --quiet`).
//! - The parent side spawns nodes with [`spawn_node`], waits on lines with
//! [`Node::wait_line`], and on exits with [`Node::wait_exit`].
//!
//! Port assignment: children bind port 0 and *report* the concrete address
//! (e.g. `LISTENING 127.0.0.1:41733`) rather than the parent pre-picking a
//! port — no bind/steal race by construction.
//!
//! Reaping: [`Node`]'s `Drop` SIGKILLs and `wait(2)`s the child, so a
//! panicking test (including a `wait_line` timeout) leaves no orphan and
//! no zombie. Tail threads exit on pipe EOF.
//!
//! Flake budget (explicit, per roadmap): every wait is bounded by
//! [`WAIT`] (10 s) against a typical cost of well under 1 s; the smoke
//! suite ran 10/10 clean at authoring time. Treat >1 failure in 100 runs
//! as a harness or runtime regression, not weather. On timeout the panic
//! message carries the node's full transcript so far.
#![allow(dead_code)] // Reusable surface: later phases use more of it than any one file.
use std::env;
use std::io::{BufRead, BufReader};
use std::process::{Child, Command, ExitStatus, Stdio};
use std::sync::mpsc::{Receiver, RecvTimeoutError};
use std::time::{Duration, Instant};
/// Env var selecting the child role in a re-exec.
const ROLE_ENV: &str = "SMARM_TWO_NODE_ROLE";
/// Upper bound for every wait in the harness. See the flake budget above.
pub const WAIT: Duration = Duration::from_secs(10);
/// In the child re-exec: run the matching role and exit. In the parent (no
/// role env set): return immediately. Call this first in every `#[test]` of
/// any file using the harness, passing the file's full role table.
pub fn maybe_child(roles: &[(&str, fn())]) {
let role = match env::var(ROLE_ENV) {
Ok(r) => r,
Err(_) => return,
};
for (name, f) in roles {
if *name == role {
f();
std::process::exit(0);
}
}
eprintln!("two_node harness: unknown role {role:?}");
std::process::exit(2);
}
/// One spawned node process with live-tailed output.
pub struct Node {
/// Role name, for panic messages.
pub role: String,
child: Option<Child>,
stdout_rx: Receiver<String>,
stderr_rx: Receiver<String>,
/// Every line consumed from stdout/stderr so far, for failure dumps.
transcript: Vec<String>,
}
fn tail(stream: impl std::io::Read + Send + 'static, prefix: &'static str) -> Receiver<String> {
let (tx, rx) = std::sync::mpsc::channel();
std::thread::spawn(move || {
for line in BufReader::new(stream).lines() {
let line = match line {
Ok(l) => l,
Err(_) => break,
};
// Receiver gone (Node dropped): stop tailing.
if tx.send(format!("{prefix}{line}")).is_err() {
break;
}
}
});
rx
}
/// Re-exec the current test binary as `role`, with any extra env vars.
pub fn spawn_node(role: &str, extra_env: &[(&str, &str)]) -> Node {
let exe = env::current_exe().expect("current_exe");
let mut cmd = Command::new(exe);
cmd.env(ROLE_ENV, role)
// --test-threads=1: exactly one test fn starts, hits maybe_child,
// and becomes the role. --nocapture: libtest must not swallow the
// role's println! lines — the parent tails them live.
.args(["--test-threads=1", "--quiet", "--nocapture"])
.stdout(Stdio::piped())
.stderr(Stdio::piped());
for (k, v) in extra_env {
cmd.env(k, v);
}
let mut child = cmd.spawn().expect("failed to spawn node process");
let stdout_rx = tail(child.stdout.take().expect("piped stdout"), "");
let stderr_rx = tail(child.stderr.take().expect("piped stderr"), "[stderr] ");
Node {
role: role.to_string(),
child: Some(child),
stdout_rx,
stderr_rx,
transcript: Vec::new(),
}
}
impl Node {
fn drain_stderr(&mut self) {
while let Ok(l) = self.stderr_rx.try_recv() {
self.transcript.push(l);
}
}
fn dump(&self) -> String {
if self.transcript.is_empty() {
"<no output>".to_string()
} else {
self.transcript.join("\n")
}
}
/// Every stdout/stderr line seen so far, in arrival order. For
/// ordering-proof assertions ("by the time X arrived, Y had not").
#[allow(dead_code)]
pub fn transcript(&self) -> &[String] {
&self.transcript
}
/// Wait until a stdout line satisfies `pred`; return it. Panics with the
/// full transcript after [`WAIT`]. `what` names the expectation in the
/// panic message.
pub fn wait_line(&mut self, what: &str, pred: impl Fn(&str) -> bool) -> String {
let deadline = Instant::now() + WAIT;
loop {
self.drain_stderr();
let left = deadline.saturating_duration_since(Instant::now());
match self.stdout_rx.recv_timeout(left) {
Ok(line) => {
self.transcript.push(line.clone());
if pred(&line) {
return line;
}
}
Err(RecvTimeoutError::Timeout) => {
// Pull in whatever stderr arrived since the last drain,
// so a role's eprintln! diagnostics survive into the dump.
self.drain_stderr();
panic!(
"node {:?}: timed out waiting for {what} after {WAIT:?}; transcript:\n{}",
self.role,
self.dump()
);
}
Err(RecvTimeoutError::Disconnected) => {
self.drain_stderr();
panic!(
"node {:?}: output closed while waiting for {what}; transcript:\n{}",
self.role,
self.dump()
);
}
}
}
}
/// Shorthand: wait for a `LISTENING <addr>` announcement, return the addr.
pub fn wait_listening(&mut self) -> String {
let line = self.wait_line("LISTENING announcement", |l| l.starts_with("LISTENING "));
line["LISTENING ".len()..].to_string()
}
/// Wait for the process to exit; panics with the transcript on timeout.
pub fn wait_exit(&mut self) -> ExitStatus {
let deadline = Instant::now() + WAIT;
loop {
let polled = match self.child.as_mut() {
Some(c) => c.try_wait(),
None => panic!("node {:?}: already reaped", self.role),
};
match polled {
Ok(Some(status)) => {
// Drain remaining output into the transcript for dumps.
self.drain_stderr();
while let Ok(l) = self.stdout_rx.try_recv() {
self.transcript.push(l);
}
self.child = None;
return status;
}
Ok(None) => {
if Instant::now() >= deadline {
self.drain_stderr();
self.kill();
panic!(
"node {:?}: did not exit within {WAIT:?}; transcript:\n{}",
self.role,
self.dump()
);
}
std::thread::sleep(Duration::from_millis(10));
}
Err(e) => panic!("node {:?}: try_wait failed: {e}", self.role),
}
}
}
/// Wait for exit and require success, dumping the transcript otherwise.
pub fn wait_exit_ok(&mut self) {
let status = self.wait_exit();
assert!(
status.success(),
"node {:?}: exited with {status}; transcript:\n{}",
self.role,
self.dump()
);
}
/// The child's OS pid, if not yet reaped.
pub fn pid(&self) -> Option<u32> {
self.child.as_ref().map(Child::id)
}
/// SIGKILL + reap now (idempotent).
pub fn kill(&mut self) {
if let Some(mut child) = self.child.take() {
let _ = child.kill();
let _ = child.wait();
}
}
}
impl Drop for Node {
fn drop(&mut self) {
self.kill();
}
}
+10 -22
View File
@@ -1,9 +1,7 @@
//! Low-level context-switch tests. These poke `init_actor_stack` and the //! Low-level context-switch tests. These poke `init_actor_stack` and the
//! naked asm shims directly — no scheduler involved. //! naked asm shims directly — no scheduler involved.
use smarm::context::{ use smarm::context::{init_actor_stack, switch_to_actor, switch_to_scheduler};
get_actor_sp, init_actor_stack, set_actor_sp, switch_to_actor, switch_to_scheduler,
};
use smarm::stack::Stack; use smarm::stack::Stack;
use std::cell::Cell; use std::cell::Cell;
@@ -31,8 +29,7 @@ fn actor_runs_and_returns_to_scheduler() {
reset_log(); reset_log();
let stack = Stack::new(64 * 1024, 4096).unwrap(); let stack = Stack::new(64 * 1024, 4096).unwrap();
let sp = init_actor_stack(stack.top(), actor_simple); let sp = init_actor_stack(stack.top(), actor_simple);
set_actor_sp(sp); let _ = unsafe { switch_to_actor(sp) };
unsafe { switch_to_actor() };
assert_eq!(get_log(), 0x1); assert_eq!(get_log(), 0x1);
} }
@@ -48,12 +45,11 @@ fn actor_yields_and_resumes() {
reset_log(); reset_log();
let stack = Stack::new(64 * 1024, 4096).unwrap(); let stack = Stack::new(64 * 1024, 4096).unwrap();
let sp = init_actor_stack(stack.top(), actor_two_steps); let sp = init_actor_stack(stack.top(), actor_two_steps);
set_actor_sp(sp);
unsafe { switch_to_actor() }; let sp = unsafe { switch_to_actor(sp) };
assert_eq!(get_log(), 0x1, "after first resume"); assert_eq!(get_log(), 0x1, "after first resume");
unsafe { switch_to_actor() }; let _ = unsafe { switch_to_actor(sp) };
assert_eq!(get_log(), 0x1 | 0x2, "after second resume"); assert_eq!(get_log(), 0x1 | 0x2, "after second resume");
} }
@@ -96,10 +92,9 @@ extern "C-unwind" fn actor_reg_check() {
fn callee_saved_registers_survive_yield() { fn callee_saved_registers_survive_yield() {
let stack = Stack::new(64 * 1024, 4096).unwrap(); let stack = Stack::new(64 * 1024, 4096).unwrap();
let sp = init_actor_stack(stack.top(), actor_reg_check); let sp = init_actor_stack(stack.top(), actor_reg_check);
set_actor_sp(sp);
unsafe { unsafe {
switch_to_actor(); let sp = switch_to_actor(sp);
switch_to_actor(); let _ = switch_to_actor(sp);
} }
assert_eq!( assert_eq!(
REG_BEFORE.get().copied().unwrap(), REG_BEFORE.get().copied().unwrap(),
@@ -138,18 +133,11 @@ fn two_actors_dont_corrupt_each_other() {
let sp_a = init_actor_stack(stack_a.top(), actor_a); let sp_a = init_actor_stack(stack_a.top(), actor_a);
let sp_b = init_actor_stack(stack_b.top(), actor_b); let sp_b = init_actor_stack(stack_b.top(), actor_b);
set_actor_sp(sp_a); let sp_a = unsafe { switch_to_actor(sp_a) };
unsafe { switch_to_actor() }; let sp_b = unsafe { switch_to_actor(sp_b) };
let sp_a = get_actor_sp();
set_actor_sp(sp_b); let _ = unsafe { switch_to_actor(sp_a) };
unsafe { switch_to_actor() }; let _ = unsafe { switch_to_actor(sp_b) };
let sp_b = get_actor_sp();
set_actor_sp(sp_a);
unsafe { switch_to_actor() };
set_actor_sp(sp_b);
unsafe { switch_to_actor() };
assert_eq!(A_VAL.with(|c| c.get()), 0xA00D); assert_eq!(A_VAL.with(|c| c.get()), 0xA00D);
assert_eq!(B_VAL.with(|c| c.get()), 0xB00D); assert_eq!(B_VAL.with(|c| c.get()), 0xB00D);
+125
View File
@@ -0,0 +1,125 @@
//! Cross-thread wake: a thread that is *not* a smarm scheduler thread must be
//! able to wake (and stop) a parked actor.
//!
//! The gap this pins down: every off-runtime wake primitive (`unpark`,
//! `unpark_at`, `request_stop`) reaches the runtime through the `RUNTIME`
//! thread-local, which is `None` on any non-scheduler thread — so a wake
//! issued from a foreign OS thread is a silent no-op and the parked actor
//! sleeps forever. Both failure modes below manifest as `Runtime::run` never
//! returning, so each test is wrapped in a watchdog: a timeout is the failure.
//!
//! The fix mirrors RFC 018's IO backend — the waker reaches the runtime
//! through a `Weak<RuntimeInner>` it already holds (the receiver captures one
//! when it parks; `Runtime::handle()` hands one to an app thread).
use std::sync::mpsc;
use std::thread;
use std::time::Duration;
const WATCHDOG: Duration = Duration::from_secs(10);
/// Give the target actor time to actually park before the foreign thread pokes
/// it, so we exercise the *wake* of a parked actor rather than the entry-side
/// stop check.
const SETTLE: Duration = Duration::from_millis(200);
fn assert_send_sync<T: Send + Sync>() {}
/// A cross-thread `send` from a plain OS thread must wake a receiver parked in
/// `recv`. Under the thread-local-only wake path the send enqueues the message
/// but never wakes the receiver, so `recv` — and therefore `run` — hangs.
#[test]
fn foreign_thread_send_wakes_parked_receiver() {
let (done_tx, done_rx) = mpsc::channel();
thread::spawn(move || {
let rt = smarm::init(smarm::Config::exact(2));
rt.run(|| {
let (tx, rx) = smarm::channel::<u32>();
// Receiver actor: parks on recv until the foreign thread sends.
let h = smarm::spawn(move || {
assert_eq!(rx.recv().expect("recv"), 42);
});
// Foreign (non-scheduler) OS thread owns the Sender and sends
// after the receiver has parked.
let sender = thread::spawn(move || {
thread::sleep(SETTLE);
tx.send(42).expect("send");
});
let _ = h.join();
sender.join().expect("sender thread");
});
let _ = done_tx.send(());
});
done_rx
.recv_timeout(WATCHDOG)
.expect("run did not return: a foreign-thread send never woke the parked receiver");
}
/// A cross-thread `request_stop` through a `RuntimeHandle` must wake and stop a
/// parked actor. The actor parks on a long sleep (only a stop can end it); the
/// handle is grabbed before `run` and driven from a foreign thread.
#[test]
fn foreign_thread_request_stop_wakes_parked_actor() {
assert_send_sync::<smarm::RuntimeHandle>();
let rt = smarm::init(smarm::Config::exact(2));
let handle = rt.handle();
// Foreign thread: learn the target pid from inside the run, let it park,
// then stop it through the handle.
let (pid_tx, pid_rx) = mpsc::channel::<smarm::Pid>();
let stopper = thread::spawn(move || {
let pid = pid_rx.recv().expect("pid");
thread::sleep(SETTLE);
handle.request_stop(pid);
});
let (done_tx, done_rx) = mpsc::channel();
thread::spawn(move || {
rt.run(move || {
let h = smarm::spawn(|| {
// Parks indefinitely; only a cooperative stop unwinds it.
smarm::sleep(Duration::from_secs(3600));
});
pid_tx.send(h.pid()).expect("send pid");
let _ = h.join();
});
let _ = done_tx.send(());
});
done_rx
.recv_timeout(WATCHDOG)
.expect("run did not return: a foreign-thread request_stop never woke the parked actor");
stopper.join().expect("stopper thread");
}
/// A `RuntimeHandle` held across (and beyond) a run must not keep the runtime
/// alive or block all-done: `run` still returns, and once the `Runtime` is
/// dropped the handle degrades to a harmless no-op (Weak lifecycle) rather than
/// panicking or touching freed memory.
#[test]
fn lingering_handle_does_not_block_all_done() {
let rt = smarm::init(smarm::Config::exact(1));
let handle = rt.handle(); // outlives the run below
let (pid_tx, pid_rx) = mpsc::channel::<smarm::Pid>();
let (done_tx, done_rx) = mpsc::channel();
let runner = thread::spawn(move || {
rt.run(move || {
let h = smarm::spawn(|| {});
pid_tx.send(h.pid()).expect("send pid");
let _ = h.join();
});
// `rt` is dropped here, at the end of this thread.
let _ = done_tx.send(());
});
done_rx
.recv_timeout(WATCHDOG)
.expect("run did not return while a RuntimeHandle was held live");
runner.join().expect("runner thread");
// Runtime is now dropped. A stop through the lingering handle must be a
// silent no-op, not a panic or use-after-free.
let dead_pid = pid_rx.recv().expect("pid");
handle.request_stop(dead_pid);
}
+20 -9
View File
@@ -88,7 +88,7 @@ impl GenServer for Lifecycle {
} }
} }
// init -> handle_call -> (drop last ref closes inbox) -> terminate. // init -> handle_call -> shutdown -> terminate.
#[test] #[test]
fn init_and_terminate_run() { fn init_and_terminate_run() {
let log = Arc::new(Mutex::new(Vec::new())); let log = Arc::new(Mutex::new(Vec::new()));
@@ -96,9 +96,9 @@ fn init_and_terminate_run() {
run(move || { run(move || {
let server = start(Lifecycle { log: log2 }); let server = start(Lifecycle { log: log2 });
server.call(()).unwrap(); server.call(()).unwrap();
// Dropping the only ref closes the inbox; the server breaks out of its // Refs are addresses: dropping one does not end the server. The
// recv loop and runs terminate. run() will not return until it has. // explicit close does, and waits for terminate.
drop(server); server.shutdown();
}); });
assert_eq!(*log.lock().unwrap(), vec!["init", "call", "terminate"]); assert_eq!(*log.lock().unwrap(), vec!["init", "call", "terminate"]);
} }
@@ -396,8 +396,11 @@ impl GenServer for Pool {
} }
// A worker spawned and watched from inside a handler delivers its Down to // A worker spawned and watched from inside a handler delivers its Down to
// handle_down. Down arms outrank the inbox, so the death is in the log by // handle_down. Nothing orders the worker's death before the follow-up
// the time the follow-up call is answered. // call (the worker sits in the shared queue while call/reply wakes ride the
// wake slot), so poll: the log must become exactly [Panic] within a bounded
// number of yields. Down-outranks-inbox ordering is covered by
// `watch_dead_pid_is_noproc_down`.
#[test] #[test]
fn worker_pool_down_reaches_handle_down() { fn worker_pool_down_reaches_handle_down() {
let got = Arc::new(Mutex::new(Vec::new())); let got = Arc::new(Mutex::new(Vec::new()));
@@ -409,7 +412,15 @@ fn worker_pool_down_reaches_handle_down() {
}); });
server.cast(PoolCast::SpawnDoomedWorker).unwrap(); server.cast(PoolCast::SpawnDoomedWorker).unwrap();
let _ = server.call(()).unwrap(); // sync point: cast handled, worker live let _ = server.call(()).unwrap(); // sync point: cast handled, worker live
*got2.lock().unwrap() = server.call(()).unwrap(); let mut log = Vec::new();
for _ in 0..10_000 {
log = server.call(()).unwrap();
if !log.is_empty() {
break;
}
smarm::yield_now();
}
*got2.lock().unwrap() = log;
}); });
assert_eq!(*got.lock().unwrap(), vec![DownReason::Panic]); assert_eq!(*got.lock().unwrap(), vec![DownReason::Panic]);
} }
@@ -702,8 +713,8 @@ fn no_timer_survives_exit() {
let _ = server.call(()).unwrap(); // sync: periodic armed let _ = server.call(()).unwrap(); // sync: periodic armed
smarm::sleep(Duration::from_millis(45)); // a couple of ticks smarm::sleep(Duration::from_millis(45)); // a couple of ticks
let mon = smarm::monitor(server.pid()); let mon = smarm::monitor(server.pid());
drop(server); // inbox closes → loop exits → guard drains timers server.shutdown(); // loop exits → guard drains timers
// Clean Down ⇒ the loop returned without the no-leak assert aborting. // Clean Down ⇒ the loop returned without the no-leak assert aborting.
assert!(mon.rx.recv().is_ok()); assert!(mon.rx.recv().is_ok());
let at_exit = f_read.lock().unwrap().len(); let at_exit = f_read.lock().unwrap().len();
smarm::sleep(Duration::from_millis(90)); // would be several more ticks smarm::sleep(Duration::from_millis(90)); // would be several more ticks
+243
View File
@@ -0,0 +1,243 @@
//! gen_server lifetime is the actor's, not its refs' (OTP: a pid is an
//! address, a process lives until it stops, is shut down, or is killed).
//!
//! - Dropping the last `GenServerRef` does NOT end the server. It ends via
//! `StopHandle::stop`, `request_shutdown` / `GenServerRef::shutdown`,
//! `request_stop`, or a handler panic.
//! - `GenServerBuilder::named(N).run()` runs the loop inline as the *current*
//! actor, so a server is a direct `ChildSpec` child: the supervisor's
//! shutdown reaches it as `handle_shutdown`, a restart re-binds the name,
//! and by-name `call`/`cast` reach whichever incarnation is live.
use smarm::gen_server::{
self, GenServer, GenServerBuilder, GenServerCtx, GenServerName, ShutdownAction, StopHandle,
};
use smarm::registry::RegisterError;
use smarm::supervisor::{ChildSpec, OneForOne, Restart, Shutdown};
use smarm::{monitor, request_shutdown, run, sleep, spawn, DownReason};
use std::sync::atomic::{AtomicUsize, Ordering};
use std::sync::{Arc, Mutex};
use std::time::Duration;
#[derive(Default, Clone)]
struct Log(Arc<Mutex<Vec<String>>>);
impl Log {
fn push(&self, s: impl Into<String>) {
self.0.lock().unwrap().push(s.into());
}
fn get(&self) -> Vec<String> {
self.0.lock().unwrap().clone()
}
}
struct Counter {
log: Log,
n: u64,
trap: bool,
stop: Option<StopHandle<Counter>>,
}
enum Call {
Get,
}
enum Cast {
Inc,
Stop,
}
impl GenServer for Counter {
type Call = Call;
type Reply = u64;
type Cast = Cast;
type Info = ();
type Timer = ();
fn init(&mut self, ctx: &GenServerCtx<Self>) {
if self.trap {
ctx.trap_exit();
}
self.stop = Some(ctx.stop_handle());
self.log.push("init");
}
fn handle_call(&mut self, Call::Get: Call) -> u64 {
self.n
}
fn handle_cast(&mut self, c: Cast) {
match c {
Cast::Inc => self.n += 1,
Cast::Stop => self.stop.as_ref().unwrap().stop(),
}
}
fn handle_shutdown(&mut self) -> ShutdownAction {
self.log.push("handle_shutdown");
ShutdownAction::Exit
}
fn terminate(&mut self) {
self.log.push("terminate");
}
}
fn counter(log: &Log, trap: bool) -> Counter {
Counter {
log: log.clone(),
n: 0,
trap,
stop: None,
}
}
// ---------------------------------------------------------------------------
// Refs are addresses: dropping the last one does not end the server.
// ---------------------------------------------------------------------------
#[test]
fn dropping_last_ref_does_not_end_server() {
let log = Log::default();
let l = log.clone();
run(move || {
let srv = gen_server::start(counter(&l, true));
let pid = srv.pid();
srv.cast(Cast::Inc).unwrap();
assert_eq!(srv.call(Call::Get).unwrap(), 1);
let mon = monitor(pid);
drop(srv);
sleep(Duration::from_millis(30));
assert!(
mon.rx.try_recv().unwrap().is_none(),
"server must outlive its last ref"
);
assert_eq!(l.get(), vec!["init"], "terminate must not have run");
// Explicit teardown still works, and is what ends it.
request_shutdown(pid);
let down = mon.rx.recv().unwrap();
assert_eq!(down.reason, DownReason::Exit);
});
assert_eq!(log.get(), vec!["init", "handle_shutdown", "terminate"]);
}
#[test]
fn ref_shutdown_is_the_explicit_close() {
let log = Log::default();
let l = log.clone();
run(move || {
let srv = gen_server::start(counter(&l, true));
srv.call(Call::Get).unwrap(); // sync: init (and trap_exit) has run
srv.shutdown(); // graceful, waits
assert_eq!(l.get(), vec!["init", "handle_shutdown", "terminate"]);
});
}
#[test]
fn forgotten_server_is_shut_down_at_root_exit() {
// A ref-less server is not a hung run: root exit shuts it down.
let log = Log::default();
let l = log.clone();
run(move || {
let srv = gen_server::start(counter(&l, false));
drop(srv);
});
assert_eq!(log.get(), vec!["init", "terminate"]);
}
// ---------------------------------------------------------------------------
// Inline run: a gen_server as a direct ChildSpec child.
// ---------------------------------------------------------------------------
const COUNTER: GenServerName<Counter> = GenServerName::new("lifetime-counter");
#[test]
fn named_run_is_a_direct_supervised_child_and_gets_shutdown() {
let log = Log::default();
let l = log.clone();
run(move || {
let l2 = l.clone();
let sup = spawn(move || {
let l3 = l2.clone();
OneForOne::new()
.child(
ChildSpec::new(Restart::Permanent, move || {
GenServerBuilder::new(counter(&l3, true))
.named(COUNTER)
.run()
.expect("name free");
})
.shutdown(Shutdown::Infinity),
)
.run();
});
sleep(Duration::from_millis(10));
gen_server::cast(COUNTER, Cast::Inc).unwrap();
assert_eq!(gen_server::call(COUNTER, Call::Get).unwrap(), 1);
request_shutdown(sup.pid());
sup.join()
.expect("ordered shutdown, supervisor returns normally");
assert_eq!(l.get(), vec!["init", "handle_shutdown", "terminate"]);
assert!(gen_server::whereis_server(COUNTER).is_none());
});
}
#[test]
fn named_run_child_restarts_and_rebinds_name() {
let log = Log::default();
let inits = Arc::new(AtomicUsize::new(0));
let l = log.clone();
let i = inits.clone();
run(move || {
let l2 = l.clone();
let i2 = i.clone();
let sup = spawn(move || {
let l3 = l2.clone();
let i3 = i2.clone();
OneForOne::new()
.child(ChildSpec::new(Restart::Permanent, move || {
i3.fetch_add(1, Ordering::SeqCst);
GenServerBuilder::new(counter(&l3, false))
.named(COUNTER)
.run()
.expect("name free on (re)start");
}))
.run();
});
sleep(Duration::from_millis(10));
gen_server::cast(COUNTER, Cast::Inc).unwrap();
assert_eq!(gen_server::call(COUNTER, Call::Get).unwrap(), 1);
// Normal self-exit → Permanent restarts it, fresh state, same name.
gen_server::cast(COUNTER, Cast::Stop).unwrap();
sleep(Duration::from_millis(30));
assert_eq!(i.load(Ordering::SeqCst), 2, "restarted once");
assert_eq!(gen_server::call(COUNTER, Call::Get).unwrap(), 0);
request_shutdown(sup.pid());
sup.join().unwrap();
});
assert_eq!(log.get(), vec!["init", "terminate", "init", "terminate"]);
}
#[test]
fn named_run_name_clash_fails_before_init() {
let log = Log::default();
let l = log.clone();
run(move || {
let first = GenServerBuilder::new(counter(&l, false))
.named(COUNTER)
.start()
.unwrap();
let l2 = l.clone();
let res = Arc::new(Mutex::new(None));
let r2 = res.clone();
let first_pid = first.pid();
spawn(move || {
let r = GenServerBuilder::new(counter(&l2, false))
.named(COUNTER)
.run();
*r2.lock().unwrap() = Some(r);
})
.join()
.unwrap();
assert_eq!(
*res.lock().unwrap(),
Some(Err(RegisterError::NameTaken { holder: first_pid }))
);
assert_eq!(l.get(), vec!["init"], "clashing server never ran init");
first.shutdown();
});
}
+222
View File
@@ -0,0 +1,222 @@
//! gen_server graceful shutdown.
//!
//! - A server that does not opt in (`ctx.trap_exit()` in `init`) is stopped
//! outright by `request_shutdown`, exactly as by `request_stop`.
//! - A trapping server receives the request as `handle_shutdown`. The default
//! returns `ShutdownAction::Exit`: the loop breaks and `terminate` runs on
//! the normal (non-unwind) path, so it may block. `Continue` keeps the loop
//! dispatching; the state later ends itself with a `StopHandle` — the only
//! way for a gen_server to exit *normally* on its own (`request_stop` on
//! self is an abnormal `Stopped`, which `Transient` restarts).
//! - Other exit signals (linked peers dying) reach a trapping server via
//! `handle_exit`.
use smarm::gen_server::{
start, GenServer, GenServerBuilder, GenServerCtx, GenServerRef, ShutdownAction, StopHandle,
};
use smarm::supervisor::{ChildSpec, OneForOne, Restart};
use smarm::{link, monitor, request_shutdown, run, self_pid, sleep, spawn, DownReason, ExitSignal};
use std::sync::atomic::{AtomicBool, AtomicUsize, Ordering};
use std::sync::{Arc, Mutex};
use std::time::Duration;
#[derive(Default, Clone)]
struct Log {
events: Arc<Mutex<Vec<&'static str>>>,
}
impl Log {
fn push(&self, e: &'static str) {
self.events.lock().unwrap().push(e);
}
fn get(&self) -> Vec<&'static str> {
self.events.lock().unwrap().clone()
}
}
/// A server with configurable shutdown behaviour.
struct Srv {
log: Log,
trap: bool,
action: ShutdownAction,
stop: Option<StopHandle<Srv>>,
exits: Arc<Mutex<Vec<ExitSignal>>>,
}
impl Srv {
fn new(log: &Log, trap: bool, action: ShutdownAction) -> Self {
Srv {
log: log.clone(),
trap,
action,
stop: None,
exits: Default::default(),
}
}
}
enum Cast {
Note(&'static str),
StopNow,
}
impl GenServer for Srv {
type Call = ();
type Reply = ();
type Cast = Cast;
type Info = ();
type Timer = ();
fn init(&mut self, ctx: &GenServerCtx<Self>) {
if self.trap {
ctx.trap_exit();
}
self.stop = Some(ctx.stop_handle());
}
fn handle_call(&mut self, _: ()) {}
fn handle_cast(&mut self, c: Cast) {
match c {
Cast::Note(s) => self.log.push(s),
Cast::StopNow => self.stop.as_ref().unwrap().stop(),
}
}
fn handle_shutdown(&mut self) -> ShutdownAction {
self.log.push("handle_shutdown");
self.action
}
fn handle_exit(&mut self, sig: ExitSignal) {
self.log.push("handle_exit");
self.exits.lock().unwrap().push(sig);
}
fn terminate(&mut self) {
// Allowed to block on the graceful path.
if self.trap {
sleep(Duration::from_millis(10));
}
self.log.push("terminate");
}
}
fn spawn_settled<G: GenServer>(state: G) -> GenServerRef<G> {
let r = start(state);
sleep(Duration::from_millis(20)); // let init (trap_exit) run
r
}
#[test]
fn non_trapping_server_is_stopped_outright() {
let log = Log::default();
let l = log.clone();
run(move || {
let r = spawn_settled(Srv::new(&l, false, ShutdownAction::Exit));
let mon = monitor(r.pid());
request_shutdown(r.pid());
let d = mon.rx.recv().unwrap();
assert_eq!(d.reason, DownReason::Stopped);
});
assert_eq!(log.get(), vec!["terminate"]);
}
#[test]
fn trapping_server_exits_normally_via_handle_shutdown() {
let log = Log::default();
let l = log.clone();
run(move || {
let r = spawn_settled(Srv::new(&l, true, ShutdownAction::Exit));
let mon = monitor(r.pid());
request_shutdown(r.pid());
let d = mon.rx.recv().unwrap();
assert_eq!(d.reason, DownReason::Exit);
});
assert_eq!(log.get(), vec!["handle_shutdown", "terminate"]);
}
#[test]
fn continue_keeps_dispatching_until_stop_handle() {
let log = Log::default();
let l = log.clone();
run(move || {
let r = spawn_settled(Srv::new(&l, true, ShutdownAction::Continue));
let mon = monitor(r.pid());
request_shutdown(r.pid());
sleep(Duration::from_millis(20));
r.cast(Cast::Note("after-shutdown-request")).unwrap();
r.cast(Cast::StopNow).unwrap();
let d = mon.rx.recv().unwrap();
assert_eq!(d.reason, DownReason::Exit);
});
assert_eq!(
log.get(),
vec!["handle_shutdown", "after-shutdown-request", "terminate"]
);
}
#[test]
fn stop_handle_is_a_normal_exit_that_transient_does_not_restart() {
let starts = Arc::new(AtomicUsize::new(0));
let s = starts.clone();
run(move || {
let s2 = s.clone();
let sup = spawn(move || {
let s3 = s2.clone();
OneForOne::new()
.child(ChildSpec::new(Restart::Transient, move || {
s3.fetch_add(1, Ordering::SeqCst);
let log = Log::default();
let r = GenServerBuilder::new(Srv::new(&log, false, ShutdownAction::Exit))
.under(self_pid())
.start();
r.cast(Cast::StopNow).unwrap();
// Block until the server is gone; a bare spawn parent
// returning would not itself end the server.
let mon = monitor(r.pid());
let _ = mon.rx.recv();
}))
.run();
});
sup.join().unwrap(); // returns only if the child was not restarted forever
});
assert_eq!(starts.load(Ordering::SeqCst), 1);
}
#[test]
fn linked_peer_death_reaches_handle_exit() {
let log = Log::default();
let l = log.clone();
let alive = Arc::new(AtomicBool::new(false));
let a = alive.clone();
run(move || {
let r = spawn_settled(Srv::new(&l, true, ShutdownAction::Exit));
let pid = r.pid();
let peer = spawn(move || {
link(pid);
panic!("peer dies");
});
let _ = peer.join();
sleep(Duration::from_millis(20));
r.cast(Cast::Note("still-serving")).unwrap();
sleep(Duration::from_millis(20));
a.store(true, Ordering::SeqCst);
r.shutdown(); // graceful; waits for terminate
});
assert!(alive.load(Ordering::SeqCst));
assert_eq!(
log.get(),
vec![
"handle_exit",
"still-serving",
"handle_shutdown",
"terminate"
]
);
}
#[test]
fn gen_server_ref_shutdown_is_graceful_for_a_trapping_server() {
let log = Log::default();
let l = log.clone();
run(move || {
let r = spawn_settled(Srv::new(&l, true, ShutdownAction::Exit));
r.shutdown();
});
assert_eq!(log.get(), vec!["handle_shutdown", "terminate"]);
}
+128
View File
@@ -0,0 +1,128 @@
//! gen_statem lifetime parity with gen_server: a machine lives until it
//! stops, is shut down, or is killed — its refs are addresses. And
//! `gen_statem::run_named` runs a machine inline as the current actor, so it
//! is a direct `ChildSpec` child addressed by name.
use smarm::gen_statem::{self, GenStatemName, Reply};
use smarm::supervisor::{ChildSpec, OneForOne, Restart, Shutdown};
use smarm::{monitor, request_shutdown, run, sleep, spawn, DownReason};
use std::sync::{Arc, Mutex};
use std::time::Duration;
#[derive(Default, Clone)]
struct Log(Arc<Mutex<Vec<&'static str>>>);
impl Log {
fn push(&self, e: &'static str) {
self.0.lock().unwrap().push(e);
}
fn get(&self) -> Vec<&'static str> {
self.0.lock().unwrap().clone()
}
}
#[derive(Clone, Copy, PartialEq, Eq, Debug)]
enum S {
On,
}
struct D {
log: Log,
trap: bool,
n: u64,
}
enum Cast {
Inc,
StopNow,
}
enum Call {
Get(Reply<u64>),
}
smarm::gen_statem! {
machine: Sm { state: S, data: D };
event: Ev { cast: Cast, call: Call, info: () };
context(data, prev, cx);
enter {
S::On => { data.log.push("enter"); if data.trap { cx.trap_exit() } },
}
on S::On => {
cast Cast::Inc => { data.n += 1; prev },
cast Cast::StopNow => stop,
call Call::Get(r) => { r.reply(data.n); prev },
shutdown => { data.log.push("shutdown"); cx.stop(); prev },
state_timeout => unhandled,
timeout _ => unhandled,
}
terminate { data.log.push("terminate"); }
}
fn d(log: &Log, trap: bool) -> D {
D {
log: log.clone(),
trap,
n: 0,
}
}
#[test]
fn dropping_last_ref_does_not_end_machine() {
let log = Log::default();
let l = log.clone();
run(move || {
let m = Sm::start(S::On, d(&l, true));
let pid = m.pid();
m.send(Ev::Cast(Cast::Inc)).unwrap();
assert_eq!(m.call(|r| Ev::Call(Call::Get(r))).unwrap(), 1);
let mon = monitor(pid);
drop(m);
sleep(Duration::from_millis(30));
assert!(
mon.rx.try_recv().unwrap().is_none(),
"machine must outlive its refs"
);
assert_eq!(l.get(), vec!["enter"]);
request_shutdown(pid);
assert_eq!(mon.rx.recv().unwrap().reason, DownReason::Exit);
});
assert_eq!(log.get(), vec!["enter", "shutdown", "terminate"]);
}
const SM: GenStatemName<Sm> = GenStatemName::new("lifetime-sm");
#[test]
fn run_named_is_a_direct_supervised_child() {
let log = Log::default();
let l = log.clone();
run(move || {
let l2 = l.clone();
let sup = spawn(move || {
let l3 = l2.clone();
OneForOne::new()
.child(
ChildSpec::new(Restart::Permanent, move || {
gen_statem::run_named(SM, Sm::new(S::On, d(&l3, true))).expect("name free");
})
.shutdown(Shutdown::Infinity),
)
.run();
});
sleep(Duration::from_millis(10));
gen_statem::send(SM, Ev::Cast(Cast::Inc)).unwrap();
assert_eq!(gen_statem::call(SM, |r| Ev::Call(Call::Get(r))).unwrap(), 1);
// Normal self-exit → Permanent restart → fresh data, same name.
gen_statem::send(SM, Ev::Cast(Cast::StopNow)).unwrap();
sleep(Duration::from_millis(30));
assert_eq!(gen_statem::call(SM, |r| Ev::Call(Call::Get(r))).unwrap(), 0);
request_shutdown(sup.pid());
sup.join().unwrap();
assert!(gen_statem::whereis_machine(SM).is_none());
});
assert_eq!(
log.get(),
vec!["enter", "terminate", "enter", "shutdown", "terminate"]
);
}
+219
View File
@@ -0,0 +1,219 @@
//! gen_statem graceful shutdown — the gen_server surface, in state-machine
//! clothes. Where gen_server routes a shutdown request to a `handle_shutdown`
//! method, a gen_statem gets it as an **event** so it can be routed by state:
//!
//! - A machine that does not opt in (`cx.trap_exit()` in the initial `enter`)
//! is stopped outright by `request_shutdown`, exactly as by `request_stop`.
//! - A trapping machine sees the request as a `shutdown` row (a unit event
//! like `state_timeout`). The macro's default, when a state writes no
//! `shutdown` row, is `stop` — the loop breaks and `terminate` runs on the
//! normal path. A row may instead transition (e.g. into a Draining state)
//! and `stop` later from any row via the `stop` tail keyword.
//! - Linked-peer deaths reach a trapping machine as `exit <pat>` rows; an
//! unmatched exit is silently dropped, like an unmatched info.
//! - `terminate { … }` is an optional macro block, run on every exit path.
use smarm::gen_statem;
use smarm::gen_statem::{GenStatemRef, Reply};
use smarm::{link, monitor, request_shutdown, run, sleep, spawn, DownReason, ExitSignal};
use std::sync::{Arc, Mutex};
use std::time::Duration;
#[derive(Default, Clone)]
struct Log(Arc<Mutex<Vec<&'static str>>>);
impl Log {
fn push(&self, e: &'static str) {
self.0.lock().unwrap().push(e);
}
fn get(&self) -> Vec<&'static str> {
self.0.lock().unwrap().clone()
}
}
#[derive(Clone, Copy, PartialEq, Eq, Debug)]
enum S {
Idle,
Draining,
}
struct D {
log: Log,
trap: bool,
exits: Vec<ExitSignal>,
}
enum Cast {
Note(&'static str),
StopNow,
}
enum Call {
Exits(Reply<usize>),
}
gen_statem! {
machine: Sm { state: S, data: D };
event: Ev { cast: Cast, call: Call, info: () };
context(data, prev, cx);
enter {
S::Idle => if data.trap { cx.trap_exit() },
S::Draining => { data.log.push("draining"); cx.state_timeout(Duration::from_millis(30)); },
}
on S::Idle => {
// Shutdown in Idle: go drain first, stop later.
shutdown => S::Draining,
cast Cast::StopNow => stop,
state_timeout => unhandled,
}
on S::Draining => {
// Drained: end the machine normally.
state_timeout => { data.log.push("drained"); cx.stop(); prev },
// A second request while draining is ignored.
shutdown => unhandled,
cast Cast::StopNow => stop,
}
on _ => {
cast Cast::Note(s) => { data.log.push(s); prev },
call Call::Exits(r) => { r.reply(data.exits.len()); prev },
exit sig => { data.log.push("exit"); data.exits.push(sig); prev },
timeout _ => unhandled,
}
terminate {
data.log.push("terminate");
}
}
/// A machine with no `shutdown` rows at all: the macro default applies.
#[derive(Clone, Copy, PartialEq, Eq, Debug)]
enum P {
On,
}
struct PD {
log: Log,
}
enum PCast {}
enum PCall {}
gen_statem! {
machine: Plain { state: P, data: PD };
event: PEv { cast: PCast, call: PCall, info: () };
context(data, prev, cx);
enter { P::On => cx.trap_exit(), }
on P::On => {
cast _ => unhandled,
call _ => unhandled,
state_timeout => unhandled,
timeout _ => unhandled,
}
terminate { data.log.push("terminate"); }
}
fn settled(log: &Log, trap: bool) -> GenStatemRef<Sm> {
let r = Sm::start(
S::Idle,
D {
log: log.clone(),
trap,
exits: Vec::new(),
},
);
sleep(Duration::from_millis(20)); // let on_start (trap_exit) run
r
}
#[test]
fn non_trapping_machine_is_stopped_outright() {
let log = Log::default();
let l = log.clone();
run(move || {
let r = settled(&l, false);
let mon = monitor(r.pid());
request_shutdown(r.pid());
assert_eq!(mon.rx.recv().unwrap().reason, DownReason::Stopped);
});
assert_eq!(log.get(), vec!["terminate"]);
}
#[test]
fn shutdown_row_routes_by_state_and_stop_tail_exits_normally() {
let log = Log::default();
let l = log.clone();
run(move || {
let r = settled(&l, true);
let mon = monitor(r.pid());
request_shutdown(r.pid());
// The second request lands in Draining and is `unhandled` (ignored).
sleep(Duration::from_millis(5));
request_shutdown(r.pid());
assert_eq!(mon.rx.recv().unwrap().reason, DownReason::Exit);
});
assert_eq!(log.get(), vec!["draining", "drained", "terminate"]);
}
#[test]
fn default_shutdown_is_stop() {
let log = Log::default();
let l = log.clone();
run(move || {
let r = Plain::start(P::On, PD { log: l });
sleep(Duration::from_millis(20));
let mon = monitor(r.pid());
request_shutdown(r.pid());
assert_eq!(mon.rx.recv().unwrap().reason, DownReason::Exit);
});
assert_eq!(log.get(), vec!["terminate"]);
}
#[test]
fn stop_tail_from_a_cast_is_a_normal_exit() {
let log = Log::default();
let l = log.clone();
run(move || {
let r = settled(&l, false);
let mon = monitor(r.pid());
r.send(Ev::Cast(Cast::Note("a"))).unwrap();
r.send(Ev::Cast(Cast::StopNow)).unwrap();
r.send(Ev::Cast(Cast::Note("after-stop"))).unwrap(); // never dispatched
assert_eq!(mon.rx.recv().unwrap().reason, DownReason::Exit);
});
assert_eq!(log.get(), vec!["a", "terminate"]);
}
#[test]
fn linked_peer_death_reaches_exit_row() {
let log = Log::default();
let l = log.clone();
run(move || {
let r = settled(&l, true);
let pid = r.pid();
let peer = spawn(move || {
link(pid);
panic!("peer dies");
});
let _ = peer.join();
sleep(Duration::from_millis(20));
r.send(Ev::Cast(Cast::Note("still-running"))).unwrap();
assert_eq!(r.call(|r| Ev::Call(Call::Exits(r))).unwrap(), 1);
r.shutdown();
});
assert_eq!(
log.get(),
vec!["exit", "still-running", "draining", "drained", "terminate"]
);
}
#[test]
fn ref_shutdown_is_graceful_and_waits() {
let log = Log::default();
let l = log.clone();
run(move || {
let r = settled(&l, true);
r.shutdown();
// terminate has run by the time shutdown() returns.
assert_eq!(l.get(), vec!["draining", "drained", "terminate"]);
});
}
+60
View File
@@ -161,3 +161,63 @@ fn demonitor_after_fire_is_none() {
let _ = h.join(); let _ = h.join();
}); });
} }
// ---------------------------------------------------------------------------
// spawn_monitor: registration precedes publish, so a child that dies before
// the parent gets another instruction in still reports its real reason.
// ---------------------------------------------------------------------------
#[test]
fn spawn_monitor_never_reports_noproc_multi_thread() {
// 4 schedulers, 200 instantly-dying children. With spawn+monitor this
// observes NoProc a few percent of the time (the child finishes on
// another scheduler before monitor() registers); with spawn_monitor it
// must be Exit every time.
let noproc = Arc::new(AtomicUsize::new(0));
let exit = Arc::new(AtomicUsize::new(0));
let (n, e) = (noproc.clone(), exit.clone());
smarm::init(smarm::Config::exact(4)).run(move || {
let mut hs = Vec::new();
for _ in 0..200 {
let (h, m) = smarm::spawn_monitor(|| {});
let d = m.rx.recv().expect("Down");
assert_eq!(d.pid, h.pid());
match d.reason {
DownReason::Exit => e.fetch_add(1, Ordering::Relaxed),
DownReason::NoProc => n.fetch_add(1, Ordering::Relaxed),
other => panic!("unexpected {other:?}"),
};
hs.push(h);
}
for h in hs {
h.join().unwrap();
}
});
assert_eq!(noproc.load(Ordering::Relaxed), 0);
assert_eq!(exit.load(Ordering::Relaxed), 200);
}
#[test]
fn spawn_monitor_sees_panic_and_demonitor_works() {
let ok = Arc::new(AtomicBool::new(false));
let o = ok.clone();
run(move || {
let (h, m) = smarm::spawn_monitor(|| panic!("boom"));
let d = m.rx.recv().expect("Down");
assert_eq!(d.pid, h.pid());
assert!(matches!(d.reason, DownReason::Panic));
let _ = h.join();
// demonitor before the child runs: no Down ever arrives.
let (h2, m2) = smarm::spawn_monitor(|| {});
demonitor(&m2);
h2.join().unwrap();
// Last sender gone with nothing sent: the channel is closed and empty.
assert!(
!matches!(m2.rx.try_recv(), Ok(Some(_))),
"demonitored spawn_monitor still delivered"
);
o.store(true, Ordering::SeqCst);
});
assert!(ok.load(Ordering::SeqCst));
}
+13 -20
View File
@@ -1,5 +1,6 @@
//! Process-group tests that run under the scheduler: `join` installs a real //! Process-group tests that run under the scheduler: `join` installs a real
//! monitor on a live actor, and a real death drives eviction on next contact. //! monitor on a live actor, and a real death drives eviction (the reaper
//! actor sweeps it; the read path hides it in the meantime).
//! (Pure structural invariants live in the `pg` unit tests.) //! (Pure structural invariants live in the `pg` unit tests.)
use smarm::{channel, members, pick, run, spawn}; use smarm::{channel, members, pick, run, spawn};
@@ -35,19 +36,15 @@ fn a_dead_actor_vanishes_from_every_group_it_joined() {
assert_eq!(members("g1"), vec![pid]); assert_eq!(members("g1"), vec![pid]);
assert_eq!(members("g2"), vec![pid]); assert_eq!(members("g2"), vec![pid]);
// Release and reap the actor. finalize_actor queues the Down to our // Release the actor. finalize_actor queues the Down to the reaper and
// monitors before unparking joiners, so by the time join() returns the // marks the slot dead before unparking joiners, so by the time join()
// Down is already waiting in the membership channel. // returns every read hides the pid whether or not the reaper has run.
tx.send(()).unwrap(); tx.send(()).unwrap();
h.join().unwrap(); h.join().unwrap();
// Drain-on-contact: touching g1 detects the death and sweeps the pid // Gone from every group it joined, not just one.
// out of every group (g2 included), not just g1. assert!(members("g1").is_empty(), "gone from g1");
assert!(members("g1").is_empty(), "evicted from the touched group"); assert!(members("g2").is_empty(), "and from g2");
assert!(
members("g2").is_empty(),
"and swept from the untouched group"
);
assert_eq!(pick("g1"), None); assert_eq!(pick("g1"), None);
}); });
} }
@@ -86,11 +83,7 @@ fn live_members_survive_a_peers_death() {
tx_a.send(()).unwrap(); tx_a.send(()).unwrap();
a.join().unwrap(); a.join().unwrap();
assert_eq!( assert_eq!(members("svc"), vec![b.pid()], "only the dead peer is gone");
members("svc"),
vec![b.pid()],
"only the dead peer is reaped"
);
assert_eq!(pick("svc"), Some(b.pid())); assert_eq!(pick("svc"), Some(b.pid()));
tx_b.send(()).unwrap(); tx_b.send(()).unwrap();
@@ -123,18 +116,18 @@ fn leave_drops_a_membership_without_affecting_others() {
} }
#[test] #[test]
fn joining_an_already_dead_pid_is_evicted_on_next_contact() { fn joining_an_already_dead_pid_never_shows_in_a_read() {
run(|| { run(|| {
let h = spawn(|| {}); let h = spawn(|| {});
let pid = h.pid(); let pid = h.pid();
h.join().unwrap(); // actor is finalized before we join it to anything h.join().unwrap(); // actor is finalized before we join it to anything
// monitor() on a gone pid queues a NoProc Down immediately, so the // join() on a gone pid queues a NoProc Down to the reaper immediately;
// membership is reaped the next time the group is touched. // reads never show it either way (slot-liveness backstop).
join("late", pid); join("late", pid);
assert!( assert!(
members("late").is_empty(), members("late").is_empty(),
"dead-at-join member is reaped on read" "dead-at-join member never reads as live"
); );
assert_eq!(pick("late"), None); assert_eq!(pick("late"), None);
}); });
+43
View File
@@ -258,3 +258,46 @@ fn send_dyn_to_dead_pid_is_dead() {
assert!(matches!(send_dyn::<u64>(p, 1u64), Err(SendError::Dead(_)))); assert!(matches!(send_dyn::<u64>(p, 1u64), Err(SendError::Dead(_))));
}); });
} }
// --- one channel per message type per actor -----------------------------------
/// Registering a second name of the same message type on one actor, with a
/// *fresh* channel, would silently replace and close the first — so it
/// panics (found in RFC 010 c9). The sanctioned shapes stay quiet: bind both
/// names to a clone of one sender, or use two actors.
#[test]
#[should_panic(expected = "already publishes a live channel")]
fn second_live_channel_of_same_type_on_one_actor_panics() {
run(|| {
let (tx1, _rx1) = channel::<u64>();
let (tx2, _rx2) = channel::<u64>();
register(Name::<u64>::new("dup-a"), tx1).unwrap();
register(Name::<u64>::new("dup-b"), tx2).unwrap(); // panics
});
}
#[test]
fn two_names_on_one_cloned_sender_is_fine() {
run(|| {
let (tx, rx) = channel::<u64>();
register(Name::<u64>::new("twin-a"), tx.clone()).unwrap();
register(Name::<u64>::new("twin-b"), tx).unwrap();
send(Name::<u64>::new("twin-a"), 1).unwrap();
send(Name::<u64>::new("twin-b"), 2).unwrap();
assert_eq!(rx.recv().unwrap(), 1);
assert_eq!(rx.recv().unwrap(), 2);
});
}
#[test]
fn replacing_a_channel_whose_receiver_is_gone_is_fine() {
run(|| {
let (tx1, rx1) = channel::<u64>();
register(Name::<u64>::new("reborn"), tx1).unwrap();
drop(rx1); // old inbox gone: replacement is the honest thing to do
let (tx2, rx2) = channel::<u64>();
register(Name::<u64>::new("reborn-2"), tx2).unwrap();
send(Name::<u64>::new("reborn-2"), 9).unwrap();
assert_eq!(rx2.recv().unwrap(), 9);
});
}
+288
View File
@@ -0,0 +1,288 @@
//! Root exit — the run's initial actor returning means "the program is done".
//!
//! When the root finalizes, the runtime delivers `request_shutdown` to every
//! **forest root**: each live actor whose parent is the run itself (a plain
//! `spawn` from the root closure) or is already dead. Nothing below a live
//! parent is touched directly — a supervisor gets one request and runs its
//! own ordered shutdown per child `Shutdown` policy.
//!
//! - Non-trapping actors are stopped outright, exactly as by
//! `request_shutdown` — a `spawn(|| { sleep(..); work() })` the root did
//! not `join` does NOT get to finish. Join it, supervise it, or trap.
//! - Trapping actors get `handle_shutdown` / an `ExitSignal{Shutdown}` and
//! may keep running (`Continue`, drain, then stop themselves) — timers and
//! all; the run ends when they do. There is no second, forcing sweep.
//! - A periodic-timer daemon (the classic wedge) never blocks `run()`.
use smarm::gen_server::{start, GenServer, GenServerCtx, ShutdownAction, StopHandle, TimerHandle};
use smarm::supervisor::{ChildSpec, OneForOne, Restart, Shutdown};
use smarm::{run, sleep, spawn, trap_exit, DownReason};
use std::sync::atomic::{AtomicBool, AtomicUsize, Ordering};
use std::sync::{Arc, Mutex};
use std::time::{Duration, Instant};
fn assert_prompt(start: Instant, what: &str) {
assert!(
start.elapsed() < Duration::from_secs(2),
"{what}: run() took {:?}",
start.elapsed()
);
}
// ---------------------------------------------------------------------------
// Bare (non-gen_server) actors
// ---------------------------------------------------------------------------
/// A non-trapping sleeper the root did not join is stopped, not waited for.
#[test]
fn unjoined_non_trapping_sleeper_is_stopped() {
let finished = Arc::new(AtomicBool::new(false));
let f = finished.clone();
let t = Instant::now();
run(move || {
spawn(move || {
sleep(Duration::from_secs(5));
f.store(true, Ordering::SeqCst);
});
});
assert_prompt(t, "sleeper");
assert!(
!finished.load(Ordering::SeqCst),
"sleeper should have been stopped"
);
}
/// A trapping bare actor sees `Shutdown` from the root's exit and may keep
/// working — here it sleeps (a timer!) after the signal, then returns. The
/// run waits for it: no forcing sweep.
#[test]
fn trapping_actor_may_finish_after_shutdown_signal() {
let finished = Arc::new(AtomicBool::new(false));
let f = finished.clone();
run(move || {
spawn(move || {
let inbox = trap_exit();
let sig = inbox.recv().expect("shutdown signal");
assert_eq!(sig.reason, DownReason::Shutdown);
sleep(Duration::from_millis(100));
f.store(true, Ordering::SeqCst);
});
// Trapping is a runtime opt-in: give the actor a chance to run
// `trap_exit()`; a not-yet-run actor is non-trapping and is stopped.
sleep(Duration::from_millis(20));
});
assert!(
finished.load(Ordering::SeqCst),
"trapping actor must be allowed to finish"
);
}
/// The classic wedge: a lazily spawned daemon that never returns on its own.
#[test]
fn parked_forever_daemon_does_not_block_run() {
let t = Instant::now();
run(|| {
let (_tx, rx) = smarm::channel::<()>();
spawn(move || {
let _ = rx.recv(); // parked forever: sender is held by the root, which returns
});
// Leak the sender into the daemon's own scope so nothing else drops it.
std::mem::forget(_tx);
});
assert_prompt(t, "daemon");
}
// ---------------------------------------------------------------------------
// gen_servers
// ---------------------------------------------------------------------------
/// A non-trapping ticker with a periodic timer: the timer wheel is never empty,
/// and root exit must still end the run.
struct Ticker {
ticks: Arc<AtomicUsize>,
}
impl GenServer for Ticker {
type Call = ();
type Reply = ();
type Cast = ();
type Info = ();
type Timer = ();
fn init(&mut self, ctx: &GenServerCtx<Self>) {
ctx.timer().tick_every(Duration::from_millis(5), ());
}
fn handle_call(&mut self, _: ()) {}
fn handle_cast(&mut self, _: ()) {}
fn handle_timer(&mut self, _: ()) {
self.ticks.fetch_add(1, Ordering::SeqCst);
}
}
#[test]
fn periodic_timer_daemon_does_not_block_run() {
let ticks = Arc::new(AtomicUsize::new(0));
let tk = ticks.clone();
let t = Instant::now();
run(move || {
let _r = start(Ticker { ticks: tk });
sleep(Duration::from_millis(50));
});
assert_prompt(t, "ticker");
assert!(
ticks.load(Ordering::SeqCst) >= 3,
"ticker should have ticked"
);
}
/// A trapping server that answers `Continue`, keeps ticking on its own timer
/// (draining), and stops itself later. Root exit must not cut it short.
struct Drainer {
log: Arc<Mutex<Vec<&'static str>>>,
shutdowns: Arc<AtomicUsize>,
ticks_after_shutdown: usize,
stop: Option<StopHandle<Drainer>>,
timer: Option<TimerHandle<Drainer>>,
draining: bool,
}
impl GenServer for Drainer {
type Call = ();
type Reply = ();
type Cast = ();
type Info = ();
type Timer = ();
fn init(&mut self, ctx: &GenServerCtx<Self>) {
ctx.trap_exit();
self.stop = Some(ctx.stop_handle());
self.timer = Some(ctx.timer());
}
fn handle_call(&mut self, _: ()) {}
fn handle_cast(&mut self, _: ()) {}
fn handle_shutdown(&mut self) -> ShutdownAction {
self.shutdowns.fetch_add(1, Ordering::SeqCst);
self.log.lock().unwrap().push("handle_shutdown");
self.draining = true;
self.timer
.as_ref()
.unwrap()
.tick_every(Duration::from_millis(10), ());
ShutdownAction::Continue
}
fn handle_timer(&mut self, _: ()) {
if !self.draining {
return;
}
self.ticks_after_shutdown += 1;
if self.ticks_after_shutdown == 3 {
self.log.lock().unwrap().push("drained");
self.stop.as_ref().unwrap().stop();
}
}
fn terminate(&mut self) {
self.log.lock().unwrap().push("terminate");
}
}
fn drainer(log: &Arc<Mutex<Vec<&'static str>>>, shutdowns: &Arc<AtomicUsize>) -> Drainer {
Drainer {
log: log.clone(),
shutdowns: shutdowns.clone(),
ticks_after_shutdown: 0,
stop: None,
timer: None,
draining: false,
}
}
#[test]
fn trapping_server_drains_with_timers_after_root_exit() {
let log = Arc::new(Mutex::new(Vec::new()));
let shutdowns = Arc::new(AtomicUsize::new(0));
let (l, s) = (log.clone(), shutdowns.clone());
run(move || {
let r = start(drainer(&l, &s));
// A gen_server's lifetime is governed by its refs: dropping the last
// one closes the inbox and ends the loop cleanly, which would cut the
// drain short for a reason unrelated to root exit. Pin it the way a
// registered name would.
std::mem::forget(r);
sleep(Duration::from_millis(20)); // let init (trap_exit) run
});
assert_eq!(
*log.lock().unwrap(),
vec!["handle_shutdown", "drained", "terminate"]
);
assert_eq!(shutdowns.load(Ordering::SeqCst), 1);
}
// ---------------------------------------------------------------------------
// Supervision trees: only forest roots are addressed
// ---------------------------------------------------------------------------
/// The supervisor gets ONE request and runs its ordered shutdown; a trapping
/// child under it sees exactly one `Shutdown` — from the supervisor, not a
/// second one from the runtime — and is allowed to finish its drain (a sleep,
/// i.e. a timer) under `Shutdown::Infinity`.
#[test]
fn supervised_children_are_shut_down_only_via_their_supervisor() {
let signals = Arc::new(AtomicUsize::new(0));
let drained = Arc::new(AtomicBool::new(false));
let sup_returned = Arc::new(AtomicBool::new(false));
let (sg, dr, sr) = (signals.clone(), drained.clone(), sup_returned.clone());
run(move || {
spawn(move || {
OneForOne::new()
.child(
ChildSpec::new(Restart::Permanent, move || {
let inbox = trap_exit();
while let Ok(sig) = inbox.recv() {
if sig.reason == DownReason::Shutdown {
sg.fetch_add(1, Ordering::SeqCst);
sleep(Duration::from_millis(100));
// A late second signal would land here.
while let Ok(Some(sig)) = inbox.try_recv() {
if sig.reason == DownReason::Shutdown {
sg.fetch_add(1, Ordering::SeqCst);
}
}
dr.store(true, Ordering::SeqCst);
return;
}
}
})
.shutdown(Shutdown::Infinity),
)
.run();
sr.store(true, Ordering::SeqCst);
});
sleep(Duration::from_millis(30)); // let the tree settle
});
assert!(
sup_returned.load(Ordering::SeqCst),
"supervisor should return normally"
);
assert!(
drained.load(Ordering::SeqCst),
"child should finish its drain"
);
assert_eq!(signals.load(Ordering::SeqCst), 1);
}
/// A supervised non-trapping child under `Shutdown::Timeout` is stopped by
/// the supervisor's policy, and the run ends promptly.
#[test]
fn supervisor_tree_is_torn_down_promptly_on_root_exit() {
let t = Instant::now();
run(|| {
spawn(|| {
OneForOne::new()
.child(
ChildSpec::new(Restart::Permanent, || loop {
sleep(Duration::from_millis(5));
})
.shutdown(Shutdown::Timeout(Duration::from_millis(50))),
)
.run();
});
sleep(Duration::from_millis(30));
});
assert_prompt(t, "tree");
}
+36
View File
@@ -0,0 +1,36 @@
//! Under `smarm-trace`, every actor the root-exit sweep reaches is recorded
//! as a `root_sweep` event — the way to *see* unsupervised leftovers. One test
//! per binary: the trace file is process-global.
#![cfg(feature = "smarm-trace")]
use smarm::gen_server::{self, GenServer, GenServerCtx};
use smarm::run;
struct Quiet;
impl GenServer for Quiet {
type Call = ();
type Reply = ();
type Cast = ();
type Info = ();
type Timer = ();
fn init(&mut self, _: &GenServerCtx<Self>) {}
fn handle_call(&mut self, _: ()) {}
fn handle_cast(&mut self, _: ()) {}
}
#[test]
fn forgotten_server_shows_up_as_root_sweep() {
let path = std::env::temp_dir().join(format!("smarm_root_sweep_{}.json", std::process::id()));
std::env::set_var("SMARM_TRACE_FILE", &path);
run(|| {
let srv = gen_server::start(Quiet);
srv.call(()).unwrap();
drop(srv); // forgotten: nobody supervises it, nobody holds it
});
let trace = std::fs::read_to_string(&path).expect("trace file written");
let _ = std::fs::remove_file(&path);
assert!(
trace.contains("root_sweep stopped"),
"expected a root_sweep line for the non-trapping leftover; got:\n{trace}"
);
}
+122
View File
@@ -0,0 +1,122 @@
//! Graceful shutdown — `request_shutdown` (OTP `exit(Pid, shutdown)`).
//!
//! `request_stop` is `exit(Pid, kill)`: an uncatchable unwind at the target's
//! next observation point. `request_shutdown` is the polite form:
//! - a target that is NOT trapping exits is stopped exactly as by
//! `request_stop` (OTP's rule: don't trap, you die);
//! - a target that IS trapping receives an `ExitSignal { reason: Shutdown }`
//! on its trap inbox and keeps running — it is expected to wind down and
//! exit normally on its own.
use smarm::{monitor, request_shutdown, run, self_pid, sleep, spawn, trap_exit, DownReason, Pid};
use std::sync::atomic::{AtomicBool, Ordering};
use std::sync::{mpsc, Arc};
use std::thread;
use std::time::Duration;
const WATCHDOG: Duration = Duration::from_secs(10);
#[test]
fn request_shutdown_stops_a_non_trapping_actor() {
run(|| {
let h = spawn(|| sleep(Duration::from_secs(3600)));
let mon = monitor(h.pid());
request_shutdown(h.pid());
let down = mon.rx.recv().expect("down");
assert_eq!(down.reason, DownReason::Stopped);
});
}
#[test]
fn request_shutdown_is_a_message_to_a_trapping_actor() {
let unwound = Arc::new(AtomicBool::new(false));
let u = unwound.clone();
run(move || {
struct Unwound(Arc<AtomicBool>);
impl Drop for Unwound {
fn drop(&mut self) {
if std::thread::panicking() {
self.0.store(true, Ordering::SeqCst);
}
}
}
let (tx, rx) = smarm::channel::<(Pid, DownReason)>();
let (ready_tx, ready_rx) = smarm::channel::<()>();
let h = spawn(move || {
let _g = Unwound(u);
let inbox = trap_exit();
let _ = ready_tx.send(());
let sig = inbox.recv().expect("exit signal");
let _ = tx.send((sig.from, sig.reason));
// Keep doing work after the request: shutdown is advisory.
sleep(Duration::from_millis(20));
});
// Trapping is set by the target itself; a request that beats it is a
// plain stop (same window as OTP's exit-before-process_flag).
ready_rx.recv().expect("ready");
let me = self_pid();
let mon = monitor(h.pid());
request_shutdown(h.pid());
let (from, reason) = rx.recv().expect("relayed");
assert_eq!(from, me);
assert_eq!(reason, DownReason::Shutdown);
let down = mon.rx.recv().expect("down");
assert_eq!(
down.reason,
DownReason::Exit,
"target exited normally, not stopped"
);
});
assert!(
!unwound.load(Ordering::SeqCst),
"trapping target must not be unwound"
);
}
#[test]
fn request_shutdown_on_dead_pid_is_a_no_op() {
run(|| {
let h = spawn(|| {});
let pid = h.pid();
let _ = h.join();
request_shutdown(pid); // must not panic
});
}
#[test]
fn handle_request_shutdown_from_foreign_thread() {
let rt = smarm::init(smarm::Config::exact(2));
let handle = rt.handle();
let (pid_tx, pid_rx) = mpsc::channel::<Pid>();
let requester = thread::spawn(move || {
let pid = pid_rx.recv().expect("pid");
thread::sleep(Duration::from_millis(50));
handle.request_shutdown(pid);
});
let (done_tx, done_rx) = mpsc::channel();
thread::spawn(move || {
rt.run(move || {
let (tx, rx) = smarm::channel::<DownReason>();
let (ready_tx, ready_rx) = smarm::channel::<()>();
let h = spawn(move || {
let inbox = trap_exit();
let _ = ready_tx.send(());
let sig = inbox.recv().expect("exit signal");
let _ = tx.send(sig.reason);
});
ready_rx.recv().expect("ready");
pid_tx.send(h.pid()).expect("send pid");
let reason = rx.recv().expect("relayed");
assert_eq!(reason, DownReason::Shutdown);
let _ = h.join();
});
let _ = done_tx.send(());
});
done_rx
.recv_timeout(WATCHDOG)
.expect("run did not return: foreign-thread request_shutdown never reached the target");
requester.join().expect("requester thread");
}
+303
View File
@@ -0,0 +1,303 @@
//! Supervisor shutdown — the OTP child-spec `shutdown` policy.
//!
//! A supervisor traps exits. A `request_shutdown` reaching it (from its parent
//! supervisor, or from the app via `request_shutdown`/`RuntimeHandle`) runs
//! the ordered shutdown: children are stopped in reverse start order, each
//! per its `Shutdown` policy — `request_shutdown`, wait up to the timeout for
//! its termination signal, `request_stop` if it overstays — and then `run()`
//! returns normally. Every supervisor-initiated child stop (ordered shutdown,
//! OneForAll/RestForOne sibling cycling) goes through the same policy.
//!
//! A *hard* `request_stop` on a supervisor unwinds it; a drop guard then
//! hard-stops its live children so the subtree is never orphaned.
use smarm::supervisor::{ChildSpec, OneForOne, Restart, Shutdown, Strategy};
use smarm::{
monitor, request_shutdown, request_stop, run, sleep, spawn, trap_exit, DownReason, JoinHandle,
};
use std::sync::atomic::{AtomicBool, AtomicUsize, Ordering};
use std::sync::{Arc, Mutex};
use std::time::{Duration, Instant};
/// A child that traps exits, records the order it was shut down in, and exits
/// normally on the request (after `delay`). Ignores the request if `comply`
/// is false — a straggler that must be hard-stopped.
fn polite_child(
tag: usize,
log: &Arc<Mutex<Vec<usize>>>,
delay: Duration,
comply: bool,
) -> impl Fn() + Send + Sync + 'static {
let log = log.clone();
move || {
let inbox = trap_exit();
loop {
let sig = match inbox.recv() {
Ok(s) => s,
Err(_) => return,
};
if sig.reason == DownReason::Shutdown {
log.lock().unwrap().push(tag);
if comply {
sleep(delay);
return;
}
// Not complying: keep running until hard-stopped.
loop {
sleep(Duration::from_millis(5));
}
}
}
}
}
/// Spawn `sup`, let its children reach `trap_exit`, return the handle.
fn spawn_settled(sup: OneForOne) -> JoinHandle {
let h = spawn(move || sup.run());
sleep(Duration::from_millis(30));
h
}
#[test]
fn shutdown_stops_children_in_reverse_order_and_returns_normally() {
let log = Arc::new(Mutex::new(Vec::new()));
let l = log.clone();
run(move || {
let sup = OneForOne::new()
.child(ChildSpec::new(
Restart::Permanent,
polite_child(1, &l, Duration::ZERO, true),
))
.child(ChildSpec::new(
Restart::Permanent,
polite_child(2, &l, Duration::ZERO, true),
))
.child(ChildSpec::new(
Restart::Permanent,
polite_child(3, &l, Duration::ZERO, true),
));
let h = spawn_settled(sup);
let mon = monitor(h.pid());
request_shutdown(h.pid());
let down = mon.rx.recv().expect("down");
assert_eq!(
down.reason,
DownReason::Exit,
"supervisor exits normally after shutdown"
);
});
assert_eq!(*log.lock().unwrap(), vec![3, 2, 1]);
}
#[test]
fn non_trapping_child_is_simply_stopped() {
let dropped = Arc::new(AtomicBool::new(false));
let d = dropped.clone();
run(move || {
struct G(Arc<AtomicBool>);
impl Drop for G {
fn drop(&mut self) {
self.0.store(true, Ordering::SeqCst);
}
}
let sup = OneForOne::new().child(ChildSpec::new(Restart::Permanent, move || {
let _g = G(d.clone());
loop {
sleep(Duration::from_millis(5));
}
}));
let h = spawn_settled(sup);
request_shutdown(h.pid());
h.join().expect("sup");
});
assert!(dropped.load(Ordering::SeqCst));
}
#[test]
fn straggler_is_hard_stopped_after_timeout() {
let log = Arc::new(Mutex::new(Vec::new()));
let l = log.clone();
run(move || {
let sup = OneForOne::new().child(
ChildSpec::new(
Restart::Permanent,
polite_child(1, &l, Duration::ZERO, false),
)
.shutdown(Shutdown::Timeout(Duration::from_millis(50))),
);
let h = spawn_settled(sup);
let t0 = Instant::now();
request_shutdown(h.pid());
h.join().expect("sup");
let took = t0.elapsed();
assert!(
took >= Duration::from_millis(50),
"returned before the grace period: {took:?}"
);
assert!(
took < Duration::from_secs(2),
"did not fall back to a hard stop: {took:?}"
);
});
assert_eq!(
*log.lock().unwrap(),
vec![1],
"the straggler did receive the request"
);
}
#[test]
fn infinity_waits_for_a_slow_but_compliant_child() {
let log = Arc::new(Mutex::new(Vec::new()));
let l = log.clone();
let finished = Arc::new(AtomicBool::new(false));
let f = finished.clone();
run(move || {
let f2 = f.clone();
let l2 = l.clone();
let sup = OneForOne::new().child(
ChildSpec::new(Restart::Permanent, move || {
let inbox = trap_exit();
let _ = inbox.recv();
l2.lock().unwrap().push(1);
sleep(Duration::from_millis(150));
f2.store(true, Ordering::SeqCst); // only reached if not hard-stopped
})
.shutdown(Shutdown::Infinity),
);
let h = spawn_settled(sup);
request_shutdown(h.pid());
h.join().expect("sup");
});
assert!(
finished.load(Ordering::SeqCst),
"Infinity must not hard-stop a compliant child"
);
}
#[test]
fn brutal_kill_skips_the_request() {
let log = Arc::new(Mutex::new(Vec::new()));
let l = log.clone();
run(move || {
let sup = OneForOne::new().child(
ChildSpec::new(
Restart::Permanent,
polite_child(1, &l, Duration::ZERO, true),
)
.shutdown(Shutdown::BrutalKill),
);
let h = spawn_settled(sup);
request_shutdown(h.pid());
h.join().expect("sup");
});
assert!(
log.lock().unwrap().is_empty(),
"a BrutalKill child never sees the request"
);
}
#[test]
fn hard_stop_of_supervisor_does_not_orphan_children() {
let alive = Arc::new(AtomicUsize::new(0));
let a = alive.clone();
run(move || {
struct Alive(Arc<AtomicUsize>);
impl Drop for Alive {
fn drop(&mut self) {
self.0.fetch_sub(1, Ordering::SeqCst);
}
}
let mk = |a: Arc<AtomicUsize>| {
move || {
a.fetch_add(1, Ordering::SeqCst);
let _g = Alive(a.clone());
loop {
sleep(Duration::from_millis(5));
}
}
};
let sup = OneForOne::new()
.child(ChildSpec::new(Restart::Permanent, mk(a.clone())))
.child(ChildSpec::new(Restart::Permanent, mk(a.clone())));
let h = spawn_settled(sup);
assert_eq!(a.load(Ordering::SeqCst), 2);
let mon = monitor(h.pid());
request_stop(h.pid());
let _ = mon.rx.recv();
sleep(Duration::from_millis(50));
assert_eq!(
a.load(Ordering::SeqCst),
0,
"children orphaned by a hard supervisor stop"
);
});
}
#[test]
fn nested_shutdown_reaches_grandchildren() {
let log = Arc::new(Mutex::new(Vec::new()));
let l = log.clone();
run(move || {
let l_inner = l.clone();
let inner = move || {
OneForOne::new()
.child(ChildSpec::new(
Restart::Permanent,
polite_child(10, &l_inner, Duration::ZERO, true),
))
.child(ChildSpec::new(
Restart::Permanent,
polite_child(11, &l_inner, Duration::ZERO, true),
))
.run()
};
let sup = OneForOne::new()
.child(ChildSpec::new(
Restart::Permanent,
polite_child(1, &l, Duration::ZERO, true),
))
.child(ChildSpec::new(Restart::Permanent, inner).shutdown(Shutdown::Infinity));
let h = spawn_settled(sup);
request_shutdown(h.pid());
h.join().expect("sup");
});
assert_eq!(*log.lock().unwrap(), vec![11, 10, 1]);
}
#[test]
fn sibling_cycling_uses_graceful_shutdown() {
// OneForAll: when child A dies, sibling B (trapping) must receive a
// Shutdown request rather than a bare stop.
let log = Arc::new(Mutex::new(Vec::new()));
let l = log.clone();
let a_runs = Arc::new(AtomicUsize::new(0));
let ar = a_runs.clone();
run(move || {
let ar2 = ar.clone();
let sup = OneForOne::new()
.strategy(Strategy::OneForAll)
.intensity(5, Duration::from_secs(60))
.child(ChildSpec::new(Restart::Transient, move || {
let n = ar2.fetch_add(1, Ordering::SeqCst) + 1;
sleep(Duration::from_millis(30));
if n == 1 {
panic!("first run dies");
}
// Second run: park until shut down.
let inbox = trap_exit();
let _ = inbox.recv();
}))
.child(ChildSpec::new(
Restart::Permanent,
polite_child(2, &l, Duration::ZERO, true),
));
let h = spawn(move || sup.run());
sleep(Duration::from_millis(150));
request_shutdown(h.pid());
h.join().expect("sup");
});
// B was shut down once by the cycle and once by the final shutdown.
assert_eq!(*log.lock().unwrap(), vec![2, 2]);
assert_eq!(a_runs.load(Ordering::SeqCst), 2);
}
+3 -2
View File
@@ -80,10 +80,11 @@ fn slot_off_means_zero_slot_traffic() {
assert_eq!(rt.stats().slot_hits(), 0); assert_eq!(rt.stats().slot_hits(), 0);
assert_eq!(rt.stats().slot_displacements(), 0); assert_eq!(rt.stats().slot_displacements(), 0);
// …and off by default (RFC 005: default off until the shootout accepts). // …and ON by default (RFC 005 accepted after the shootout — history.md
// finding 17/18: ping-pong 12.5× at 20T, ~9% at 1T, 100% slot hits).
let rt = init(Config::exact(1)); let rt = init(Config::exact(1));
rt.run(ping_pong(2, 100)); rt.run(ping_pong(2, 100));
assert_eq!(rt.stats().slot_hits(), 0, "wake_slot must default to OFF"); assert!(rt.stats().slot_hits() > 0, "wake_slot must default to ON");
} }
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------