8 Commits
Author SHA1 Message Date
Claude (sandbox) 741c10337b release: v0.7.0
Bump crate version to 0.7.0.
2026-08-21 13:10:25 +02:00
Claude (sandbox) e570138da5 docs(roadmap): supervisor start order is not start readiness
Filed from the urus v0.3 endpoint work. start_child spawns and moves on,
so a later sibling can whereis an earlier named child before that child's
actor has run. Notes why blocking spawn is not the fix ('has begun
executing' != 'has bound its name', plus a per-accept round-trip tax and
every spawn becoming a context-switch point), that OTP has the same async
spawn and synchronises one level up in gen_server:start_link, the
readiness-ack shape if scheduled, and the structural workaround urus uses
today (registrar spawns its own consumers).
2026-08-20 13:20:42 +00:00
Claude (sandbox) 415effb2e9 feat(gen_server,gen_statem): lifetime is the actor's — refs are addresses; inline named run
Root cause behind the "pin the endpoint" gotcha and the trapping-wrapper
pattern: a gen_server had two lifetime authorities — its refs (last one
dropped → inbox closes → exit) and, when supervised, its supervisor. OTP has
one: a process lives until it stops, is shut down, or is killed; a pid is an
address. Root exit now shutting down every forest root removes the reason the
ref-governed idiom existed (a forgotten server no longer hangs the run), so
adopt the one rule:

- The server/machine loop holds one inbox sender for its life; the inbox
  never closes. GenServerRef / GenStatemRef are addresses. Explicit close is
  `shutdown()`; a forgotten one is swept at root exit.
- `NamedGenServerBuilder::run()` / `gen_statem::run_named(name, m)` run the
  loop inline as the current actor: a server is a direct ChildSpec child,
  gets the supervisor's shutdown as handle_shutdown / a shutdown row, re-binds
  its name on restart, and is addressed by name. The wrapper in
  examples/graceful_shutdown.rs is gone.
- gen_statem gains GenStatemName + whereis_machine/send/call/shutdown by name
  (parity with gen_server); the macro gets `Sm::new`.
- Root-exit sweep records `Event::RootSweep { target, trapping }` under
  smarm-trace ("root_sweep shutdown|stopped"): unsupervised leftovers are
  visible rather than silently owned-by-refs.
- Named start() name-clash path stops the spawned actor instead of relying
  on ref drop.

Tests: tests/gen_server_lifetime.rs, tests/gen_statem_lifetime.rs,
tests/root_sweep_trace.rs (feature-gated); three existing tests that used
drop-closes-inbox now use shutdown(). Docs/README/ROADMAP/Deep Dive updated.
2026-08-19 17:47:31 +00:00
Claude (sandbox) 849a424c8e docs,examples: graceful shutdown — new examples/graceful_shutdown.rs, README 'Stopping actors', named_genserver uses shutdown(), Deep Dive terminate note, ROADMAP open items 2026-08-19 16:28:42 +00:00
Claude (sandbox) 6ceb138f5f feat(gen_statem): graceful-shutdown parity — trap_exit, shutdown/exit rows, stop, terminate
Mirrors the gen_server surface in gen_statem's event model:
- Cx::trap_exit() (in the initial enter): a shutdown request then arrives
  as the Shutdown event, routed by state through `shutdown` rows (default
  for a state with no row: stop); linked-peer deaths as `exit <pat>` rows
  (default: drop). Non-trapping machines are stopped outright, as before.
- Cx::stop(): normal self-exit after the current event; `stop` tail keyword
  is sugar for { cx.stop(); prev }.
- Machine::terminate (optional `terminate { … }` macro block), run from a
  Drop guard on every exit path; the guard also drains armed timers.
- Machine::shutdown_ev / exit_ev (defaults None) so hand-written machines
  keep compiling; GenStatemRef::shutdown() is graceful and waits.
- Loop selects exits > timers > inbox.

Tests: tests/gen_statem_shutdown.rs.
2026-08-19 16:25:15 +00:00
Claude (sandbox) 250f31265b feat(runtime): root exit is graceful shutdown of the forest roots
The RFC 014 root-exit sweep hard-stopped every live slot once nothing was
runnable. That deferral privileged queued work over parked-with-a-pending-
wake work (a sleeper was killed, a queued cast was drained) and any attempt
to widen the notion of pending wake (timers, fd readiness) re-wedges the
run on the periodic-timer daemon the sweep exists to end.

Root exit now means "the program is done": finalize_actor delivers
request_shutdown to every forest root — each live actor whose parent is
the run (ROOT_PID) or is dead — synchronously, before the live-count
decrement. Supervisors cascade per child Shutdown policy; trapping actors
may Continue/drain with working timers and end the run when they stop
themselves; non-trapping actors are stopped outright. No forcing sweep.

Removes root_exited/root_swept, Pop::RootDrain and the idle-verdict
condition; adds tests/root_exit.rs.
2026-08-19 07:11:46 +00:00
claude 9c8f59ca53 feat(scheduler,supervisor,gen_server): graceful shutdown — request_shutdown, child Shutdown policy, handle_shutdown
Lift OTP's `exit(Pid, shutdown)` + child-spec `shutdown` wholesale.

scheduler / runtime
- `request_shutdown(pid)`: the polite stop. A target trapping exits gets an
  `ExitSignal { reason: DownReason::Shutdown }` on its trap inbox and keeps
  running; a non-trapping target is stopped as by `request_stop`, which is
  now documented as the hard stop (`exit(Pid, kill)`). Dead pid: no-op.
- `RuntimeHandle::request_shutdown` for the off-runtime (signal thread) path;
  `from == ROOT_PID` there.
- `DownReason::Shutdown` — appears only in ExitSignal, never in Down (a
  complying target exits *normally*).

supervisor
- `ChildSpec::shutdown(Shutdown::{BrutalKill, Timeout(d), Infinity})`,
  default Timeout(5s). Every supervisor-initiated stop (ordered shutdown and
  OneForAll/RestForOne sibling cycling) is: request_shutdown → await the
  child's Signal up to the grace → request_stop → await. Sequential, reverse
  start order.
- The supervisor traps exits; a Shutdown ExitSignal runs the ordered
  shutdown and `run()` returns normally, so `request_shutdown(root_sup)`
  tears a whole tree down top-down with each child's grace period.
- FIX: a hard `request_stop` on a supervisor previously orphaned its
  children (the ordered shutdown lived after the loop, and the unwind
  skipped it). `Live` (the by_pid map) now carries a drop guard that
  fire-and-forget hard-stops live children when unwinding.

gen_server
- `GenServerCtx::trap_exit()` opt-in in `init`; the trap inbox becomes arm 0
  of the loop's select. Shutdown ExitSignal → `handle_shutdown() ->
  ShutdownAction::{Exit, Continue}` (default Exit: loop breaks, `terminate`
  runs on the normal path and may block). Other ExitSignals →
  `handle_exit(sig)`.
- `GenServerCtx::stop_handle() -> StopHandle`, `stop()` ends the server
  after the current message with a *normal* exit — the missing
  `{stop, normal, State}`; `request_stop(self_pid())` was the only self-exit
  and it is abnormal (Transient restarts it).
- `GenServerRef::shutdown()` / `gen_server::shutdown(name)` now go through
  `request_shutdown`.

Tests: tests/shutdown.rs, tests/supervisor_shutdown.rs,
tests/gen_server_shutdown.rs. Full suite green; fmt + clippy --lib clean.
2026-08-19 06:34:48 +00:00
Claude (sandbox) 1002777ef3 feat(channel,runtime): off-runtime cross-thread wake for parked receivers and stop
Wakes issued from a non-scheduler OS thread were silent no-ops. Every
off-runtime wake primitive (unpark, unpark_at, request_stop) reaches the
runtime through the RUNTIME thread-local, which is unset on any foreign
thread — so a send from a plain std::thread enqueued its message but never
woke the parked receiver, and there was no way to drive a stop into a
runtime from an application thread (e.g. an OS-signal handler). The former
strands a parked recv forever; the latter is why a downstream server must
poll a shutdown flag instead of parking on it.

Generalize RFC 018's rule — a producer reaches the runtime through a Weak it
holds — from the IO backend to channel senders and to a new handle:

- A receiver captures a Weak<RuntimeInner> (provably live at that moment)
  alongside its (pid, epoch) when it parks. send() and the last-sender drop
  wake through scheduler::unpark_at_via, which takes the thread-local path
  when on a scheduler thread (preemption-gated, slot-eligible) and the
  captured Weak otherwise — the same cross-context wake the IO threads do.
- Runtime::handle() returns a Send + Sync RuntimeHandle carrying that Weak;
  RuntimeHandle::request_stop drives a cooperative stop from any thread and
  is a no-op once the runtime is dropped.

The in-runtime wake paths (recv/select timers) are unchanged; only the sites
reachable from a foreign thread route through the Weak. RuntimeHandle exposes
request_stop only: send-wake needs no user-facing handle, and off-runtime
unpark is covered because request_stop drives unpark on the upgraded inner.

tests/cross_thread_wake.rs: a foreign-thread send wakes a parked receiver; a
foreign-thread request_stop wakes and stops a parked actor; a RuntimeHandle
held across and beyond run never blocks all-done and degrades to a no-op.
2026-08-19 05:50:54 +00:00
26 changed files with 3320 additions and 282 deletions
+1 -1
View File
@@ -1,6 +1,6 @@
[package] [package]
name = "smarm" name = "smarm"
version = "0.6.1" version = "0.7.0"
edition = "2021" edition = "2021"
rust-version = "1.95" rust-version = "1.95"
+23
View File
@@ -54,6 +54,29 @@ run(|| {
}); });
``` ```
## Stopping actors
Two strengths, as in OTP. `request_stop(pid)` is `exit(Pid, kill)`: a cooperative
hard stop, unwinding at the actor's next observation point. `request_shutdown(pid)`
is `exit(Pid, shutdown)`: an actor that traps exits (`trap_exit()`, or
`ctx.trap_exit()` in a gen_server / `cx.trap_exit()` in a gen_statem) receives it
as a signal — `handle_shutdown` / a `shutdown` row — and may drain before stopping
itself; one that does not trap is stopped outright. Supervisors trap:
`request_shutdown(sup)` tears the tree down top-down, each child per its
`ChildSpec` `Shutdown` policy (`Timeout(d)`, `Infinity`, `BrutalKill`). The run's
root actor returning means "the program is done": every top-level actor gets a
`request_shutdown`, and `run()` returns when they are gone. From outside the
runtime (a signal thread), `Runtime::handle().request_shutdown(pid)` does the
same. `examples/graceful_shutdown.rs` shows all of it.
A gen_server or gen_statem lives until it stops, is shut down, or is killed; its
refs are addresses — dropping them never ends it (a forgotten one is swept at
root exit; with `--features smarm-trace` each such sweep is a `root_sweep`
trace line). The supervised shape is `GenServerBuilder::named(N).run()` /
`gen_statem::run_named(N, m)`: the server runs inline as the `ChildSpec` child
itself, so the supervisor's shutdown reaches it directly, a restart re-binds the
name, and the program addresses it by name.
## Layout ## Layout
``` ```
+48 -4
View File
@@ -77,10 +77,10 @@ Delivered surface:
`ServerBuilder::start` untouched; free `call` / `cast` / `whereis_server`; `ServerBuilder::start` untouched; free `call` / `cast` / `whereis_server`;
`ServerRef::shutdown` + free `shutdown` as the sys-style synchronous stop. `ServerRef::shutdown` + free `shutdown` as the sys-style synchronous stop.
- **Root-exit teardown** (final phase): the run's initial actor is the root; - **Root-exit teardown** (final phase): the run's initial actor is the root;
when it exits, the scheduler's idle verdict stops the parked-forever remainder when it exits the run winds down. *(Reworked with the graceful-shutdown work:
(deferred past the queue drain, so actors with in-flight work finish rather root exit now delivers `request_shutdown` to every forest root — see
than unwinding on the stop). Closes the "app actor blocks AllDone" stall — see "Root exit" below and `tests/root_exit.rs`.)* Closes the "app actor blocks
Look into, below. AllDone" stall — see Look into, below.
Extends — does not retire — the "select exists; a unified per-process mailbox Extends — does not retire — the "select exists; a unified per-process mailbox
still does not" invariant: 014 adds addressable *delivery*, not a unified inbox; still does not" invariant: 014 adds addressable *delivery*, not a unified inbox;
@@ -260,8 +260,52 @@ path the atomic-bool workaround stood in for. Re-check the urus crud repro to
confirm the workaround can be retired (the teardown is cooperative — an actor in confirm the workaround can be retired (the teardown is cooperative — an actor in
a tight loop with no observation point still can't be stopped). a tight loop with no observation point still can't be stopped).
**Update (graceful shutdown):** the RFC 014 sweep was a hard `request_stop` of
every live slot, deferred until nothing was runnable — which killed a sleeping
actor (timer pending) but drained a queued one, for no principled reason. It is
now the OTP semantics: root exit = "the program is done" = `request_shutdown`
to every **forest root** (live actor whose parent is the run or is dead), run
synchronously on the root's finalize path. Supervisors cascade with their child
`Shutdown` policies; trapping actors may `Continue`/drain (timers keep working)
and end the run when they stop themselves; non-trapping actors are stopped
outright — `join` what you need finished. No forcing sweep follows.
--- ---
### Open items from the graceful-shutdown work (not scheduled)
- ~~gen_server / gen_statem as a direct supervised child.~~ Done: lifetime is
the actor's (refs are addresses, the loop holds an inbox sender);
`NamedGenServerBuilder::run` / `gen_statem::run_named` run the loop inline as
the `ChildSpec` child; the root-exit sweep traces each leftover as
`root_sweep` under `smarm-trace`.
- Supervisor `Live` drop-guard sweep is `request_stop` (kill propagates as
kill); OTP would deliver a trappable `killed`. Chosen for boundedness.
- A root-exit shutdown reaches only actors live *at that instant*; a
non-trapping forest root that spawns before it unwinds leaves that spawn
to itself (Erlang: an unlinked spawn is nobody's child).
- **Supervisor start *order* is not start *readiness*.** `start_child`
spawns and moves straight on, so an earlier child is merely *scheduled*,
not initialised, when a later sibling starts. A later child that resolves
an earlier one by name (`whereis_server`) can therefore miss it — the
classic "named registry sibling, then its consumers" tree. Ordered
`OneForOne`/`RestForOne` shutdown is unaffected (reverse order is honoured
and each stop *is* awaited); this is a start-side gap only.
Making `spawn` itself block does NOT fix it — it would only shrink the
window to "child has begun executing", while the property callers need is
"child has bound its name / opened its socket", which only the child can
declare. It would also tax the hot path (one round-trip per accepted
connection) and turn every spawn into a context-switch point. OTP has the
same async `spawn` and puts the synchronisation one level up:
`gen_server:start_link` blocks the caller until `init/1` returns.
Fix shape when scheduled: a readiness ack in the supervisor's child-start
path (`ChildSpec` variant whose factory receives a ready-signal;
`NamedGenServerBuilder::run` acks after its name bind, gen_server default
acks after `init`; plain closures ack at spawn as today, i.e. opt-in with
no cost to existing children). Until then the workaround is structural:
have the registrar spawn its own consumers so the ordering is program
order inside one actor, not a cross-actor guarantee (urus v0.3 endpoint
does exactly this).
## Invariants & gotchas (respect these across all cycles) ## Invariants & gotchas (respect these across all cycles)
- **Shared mutex is non-reentrant.** `Sender::send` can call `unpark` → - **Shared mutex is non-reentrant.** `Sender::send` can call `unpark` →
+229
View File
@@ -0,0 +1,229 @@
# urus / smarm handoff — updated 2026-08-19 (session 3)
## TL;DR for the next session
**smarm is done for now** (5 unpushed commits on local `master`, see below).
**Next = urus v0.3 endpoint refactor.** You should NOT need to read smarm
scheduler internals; the contract you build on is fully described here and in
`smarm_full/examples/graceful_shutdown.rs` (read that file first — it is the
exact shape urus's tree will take) plus `smarm_full/tests/root_exit.rs`.
### The smarm contract urus builds on (all on local master, verified by tests)
- `request_stop(pid)` = kill (cooperative hard stop). `request_shutdown(pid)` =
polite: trapping target gets `ExitSignal{reason: Shutdown}`, non-trapping is
stopped outright. `RuntimeHandle::{request_stop,request_shutdown}` do the same
from any OS thread (signal handler); grab `rt.handle()` before `rt.run`.
- Supervisor traps; `request_shutdown(sup)` = ordered reverse-start shutdown,
per-child `ChildSpec::shutdown(Shutdown::{Timeout(d)|Infinity|BrutalKill})`
(default Timeout(5s)); sup then returns normally. `request_stop(sup)`
hard-stops children too (no orphans).
- gen_server: `ctx.trap_exit()` in init; `handle_shutdown() -> Exit|Continue`;
`handle_exit(sig)`; `ctx.stop_handle().stop()` = normal self-exit;
`terminate()` may block only on the graceful path (Exit / stop / inbox close).
`GenServerRef::shutdown()` is graceful and waits.
- gen_statem: same in event clothes — `cx.trap_exit()` in initial enter,
`shutdown` rows (default `stop`), `exit sig` rows, `cx.stop()` / `stop` tail,
optional `terminate { }` block. `GenStatemRef::shutdown()`.
- **Root exit = program done**: when the root actor returns, the runtime
`request_shutdown`s every *forest root* (live actor whose parent is the run
or dead). Supervisors cascade; trapping actors may drain (timers keep
working) and end the run when they stop; non-trapping are stopped; **no
forcing sweep** (`join` what must finish). The old "wait until nothing
runnable then kill all" deferral is gone.
- **Gotcha for urus:** a gen_server's lifetime is governed by its refs — drop
the last `GenServerRef` and the inbox closes → clean exit *even mid-drain*.
The endpoint must be pinned (named, or its ref held by the supervisor
wrapper) or it will terminate the moment the root drops its ref.
- **Known gap (ROADMAP open item):** a gen_server can't be a direct `ChildSpec`
child; use the trapping wrapper pattern in `examples/graceful_shutdown.rs::
drainer_child` (starts `under(self_pid())`, forwards shutdown, waits). Doing
an inline `GenServerBuilder::run()` first may be worth a short smarm detour —
decide with Markk.
### smarm commits this session (local master, NOT pushed, NOT tagged)
`250f312` root-exit = graceful shutdown of forest roots (tests/root_exit.rs)
`6ceb138` gen_statem shutdown parity (tests/gen_statem_shutdown.rs)
`849a424` docs + examples/graceful_shutdown.rs + README "Stopping actors"
On top of `1002777` (cross-thread wake) and `9c8f59c` (graceful shutdown).
Cargo.toml still `0.6.1`. Release cut (push, tag — v0.7 is justified by the
API surface — version bump) is Markk's. Full suite, doc tests, examples,
`cargo fmt`, `cargo clippy --lib` all clean. (`clippy --tests` has pre-existing
unwrap lints in tests/fd_select.rs, untouched.)
### Decisions taken this session (Markk)
- Root exit means "program done" (Go/tokio/OTP), not "wait for pending work";
the previously agreed "sleep(50ms) must finish" test was dropped as encoding
the wrong contract (a timer-wheel gate would re-wedge periodic-timer daemons).
- No behaviour-preserving deferral, no forcing second sweep.
- Examples/docs done in the same session; urus next session.
---
# Previous handoff (still accurate where not superseded above)
## Next-session goal
Phase 1 is **done and committed**; Phase 2 is next:
1. **smarm v0.6.2** — cross-thread wake root fix. **DONE**, committed on `master`
as `1002777`. Not yet tagged, not yet version-bumped (Cargo.toml still reads
`0.6.1`), and **not yet pushed to origin** — it exists only in the delivered
snapshot zip and the local sandbox clone. Cutting the release (push + tag
`v0.6.2` + bump `0.6.1`→`0.6.2`) is Markk's step.
2. **urus v0.3** — endpoint refactor. Working against `smarm = { path = "../smarm_full" }`
with `git update-index --skip-worktree Cargo.toml` (Markk approved); release commit
swaps back to the tag once Markk cuts it. Phase 2 plan below is STALE where it says
drain-in-terminate; the endpoint is a trapping GenServer: `handle_shutdown` →
`Continue`, enter Draining, `StopHandle::stop()` when the conn set empties.
Decisions below are locked unless marked *(confirm)*.
## Reconstruction (the sandbox resets between sessions)
A fresh sandbox has an empty home and **no Rust toolchain**. To restore:
- Install rustup/cargo. smarm reformats under **rustc 1.97.1**; urus `rust-version`
is 1.95. Use 1.97.1.
- urus: `git clone https://git.kalsbeek.dev/Markk116/urus` — `origin` is registered
and public-read. master `8bdec97` = the v0.2.x line. (Zips in outputs are stale;
prefer the remote now.)
- smarm: `git clone https://git.kalsbeek.dev/Markk116/smarm`. Latest tag **v0.6.1**
(`ca1c983`). The cross-thread wake fix is committed as `1002777` on top of the
post-v0.6.1 README commit `8f2d513` (= origin/master). **It is NOT on origin
yet** — a fresh clone won't have it until Markk pushes. Restore it from the
snapshot zip if working before the push. **v0.6.2 is not yet tagged.**
- urus pins smarm by git **tag** in `Cargo.toml` (currently `v0.6.0`). A trivial
first commit bumps it to `v0.6.1` (also picks up `try_spawn` + monitor
terminal-outcome fixes).
## Why (context — the finding that drives the plan)
urus's shutdown machinery (the `AtomicBool` listener flag + the `SHUTDOWN_POLL`
loop in `serve.rs`) is scaffolding around two smarm properties. Their statuses
differ, which is the whole point:
- **Issue A — lossy stop vs a QUEUED actor: ALREADY FIXED in smarm.** Commit
`7bab4d2` added an entry-side `check_cancelled()` in `park_current`. A
`request_stop` against a listener parked in `wait_readable_timeout` now unwinds
cleanly (it parks via `try_select_timeout → park_current`). urus's flag + its
stale "smarm's lossy stop-while-QUEUED window" comment can be deleted.
- **Issue B — foreign-thread wake is a no-op: FIXED in `1002777` (was present
through v0.6.1).** The gap: `unpark`/`unpark_at`/`request_stop` all route through
`try_with_runtime`, which reads a thread-local that is `None` on any non-scheduler
thread, so no cross-thread wake worked — a signal handler / OS thread could not
wake *or* stop a parked actor, which is why `serve.rs` polls the shutdown signal
instead of parking on it. Now closed (see Phase 1 below): urus's `SHUTDOWN_POLL`
loop can be deleted and its `Handle::shutdown` can park on a handle-driven stop.
## Phase 1 — smarm cross-thread wake (root fix) — DONE (`1002777`)
Shipped as one commit generalizing RFC 018 (a producer reaches the runtime through
a `Weak` it holds) from the IO backend to channel senders and a new handle:
- **`Runtime::handle() -> RuntimeHandle`** (`Send + Sync`), holding a
`Weak<RuntimeInner>`. Grab it before `rt.run` and hand it to the signal thread.
- **`RuntimeHandle::request_stop<A>(Pid<A>)`** — upgrades the Weak and calls
`request_stop_inner` on the inner; no-op if the runtime is gone. This is the
signal-handler-drives-shutdown path; it cascades the ordered stop down the tree
exactly like an in-runtime `request_stop`.
- **Send-wake:** the receiver captures `scheduler::runtime_weak()` into its
`parked_receiver` tuple **at park time** (not at channel creation — the resolved
sub-decision; a parked receiver is a live actor so the Weak is provably upgradable,
and it scopes the capture to when a wake is possible). `send()` and last-sender
`drop` wake via `scheduler::unpark_at_via(pid, epoch, &weak)`: thread-local path
when on a scheduler thread (preempt-gated, slot-eligible), captured Weak otherwise.
In-runtime timer wakes (recv/select) were left on `scheduler::unpark_at`.
**API scope decision (signed off):** `RuntimeHandle` exposes **`request_stop` only**.
No public `unpark`/`unpark_at` on the handle — send-wake needs no user-facing handle,
and "unpark off-runtime" is covered because `request_stop` drives `unpark` on the
upgraded inner. No `is_alive()`. Both are one-line additions if a consumer appears.
**No RFC written** — pattern was already established (RFC 018), agreed not needed.
Tests: `tests/cross_thread_wake.rs` (foreign-thread send wakes a parked receiver;
foreign-thread `request_stop` wakes+stops a parked actor; a lingering handle never
blocks all-done and degrades to a no-op once the runtime drops). Full suite green;
`cargo fmt` + `cargo clippy --lib` clean.
**Remaining release step (Markk):** push `master`, tag `v0.6.2`, bump Cargo.toml
`0.6.1`→`0.6.2`. Left paired with the tag as the release cut, not done in `1002777`.
## Phase 1b — smarm graceful shutdown (OTP lift) — DONE (`9c8f59c`, on top of `1002777`)
Decided this session (Markk): B — fix at the smarm level rather than a two-stop
split in urus. No RFC (Markk: "just implement it"). Shipped, tested, committed on
the local `master`, **not pushed, not tagged**. It should ship as the same
release as 1002777 (v0.6.2, or v0.7 given the API surface — Markk's call).
- `request_shutdown(pid)` / `RuntimeHandle::request_shutdown` = `exit(Pid, shutdown)`;
`request_stop` = `exit(Pid, kill)`. Trapping target gets `ExitSignal{reason:
DownReason::Shutdown}`; non-trapping is stopped outright.
- `ChildSpec::shutdown(Shutdown::{BrutalKill, Timeout(d), Infinity})`, default 5s.
Supervisor traps exits; `request_shutdown(sup)` = ordered top-down shutdown,
returns normally. **Also fixed**: `request_stop(sup)` used to ORPHAN children
(probe-verified; the handoff's "cascade" claim was wrong) — `Live` drop guard now
hard-stops them.
- gen_server: `ctx.trap_exit()`, `handle_shutdown() -> ShutdownAction::{Exit,
Continue}`, `handle_exit(ExitSignal)`, `ctx.stop_handle().stop()` = normal
self-exit (`{stop, normal}`; previously impossible — only abnormal `Stopped`).
`GenServerRef::shutdown()` is graceful now.
- Root finding that forced this: gen_server `terminate()` runs from a Drop guard,
mid-unwind on the stop path; any park in it = double panic = abort. So
"drain-in-terminate()" (the old Phase 2 plan) was never viable.
### ~~Next-session smarm work~~ DONE this session (see TL;DR)
1. **Root-exit sweep**: make `Pop::RootDrain` also require an empty timer wheel
(and it already requires nothing runnable; io_out is only checked for AllDone —
check whether it should gate RootDrain too). TDD: an actor in `sleep(50ms)` when
the root returns must finish, not be swept. Then a `Reservoir`-style test that a
*truly* parked-forever daemon still gets swept.
2. **gen_statem parity**: `ctx.trap_exit()`, `handle_shutdown -> ShutdownAction`,
`handle_exit`, stop handle. Mirror gen_server; mechanical.
3. **Examples review**: `examples/*.rs` predate all of this. Rework where they show
shutdown/teardown to use `request_shutdown`, `Shutdown` policies, and
`StopHandle`; `named_genserver.rs` first (uses `shutdown`). Also
`docs/smarm - Deep Dive.html` says terminate() must be non-blocking — now only
true on the unwind paths; and README could use a "Stopping actors" paragraph
(request_stop = kill, request_shutdown = shutdown, Shutdown policy).
### Open smarm items found on the way (noted, not scheduled)
- (root-exit sweep and gen_statem parity moved up to the scheduled list.)
- Sweep in the supervisor `Live` drop guard is `request_stop` (kill propagates as
kill); OTP would deliver a trappable `killed`. Chosen for boundedness.
## Phase 2 — urus v0.3: endpoint refactor (after v0.6.2 is tagged)
Target = the spec's original shape (`urus-spec.md` §2.1/§6: `listener_sup` under the
**user's** root supervisor). Deviation to unwind: `serve` owning `rt.run`.
- App owns the runtime: `smarm::init(cfg).run(|| root_sup.run())`, root e.g.
`RestForOne[ app actors…, urus::endpoint(config, pipeline) ]`. This is what kills
the `Arc<OnceLock>` idiom for the right reason (app state born in-runtime as a
supervised, ordered child).
- `urus::endpoint` = one GenServer child owning the registry + an **internal**
listener sub-supervisor + drain-in-`terminate()`. Listeners stay internal, not
app-visible peers.
- Shutdown = `request_stop` the root supervisor (or via the runtime handle from a
signal thread) → cascades down → `endpoint.terminate()` runs the drain
(`drain_timeout`, force-stop sweep).
- **DELETE:** the `AtomicBool` listener flag (A fixed) and the `SHUTDOWN_POLL` loop
+ its apologetic comment (B fixed → park, don't poll).
- Nuance: `request_stop` → `Signal::Stopped` is *abnormal* → `Transient` restarts.
Stop-without-restart = stop the **supervisor**, not the children.
- *(confirm)* Keep `serve`/`serve_with`/`serve_with_shutdown` as thin wrappers that
build the one-child tree internally, so the simple case stays one line.
- *(confirm)* Keep `Handle`/`ShutdownSignal`? Now that cross-thread wake works,
`Handle::shutdown` can map to handle-driven `request_stop` on the endpoint.
- Breaking → cut **urus v0.3**; bump the smarm pin to `v0.6.2` here.
## Working norms
- Every bash call: `export PATH=$HOME/.cargo/bin:$PATH` (once the toolchain's in).
- **TDD**: failing test first, then implement; keep suites green.
- **Hammer ritual** for ANY connection-lifecycle change (the urus #2 shutdown work
qualifies): 35× subset (`shutdown timeout reaped slowloris streaming chunked sse
stalled ws_ channels session`) + 3× full + 1× trace. `scripts/hammer.sh` does NOT
pass feature flags — loop manually with `--features phoenix`. Subset filter must
NOT use `--test integration` (session tests live in lib).
- Example smoke tests: hold the server's stdin open (`mkfifo` + `sleep > fifo`) or
the Enter-to-shutdown thread fires on EOF instantly.
- Background procs are reaped BETWEEN bash calls; `pkill -f` matches your own shell.
- Artefact store (specs): `curl -H "Authorization: Bearer sk-llmingest-2e45d80c63db24c6781e761eb2a9a58e83d9f48ef77a42185bad311d07c80e68" https://artefacts.kalsbeek.dev/artifacts/<name>`
— `urus-spec.md`, `urus-bench-spec.md`, `rfc_008-implementation-notes.md`, …
- smarm feature flags: `smarm-trace`, `smarm-causal` (urus re-exports both).
## Local cross-repo testing — KEEP OUT OF COMMITS
To test urus #2 against un-tagged smarm 0.6.2, point urus's `Cargo.toml` smarm dep
at a local path (`smarm = { path = "../smarm" }`) instead of the git tag.
- Must NOT land in commits. Guard: `git update-index --skip-worktree Cargo.toml`
after editing (undo with `--no-skip-worktree`), or stash before committing.
- The committed `Cargo.toml` stays pinned to the git tag; restore the tag (bumped to
`v0.6.2`) for the release commit.
+4 -3
View File
@@ -1620,8 +1620,9 @@
wait: <code>select</code> priority is <strong>Down arms › Watcher arm › info channels (declaration order) › wait: <code>select</code> priority is <strong>Down arms › Watcher arm › info channels (declaration order) ›
inbox</strong>, rebuilt each turn. A hot inbox can't starve a death notice or a system message; inbox</strong>, rebuilt each turn. A hot inbox can't starve a death notice or a system message;
conversely a hot info channel <em>can</em> starve the inbox — deliberately. A closed info arm is conversely a hot info channel <em>can</em> starve the inbox — deliberately. A closed info arm is
silently dropped from the set; a closed <em>inbox</em> (every <code>ServerRef</code> gone) is graceful silently dropped from the set. The inbox never closes — the loop holds one sender for its whole
shutdown.</p> life, so a <code>GenServerRef</code> is an address, not an owner: the server ends only by
<code>StopHandle::stop</code>, a shutdown, a hard stop, or a panic.</p>
<h3>Death needs no monitor</h3> <h3>Death needs no monitor</h3>
<p>Server death detection falls out of channel closure. Already dead → the inbox is closed and <p>Server death detection falls out of channel closure. Already dead → the inbox is closed and
@@ -1700,7 +1701,7 @@
</div> </div>
<div class="module-card"> <div class="module-card">
<div class="module-name" style="color:var(--red)">Panics in <code>terminate()</code></div> <div class="module-name" style="color:var(--red)">Panics in <code>terminate()</code></div>
<p>gen_server's <code>terminate()</code> runs from a drop guard, possibly mid-unwind. A panic inside it during an unwind is a double panic → process abort, no supervision tree to save you. Keep it cheap, non-blocking, non-panicking.</p> <p>gen_server's <code>terminate()</code> runs from a drop guard, possibly mid-unwind. A panic inside it during an unwind is a double panic → process abort, no supervision tree to save you. On the panic and hard-stop paths keep it cheap, non-blocking, non-panicking. Only the graceful path (<code>handle_shutdown → Exit</code>, <code>StopHandle::stop</code>) runs it outside an unwind, where it may do real work.</p>
</div> </div>
<div class="module-card"> <div class="module-card">
<div class="module-name" style="color:var(--yellow)">Cold locks are leaf locks</div> <div class="module-name" style="color:var(--yellow)">Cold locks are leaf locks</div>
+139
View File
@@ -0,0 +1,139 @@
//! Graceful shutdown, end to end: a supervised app tree, a server that
//! drains before it exits, and the two ways the whole thing winds down.
//!
//! Stopping an actor comes in two strengths, as in OTP:
//! - `request_stop(pid)` = `exit(Pid, kill)`: cooperative hard stop,
//! unwinds at the next observation point.
//! - `request_shutdown(pid)` = `exit(Pid, shutdown)`: a trapping target gets
//! an `ExitSignal { reason: Shutdown }` and winds
//! down on its own terms; a non-trapping one is
//! stopped outright.
//!
//! A supervisor traps exits. `request_shutdown(sup)` runs its ordered
//! shutdown — children in reverse start order, each per its `ChildSpec`
//! `Shutdown` policy (`Timeout(d)` default 5s, `Infinity`, `BrutalKill`) —
//! and the supervisor then returns normally.
//!
//! Two triggers are shown:
//! 1. **Root exit.** The run's root actor returning means "the program is
//! done": the runtime delivers `request_shutdown` to every top-level actor
//! (here: the supervisor). Trapping actors may keep running to drain and
//! end the run when they stop themselves; non-trapping ones are stopped.
//! 2. **An outside thread** (e.g. a signal handler) driving it via
//! `RuntimeHandle::request_shutdown` on the supervisor — the root then
//! just waits for the tree to come down.
use smarm::gen_server::{
GenServer, GenServerBuilder, GenServerCtx, GenServerName, ShutdownAction, StopHandle,
TimerHandle,
};
use smarm::supervisor::{ChildSpec, OneForOne, Restart, Shutdown};
use smarm::{sleep, spawn};
use std::thread;
use std::time::Duration;
/// A server with in-flight work: on shutdown it stops accepting, finishes what
/// it has (simulated with a ticking timer), then ends itself.
struct Drainer {
pending: u32,
stop: Option<StopHandle<Drainer>>,
timer: Option<TimerHandle<Drainer>>,
}
impl GenServer for Drainer {
type Call = ();
type Reply = ();
type Cast = ();
type Info = ();
type Timer = ();
fn init(&mut self, ctx: &GenServerCtx<Self>) {
ctx.trap_exit(); // opt in: shutdown arrives as handle_shutdown
self.stop = Some(ctx.stop_handle());
self.timer = Some(ctx.timer());
}
fn handle_call(&mut self, _: ()) {}
fn handle_cast(&mut self, _: ()) {}
fn handle_shutdown(&mut self) -> ShutdownAction {
println!(
"drainer: shutdown requested, {} items pending",
self.pending
);
self.timer
.as_ref()
.unwrap()
.tick_every(Duration::from_millis(20), ());
ShutdownAction::Continue // keep serving until drained
}
fn handle_timer(&mut self, _: ()) {
self.pending -= 1;
if self.pending == 0 {
println!("drainer: drained, stopping");
self.stop.as_ref().unwrap().stop(); // normal exit
}
}
fn terminate(&mut self) {
// Graceful path: this runs on the normal path and may block.
println!("drainer: terminate");
}
}
/// The server's name: how the rest of the app reaches it (and the only handle
/// that survives a restart).
const DRAINER: GenServerName<Drainer> = GenServerName::new("drainer");
fn app_tree() -> OneForOne {
OneForOne::new()
.child(
ChildSpec::new(Restart::Permanent, || {
// A plain worker that does not trap: stopped outright on shutdown.
loop {
sleep(Duration::from_millis(10));
}
})
.shutdown(Shutdown::Timeout(Duration::from_millis(100))),
)
// A gen_server is a direct child: `named(N).run()` runs the loop as
// the child actor itself, so the supervisor's shutdown arrives as
// `handle_shutdown` and a restart re-binds the name.
.child(
ChildSpec::new(Restart::Permanent, || {
GenServerBuilder::new(Drainer {
pending: 3,
stop: None,
timer: None,
})
.named(DRAINER)
.run()
.expect("drainer name is free");
})
.shutdown(Shutdown::Infinity),
)
}
fn main() {
println!("--- 1. root exit drives the shutdown ---");
smarm::run(|| {
spawn(|| app_tree().run());
sleep(Duration::from_millis(50)); // the app "runs" for a while
// Returning here asks the supervisor to shut down; the run ends when
// the tree — drainer included — is gone.
});
println!("--- 2. an outside thread drives the shutdown ---");
let rt = smarm::init(smarm::Config::default());
let handle = rt.handle(); // Send + Sync; grab it before run
rt.run(move || {
let sup = spawn(|| app_tree().run());
let sup_pid = sup.pid();
// Stand-in for a SIGTERM handler thread.
thread::spawn(move || {
thread::sleep(Duration::from_millis(50));
println!("signal thread: requesting shutdown");
handle.request_shutdown(sup_pid);
});
sup.join()
.expect("supervisor returns normally after ordered shutdown");
println!("supervisor down; root returns");
});
}
+6
View File
@@ -66,6 +66,12 @@ fn main() {
let svc: Option<GenServerRef<Counter>> = whereis_server(COUNTER); let svc: Option<GenServerRef<Counter>> = whereis_server(COUNTER);
if let Some(svc) = svc { if let Some(svc) = svc {
let _ = svc.call(Query::Get); let _ = svc.call(Query::Get);
// A named server is pinned alive by the registry, so dropping refs
// does not end it. Stop it explicitly: `shutdown()` asks politely
// (a trapping server drains first; this one is stopped outright)
// and waits until it is gone. Left running, the root's return
// would shut it down the same way — see examples/graceful_shutdown.rs.
svc.shutdown();
} }
}); });
} }
+32 -19
View File
@@ -90,8 +90,9 @@
use crate::pid::Pid; use crate::pid::Pid;
use crate::raw_mutex::RawMutex; use crate::raw_mutex::RawMutex;
use crate::runtime::RuntimeInner;
use std::collections::VecDeque; use std::collections::VecDeque;
use std::sync::Arc; use std::sync::{Arc, Weak};
/// Create a new channel and return its `(Sender, Receiver)` halves. /// Create a new channel and return its `(Sender, Receiver)` halves.
/// ///
@@ -114,12 +115,15 @@ pub fn channel<T>() -> (Sender<T>, Receiver<T>) {
struct Inner<T> { struct Inner<T> {
queue: VecDeque<T>, queue: VecDeque<T>,
/// The parked receiver's `(pid, park-epoch)`, if one is currently /// The parked receiver's `(pid, park-epoch, runtime)`, if one is currently
/// waiting. The epoch identifies exactly which wait this is, so a waker /// waiting. The epoch identifies exactly which wait this is, so a waker
/// left over from a wait that already ended (a losing `select` arm, a /// left over from a wait that already ended (a losing `select` arm, a
/// `recv_timeout` whose timer fired after it was already satisfied) is /// `recv_timeout` whose timer fired after it was already satisfied) is
/// inert and does nothing when it fires. /// inert and does nothing when it fires. The `Weak<RuntimeInner>` is the
parked_receiver: Option<(Pid, u32)>, /// receiver's runtime, captured while it parked (so provably alive then);
/// it lets a sender on a foreign OS thread wake the receiver without the
/// `RUNTIME` thread-local, which is unset off a scheduler thread.
parked_receiver: Option<(Pid, u32, Weak<RuntimeInner>)>,
senders: usize, senders: usize,
receiver_alive: bool, receiver_alive: bool,
} }
@@ -206,8 +210,8 @@ impl<T> Drop for Sender<T> {
None None
} }
}; };
if let Some((pid, epoch)) = unpark { if let Some((pid, epoch, rt)) = unpark {
crate::scheduler::unpark_at(pid, epoch); crate::scheduler::unpark_at_via(pid, epoch, &rt);
} }
} }
} }
@@ -254,13 +258,13 @@ impl<T> Sender<T> {
g.queue.push_back(value); g.queue.push_back(value);
g.parked_receiver.take() g.parked_receiver.take()
}; };
if let Some((pid, epoch)) = unpark { if let Some((pid, epoch, rt)) = unpark {
crate::te!(crate::trace::Event::Send { crate::te!(crate::trace::Event::Send {
sender: crate::actor::current_pid() sender: crate::actor::current_pid()
.unwrap_or(crate::pid::Pid::new(u32::MAX, u32::MAX)), .unwrap_or(crate::pid::Pid::new(u32::MAX, u32::MAX)),
receiver: Some(pid) receiver: Some(pid)
}); });
crate::scheduler::unpark_at(pid, epoch); crate::scheduler::unpark_at_via(pid, epoch, &rt);
} else { } else {
crate::te!(crate::trace::Event::Send { crate::te!(crate::trace::Event::Send {
sender: crate::actor::current_pid() sender: crate::actor::current_pid()
@@ -293,13 +297,17 @@ impl<T> Receiver<T> {
None => panic!("smarm: recv() called outside an actor"), None => panic!("smarm: recv() called outside an actor"),
}; };
debug_assert!( debug_assert!(
g.parked_receiver.is_none_or(|(p, _)| p == me), g.parked_receiver.as_ref().is_none_or(|(p, _, _)| *p == me),
"channel has more than one receiver" "channel has more than one receiver"
); );
// begin_wait is lock-free, so it's legal under the Channel lock; // begin_wait is lock-free, so it's legal under the Channel lock;
// registering in the same critical section makes the epoch // registering in the same critical section makes the epoch
// atomic with the senders' view of the registration. // atomic with the senders' view of the registration.
g.parked_receiver = Some((me, crate::scheduler::begin_wait())); g.parked_receiver = Some((
me,
crate::scheduler::begin_wait(),
crate::scheduler::runtime_weak(),
));
crate::te!(crate::trace::Event::RecvPark(me)); crate::te!(crate::trace::Event::RecvPark(me));
} }
// Release the lock before parking: the unparker will need it. // Release the lock before parking: the unparker will need it.
@@ -347,11 +355,11 @@ impl<T> Receiver<T> {
return Err(RecvTimeoutError::Disconnected); return Err(RecvTimeoutError::Disconnected);
} }
debug_assert!( debug_assert!(
g.parked_receiver.is_none_or(|(p, _)| p == me), g.parked_receiver.as_ref().is_none_or(|(p, _, _)| *p == me),
"channel has more than one receiver" "channel has more than one receiver"
); );
epoch = crate::scheduler::begin_wait(); epoch = crate::scheduler::begin_wait();
g.parked_receiver = Some((me, epoch)); g.parked_receiver = Some((me, epoch, crate::scheduler::runtime_weak()));
crate::te!(crate::trace::Event::RecvPark(me)); crate::te!(crate::trace::Event::RecvPark(me));
} }
@@ -423,10 +431,14 @@ impl<T> Receiver<T> {
None => panic!("smarm: recv_match() called outside an actor"), None => panic!("smarm: recv_match() called outside an actor"),
}; };
debug_assert!( debug_assert!(
g.parked_receiver.is_none_or(|(p, _)| p == me), g.parked_receiver.as_ref().is_none_or(|(p, _, _)| *p == me),
"channel has more than one receiver" "channel has more than one receiver"
); );
g.parked_receiver = Some((me, crate::scheduler::begin_wait())); g.parked_receiver = Some((
me,
crate::scheduler::begin_wait(),
crate::scheduler::runtime_weak(),
));
crate::te!(crate::trace::Event::RecvPark(me)); crate::te!(crate::trace::Event::RecvPark(me));
} }
// Release the lock before parking: the unparker will need it. // Release the lock before parking: the unparker will need it.
@@ -497,11 +509,12 @@ impl<T: Send + 'static> crate::timer::TimerTarget for RawMutex<Inner<T>> {
// keeps the registration bookkeeping exact.) // keeps the registration bookkeeping exact.)
let unpark = { let unpark = {
let mut g = self.lock(); let mut g = self.lock();
if g.parked_receiver == Some((pid, epoch)) { match g.parked_receiver {
Some((p, e, _)) if p == pid && e == epoch => {
g.parked_receiver = None; g.parked_receiver = None;
true true
} else { }
false _ => false,
} }
}; };
// Unpark outside the channel lock: it may take the run-queue lock; // Unpark outside the channel lock: it may take the run-queue lock;
@@ -562,10 +575,10 @@ impl<T> Selectable for Receiver<T> {
return Ok(false); return Ok(false);
} }
debug_assert!( debug_assert!(
g.parked_receiver.is_none_or(|(p, _)| p == pid), g.parked_receiver.as_ref().is_none_or(|(p, _, _)| *p == pid),
"channel has more than one receiver" "channel has more than one receiver"
); );
g.parked_receiver = Some((pid, epoch)); g.parked_receiver = Some((pid, epoch, crate::scheduler::runtime_weak()));
Ok(true) Ok(true)
} }
+220 -40
View File
@@ -127,16 +127,51 @@
//! - [`GenServer::init`] runs once before the first message. Use it to start //! - [`GenServer::init`] runs once before the first message. Use it to start
//! timers or set up monitors; see the [`GenServerCtx`] it receives. //! timers or set up monitors; see the [`GenServerCtx`] it receives.
//! - [`GenServer::terminate`] runs when the server is about to exit. It fires //! - [`GenServer::terminate`] runs when the server is about to exit. It fires
//! on every exit path (all `GenServerRef`s dropped, a handler panic, or an //! on every exit path (a self-stop, a graceful shutdown, a cooperative hard
//! explicit [`GenServerRef::shutdown`]), not only on clean shutdown. Keep it //! stop, or a handler panic), not only on clean shutdown.
//! short and non-blocking: if `terminate` panics while the server is already //! Keep it non-panicking: on the panic and hard-stop paths it runs
//! unwinding from a handler panic, the process aborts. //! mid-unwind, where a second panic aborts the process and where it must
//! not block (any park re-observes the stop). Only on the graceful path
//! (see below) may it do real work.
//! //!
//! ## When the server stops //! ## When the server stops
//! //!
//! The server runs as long as at least one [`GenServerRef`] exists. When the last //! A server lives until it stops, is shut down, or is killed — as an OTP
//! one is dropped, the inbox closes and the loop exits gracefully. To stop a //! process does. A [`GenServerRef`] is an *address*: cloning and dropping it
//! server explicitly and wait for it to finish, call [`GenServerRef::shutdown`]. //! never changes the server's lifetime, and a ref that nobody holds is not a
//! leak — a forgotten server idles until the run ends, when the root-exit
//! shutdown (see [`Runtime::run`](crate::Runtime::run)) takes it down with
//! every other unsupervised actor. Anything meant to live long should be
//! supervised (see *Supervised servers* below); [`start`] / [`start_under`]
//! are for scripts, tests and short-lived helpers, and the explicit close is
//! [`GenServerRef::shutdown`].
//!
//! A server can end itself: clone a [`StopHandle`] from
//! [`GenServerCtx::stop_handle`] in `init` and call [`StopHandle::stop`] from
//! any handler — the loop breaks after the current message and exits
//! *normally* (OTP's `{stop, normal}`). This is distinct from
//! `request_stop(self_pid())`, which is an abnormal `Stopped` and gets a
//! `Transient` child restarted.
//!
//! ## Graceful shutdown
//!
//! From outside, [`GenServerRef::shutdown`] (or a plain
//! [`request_shutdown`](crate::request_shutdown), which is what a supervisor
//! sends) asks the server to stop. What happens next is the server's choice:
//!
//! - By default a server does not trap exits, and the request stops it
//! outright at its next observation point — `terminate` runs mid-unwind.
//! - A server that calls [`GenServerCtx::trap_exit`] in `init` receives the
//! request as [`GenServer::handle_shutdown`]. Return
//! [`ShutdownAction::Exit`] (the default) to have the loop break and
//! `terminate` run on the normal path, where it may block; return
//! [`ShutdownAction::Continue`] to keep serving — e.g. to drain in-flight
//! work — and end the server later with a [`StopHandle`]. The supervisor's
//! [`Shutdown`](crate::supervisor::Shutdown) policy bounds how long that
//! may take before it falls back to a hard stop.
//!
//! A trapping server also receives the deaths of its linked peers as
//! [`GenServer::handle_exit`] messages instead of dying with them.
//! //!
//! If the server panics inside a handler, the panic unwinds the server thread. //! If the server panics inside a handler, the panic unwinds the server thread.
//! Any caller currently waiting in `call` sees `Err(ServerDown)`: the reply //! Any caller currently waiting in `call` sees `Err(ServerDown)`: the reply
@@ -167,9 +202,23 @@
//! Names and registration: a server can be given a static name so other //! Names and registration: a server can be given a static name so other
//! actors can reach it without holding a `GenServerRef`. Use //! actors can reach it without holding a `GenServerRef`. Use
//! [`GenServerBuilder::named`] to register on start, and the free functions //! [`GenServerBuilder::named`] to register on start, and the free functions
//! [`call`], [`cast`], and [`whereis_server`] to address it by name. Registered //! [`call`], [`cast`], and [`whereis_server`] to address it by name.
//! servers are a natural fit for supervision; see `supervisor` for how to //!
//! build a tree that restarts servers on failure. //! ## Supervised servers
//!
//! The supervised shape is [`NamedGenServerBuilder::run`]: it runs the loop
//! **inline, as the current actor**, so the closure of a
//! [`ChildSpec`](crate::supervisor::ChildSpec) *is* the server — the
//! supervisor's shutdown arrives as [`GenServer::handle_shutdown`], a restart
//! runs the factory again and re-binds the name, and the rest of the program
//! addresses it by name (a held ref would go stale on restart anyway).
//!
//! ```ignore
//! const COUNTER: GenServerName<Counter> = GenServerName::new("counter");
//! OneForOne::new().child(ChildSpec::new(Restart::Permanent, || {
//! GenServerBuilder::new(Counter::default()).named(COUNTER).run().unwrap();
//! }));
//! ```
//! //!
//! ## Limitations //! ## Limitations
//! //!
@@ -181,10 +230,12 @@
use crate::channel::{ use crate::channel::{
channel, select, select_timeout, Receiver, RecvTimeoutError, Selectable, Sender, channel, select, select_timeout, Receiver, RecvTimeoutError, Selectable, Sender,
}; };
use crate::link::ExitSignal;
use crate::monitor::DownReason;
use crate::monitor::{demonitor, monitor, Down, Monitor}; use crate::monitor::{demonitor, monitor, Down, Monitor};
use crate::pid::Pid; use crate::pid::Pid;
use crate::registry::{register_with, resolve_named_sender, RegisterError}; use crate::registry::{register_with, resolve_named_sender, RegisterError};
use crate::scheduler::{cancel_timer, request_stop, send_after_to}; use crate::scheduler::{cancel_timer, request_shutdown, send_after_to};
use crate::timer::TimerId; use crate::timer::TimerId;
use std::cell::Cell; use std::cell::Cell;
use std::collections::HashMap; use std::collections::HashMap;
@@ -254,10 +305,34 @@ pub trait GenServer: Send + 'static {
/// Default: no-op. /// Default: no-op.
fn handle_idle(&mut self) {} fn handle_idle(&mut self) {}
/// A graceful shutdown request (a [`request_shutdown`](crate::request_shutdown)
/// reaching this server), delivered only if `init` called
/// [`GenServerCtx::trap_exit`]. Return [`ShutdownAction::Exit`] to stop
/// now (the default), or [`ShutdownAction::Continue`] to keep serving and
/// end the server later with a [`StopHandle`].
fn handle_shutdown(&mut self) -> ShutdownAction {
ShutdownAction::Exit
}
/// A linked peer's abnormal death (an [`ExitSignal`] that is not a
/// shutdown request), delivered only if `init` called
/// [`GenServerCtx::trap_exit`]. Default: drop it.
fn handle_exit(&mut self, _sig: ExitSignal) {}
/// Runs as the server actor exits, on any exit path (see module docs). /// Runs as the server actor exits, on any exit path (see module docs).
fn terminate(&mut self) {} fn terminate(&mut self) {}
} }
/// What a server does with a shutdown request; see [`GenServer::handle_shutdown`].
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum ShutdownAction {
/// Break the loop now. `terminate` runs on the normal path and may block.
Exit,
/// Keep dispatching. The server is expected to end itself with a
/// [`StopHandle`] once it is done winding down.
Continue,
}
/// What travels the server's single inbox channel: a synchronous call (with a /// What travels the server's single inbox channel: a synchronous call (with a
/// reply sender) or an asynchronous cast. Private — callers use [`GenServerRef`]. /// reply sender) or an asynchronous cast. Private — callers use [`GenServerRef`].
enum Envelope<G: GenServer> { enum Envelope<G: GenServer> {
@@ -265,9 +340,9 @@ enum Envelope<G: GenServer> {
Cast(G::Cast), Cast(G::Cast),
} }
/// A clonable handle to a running server. Cloning yields another sender to the /// A clonable handle to a running server: an *address*, not an owner. Cloning
/// same inbox; the server lives until the last `GenServerRef` is dropped, at which /// yields another sender to the same inbox; dropping refs never ends the
/// point its inbox closes and the loop exits normally. /// server (see the module docs, *When the server stops*).
pub struct GenServerRef<G: GenServer> { pub struct GenServerRef<G: GenServer> {
tx: Sender<Envelope<G>>, tx: Sender<Envelope<G>>,
pid: Pid, pid: Pid,
@@ -362,18 +437,22 @@ impl<G: GenServer> GenServerRef<G> {
/// Stop the server and block until it has fully exited. /// Stop the server and block until it has fully exited.
/// ///
/// Sends a cooperative stop signal to the server actor and waits for it to /// Asks the server to shut down (a [`request_shutdown`](crate::request_shutdown))
/// exit, so [`GenServer::terminate`] has run by the time this returns. /// and waits for it to exit, so [`GenServer::terminate`] has run by the
/// Returns immediately if the server is already gone. /// time this returns. A trapping server gets to wind down via
/// [`GenServer::handle_shutdown`]; any other is stopped outright. Returns
/// immediately if the server is already gone. Waits as long as the server
/// takes — the caller, not the server, decides whether that is acceptable;
/// a supervisor uses its child's [`Shutdown`](crate::supervisor::Shutdown)
/// policy to bound it.
/// ///
/// This is the right teardown for a server kept alive by a registered /// This is the explicit close: dropping refs never ends a server. Like
/// [`GenServerName`], where dropping every external `GenServerRef` is not enough /// all cooperative cancellation, it is best-effort:
/// to close the inbox. Like all cooperative cancellation, it is best-effort:
/// a server wedged in a tight loop with no observation point cannot be /// a server wedged in a tight loop with no observation point cannot be
/// stopped this way. Panics if called outside `Runtime::run()`. /// stopped this way. Panics if called outside `Runtime::run()`.
pub fn shutdown(&self) { pub fn shutdown(&self) {
let mon = monitor(self.pid); let mon = monitor(self.pid);
request_stop(self.pid); request_shutdown(self.pid);
// The Down lands when the server finalizes; an already-dead target makes // The Down lands when the server finalizes; an already-dead target makes
// `monitor` deliver NoProc immediately, so this never blocks forever. // `monitor` deliver NoProc immediately, so this never blocks forever.
let _ = mon.rx.recv(); let _ = mon.rx.recv();
@@ -396,6 +475,8 @@ enum Sys<G: GenServer> {
/// the payload factory, dispatches it to [`GenServer::handle_timer`], and /// the payload factory, dispatches it to [`GenServer::handle_timer`], and
/// re-arms the next tick before returning. /// re-arms the next tick before returning.
Tick(crate::timer::TimerId), Tick(crate::timer::TimerId),
/// The state asked to end the server (via [`StopHandle::stop`]).
Stop,
} }
/// The server loop's runtime hook, passed to [`GenServer::init`]. Hands out the /// The server loop's runtime hook, passed to [`GenServer::init`]. Hands out the
@@ -411,9 +492,29 @@ pub struct GenServerCtx<G: GenServer> {
/// because `init` holds only `&ctx`; not `Send`, but `GenServerCtx` is only ever /// because `init` holds only `&ctx`; not `Send`, but `GenServerCtx` is only ever
/// borrowed on the actor's own stack during `init`, never sent. /// borrowed on the actor's own stack during `init`, never sent.
idle: Cell<Option<Duration>>, idle: Cell<Option<Duration>>,
/// Whether the loop should trap exits (set via [`trap_exit`](Self::trap_exit)
/// during `init`, read by the loop after).
trap: Cell<bool>,
} }
impl<G: GenServer> GenServerCtx<G> { impl<G: GenServer> GenServerCtx<G> {
/// Trap exits for the server's lifetime: shutdown requests then arrive as
/// [`GenServer::handle_shutdown`] and linked-peer deaths as
/// [`GenServer::handle_exit`], instead of stopping the server outright.
/// Call this once during `init`.
pub fn trap_exit(&self) {
self.trap.set(true);
}
/// A clonable handle that lets the state end the server from any handler
/// (a normal exit; see the module docs). Store it on the state during
/// `init`.
pub fn stop_handle(&self) -> StopHandle<G> {
StopHandle {
sys_tx: self.sys_tx.clone(),
}
}
/// A clonable handle to the loop's monitor intake. Store it in the state /// A clonable handle to the loop's monitor intake. Store it in the state
/// during `init` to watch monitors from later handlers. /// during `init` to watch monitors from later handlers.
pub fn watcher(&self) -> Watcher<G> { pub fn watcher(&self) -> Watcher<G> {
@@ -452,6 +553,30 @@ impl<G: GenServer> GenServerCtx<G> {
} }
} }
/// Lets a server's state end the server, cloned from
/// [`GenServerCtx::stop_handle`] during `init`. [`stop`](Self::stop) makes the
/// loop break after the current message and exit normally; `terminate` runs on
/// the normal path.
pub struct StopHandle<G: GenServer> {
sys_tx: Sender<Sys<G>>,
}
impl<G: GenServer> Clone for StopHandle<G> {
fn clone(&self) -> Self {
StopHandle {
sys_tx: self.sys_tx.clone(),
}
}
}
impl<G: GenServer> StopHandle<G> {
/// End the server after the current message. Idempotent; a no-op once the
/// server is gone.
pub fn stop(&self) {
let _ = self.sys_tx.send(Sys::Stop);
}
}
/// Per-server timer bookkeeping, shared between the loop and every /// Per-server timer bookkeeping, shared between the loop and every
/// [`TimerHandle`] clone. A gen_server actor is single-threaded — handlers /// [`TimerHandle`] clone. A gen_server actor is single-threaded — handlers
/// and the loop never run concurrently — so this `Mutex` is always /// and the loop never run concurrently — so this `Mutex` is always
@@ -704,9 +829,10 @@ impl<G: GenServer> GenServerBuilder<G> {
self self
} }
/// Spawn the server actor and hand back its [`GenServerRef`]. The server's /// Spawn the server actor and hand back its [`GenServerRef`] (an address;
/// lifetime is governed by its refs, not by joining, so the backing join /// the server's lifetime is its own, see the module docs). The backing
/// handle is dropped. /// join handle is dropped. For a supervised server use
/// [`named`](Self::named) + [`NamedGenServerBuilder::run`] instead.
pub fn start(self) -> GenServerRef<G> { pub fn start(self) -> GenServerRef<G> {
self.spawn_server() self.spawn_server()
} }
@@ -734,13 +860,14 @@ impl<G: GenServer> GenServerBuilder<G> {
supervisor, supervisor,
stack_opts, stack_opts,
} = self; } = self;
let keep = tx.clone();
let handle = match supervisor { let handle = match supervisor {
Some(sup) => crate::scheduler::spawn_under_with(sup, stack_opts, move || { Some(sup) => crate::scheduler::spawn_under_with(sup, stack_opts, move || {
server_loop::<G>(rx, state, infos) server_loop::<G>(keep, rx, state, infos)
}),
None => crate::scheduler::spawn_with(stack_opts, move || {
server_loop::<G>(keep, rx, state, infos)
}), }),
None => {
crate::scheduler::spawn_with(stack_opts, move || server_loop::<G>(rx, state, infos))
}
}; };
GenServerRef { GenServerRef {
tx, tx,
@@ -828,19 +955,38 @@ impl<G: GenServer> NamedGenServerBuilder<G> {
/// The inbox sender is published under the name **from the parent side, /// The inbox sender is published under the name **from the parent side,
/// before this returns**, so a by-name `call` / `cast` resolves the instant /// before this returns**, so a by-name `call` / `cast` resolves the instant
/// `start()` returns — no race with the server body. On a name clash the /// `start()` returns — no race with the server body. On a name clash the
/// just-spawned server is wound down (its only ref is dropped, closing the /// just-spawned server is stopped, so a failed bind leaks no actor.
/// inbox), so a failed bind leaks no actor.
pub fn start(self) -> Result<GenServerRef<G>, RegisterError> { pub fn start(self) -> Result<GenServerRef<G>, RegisterError> {
let NamedGenServerBuilder { builder, name } = self; let NamedGenServerBuilder { builder, name } = self;
let server = builder.spawn_server(); let server = builder.spawn_server();
match register_with::<Envelope<G>>(server.pid, name, server.tx.clone()) { match register_with::<Envelope<G>>(server.pid, name, server.tx.clone()) {
Ok(()) => Ok(server), Ok(()) => Ok(server),
Err(e) => { Err(e) => {
drop(server); // inbox closes → loop exits gracefully crate::scheduler::request_stop(server.pid); // never ran init
Err(e) Err(e)
} }
} }
} }
/// Run the server **inline, as the current actor**, bound to its name.
/// This is the supervised shape: the closure of a
/// [`ChildSpec`](crate::supervisor::ChildSpec) *is* the server, so the
/// supervisor's shutdown reaches it as [`GenServer::handle_shutdown`], a
/// restart runs the factory again and re-binds the name, and clients
/// address it by name ([`call`], [`cast`], [`whereis_server`]). Returns
/// when the server exits; [`RegisterError::NameTaken`] (before `init`) if
/// the name is held by a different live actor.
///
/// `under` / `stack_opts` are spawn options and do not apply here — the
/// actor already exists.
pub fn run(self) -> Result<(), RegisterError> {
let NamedGenServerBuilder { builder, name } = self;
let GenServerBuilder { state, infos, .. } = builder;
let (tx, rx) = channel::<Envelope<G>>();
register_with::<Envelope<G>>(crate::scheduler::self_pid(), name, tx.clone())?;
server_loop::<G>(tx, rx, state, infos);
Ok(())
}
} }
/// Resolve a [`GenServerName`] to a [`GenServerRef`] when you want a handle to hold or /// Resolve a [`GenServerName`] to a [`GenServerRef`] when you want a handle to hold or
@@ -899,10 +1045,17 @@ pub fn start_under<G: GenServer>(supervisor: Pid, state: G) -> GenServerRef<G> {
} }
fn server_loop<G: GenServer>( fn server_loop<G: GenServer>(
keep: Sender<Envelope<G>>,
rx: Receiver<Envelope<G>>, rx: Receiver<Envelope<G>>,
state: G, state: G,
mut infos: Vec<Receiver<G::Info>>, mut infos: Vec<Receiver<G::Info>>,
) { ) {
// The loop holds one inbox sender for its whole life: the inbox never
// closes, so refs are addresses and the server's lifetime is the actor's
// (stop handle, shutdown, stop, panic). The `Disconnected` arms below are
// defensive only.
let _keep = keep;
// Drop guard — owns the server state and the timer registry. // Drop guard — owns the server state and the timer registry.
// //
// Why a guard rather than code after the loop: // Why a guard rather than code after the loop:
@@ -974,9 +1127,14 @@ fn server_loop<G: GenServer>(
sys_tx, sys_tx,
reg: reg.clone(), reg: reg.clone(),
idle: Cell::new(None), idle: Cell::new(None),
trap: Cell::new(false),
}; };
guard.0.init(&ctx); guard.0.init(&ctx);
let idle = ctx.idle.get(); let idle = ctx.idle.get();
// Trapping is opted into during init and fixed for the loop's life. The
// inbox is armed only when set: an untrapped server keeps the fast path,
// and a shutdown request simply stops it as `request_stop` would.
let exits: Option<Receiver<ExitSignal>> = ctx.trap.get().then(crate::link::trap_exit);
drop(ctx); drop(ctx);
let mut monitors: Vec<Monitor> = Vec::new(); let mut monitors: Vec<Monitor> = Vec::new();
@@ -992,7 +1150,7 @@ fn server_loop<G: GenServer>(
}; };
loop { loop {
if monitors.is_empty() && !sys_open && infos.is_empty() { if exits.is_none() && monitors.is_empty() && !sys_open && infos.is_empty() {
// Fast path: no extra arms, no select overhead — park directly on // Fast path: no extra arms, no select overhead — park directly on
// the inbox. Mirrors the inbox arm of the select path below; any // the inbox. Mirrors the inbox arm of the select path below; any
// change there must be applied here too. // change there must be applied here too.
@@ -1009,7 +1167,7 @@ fn server_loop<G: GenServer>(
guard.0.handle_idle(); guard.0.handle_idle();
reset_idle(&mut idle_deadline); reset_idle(&mut idle_deadline);
} }
// All ServerRefs dropped → inbox closed → shutdown. // Defensive: the loop holds a sender, so unreachable.
Err(RecvTimeoutError::Disconnected) => break, Err(RecvTimeoutError::Disconnected) => break,
} }
} }
@@ -1020,16 +1178,21 @@ fn server_loop<G: GenServer>(
} }
} else { } else {
// Slow path: one or more extra arms live — build the arm slice and // Slow path: one or more extra arms live — build the arm slice and
// select. Arm order encodes priority: downs → system → infos → // select. Arm order encodes priority: exits → downs → system →
// inbox. The slice is rebuilt each iteration because the monitor // infos → inbox (a shutdown request is noticed under any load).
// and info sets shrink/grow. Mirrors the fast-path inbox park // The slice is rebuilt each iteration because the monitor and info
// above; keep them in sync. // sets shrink/grow. Mirrors the fast-path inbox park above; keep
let nd = monitors.len(); // monitor band: [0, nd) // them in sync.
let ne = exits.is_some() as usize; // exit arm: [0, ne)
let nd = ne + monitors.len(); // monitor band: [ne, nd)
let nw = sys_open as usize; // system arm: [nd, nd+nw) let nw = sys_open as usize; // system arm: [nd, nd+nw)
// info band: [nd+nw, nd+nw+ni) // info band: [nd+nw, nd+nw+ni)
// inbox arm: [nd+nw+ni] // inbox arm: [nd+nw+ni]
let sel = { let sel = {
let mut arms: Vec<&dyn Selectable> = Vec::with_capacity(nd + nw + infos.len() + 1); let mut arms: Vec<&dyn Selectable> = Vec::with_capacity(nd + nw + infos.len() + 1);
if let Some(e) = &exits {
arms.push(e);
}
for m in &monitors { for m in &monitors {
arms.push(&m.rx); arms.push(&m.rx);
} }
@@ -1055,10 +1218,25 @@ fn server_loop<G: GenServer>(
continue; continue;
} }
}; };
if i < nd { if i < ne {
// Exit arm: a shutdown request or a linked peer's death.
// The inbox lives for the loop's life, so it never closes.
let sig = exits.as_ref().and_then(|e| e.try_recv().ok().flatten());
if let Some(sig) = sig {
if sig.reason == DownReason::Shutdown {
match guard.0.handle_shutdown() {
ShutdownAction::Exit => break,
ShutdownAction::Continue => {}
}
} else {
guard.0.handle_exit(sig);
}
reset_idle(&mut idle_deadline);
}
} else if i < nd {
// Monitor band: a Down retires its arm either way (one-shot) // Monitor band: a Down retires its arm either way (one-shot)
// or closes without delivering (defensive; shouldn't happen). // or closes without delivering (defensive; shouldn't happen).
let m = monitors.remove(i); let m = monitors.remove(i - ne);
if let Ok(Some(down)) = m.rx.try_recv() { if let Ok(Some(down)) = m.rx.try_recv() {
guard.0.handle_down(down); guard.0.handle_down(down);
reset_idle(&mut idle_deadline); reset_idle(&mut idle_deadline);
@@ -1067,6 +1245,8 @@ fn server_loop<G: GenServer>(
match sys_rx.try_recv() { match sys_rx.try_recv() {
// Control intake, not a dispatched message: no idle reset. // Control intake, not a dispatched message: no idle reset.
Ok(Some(Sys::Watch(m))) => monitors.push(m), Ok(Some(Sys::Watch(m))) => monitors.push(m),
// The state ended the server: a normal exit.
Ok(Some(Sys::Stop)) => break,
Ok(Some(Sys::Timer(id, msg))) => { Ok(Some(Sys::Timer(id, msg))) => {
// The one-shot fired: retire its registry entry so the // The one-shot fired: retire its registry entry so the
// live set tracks only still-pending timers, then // live set tracks only still-pending timers, then
+469 -66
View File
@@ -68,10 +68,38 @@
//! inbox or timer event; a replayed event may postpone again (it re-queues for //! inbox or timer event; a replayed event may postpone again (it re-queues for
//! the next transition). See the macro docs for the row surface and [`Step`] //! the next transition). See the macro docs for the row surface and [`Step`]
//! for how a postpone surfaces to the loop. //! for how a postpone surfaces to the loop.
//!
//! ## Stopping, shutdown, and exits
//!
//! A machine lives until it stops, is shut down, or is killed; a
//! [`GenStatemRef`] is an address, and dropping refs never ends it (the
//! gen_server rule — see its *When the server stops*). The supervised shape
//! is [`run_named`], which runs the machine inline as the current actor so it
//! is a direct `ChildSpec` child addressed by [`GenStatemName`]. A machine
//! can end itself: any row
//! body may call [`cx.stop()`](Cx::stop) (or use the `stop` tail keyword,
//! sugar for `{ cx.stop(); prev }`) — the loop breaks after that event and the
//! actor exits *normally* (OTP's `{stop, normal}`; a `Transient` child is not
//! restarted). [`terminate`](Machine::terminate) — the optional `terminate
//! { … }` macro block — runs on every exit path.
//!
//! From outside, [`GenStatemRef::shutdown`] (or a plain
//! [`request_shutdown`](crate::request_shutdown), which is what a supervisor
//! sends) asks the machine to stop. By default a machine does not trap exits
//! and the request stops it outright. A machine that calls
//! [`cx.trap_exit()`](Cx::trap_exit) in its initial `enter` instead receives
//! it as the **`shutdown` event**, routed by state like any other — a
//! `Connected` state may transition into `Draining` and stop later from a
//! timeout row, while a state with no `shutdown` row takes the macro's
//! default, `stop`. Linked-peer deaths reach a trapping machine as
//! `exit <pat>` events; an unmatched one is dropped like an unmatched info.
use crate::channel::{channel, select, Receiver, Sender}; use crate::channel::{channel, select, Receiver, Selectable, Sender};
use crate::link::ExitSignal;
use crate::monitor::{monitor, DownReason};
use crate::pid::Pid; use crate::pid::Pid;
use crate::scheduler::{cancel_timer, send_after_to}; use crate::registry::{register_with, resolve_named_sender, RegisterError};
use crate::scheduler::{cancel_timer, request_shutdown, send_after_to};
use crate::timer::TimerId; use crate::timer::TimerId;
use std::collections::{HashMap, VecDeque}; use std::collections::{HashMap, VecDeque};
use std::marker::PhantomData; use std::marker::PhantomData;
@@ -119,6 +147,33 @@ pub trait Machine: Send + 'static {
/// state's `enter`, and returns [`Step::Transitioned`]; a stay or unmatched /// state's `enter`, and returns [`Step::Transitioned`]; a stay or unmatched
/// event returns [`Step::Stayed`]. /// event returns [`Step::Stayed`].
fn handle(&mut self, ev: Self::Ev, cx: &mut Cx<Self::Ev>) -> Step<Self::Ev>; fn handle(&mut self, ev: Self::Ev, cx: &mut Cx<Self::Ev>) -> Step<Self::Ev>;
/// Wrap a graceful shutdown request into this machine's event, so a
/// trapping machine (see [`Cx::trap_exit`]) can route it **by state**. The
/// macro generates it as `Ev::Shutdown` and matches it in `shutdown` rows;
/// its default for a state that writes no such row is `stop`. A
/// hand-written machine that returns `None` (the default) is simply
/// stopped — the loop breaks and `terminate` runs on the normal path.
fn shutdown_ev() -> Option<Self::Ev> {
None
}
/// Wrap a linked peer's death (an [`ExitSignal`] that is not a shutdown
/// request, delivered only when trapping) into this machine's event. The
/// macro generates it as `Ev::Exit(sig)` and matches it in `exit <pat>`
/// rows; an unmatched exit is silently dropped, like an unmatched info.
/// A hand-written machine that returns `None` (the default) drops it.
fn exit_ev(_sig: ExitSignal) -> Option<Self::Ev> {
None
}
/// Runs as the machine actor exits, on any exit path (a `stop`, a
/// graceful shutdown, a handler panic, a hard stop). Like
/// `gen_server::terminate`: on the panic and hard-stop paths it runs
/// mid-unwind — do not panic or park there; only on the normal path
/// (`stop`, shutdown rows) may it do real work. The macro's optional
/// `terminate { … }` block generates it.
fn terminate(&mut self) {}
} }
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -247,6 +302,11 @@ impl Timers {
pub struct Cx<Ev> { pub struct Cx<Ev> {
sys_tx: Sender<Sys>, sys_tx: Sender<Sys>,
reg: Arc<Mutex<Timers>>, reg: Arc<Mutex<Timers>>,
/// Set by [`trap_exit`](Self::trap_exit) during `on_start`; read once by
/// the loop right after, fixed for the machine's life.
trap: bool,
/// Set by [`stop`](Self::stop); the loop breaks after the current event.
stop: bool,
_ev: PhantomData<fn() -> Ev>, _ev: PhantomData<fn() -> Ev>,
} }
@@ -255,10 +315,30 @@ impl<Ev> Cx<Ev> {
Cx { Cx {
sys_tx, sys_tx,
reg, reg,
trap: false,
stop: false,
_ev: PhantomData, _ev: PhantomData,
} }
} }
/// Trap exits for the machine's lifetime: a shutdown request then arrives
/// as the `shutdown` event (routed by state) and linked-peer deaths as
/// `exit` events, instead of stopping the machine outright. Call it in the
/// initial state's `enter` (i.e. during `on_start`); later calls have no
/// effect.
pub fn trap_exit(&mut self) {
self.trap = true;
}
/// End the machine after the current event: the loop breaks and the actor
/// exits *normally* (OTP's `{stop, normal}`); `terminate` runs on the
/// normal path. Anything still queued or postponed is dropped. The `stop`
/// tail keyword in a macro row is sugar for `{ cx.stop(); prev }`.
/// Mirrors gen_server's [`StopHandle`](crate::gen_server::StopHandle).
pub fn stop(&mut self) {
self.stop = true;
}
/// Arm the **state timeout**: fire a `state_timeout` event after `after` in /// Arm the **state timeout**: fire a `state_timeout` event after `after` in
/// the current state. Auto-reset on any state change (the loop cancels and /// the current state. Auto-reset on any state change (the loop cancels and
/// clears it on every real transition), so it measures quiet time *within* a /// clears it on every real transition), so it measures quiet time *within* a
@@ -386,8 +466,7 @@ pub enum SendError {
} }
/// A clonable handle to a running machine. Cloning yields another sender to the /// A clonable handle to a running machine. Cloning yields another sender to the
/// same inbox; the machine lives until the last `GenStatemRef` is dropped, at which /// same inbox. An address, not an owner: dropping refs never ends the machine.
/// point its inbox closes and the loop exits.
pub struct GenStatemRef<M: Machine> { pub struct GenStatemRef<M: Machine> {
tx: Sender<M::Ev>, tx: Sender<M::Ev>,
pid: Pid, pid: Pid,
@@ -433,6 +512,19 @@ impl<M: Machine> GenStatemRef<M> {
self.send(ev).map_err(|_| CallError::Down)?; self.send(ev).map_err(|_| CallError::Down)?;
rx.recv().map_err(|_| CallError::Down) rx.recv().map_err(|_| CallError::Down)
} }
/// Ask the machine to shut down and block until it has fully exited, so
/// [`Machine::terminate`] has run by the time this returns. A trapping
/// machine winds down through its `shutdown` rows; any other is stopped
/// outright. Returns immediately if the machine is already gone. Waits as
/// long as the machine takes — a supervisor bounds that with its child's
/// [`Shutdown`](crate::supervisor::Shutdown) policy. Mirrors
/// [`GenServerRef::shutdown`](crate::gen_server::GenServerRef::shutdown).
pub fn shutdown(&self) {
let mon = monitor(self.pid);
request_shutdown(self.pid);
let _ = mon.rx.recv();
}
} }
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -441,7 +533,8 @@ impl<M: Machine> GenStatemRef<M> {
/// Spawn `machine` as an actor and hand back its [`GenStatemRef`]. Shape mirrors /// Spawn `machine` as an actor and hand back its [`GenStatemRef`]. Shape mirrors
/// `gen_server::start`: make the inbox, spawn the loop, return the ref; the /// `gen_server::start`: make the inbox, spawn the loop, return the ref; the
/// backing join handle is dropped (lifetime is governed by refs, not joining). /// backing join handle is dropped (the machine's lifetime is its own). For a
/// supervised machine use [`run_named`].
/// ///
/// Panics if called outside `Runtime::run()`. /// Panics if called outside `Runtime::run()`.
pub fn spawn<M: Machine>(machine: M) -> GenStatemRef<M> { pub fn spawn<M: Machine>(machine: M) -> GenStatemRef<M> {
@@ -456,40 +549,188 @@ pub fn spawn<M: Machine>(machine: M) -> GenStatemRef<M> {
/// Panics if called outside `Runtime::run()`. /// Panics if called outside `Runtime::run()`.
pub fn spawn_with<M: Machine>(opts: crate::scheduler::SpawnOpts, machine: M) -> GenStatemRef<M> { pub fn spawn_with<M: Machine>(opts: crate::scheduler::SpawnOpts, machine: M) -> GenStatemRef<M> {
let (tx, rx) = channel::<M::Ev>(); let (tx, rx) = channel::<M::Ev>();
let handle = crate::scheduler::spawn_with(opts, move || statem_loop(rx, machine)); let keep = tx.clone();
let handle = crate::scheduler::spawn_with(opts, move || statem_loop(keep, rx, machine));
GenStatemRef { GenStatemRef {
tx, tx,
pid: handle.pid(), pid: handle.pid(),
} }
} }
/// A typed, static name for a gen_statem, used to address a machine through
/// the registry without holding a [`GenStatemRef`]. Mirrors
/// [`GenServerName`](crate::gen_server::GenServerName): declare it as a
/// constant and bind it with [`run_named`].
pub struct GenStatemName<M> {
name: &'static str,
_marker: PhantomData<fn() -> M>,
}
impl<M> GenStatemName<M> {
/// Bind a static string as a machine name.
#[inline]
pub const fn new(name: &'static str) -> Self {
Self {
name,
_marker: PhantomData,
}
}
/// The underlying registry key.
#[inline]
pub const fn as_str(self) -> &'static str {
self.name
}
}
impl<M> Copy for GenStatemName<M> {}
impl<M> Clone for GenStatemName<M> {
fn clone(&self) -> Self {
*self
}
}
/// Run `machine` **inline, as the current actor**, bound to `name`. The
/// supervised shape: the closure of a
/// [`ChildSpec`](crate::supervisor::ChildSpec) *is* the machine, so the
/// supervisor's shutdown reaches it as a `shutdown` row, a restart runs the
/// factory again and re-binds the name, and clients address it by name
/// ([`send`], [`call`], [`whereis_machine`]). Returns when the machine exits;
/// [`RegisterError::NameTaken`] (before `on_start`) if the name is held by a
/// different live actor. Mirrors
/// [`NamedGenServerBuilder::run`](crate::gen_server::NamedGenServerBuilder::run).
pub fn run_named<M: Machine>(name: GenStatemName<M>, machine: M) -> Result<(), RegisterError> {
let (tx, rx) = channel::<M::Ev>();
register_with::<M::Ev>(crate::scheduler::self_pid(), name.as_str(), tx.clone())?;
statem_loop(tx, rx, machine);
Ok(())
}
/// Resolve a [`GenStatemName`] to a [`GenStatemRef`]; `None` if no live
/// machine holds the name. Panics if called outside `Runtime::run()`.
pub fn whereis_machine<M: Machine>(name: GenStatemName<M>) -> Option<GenStatemRef<M>> {
resolve_named_sender::<M::Ev>(name.as_str()).map(|(pid, tx)| GenStatemRef { tx, pid })
}
/// Push an event to the machine registered under `name`, resolving per send.
/// [`SendError::Down`] if no live machine holds the name.
pub fn send<M: Machine>(name: GenStatemName<M>, ev: M::Ev) -> Result<(), SendError> {
match whereis_machine(name) {
Some(m) => m.send(ev),
None => Err(SendError::Down),
}
}
/// Synchronous request-reply to the machine registered under `name`,
/// resolving per call (a machine restarted under the same name is reached
/// transparently). [`CallError::Down`] if no live machine holds the name.
pub fn call<M, T, F>(name: GenStatemName<M>, make: F) -> Result<T, CallError>
where
M: Machine,
T: Send + 'static,
F: FnOnce(Reply<T>) -> M::Ev,
{
match whereis_machine(name) {
Some(m) => m.call(make),
None => Err(CallError::Down),
}
}
/// Shut down the machine registered under `name` and wait for it (see
/// [`GenStatemRef::shutdown`]). A no-op if no live machine holds the name.
pub fn shutdown<M: Machine>(name: GenStatemName<M>) {
if let Some(m) = whereis_machine(name) {
m.shutdown();
}
}
/// The machine actor body: `on_start`, then one `handle` per event until the /// The machine actor body: `on_start`, then one `handle` per event until the
/// inbox closes (all refs dropped → graceful shutdown). /// row resolves to `stop`, a shutdown row stops it, or the actor is stopped
/// from outside.
/// ///
/// Two intake sources are selected each iteration with the **timer arm above /// Intake arms are selected each iteration in priority order — **exits**
/// the inbox**, so a timeout fire is never starved by inbox traffic: `sys_rx` /// (only when trapping) above **timers** above the **inbox** — so a shutdown
/// carries timer fires armed through `cx`, `rx` is the user inbox. A fire is /// request or a timeout fire is never starved by inbox traffic. `sys_rx`
/// turned into the matching internal event (`state_timeout` / `timeout(name)`) /// carries timer fires armed through `cx`; a fire is turned into the matching
/// and run through the same `handle` dispatch as an inbox event — the /// internal event (`state_timeout` / `timeout(name)`) and run through the same
/// gen_statem model, where timeouts surface as ordinary events. /// `handle` dispatch as an inbox event — the gen_statem model, where timeouts
/// (and, when trapping, shutdown and exits) surface as ordinary events.
/// ///
/// The loop owns the **postpone queue**: a `handle` that defers its event hands /// The loop owns the **postpone queue**: a `handle` that defers its event hands
/// it back ([`Step::Postponed`]) for the queue; a `handle` that transitions /// it back ([`Step::Postponed`]) for the queue; a `handle` that transitions
/// ([`Step::Transitioned`]) triggers a [`replay`] of the queue in the new /// ([`Step::Transitioned`]) triggers a [`replay`] of the queue in the new
/// state, ahead of the next intake. /// state, ahead of the next intake.
fn statem_loop<M: Machine>(rx: Receiver<M::Ev>, mut machine: M) { fn statem_loop<M: Machine>(keep: Sender<M::Ev>, rx: Receiver<M::Ev>, machine: M) {
// One inbox sender lives with the loop: the inbox never closes, refs are
// addresses, the machine's lifetime is the actor's (stop row, shutdown,
// stop, panic). The `Disconnected` inbox arm below is defensive only.
let _keep = keep;
// Drop guard — owns the machine and the timer registry, so `terminate`
// fires on every exit path (clean close, `stop`, a handler panic, a hard
// stop) and the timer drain is sequenced before it. Same shape as
// gen_server's guard; see the rationale there.
struct Terminate<M: Machine>(M, Arc<Mutex<Timers>>);
impl<M: Machine> Drop for Terminate<M> {
fn drop(&mut self) {
{
let mut reg = match self.1.lock() {
Ok(g) => g,
Err(e) => panic!("smarm: gen_statem reg lock poisoned (core corrupt): {e}"),
};
if let Some((_, sub)) = reg.state.take() {
cancel_timer(sub);
}
for (_, (_, sub)) in reg.named.drain() {
cancel_timer(sub);
}
}
self.0.terminate();
}
}
let (sys_tx, sys_rx) = channel::<Sys>(); let (sys_tx, sys_rx) = channel::<Sys>();
let reg = Arc::new(Mutex::new(Timers::new())); let reg = Arc::new(Mutex::new(Timers::new()));
let mut guard = Terminate(machine, reg.clone());
// The loop owns `cx` (and through it a `sys_tx` clone) for its whole life, // The loop owns `cx` (and through it a `sys_tx` clone) for its whole life,
// so the sys arm never closes from under us — no auto-close dance needed. // so the sys arm never closes from under us — no auto-close dance needed.
let mut cx = Cx::new(sys_tx, reg.clone()); let mut cx = Cx::new(sys_tx, reg.clone());
// Events deferred by `postpone` rows, replayed FIFO on the next transition. // Events deferred by `postpone` rows, replayed FIFO on the next transition.
let mut postpone: VecDeque<M::Ev> = VecDeque::new(); let mut postpone: VecDeque<M::Ev> = VecDeque::new();
machine.on_start(&mut cx); guard.0.on_start(&mut cx);
// Trapping is opted into during on_start and fixed for the loop's life.
// The inbox is armed only when set: an untrapped machine keeps the
// two-arm select, and a shutdown request simply stops it as
// `request_stop` would.
let exits: Option<Receiver<ExitSignal>> = cx.trap.then(crate::link::trap_exit);
loop { loop {
// Timer arm first: a ready fire is taken in preference to the inbox. // Arm order encodes priority: exits → timers → inbox.
let i = select(&[&sys_rx, &rx]); let ne = exits.is_some() as usize;
if i == 0 { let i = {
let mut arms: Vec<&dyn Selectable> = Vec::with_capacity(3);
if let Some(e) = &exits {
arms.push(e);
}
arms.push(&sys_rx);
arms.push(&rx);
select(&arms)
};
if i < ne {
// Exit arm: a shutdown request or a linked peer's death. The trap
// inbox lives for the loop's life, so it never closes.
let sig = exits.as_ref().and_then(|e| e.try_recv().ok().flatten());
match sig {
Some(sig) if sig.reason == DownReason::Shutdown => match M::shutdown_ev() {
Some(ev) => dispatch(&mut guard.0, &mut cx, &mut postpone, ev),
None => cx.stop(),
},
Some(sig) => {
if let Some(ev) = M::exit_ev(sig) {
dispatch(&mut guard.0, &mut cx, &mut postpone, ev);
}
}
None => {}
}
} else if i == ne {
match sys_rx.try_recv() { match sys_rx.try_recv() {
Ok(Some(fire)) => { Ok(Some(fire)) => {
// Confirm the fire is still the live one before dispatching: // Confirm the fire is still the live one before dispatching:
@@ -528,7 +769,7 @@ fn statem_loop<M: Machine>(rx: Receiver<M::Ev>, mut machine: M) {
} }
}; };
if let Some(ev) = ev { if let Some(ev) = ev {
dispatch(&mut machine, &mut cx, &mut postpone, ev); dispatch(&mut guard.0, &mut cx, &mut postpone, ev);
} }
} }
// Single-receiver: nothing can drain the arm between select's // Single-receiver: nothing can drain the arm between select's
@@ -540,12 +781,16 @@ fn statem_loop<M: Machine>(rx: Receiver<M::Ev>, mut machine: M) {
} }
} else { } else {
match rx.try_recv() { match rx.try_recv() {
Ok(Some(ev)) => dispatch(&mut machine, &mut cx, &mut postpone, ev), Ok(Some(ev)) => dispatch(&mut guard.0, &mut cx, &mut postpone, ev),
Ok(None) => debug_assert!(false, "ready inbox was empty"), Ok(None) => debug_assert!(false, "ready inbox was empty"),
// All GenStatemRefs dropped → inbox closed → shutdown. // Defensive: the loop holds a sender, so unreachable.
Err(_) => break, Err(_) => break,
} }
} }
// A handler (or a replay) asked to stop: a normal exit.
if cx.stop {
break;
}
// Observation point so a machine fed a hot inbox stays preemptible and // Observation point so a machine fed a hot inbox stays preemptible and
// cancellable. // cancellable.
crate::check!(); crate::check!();
@@ -554,7 +799,8 @@ fn statem_loop<M: Machine>(rx: Receiver<M::Ev>, mut machine: M) {
/// Run one event through `handle` and act on its [`Step`]: stash a deferred /// Run one event through `handle` and act on its [`Step`]: stash a deferred
/// event on the postpone queue, or — on a real transition — [`replay`] the /// event on the postpone queue, or — on a real transition — [`replay`] the
/// queue in the new state. A stay/unmatched event needs nothing further. /// queue in the new state. A stay/unmatched event needs nothing further. A
/// [`Cx::stop`] raised by the handler skips the replay; the loop breaks next.
fn dispatch<M: Machine>( fn dispatch<M: Machine>(
machine: &mut M, machine: &mut M,
cx: &mut Cx<M::Ev>, cx: &mut Cx<M::Ev>,
@@ -564,14 +810,19 @@ fn dispatch<M: Machine>(
match machine.handle(ev, cx) { match machine.handle(ev, cx) {
Step::Postponed(ev) => postpone.push_back(ev), Step::Postponed(ev) => postpone.push_back(ev),
Step::Stayed => {} Step::Stayed => {}
Step::Transitioned => replay(machine, cx, postpone), Step::Transitioned => {
if !cx.stop {
replay(machine, cx, postpone)
}
}
} }
} }
/// Replay deferred events after a real transition: each goes back through /// Replay deferred events after a real transition: each goes back through
/// `handle` in FIFO order, in the now-current state. An event that postpones /// `handle` in FIFO order, in the now-current state. An event that postpones
/// again re-queues (to wait for the *next* transition); one that transitions /// again re-queues (to wait for the *next* transition); one that transitions
/// re-arms the replay, so a later state can in turn drain what is still pending. /// re-arms the replay, so a later state can in turn drain what is still pending;
/// one that raises [`Cx::stop`] ends the replay (and the machine) at once.
/// Subsequent events in a batch already see the post-transition state, since /// Subsequent events in a batch already see the post-transition state, since
/// `handle` reads the live state cell — the outer loop only re-runs to give /// `handle` reads the live state cell — the outer loop only re-runs to give
/// re-queued events another pass once a transition has occurred within a batch. /// re-queued events another pass once a transition has occurred within a batch.
@@ -590,6 +841,9 @@ fn replay<M: Machine>(machine: &mut M, cx: &mut Cx<M::Ev>, postpone: &mut VecDeq
Step::Stayed => {} Step::Stayed => {}
Step::Transitioned => transitioned = true, Step::Transitioned => transitioned = true,
} }
if cx.stop {
return;
}
} }
if !transitioned { if !transitioned {
return; return;
@@ -691,8 +945,9 @@ fn replay<M: Machine>(machine: &mut M, cx: &mut Cx<M::Ev>, postpone: &mut VecDeq
/// // The transition table. Group rows by current state with `on <pat>`. /// // The transition table. Group rows by current state with `on <pat>`.
/// // A row is: <kind> <event-pattern> [if <guard>] => <tail> , /// // A row is: <kind> <event-pattern> [if <guard>] => <tail> ,
/// // where <kind> is one of `cast`, `call`, `info`, `state_timeout` /// // where <kind> is one of `cast`, `call`, `info`, `state_timeout`
/// // (no pattern — it is a unit event), or `timeout <name-pattern>`, /// // (no pattern — it is a unit event), `timeout <name-pattern>`,
/// // and the tail is one of: /// // `shutdown` (unit; trapping machines only) or `exit <sig-pattern>`
/// // (trapping only), and the tail is one of:
/// // * a state tag `Door::Closed` (transition, or "stay" /// // * a state tag `Door::Closed` (transition, or "stay"
/// // if it equals current) /// // if it equals current)
/// // * a block ending in one `{ data.enters += 1; Door::Closed }` /// // * a block ending in one `{ data.enters += 1; Door::Closed }`
@@ -701,6 +956,9 @@ fn replay<M: Machine>(machine: &mut M, cx: &mut Cx<M::Ev>, postpone: &mut VecDeq
/// // * the keyword `postpone` (defer until next /// // * the keyword `postpone` (defer until next
/// // transition; cast/call/ /// // transition; cast/call/
/// // info only) /// // info only)
/// // * the keyword `stop` (end the machine
/// // normally; sugar for
/// // `{ cx.stop(); prev }`)
/// on Door::Open => { /// on Door::Open => {
/// cast DoorCast::Push => Door::Closed, /// cast DoorCast::Push => Door::Closed,
/// // An armed state-timeout surfaces as an ordinary event: /// // An armed state-timeout surfaces as an ordinary event:
@@ -724,6 +982,9 @@ fn replay<M: Machine>(machine: &mut M, cx: &mut Cx<M::Ev>, postpone: &mut VecDeq
/// on _ => { /// on _ => {
/// call DoorCall::GetState(r) => { r.reply(prev); prev }, /// call DoorCall::GetState(r) => { r.reply(prev); prev },
/// } /// }
///
/// // Optional: runs as the machine exits, on every exit path.
/// terminate { data.enters = 0; }
/// } /// }
/// ``` /// ```
/// ///
@@ -748,6 +1009,12 @@ fn replay<M: Machine>(machine: &mut M, cx: &mut Cx<M::Ev>, postpone: &mut VecDeq
/// ignores info writes no `info` rows at all. State-timeouts and named /// ignores info writes no `info` rows at all. State-timeouts and named
/// timeouts have **no** such default: a state that can see one must handle it /// timeouts have **no** such default: a state that can see one must handle it
/// (or `unhandled` it) or the match is non-exhaustive. /// (or `unhandled` it) or the match is non-exhaustive.
/// * **`shutdown` defaults to `stop`, `exit` to a silent drop.** Both reach a
/// machine only if its initial `enter` called `cx.trap_exit()`; a
/// non-trapping machine is stopped outright by a shutdown request. Write
/// `shutdown => …` rows only in the states that want to wind down first
/// (transition into a draining state, `stop` later); write `exit sig => …`
/// rows to react to linked-peer deaths. Neither is postponable.
/// * **Stay** = return the current tag. The `prev` you named in `context` is /// * **Stay** = return the current tag. The `prev` you named in `context` is
/// bound to the pre-handler state for exactly this — handy in any-state /// bound to the pre-handler state for exactly this — handy in any-state
/// (`on _`) rows where there is no single literal tag to write. /// (`on _`) rows where there is no single literal tag to write.
@@ -773,12 +1040,12 @@ fn replay<M: Machine>(machine: &mut M, cx: &mut Cx<M::Ev>, postpone: &mut VecDeq
/// # What it emits /// # What it emits
/// ///
/// The unified `enum $Ev` (the `Cast`/`Call`/`Info` wrappers plus the internal /// The unified `enum $Ev` (the `Cast`/`Call`/`Info` wrappers plus the internal
/// `StateTimeout` / `Timeout` events), `struct $Sm { state, data }`, /// `StateTimeout` / `Timeout` / `Shutdown` / `Exit` events), `struct $Sm {
/// `$Sm::start(init, data) -> GenStatemRef<$Sm>`, and the `Machine` impl: /// state, data }`, `$Sm::start(init, data) -> GenStatemRef<$Sm>`, and the
/// `on_start` runs the initial `enter`; `handle` is the dispatch match plus the /// `Machine` impl: `on_start` runs the initial `enter`; `handle` is the
/// stay/transition/unhandled apply-tail (the cell's sole writer, which also /// dispatch match plus the stay/transition/unhandled apply-tail (the cell's
/// auto-resets the state-timeout on every real transition); and the `enter` /// sole writer, which also auto-resets the state-timeout on every real
/// dispatch. /// transition); the `enter` dispatch; and `terminate` when the block is given.
/// ///
/// # Limitation /// # Limitation
/// ///
@@ -794,9 +1061,10 @@ macro_rules! gen_statem {
context ( $data:ident , $cur:ident , $cx:ident ) ; context ( $data:ident , $cur:ident , $cx:ident ) ;
enter { $( $est:pat => $ebody:expr ),+ $(,)? } enter { $( $est:pat => $ebody:expr ),+ $(,)? }
$( on $st:pat => { $($rows:tt)* } )+ $( on $st:pat => { $($rows:tt)* } )+
$( terminate { $($tbody:tt)* } )?
) => { ) => {
/// Unified inbox payload: the user's `cast`/`call`/`info` enums folded /// Unified inbox payload: the user's `cast`/`call`/`info` enums folded
/// together with the runtime's internal timeout events. /// together with the runtime's internal events.
enum $Ev { enum $Ev {
Cast($Cast), Cast($Cast),
Call($Call), Call($Call),
@@ -807,6 +1075,12 @@ macro_rules! gen_statem {
StateTimeout, StateTimeout,
/// A named timeout fired (matched in `timeout <pat>` rows). /// A named timeout fired (matched in `timeout <pat>` rows).
Timeout(&'static str), Timeout(&'static str),
/// A graceful shutdown request reached this (trapping) machine
/// (matched in `shutdown` rows). A state with no such row `stop`s.
Shutdown,
/// A linked peer died (trapping only; matched in `exit <pat>`
/// rows). An unmatched exit is silently dropped.
Exit($crate::ExitSignal),
} }
struct $sm { struct $sm {
@@ -815,8 +1089,14 @@ macro_rules! gen_statem {
} }
impl $sm { impl $sm {
/// The machine value, for [`gen_statem::run_named`]
/// (`$crate::gen_statem::run_named`) or `spawn`.
fn new(init: $State, data: $Data) -> $sm {
$sm { state: init, data }
}
fn start(init: $State, data: $Data) -> $crate::gen_statem::GenStatemRef<$sm> { fn start(init: $State, data: $Data) -> $crate::gen_statem::GenStatemRef<$sm> {
$crate::gen_statem::spawn($sm { state: init, data }) $crate::gen_statem::spawn($sm::new(init, data))
} }
#[allow(unused_variables)] #[allow(unused_variables)]
@@ -840,11 +1120,27 @@ macro_rules! gen_statem {
$Ev::Timeout(name) $Ev::Timeout(name)
} }
fn shutdown_ev() -> Option<$Ev> {
Some($Ev::Shutdown)
}
fn exit_ev(sig: $crate::ExitSignal) -> Option<$Ev> {
Some($Ev::Exit(sig))
}
fn on_start(&mut self, $cx: &mut $crate::gen_statem::Cx<$Ev>) { fn on_start(&mut self, $cx: &mut $crate::gen_statem::Cx<$Ev>) {
let s = self.state; let s = self.state;
self.enter(s, $cx); self.enter(s, $cx);
} }
$(
#[allow(unused_variables)]
fn terminate(&mut self) {
let $data = &mut self.data;
$($tbody)*
}
)?
#[allow(unused_variables)] #[allow(unused_variables)]
#[deny(unreachable_patterns)] // conflicting rows must fail even though #[deny(unreachable_patterns)] // conflicting rows must fail even though
// this match is external-macro-expanded // this match is external-macro-expanded
@@ -860,7 +1156,7 @@ macro_rules! gen_statem {
// untouched), then the consuming `match (state, event)` whose // untouched), then the consuming `match (state, event)` whose
// value is this `Resolution`. // value is this `Resolution`.
let next: $crate::gen_statem::Resolution<$State> = let next: $crate::gen_statem::Resolution<$State> =
$crate::gen_statem!(@arms ($Ev) ($cur, ev) [ ] [ ] $crate::gen_statem!(@arms ($Ev) ($cur, ev, $cx) [ ] [ ]
$( on $st => { $($rows)* } )+); $( on $st => { $($rows)* } )+);
match next { match next {
$crate::gen_statem::Resolution::To(s) if s == $cur => { $crate::gen_statem::Resolution::To(s) if s == $cur => {
@@ -896,7 +1192,7 @@ macro_rules! gen_statem {
// is the `Info` silent-drop — last and broadest, so per-state `info` rows // is the `Info` silent-drop — last and broadest, so per-state `info` rows
// stay reachable; cast/call/timeouts get no fallback, so a forgotten pair is // stay reachable; cast/call/timeouts get no fallback, so a forgotten pair is
// still E0004). // still E0004).
(@arms ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ]) => { (@arms ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ]) => {
{ {
// Phase 1 — postpone routing (borrow-only). A guard on a postpone // Phase 1 — postpone routing (borrow-only). A guard on a postpone
// row runs here, against by-ref bindings, so it must not depend on // row runs here, against by-ref bindings, so it must not depend on
@@ -913,140 +1209,247 @@ macro_rules! gen_statem {
// returned for it. // returned for it.
match ($ss, $se) { match ($ss, $se) {
$($arms)* $($arms)*
// Macro-injected defaults, last and broadest so per-state rows
// stay reachable; a user's own catch-all row may shadow them
// entirely, hence the allow.
#[allow(unreachable_patterns)]
(_, $Ev::Info(_)) => $crate::gen_statem::Resolution::Unhandled, (_, $Ev::Info(_)) => $crate::gen_statem::Resolution::Unhandled,
#[allow(unreachable_patterns)]
(_, $Ev::Exit(_)) => $crate::gen_statem::Resolution::Unhandled,
#[allow(unreachable_patterns)]
(_, $Ev::Shutdown) => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) },
} }
} }
}; };
// Open an on-block: remember its state pat, drain its rows, then continue. // Open an on-block: remember its state pat, drain its rows, then continue.
(@arms ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] (@arms ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ]
on $st:pat => { $($rows:tt)* } $($more:tt)* on $st:pat => { $($rows:tt)* } $($more:tt)*
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) [ $($arms)* ] [ $($post)* ] ($st) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx) [ $($arms)* ] [ $($post)* ] ($st)
{ $($rows)* } { $($more)* }) { $($rows)* } { $($more)* })
}; };
// ===== @rows: drain one on-block's rows, threading both accs ============= // ===== @rows: drain one on-block's rows, threading both accs =============
// cast, explicit refusal // cast, explicit refusal
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ cast $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* } { cast $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Cast($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ] [ $($arms)* ($st, $Ev::Cast($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ]
[ $($post)* ] [ $($post)* ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// cast, stop (end the machine normally after this event)
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ cast $ev:pat $(if $g:expr)? => stop , $($rows:tt)* } { $($more:tt)* }
) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Cast($ev)) $(if $g)? => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) }, ]
[ $($post)* ]
($st) { $($rows)* } { $($more)* })
};
// cast, postpone (defer the event; the replay in a later state handles it) // cast, postpone (defer the event; the replay in a later state handles it)
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ cast $ev:pat $(if $g:expr)? => postpone , $($rows:tt)* } { $($more:tt)* } { cast $ev:pat $(if $g:expr)? => postpone , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Cast($ev)) $(if $g)? => unreachable!("postponed event is replayed, not dispatched here"), ] [ $($arms)* ($st, $Ev::Cast($ev)) $(if $g)? => unreachable!("postponed event is replayed, not dispatched here"), ]
[ $($post)* ($st, $Ev::Cast($ev)) $(if $g)? => true, ] [ $($post)* ($st, $Ev::Cast($ev)) $(if $g)? => true, ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// cast, transition / stay / branch // cast, transition / stay / branch
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ cast $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* } { cast $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Cast($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ] [ $($arms)* ($st, $Ev::Cast($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ]
[ $($post)* ] [ $($post)* ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// call, explicit refusal // call, explicit refusal
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ call $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* } { call $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Call($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ] [ $($arms)* ($st, $Ev::Call($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ]
[ $($post)* ] [ $($post)* ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// call, stop (end the machine normally after this event)
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ call $ev:pat $(if $g:expr)? => stop , $($rows:tt)* } { $($more:tt)* }
) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Call($ev)) $(if $g)? => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) }, ]
[ $($post)* ]
($st) { $($rows)* } { $($more)* })
};
// call, postpone (the Reply rides inside the event onto the queue) // call, postpone (the Reply rides inside the event onto the queue)
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ call $ev:pat $(if $g:expr)? => postpone , $($rows:tt)* } { $($more:tt)* } { call $ev:pat $(if $g:expr)? => postpone , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Call($ev)) $(if $g)? => unreachable!("postponed event is replayed, not dispatched here"), ] [ $($arms)* ($st, $Ev::Call($ev)) $(if $g)? => unreachable!("postponed event is replayed, not dispatched here"), ]
[ $($post)* ($st, $Ev::Call($ev)) $(if $g)? => true, ] [ $($post)* ($st, $Ev::Call($ev)) $(if $g)? => true, ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// call, transition / stay / branch // call, transition / stay / branch
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ call $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* } { call $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Call($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ] [ $($arms)* ($st, $Ev::Call($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ]
[ $($post)* ] [ $($post)* ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// info, explicit refusal // info, explicit refusal
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ info $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* } { info $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Info($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ] [ $($arms)* ($st, $Ev::Info($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ]
[ $($post)* ] [ $($post)* ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// info, stop (end the machine normally after this event)
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ info $ev:pat $(if $g:expr)? => stop , $($rows:tt)* } { $($more:tt)* }
) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Info($ev)) $(if $g)? => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) }, ]
[ $($post)* ]
($st) { $($rows)* } { $($more)* })
};
// info, postpone // info, postpone
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ info $ev:pat $(if $g:expr)? => postpone , $($rows:tt)* } { $($more:tt)* } { info $ev:pat $(if $g:expr)? => postpone , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Info($ev)) $(if $g)? => unreachable!("postponed event is replayed, not dispatched here"), ] [ $($arms)* ($st, $Ev::Info($ev)) $(if $g)? => unreachable!("postponed event is replayed, not dispatched here"), ]
[ $($post)* ($st, $Ev::Info($ev)) $(if $g)? => true, ] [ $($post)* ($st, $Ev::Info($ev)) $(if $g)? => true, ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// info, transition / stay / branch // info, transition / stay / branch
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ info $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* } { info $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Info($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ] [ $($arms)* ($st, $Ev::Info($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ]
[ $($post)* ] [ $($post)* ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// state_timeout, explicit refusal (unit event — no pattern; not postponable) // state_timeout, explicit refusal (unit event — no pattern; not postponable)
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ state_timeout $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* } { state_timeout $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::StateTimeout) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ] [ $($arms)* ($st, $Ev::StateTimeout) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ]
[ $($post)* ] [ $($post)* ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// state_timeout, stop (end the machine normally after this event)
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ state_timeout $(if $g:expr)? => stop , $($rows:tt)* } { $($more:tt)* }
) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::StateTimeout) $(if $g)? => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) }, ]
[ $($post)* ]
($st) { $($rows)* } { $($more)* })
};
// state_timeout, transition / stay / branch // state_timeout, transition / stay / branch
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ state_timeout $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* } { state_timeout $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::StateTimeout) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ] [ $($arms)* ($st, $Ev::StateTimeout) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ]
[ $($post)* ] [ $($post)* ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// timeout, explicit refusal (pattern matches the name; not postponable) // timeout, explicit refusal (pattern matches the name; not postponable)
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ timeout $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* } { timeout $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Timeout($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ] [ $($arms)* ($st, $Ev::Timeout($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ]
[ $($post)* ] [ $($post)* ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// timeout, stop (end the machine normally after this event)
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ timeout $ev:pat $(if $g:expr)? => stop , $($rows:tt)* } { $($more:tt)* }
) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Timeout($ev)) $(if $g)? => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) }, ]
[ $($post)* ]
($st) { $($rows)* } { $($more)* })
};
// timeout, transition / stay / branch // timeout, transition / stay / branch
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ timeout $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* } { timeout $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se) $crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Timeout($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ] [ $($arms)* ($st, $Ev::Timeout($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ]
[ $($post)* ] [ $($post)* ]
($st) { $($rows)* } { $($more)* }) ($st) { $($rows)* } { $($more)* })
}; };
// shutdown, explicit refusal (unit event — no pattern; not postponable; a state with no shutdown row defaults to `stop`)
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ shutdown $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* }
) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Shutdown) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ]
[ $($post)* ]
($st) { $($rows)* } { $($more)* })
};
// shutdown, stop (end the machine normally after this event)
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ shutdown $(if $g:expr)? => stop , $($rows:tt)* } { $($more:tt)* }
) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Shutdown) $(if $g)? => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) }, ]
[ $($post)* ]
($st) { $($rows)* } { $($more)* })
};
// shutdown, transition / stay / branch
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ shutdown $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* }
) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Shutdown) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ]
[ $($post)* ]
($st) { $($rows)* } { $($more)* })
};
// exit, explicit refusal (pattern matches the ExitSignal; not postponable)
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ exit $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* }
) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Exit($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ]
[ $($post)* ]
($st) { $($rows)* } { $($more)* })
};
// exit, stop (end the machine normally after this event)
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ exit $ev:pat $(if $g:expr)? => stop , $($rows:tt)* } { $($more:tt)* }
) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Exit($ev)) $(if $g)? => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) }, ]
[ $($post)* ]
($st) { $($rows)* } { $($more)* })
};
// exit, transition / stay / branch
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ exit $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* }
) => {
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
[ $($arms)* ($st, $Ev::Exit($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ]
[ $($post)* ]
($st) { $($rows)* } { $($more)* })
};
// this block is drained: hand the remaining on-blocks back to @arms // this block is drained: hand the remaining on-blocks back to @arms
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat) (@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
{ } { $($more:tt)* } { } { $($more:tt)* }
) => { ) => {
$crate::gen_statem!(@arms ($Ev) ($ss, $se) [ $($arms)* ] [ $($post)* ] $($more)*) $crate::gen_statem!(@arms ($Ev) ($ss, $se, $cx) [ $($arms)* ] [ $($post)* ] $($more)*)
}; };
} }
+9 -9
View File
@@ -60,10 +60,10 @@ pub use channel::{
pub use gen_server::{ pub use gen_server::{
call, cast, shutdown, whereis_server, CallError, CallTimeoutError, CastError, GenServer, call, cast, shutdown, whereis_server, CallError, CallTimeoutError, CastError, GenServer,
GenServerBuilder, GenServerCtx, GenServerName, GenServerRef, NamedGenServerBuilder, GenServerBuilder, GenServerCtx, GenServerName, GenServerRef, NamedGenServerBuilder,
TimerHandle, Watcher, ShutdownAction, StopHandle, TimerHandle, Watcher,
}; };
pub use gen_statem::{ pub use gen_statem::{
CallError as GenStatemCallError, Cx, GenStatemRef, Machine, Reply, Resolution, CallError as GenStatemCallError, Cx, GenStatemName, GenStatemRef, Machine, Reply, Resolution,
SendError as GenStatemSendError, SendError as GenStatemSendError,
}; };
pub use introspect::{ pub use introspect::{
@@ -85,15 +85,15 @@ pub use registry::{
install, lookup_as, register, resolve_name, send, send_dyn, send_to, unregister, whereis, install, lookup_as, register, resolve_name, send, send_dyn, send_to, unregister, whereis,
NameResolution, RegisterError, SendError, NameResolution, RegisterError, SendError,
}; };
pub use runtime::{init, Config, Runtime}; pub use runtime::{init, Config, Runtime, RuntimeHandle};
pub use scheduler::{ pub use scheduler::{
block_on_io, cancel_timer, request_stop, run, self_pid, send_after, send_after_named, block_on_io, cancel_timer, request_shutdown, request_stop, run, self_pid, send_after,
send_after_named_wall, send_after_wall, sleep, sleep_wall, spawn, spawn_addr, spawn_addr_with, send_after_named, send_after_named_wall, send_after_wall, sleep, sleep_wall, spawn, spawn_addr,
spawn_under, spawn_under_with, spawn_with, try_spawn, try_spawn_under_with, wait_readable, spawn_addr_with, spawn_under, spawn_under_with, spawn_with, try_spawn, try_spawn_under_with,
wait_readable_timeout, wait_writable, wait_writable_timeout, yield_now, FdArm, JoinError, wait_readable, wait_readable_timeout, wait_writable, wait_writable_timeout, yield_now, FdArm,
JoinHandle, SpawnError, SpawnOpts, JoinError, JoinHandle, SpawnError, SpawnOpts,
}; };
pub use supervisor::{ChildSpec, OneForOne, Restart, Signal, Strategy}; pub use supervisor::{ChildSpec, OneForOne, Restart, Shutdown, Signal, Strategy};
pub use timer::TimerId; pub use timer::TimerId;
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
+5
View File
@@ -98,6 +98,11 @@ pub enum DownReason {
Panic, Panic,
/// The target was cooperatively cancelled via `request_stop`. /// The target was cooperatively cancelled via `request_stop`.
Stopped, Stopped,
/// A graceful shutdown was requested via `request_shutdown`. Only ever
/// appears in an [`ExitSignal`](crate::link::ExitSignal) delivered to a
/// trapping actor — never in a [`Down`]: a target that honours the request
/// exits *normally*, one that does not trap is `Stopped`.
Shutdown,
/// The target was already gone (finished and reclaimed, or never alive) /// The target was already gone (finished and reclaimed, or never alive)
/// at the moment `monitor()` was called. /// at the moment `monitor()` was called.
NoProc, NoProc,
+132 -54
View File
@@ -127,7 +127,7 @@ use crate::supervisor::Signal;
use crate::timer::Timers; use crate::timer::Timers;
use std::sync::atomic::{AtomicBool, AtomicPtr, AtomicU32, AtomicU64, AtomicUsize, Ordering}; use std::sync::atomic::{AtomicBool, AtomicPtr, AtomicU32, AtomicU64, AtomicUsize, Ordering};
use std::sync::{Arc, Mutex}; use std::sync::{Arc, Mutex, Weak};
use std::thread; use std::thread;
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -904,17 +904,10 @@ pub(crate) struct RuntimeInner {
pub(crate) live_actors: AtomicU32, pub(crate) live_actors: AtomicU32,
/// Packed `(index << 32 | generation)` of the run's root (initial) actor, /// Packed `(index << 32 | generation)` of the run's root (initial) actor,
/// or `u64::MAX` (the ROOT_PID sentinel) before one is set. When this actor /// or `u64::MAX` (the ROOT_PID sentinel) before one is set. When this actor
/// finalizes it flags `root_exited`; the scheduler's idle verdict then /// finalizes, `finalize_actor` runs the root-exit shutdown (see
/// stops the remaining (parked-forever) actors. Set once per `run()`, right /// `shutdown_forest_roots`). Set once per `run()`, right after the initial
/// after the initial spawn. /// spawn.
pub(crate) root_bits: AtomicU64, pub(crate) root_bits: AtomicU64,
/// Set when the root actor finalizes; read by the scheduler's idle verdict
/// to trigger the one-shot teardown sweep. Reset per `run()`.
pub(crate) root_exited: AtomicBool,
/// Guards the teardown sweep to fire at most once per run (a parked-forever
/// remainder that survives the sweep falls through to the normal idle wait
/// rather than busy-spinning). Reset per `run()`.
pub(crate) root_swept: AtomicBool,
/// Timer heap. Independent lock: never nested with any other. /// Timer heap. Independent lock: never nested with any other.
pub(crate) timers: Mutex<Timers>, pub(crate) timers: Mutex<Timers>,
/// IO subsystem. `None` between runs. Lock order: io before everything. /// IO subsystem. `None` between runs. Lock order: io before everything.
@@ -1004,8 +997,6 @@ impl RuntimeInner {
free: RawMutex::new(free), free: RawMutex::new(free),
live_actors: AtomicU32::new(0), live_actors: AtomicU32::new(0),
root_bits: AtomicU64::new(u64::MAX), root_bits: AtomicU64::new(u64::MAX),
root_exited: AtomicBool::new(false),
root_swept: AtomicBool::new(false),
timers: Mutex::new(timers), timers: Mutex::new(timers),
io: Mutex::new(None), io: Mutex::new(None),
next_monitor_id: AtomicU64::new(0), next_monitor_id: AtomicU64::new(0),
@@ -1338,10 +1329,9 @@ impl Runtime {
// requires a running runtime in the thread-local). // requires a running runtime in the thread-local).
RUNTIME.with(|r| *r.borrow_mut() = Some(self.inner.clone())); RUNTIME.with(|r| *r.borrow_mut() = Some(self.inner.clone()));
let initial_handle = crate::scheduler::spawn(f); let initial_handle = crate::scheduler::spawn(f);
// The initial actor is the run's root: when it exits, remaining actors // The initial actor is the run's root: its exit means "the program is
// are stopped so the run winds down (see finalize_actor / schedule_loop). // done" — every remaining top-level actor is asked to shut down (see
self.inner.root_exited.store(false, Ordering::Relaxed); // `finalize_actor` / `shutdown_forest_roots`).
self.inner.root_swept.store(false, Ordering::Relaxed);
self.inner.set_root(initial_handle.pid()); self.inner.set_root(initial_handle.pid());
// Launch N-1 extra scheduler threads, named `smarm-sched-{slot}` so // Launch N-1 extra scheduler threads, named `smarm-sched-{slot}` so
@@ -1457,6 +1447,71 @@ impl Runtime {
inner: self.inner.clone(), inner: self.inner.clone(),
} }
} }
/// A `Send + Sync` handle to this runtime, usable from any thread —
/// including threads that are not smarm schedulers (an OS-signal handler
/// thread, an external event source). Grab it *before* [`run`](Self::run)
/// and hand it to e.g. a signal thread; that thread can then
/// [`request_stop`](RuntimeHandle::request_stop) the runtime's top
/// supervisor to drive an ordered shutdown from outside the runtime.
///
/// The in-runtime primitives ([`scheduler::request_stop`](crate::request_stop)
/// and friends) reach the runtime through a thread-local that is unset on
/// any non-scheduler thread, so they are silent no-ops off-runtime; this
/// handle carries its own reference and closes that gap.
pub fn handle(&self) -> RuntimeHandle {
RuntimeHandle {
inner: Arc::downgrade(&self.inner),
}
}
}
// ---------------------------------------------------------------------------
// RuntimeHandle — off-runtime wake/stop
// ---------------------------------------------------------------------------
/// A `Send + Sync` handle to a [`Runtime`], obtained from
/// [`Runtime::handle`]. Lets a thread that is *not* a smarm scheduler thread
/// drive a cooperative stop into the runtime — the off-runtime counterpart to
/// [`scheduler::request_stop`](crate::request_stop).
///
/// Holds a [`Weak`] to the runtime, for the same reason the IO backend does
/// (RFC 018): a lingering handle can never keep the runtime's slot table alive
/// and can never block [`Runtime::run`] from finishing. Once the `Runtime` is
/// dropped every method is a harmless no-op — the same end state as calling
/// `request_stop` on an actor that has already exited.
#[derive(Clone)]
pub struct RuntimeHandle {
inner: Weak<RuntimeInner>,
}
impl RuntimeHandle {
/// Ask `pid` to stop cooperatively, from any thread. The off-runtime
/// equivalent of [`scheduler::request_stop`](crate::request_stop): it sets
/// the target's stop flag and wakes it, so a parked actor unwinds at its
/// next checkpoint exactly as it would for an in-runtime stop. A no-op if
/// the runtime has been dropped, or if the actor has already exited.
pub fn request_stop<A>(&self, pid: Pid<A>) {
let pid = pid.erase();
// Upgrade the Weak per call, like the IO backend does (io.rs): a live
// runtime yields the inner and we drive the same stop the in-runtime
// path would; a dropped runtime makes this a no-op.
if let Some(inner) = self.inner.upgrade() {
crate::scheduler::request_stop_inner(&inner, pid);
}
}
/// Ask `pid` to shut down gracefully, from any thread. The off-runtime
/// equivalent of [`scheduler::request_shutdown`](crate::request_shutdown);
/// the delivered [`ExitSignal`](crate::ExitSignal) carries `from ==
/// ROOT_PID`, since no actor made the request. A no-op if the runtime has
/// been dropped, or if the actor has already exited.
pub fn request_shutdown<A>(&self, pid: Pid<A>) {
let pid = pid.erase();
if let Some(inner) = self.inner.upgrade() {
crate::scheduler::request_shutdown_inner(&inner, pid, ROOT_PID);
}
}
} }
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -1851,14 +1906,12 @@ fn finalize_actor(inner: &Arc<RuntimeInner>, pid: Pid, outcome: Outcome) {
// Reclaim if no outstanding handles (re-verified inside). // Reclaim if no outstanding handles (re-verified inside).
reclaim_slot(inner, pid); reclaim_slot(inner, pid);
// Root-exit teardown is DEFERRED to the scheduler's idle verdict, not done // Root exit = the program is done. Ask every top-level survivor to shut
// here: stopping eagerly would cut off actors that still have queued work // down, right here, before the live-count decrement below: any wake this
// (they'd unwind on the stop before draining their mailbox). Flagging it // produces is then ordered before `live_actors` can be observed at its
// instead lets the run queue drain naturally first; only the parked-forever // decremented value, same as every other wakeup finalize issues.
// remainder (e.g. a server pinned alive by a registered name) is then
// stopped, once nothing runnable is left. See `schedule_loop`.
if inner.is_root(pid) { if inner.is_root(pid) {
inner.root_exited.store(true, Ordering::Release); shutdown_forest_roots(inner, pid);
} }
// The decrement is LAST: every wakeup this finalize produced (joiners, // The decrement is LAST: every wakeup this finalize produced (joiners,
@@ -1869,16 +1922,62 @@ fn finalize_actor(inner: &Arc<RuntimeInner>, pid: Pid, outcome: Outcome) {
debug_assert!(prev >= 1, "live_actors underflow — double finalize"); debug_assert!(prev >= 1, "live_actors underflow — double finalize");
} }
/// Cooperatively stop every live actor — the root-exit teardown sweep, run from /// The root-exit shutdown. Delivers [`request_shutdown`](crate::request_shutdown)
/// `schedule_loop` once the run queue is empty after the root has exited. Each /// to every **forest root**: each live actor whose recorded parent
/// [`request_stop_inner`](crate::scheduler::request_stop_inner) re-verifies the /// (`Actor::supervisor` — the spawner for a plain `spawn`, the supervisor for
/// target under its cold lock, so the racy per-slot generation read is safe: a /// `spawn_under`) is the run itself (`ROOT_PID`) or is no longer live. Actors
/// vacant, dead, or reused slot no-ops. The swept actors unpark, unwind at their /// under a live parent are not addressed — that parent is responsible for
/// next observation point, and finalize, dropping `live_actors` to zero. /// them: a supervisor traps and runs its ordered, policy-driven shutdown; a
fn stop_live_actors(inner: &Arc<RuntimeInner>) { /// bare parent that dies takes non-trapping children with it via the next
/// pass of this same rule only if it dies *now*, so a parent that outlives
/// this scan and later dies leaves its subtree to itself (Erlang semantics: an
/// unlinked spawn is nobody's child).
///
/// Semantics per target follow `request_shutdown`: a trapping actor receives
/// `ExitSignal { from: root, reason: Shutdown }` and may finish work — drain,
/// keep its timers ticking, then stop itself; a non-trapping one is stopped
/// outright. There is no second, forcing sweep: an actor that traps and never
/// stops keeps the run alive by design (put it under a supervisor with a
/// `Shutdown::Timeout` policy if that is not wanted). Runs once, on the root's
/// finalize path, so it races only against actors that are still running —
/// each `request_shutdown_inner` re-verifies its target under the cold lock,
/// so a slot that dies or is reused mid-scan is a no-op.
fn shutdown_forest_roots(inner: &Arc<RuntimeInner>, root: Pid) {
for idx in 0..inner.slots.len() as u32 { for idx in 0..inner.slots.len() as u32 {
let pid = Pid::new(idx, inner.slots[idx as usize].generation()); let slot = &inner.slots[idx as usize];
crate::scheduler::request_stop_inner(inner, pid); let pid = Pid::new(idx, slot.generation());
if pid == root {
continue;
}
// Read the parent under the cold lock (generation-verified); act
// outside it — `request_shutdown_inner` sends and may unpark.
let parent = {
let cold = slot.cold.lock();
if slot.generation() != pid.generation() {
continue;
}
match cold.actor.as_ref() {
Some(a) => a.supervisor,
None => continue,
}
};
let parent_live = inner
.slot_at(parent)
.is_some_and(|ps| ps.is_live_for(parent));
if !parent_live {
// `_probe` so the trace can say what each leftover was: under
// `smarm-trace` every swept actor is a `root_sweep` line — the
// visibility that makes "a forgotten actor costs only a slot, never
// a hung run" a checkable claim rather than a hope.
let found = crate::scheduler::request_shutdown_inner_probe(inner, pid, root);
#[cfg_attr(not(feature = "smarm-trace"), allow(unused_variables))]
if let Some(trapping) = found {
crate::te!(crate::trace::Event::RootSweep {
target: pid,
trapping
});
}
}
} }
} }
@@ -1962,9 +2061,6 @@ fn schedule_loop(inner: &Arc<RuntimeInner>, slot_idx: usize) {
Got(Pid), Got(Pid),
Idle, Idle,
AllDone, AllDone,
/// Root has exited and nothing is runnable: stop the parked-forever
/// remainder, then re-pop. Fires at most once per run.
RootDrain,
} }
// 2a. RFC 005: drain this thread's wake slot before touching the // 2a. RFC 005: drain this thread's wake slot before touching the
@@ -2014,15 +2110,6 @@ fn schedule_loop(inner: &Arc<RuntimeInner>, slot_idx: usize) {
let live = inner.live_actors.load(Ordering::Acquire); let live = inner.live_actors.load(Ordering::Acquire);
if live == 0 && io_out == 0 { if live == 0 && io_out == 0 {
Pop::AllDone Pop::AllDone
} else if inner.root_exited.load(Ordering::Acquire)
&& !inner.root_swept.swap(true, Ordering::AcqRel)
{
// Root gone and nothing runnable — the live remainder
// are parked-forever daemons (Queued actors with pending
// work drained before the queue emptied). Stop them so
// the run can end. One-shot: a survivor falls through to
// the idle wait below on the next pass.
Pop::RootDrain
} else { } else {
Pop::Idle Pop::Idle
} }
@@ -2050,13 +2137,6 @@ fn schedule_loop(inner: &Arc<RuntimeInner>, slot_idx: usize) {
inner.coord.wake_all(); inner.coord.wake_all();
return; return;
} }
Pop::RootDrain => {
// Root has exited and nothing is runnable: stop the
// parked-forever remainder, then loop back to re-pop the
// now-runnable (stopping) actors.
stop_live_actors(inner);
continue;
}
Pop::Idle => { Pop::Idle => {
// Something is still in flight. Park on our own futex // Something is still in flight. Park on our own futex
// until a producer wakes us (enqueue tail), a deadline // until a producer wakes us (enqueue tail), a deadline
@@ -2087,8 +2167,6 @@ fn schedule_loop(inner: &Arc<RuntimeInner>, slot_idx: usize) {
|| (inner.live_actors.load(Ordering::Acquire) == 0 || (inner.live_actors.load(Ordering::Acquire) == 0
&& inner.io_outstanding.load(Ordering::Acquire) == 0 && inner.io_outstanding.load(Ordering::Acquire) == 0
&& inner.io_fd_waiters.load(Ordering::Acquire) == 0) && inner.io_fd_waiters.load(Ordering::Acquire) == 0)
|| (inner.root_exited.load(Ordering::Acquire)
&& !inner.root_swept.load(Ordering::Acquire))
|| inner.coord.deadline_due() || inner.coord.deadline_due()
}); });
if tk_deadline.is_some() { if tk_deadline.is_some() {
+94 -6
View File
@@ -71,7 +71,7 @@ use crate::pid::{Name, Pid};
use crate::runtime::{self, RuntimeInner, YieldIntent, RUNTIME}; use crate::runtime::{self, RuntimeInner, YieldIntent, RUNTIME};
use crate::supervisor::Signal; use crate::supervisor::Signal;
use std::sync::atomic::Ordering; use std::sync::atomic::Ordering;
use std::sync::Arc; use std::sync::{Arc, Weak};
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// with_runtime / try_with_runtime // with_runtime / try_with_runtime
@@ -544,6 +544,29 @@ pub(crate) fn unpark_at(pid: Pid, epoch: u32) {
let _ = try_with_runtime(|inner| inner.unpark_at(pid, epoch)); let _ = try_with_runtime(|inner| inner.unpark_at(pid, epoch));
} }
// The current actor's runtime as a `Weak`, for a waker that must reach the
// runtime from a foreign thread later. A channel captures this when its
// receiver parks, so a cross-thread `send` can wake without the `RUNTIME`
// thread-local (unset off a scheduler thread). Panics outside `Runtime::run()`,
// the same contract as `begin_wait`.
pub(crate) fn runtime_weak() -> Weak<RuntimeInner> {
with_runtime(Arc::downgrade)
}
// Epoch-matched wake of `pid` from a waker that may or may not be on a
// scheduler thread. On a scheduler thread we take the thread-local path
// (preemption-gated, slot-eligible); off one that path is a silent no-op, so
// we reach the runtime through `rt` — the `Weak` the waker captured while it
// was in-runtime. Mirrors the IO backend's cross-context wake (io.rs, RFC 018).
pub(crate) fn unpark_at_via(pid: Pid, epoch: u32, rt: &Weak<RuntimeInner>) {
if try_with_runtime(|inner| inner.unpark_at(pid, epoch)).is_some() {
return;
}
if let Some(inner) = rt.upgrade() {
inner.unpark_at(pid, epoch);
}
}
// Open a new wait for the current actor and return its wait identity // Open a new wait for the current actor and return its wait identity
// ("epoch"). Call once per wait, before registering with any waker. Lock-free, // ("epoch"). Call once per wait, before registering with any waker. Lock-free,
// so it's legal to call while already holding another internal lock. // so it's legal to call while already holding another internal lock.
@@ -577,10 +600,12 @@ pub(crate) fn retire_wait() {
/// [`JoinHandle::join`] reports it as a normal, non-error exit: cooperative /// [`JoinHandle::join`] reports it as a normal, non-error exit: cooperative
/// stop is a controlled shutdown, not a failure. /// stop is a controlled shutdown, not a failure.
/// ///
/// This is exactly the mechanism `gen_server` shutdown, supervisor restarts, /// This is the *hard* stop — OTP's `exit(Pid, kill)`. It is what a supervisor
/// and structured teardown are built from: reach for [`GenServerRef::shutdown`](crate::GenServerRef::shutdown) /// falls back to when a child overstays its [`Shutdown`](crate::supervisor::Shutdown)
/// or a [`supervisor`](crate::supervisor) instead of calling this directly /// grace period. For a stop the target gets to prepare for, use
/// where those apply. /// [`request_shutdown`]; for structured teardown, reach for
/// [`GenServerRef::shutdown`](crate::GenServerRef::shutdown) or a
/// [`supervisor`](crate::supervisor) instead of calling this directly.
/// ///
/// Because it's cooperative, an actor stuck in a tight loop with no /// Because it's cooperative, an actor stuck in a tight loop with no
/// blocking call, no [`check!`](crate::check), and no allocation cannot be /// blocking call, no [`check!`](crate::check), and no allocation cannot be
@@ -592,7 +617,8 @@ pub fn request_stop<A>(pid: Pid<A>) {
} }
// The core of `request_stop`, taking the runtime directly so it can also be // The core of `request_stop`, taking the runtime directly so it can also be
// driven from inside the runtime itself (the root-exit sweep) without // driven from inside the runtime itself (the RuntimeHandle path, supervisor
// sweeps) without
// re-borrowing the thread-local. Sets the stop flag under the target's lock // re-borrowing the thread-local. Sets the stop flag under the target's lock
// (a generation mismatch, or no live actor there, makes it a no-op) and // (a generation mismatch, or no live actor there, makes it a no-op) and
// wakes the target. // wakes the target.
@@ -612,6 +638,68 @@ pub(crate) fn request_stop_inner(inner: &RuntimeInner, pid: Pid) {
} }
} }
/// Ask an actor to shut down gracefully — OTP's `exit(Pid, shutdown)`, where
/// [`request_stop`] is `exit(Pid, kill)`.
///
/// If the target has called [`trap_exit`](crate::trap_exit), it receives an
/// [`ExitSignal`](crate::ExitSignal) with reason
/// [`DownReason::Shutdown`](crate::DownReason::Shutdown) on its trap inbox and
/// keeps running: the request is advisory, and the target is expected to wind
/// down and exit normally in its own time (a supervisor bounds that time with
/// its child's [`Shutdown`](crate::supervisor::Shutdown) policy and falls back
/// to `request_stop`). A target that is not trapping is stopped exactly as by
/// `request_stop`. A dead pid is a no-op.
///
/// The signal's `from` is the calling actor, or `ROOT_PID` when driven from
/// outside the runtime (see [`RuntimeHandle::request_shutdown`](crate::RuntimeHandle::request_shutdown)).
pub fn request_shutdown<A>(pid: Pid<A>) {
let pid = pid.erase();
let from = current_pid().unwrap_or(crate::runtime::ROOT_PID);
let _ = try_with_runtime(|inner| request_shutdown_inner(inner, pid, from));
}
// The core of `request_shutdown`. Reads the target's trap sender under its
// cold lock (generation-verified), then acts outside the lock: a trap send
// may unpark the receiver, and `request_stop_inner` re-takes the lock.
pub(crate) fn request_shutdown_inner(inner: &RuntimeInner, pid: Pid, from: Pid) {
request_shutdown_inner_probe(inner, pid, from);
}
/// [`request_shutdown_inner`], reporting what it found: `Some(true)` if the
/// target was trapping (got the signal), `Some(false)` if it was stopped
/// outright, `None` if there was nothing live at `pid`.
pub(crate) fn request_shutdown_inner_probe(
inner: &RuntimeInner,
pid: Pid,
from: Pid,
) -> Option<bool> {
let trap = match inner.slot_at(pid) {
Some(slot) => {
let cold = slot.cold.lock();
if slot.generation() == pid.generation() {
cold.actor.as_ref().map(|a| a.trap.clone())
} else {
None // stale pid: nothing there to shut down
}
}
None => None,
};
match trap {
Some(Some(tx)) => {
let _ = tx.send(crate::link::ExitSignal {
from,
reason: crate::monitor::DownReason::Shutdown,
});
Some(true)
}
Some(None) => {
request_stop_inner(inner, pid);
Some(false)
}
None => None,
}
}
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// NoPreempt // NoPreempt
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
+187 -66
View File
@@ -140,7 +140,8 @@ impl Signal {
} }
} }
use crate::channel::channel; use crate::channel::{channel, RecvTimeoutError};
use crate::monitor::DownReason;
use std::collections::{HashMap, VecDeque}; use std::collections::{HashMap, VecDeque};
use std::sync::Arc; use std::sync::Arc;
use std::time::{Duration, Instant}; use std::time::{Duration, Instant};
@@ -165,15 +166,53 @@ pub enum Restart {
pub struct ChildSpec { pub struct ChildSpec {
start: Arc<dyn Fn() + Send + Sync + 'static>, start: Arc<dyn Fn() + Send + Sync + 'static>,
restart: Restart, restart: Restart,
shutdown: Shutdown,
} }
impl ChildSpec { impl ChildSpec {
/// A child with the given restart policy and the default
/// [`Shutdown::Timeout`] of 5 seconds.
pub fn new(restart: Restart, start: impl Fn() + Send + Sync + 'static) -> Self { pub fn new(restart: Restart, start: impl Fn() + Send + Sync + 'static) -> Self {
Self { Self {
start: Arc::new(start), start: Arc::new(start),
restart, restart,
shutdown: Shutdown::default(),
} }
} }
/// Set how the supervisor stops this child (see [`Shutdown`]). A child
/// that is itself a supervisor should use [`Shutdown::Infinity`] so its
/// own subtree gets its full grace periods.
pub fn shutdown(mut self, shutdown: Shutdown) -> Self {
self.shutdown = shutdown;
self
}
}
/// How a supervisor stops a child it is taking down — the OTP child-spec
/// `shutdown` value. Applies to every supervisor-initiated stop: the ordered
/// shutdown of the whole set and the sibling cycling of
/// [`Strategy::OneForAll`] / [`Strategy::RestForOne`].
///
/// A graceful stop is a [`request_shutdown`](crate::request_shutdown): a child
/// that traps exits receives the request as a message and winds down in its
/// own time; one that does not is stopped outright.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum Shutdown {
/// `request_stop` immediately; no request, no grace period.
BrutalKill,
/// `request_shutdown`, wait up to the duration for the child to exit, then
/// `request_stop` it. The default, at 5 seconds.
Timeout(Duration),
/// `request_shutdown` and wait however long the child takes. Use for a
/// child supervisor, whose subtree has its own timeouts.
Infinity,
}
impl Default for Shutdown {
fn default() -> Self {
Shutdown::Timeout(Duration::from_secs(5))
}
} }
/// How a supervisor reacts when one child terminates and a restart is due. /// How a supervisor reacts when one child terminates and a restart is due.
@@ -244,23 +283,35 @@ impl OneForOne {
} }
/// Run the supervision loop on the current actor. Returns when every child /// Run the supervision loop on the current actor. Returns when every child
/// has reached a terminal, non-restartable state, or when the restart /// has reached a terminal, non-restartable state, when the restart
/// intensity cap is tripped. /// intensity cap is tripped, or when the supervisor is asked to shut down
/// (a [`request_shutdown`](crate::request_shutdown) — from its own
/// supervisor, or from the app). On every one of those exits the survivors
/// are stopped in reverse start order, each per its
/// [`Shutdown`] policy, before this returns.
///
/// The supervisor traps exits for the length of the loop (that is how the
/// shutdown request reaches it as a message). Should the supervisor itself
/// be hard-stopped with [`request_stop`](crate::request_stop), it unwinds
/// without waiting for anything — but a drop guard hard-stops its live
/// children on the way out, so the subtree is not orphaned (a child
/// supervisor unwinds the same way, recursively).
pub fn run(self) { pub fn run(self) {
let me = crate::scheduler::self_pid(); let me = crate::scheduler::self_pid();
let (tx, rx) = channel::<Signal>(); let (tx, rx) = channel::<Signal>();
crate::scheduler::register_supervisor_channel(me, tx); crate::scheduler::register_supervisor_channel(me, tx);
let exits = crate::link::trap_exit();
// pid -> index into `self.children`, for the children currently alive. // pid -> index into `self.children`, for the children currently alive.
let mut by_pid: HashMap<Pid, usize> = HashMap::new(); let mut live = Live::default();
let mut active: usize = 0; let mut active: usize = 0;
// Sliding window of recent restart instants, for the intensity cap. // Sliding window of recent restart instants, for the intensity cap.
let mut restarts: Vec<Instant> = Vec::new(); let mut restarts: Vec<Instant> = Vec::new();
let start_child = |idx: usize, by_pid: &mut HashMap<Pid, usize>| { let start_child = |idx: usize, live: &mut Live| {
let start = self.children[idx].start.clone(); let start = self.children[idx].start.clone();
let h = crate::scheduler::spawn_under(me, move || (start)()); let h = crate::scheduler::spawn_under(me, move || (start)());
by_pid.insert(h.pid(), idx); live.insert(h.pid(), idx);
// We supervise via the signal funnel, not by joining; drop the // We supervise via the signal funnel, not by joining; drop the
// handle so the child's slot is reclaimed promptly on death (the // handle so the child's slot is reclaimed promptly on death (the
// termination Signal is delivered before reclamation regardless). // termination Signal is delivered before reclamation regardless).
@@ -268,28 +319,105 @@ impl OneForOne {
}; };
for idx in 0..self.children.len() { for idx in 0..self.children.len() {
start_child(idx, &mut by_pid); start_child(idx, &mut live);
active += 1; active += 1;
} }
// A signal that arrives while we are awaiting stop-confirmations (for a // A signal that arrives while we are awaiting stop-confirmations (for a
// child we are *not* currently stopping) is stashed here and processed // child we are *not* currently stopping) is stashed here and processed
// by the main loop before it blocks on `recv` again. // by the main loop before it blocks again.
let mut pending: VecDeque<Signal> = VecDeque::new(); let mut pending: VecDeque<Signal> = VecDeque::new();
let next_signal = |pending: &mut VecDeque<Signal>| -> Option<Signal> {
// Stop one child per its policy and wait for its termination signal.
// Signals for other pids that arrive meanwhile are stashed. Bounded by
// construction: `request_stop` (used directly, or as the fallback once
// the grace period lapses) always produces a signal.
let stop_child = |pid: Pid, idx: usize, pending: &mut VecDeque<Signal>| {
let await_one = |deadline: Option<Instant>, pending: &mut VecDeque<Signal>| -> bool {
loop {
let sig = match pending.iter().position(|s| s.pid() == pid) {
Some(i) => pending.remove(i),
None => match deadline {
None => rx.recv().ok(),
Some(dl) => {
match rx.recv_timeout(dl.saturating_duration_since(Instant::now()))
{
Ok(s) => Some(s),
Err(RecvTimeoutError::Timeout) => return false,
Err(RecvTimeoutError::Disconnected) => None,
}
}
},
};
match sig {
Some(s) if s.pid() == pid => return true,
Some(s) => pending.push_back(s),
None => return true, // funnel closed: nothing more can arrive
}
}
};
match self.children[idx].shutdown {
Shutdown::BrutalKill => {
crate::scheduler::request_stop(pid);
await_one(None, pending);
}
Shutdown::Timeout(grace) => {
crate::scheduler::request_shutdown(pid);
if !await_one(Some(Instant::now() + grace), pending) {
crate::scheduler::request_stop(pid);
await_one(None, pending);
}
}
Shutdown::Infinity => {
crate::scheduler::request_shutdown(pid);
await_one(None, pending);
}
}
};
// Stop a set of children in reverse start order, one at a time.
let stop_set =
|set: &mut Vec<(Pid, usize)>, live: &mut Live, pending: &mut VecDeque<Signal>| {
set.sort_unstable_by_key(|x| std::cmp::Reverse(x.1));
for (pid, idx) in set.iter() {
live.remove(pid);
stop_child(*pid, *idx, pending);
}
};
// Wait for the next event: a stashed signal, a child signal, or a
// shutdown request. `Ok(sig)`, or `Err(())` when we must wind down.
let next_event = |pending: &mut VecDeque<Signal>| -> Result<Signal, ()> {
loop {
if let Some(s) = pending.pop_front() { if let Some(s) = pending.pop_front() {
Some(s) return Ok(s);
} else { }
rx.recv().ok() // The trap inbox is arm 0: a shutdown request is noticed even
// under a flood of child signals.
match crate::channel::select(&[&exits, &rx]) {
0 => match exits.try_recv() {
Ok(Some(sig)) if sig.reason == DownReason::Shutdown => return Err(()),
// Any other exit signal (a linked peer's death — a
// supervisor links nothing itself, but may be linked
// to) is not ours to act on; a closed trap inbox is
// impossible while `exits` is held here.
_ => {}
},
_ => match rx.try_recv() {
Ok(Some(s)) => return Ok(s),
Ok(None) => {}
Err(_) => return Err(()), // funnel closed: nothing left to supervise
},
}
} }
}; };
while active > 0 { while active > 0 {
let sig = match next_signal(&mut pending) { let sig = match next_event(&mut pending) {
Some(s) => s, Ok(s) => s,
None => break, // mailbox closed: nothing left to supervise Err(()) => break,
}; };
let idx = match by_pid.remove(&sig.pid()) { let idx = match live.remove(&sig.pid()) {
Some(i) => i, Some(i) => i,
None => continue, // stray/duplicate signal None => continue, // stray/duplicate signal
}; };
@@ -321,76 +449,69 @@ impl OneForOne {
restarts.push(now); restarts.push(now);
// Which *live* siblings get cycled along with the failed child. // Which *live* siblings get cycled along with the failed child.
// (The failed child is already gone — removed from `by_pid` above.) // (The failed child is already gone — removed from `live` above.)
let mut to_stop: Vec<(Pid, usize)> = match self.strategy { let mut to_stop: Vec<(Pid, usize)> = match self.strategy {
Strategy::OneForOne => Vec::new(), Strategy::OneForOne => Vec::new(),
Strategy::OneForAll => by_pid.iter().map(|(p, i)| (*p, *i)).collect(), Strategy::OneForAll => live.iter().map(|(p, i)| (*p, *i)).collect(),
Strategy::RestForOne => by_pid Strategy::RestForOne => live
.iter() .iter()
.filter(|(_, i)| **i > idx) .filter(|(_, i)| **i > idx)
.map(|(p, i)| (*p, *i)) .map(|(p, i)| (*p, *i))
.collect(), .collect(),
}; };
// Stop survivors in reverse start order (highest child index first).
to_stop.sort_unstable_by_key(|x| std::cmp::Reverse(x.1));
// The set we will restart: the failed child plus every sibling we // The set we will restart: the failed child plus every sibling we
// are about to stop, restarted in start (ascending index) order. // are about to stop, restarted in start (ascending index) order.
let mut restart_set: Vec<usize> = Vec::with_capacity(to_stop.len() + 1); let mut restart_set: Vec<usize> = Vec::with_capacity(to_stop.len() + 1);
restart_set.push(idx); restart_set.push(idx);
restart_set.extend(to_stop.iter().map(|(_, i)| *i));
// Request stops, then await each survivor's termination signal // Stop the survivors (each per its policy, reverse start order),
// before restarting. `request_stop` on an already-dead pid is a // then restart the whole set in start order. Net effect on
// no-op; in that case its (already-sent) Exit signal serves as the // `active`: one child died (idx), `to_stop.len()` were stopped,
// confirmation. Any signal for a pid we are *not* awaiting is // and `restart_set.len() == 1 + to_stop.len()` are started — so
// stashed for the main loop.
let mut awaiting: Vec<Pid> = Vec::with_capacity(to_stop.len());
for (pid, cidx) in &to_stop {
by_pid.remove(pid);
restart_set.push(*cidx);
crate::scheduler::request_stop(*pid);
awaiting.push(*pid);
}
while !awaiting.is_empty() {
let s = match next_signal(&mut pending) {
Some(s) => s,
None => break, // mailbox closed mid-await; stop waiting
};
if let Some(pos) = awaiting.iter().position(|p| *p == s.pid()) {
awaiting.swap_remove(pos);
} else {
pending.push_back(s);
}
}
// Restart the whole set in start order. Net effect on `active`:
// one child died (idx), `to_stop.len()` were stopped, and
// `restart_set.len() == 1 + to_stop.len()` are started — so
// `active` is unchanged and needs no adjustment here. // `active` is unchanged and needs no adjustment here.
stop_set(&mut to_stop, &mut live, &mut pending);
restart_set.sort_unstable(); restart_set.sort_unstable();
for cidx in restart_set { for cidx in restart_set {
start_child(cidx, &mut by_pid); start_child(cidx, &mut live);
} }
} }
// Ordered shutdown: stop any survivors in reverse start order and await // Ordered shutdown: stop any survivors in reverse start order, each per
// their termination. On the normal `active == 0` exit `by_pid` is empty // its policy. On the normal `active == 0` exit `live` is empty and this
// and this is a no-op; on a cap-trip or mailbox-closed break it tears // is a no-op; on a shutdown request, a cap-trip, or a closed funnel it
// the remaining children down deterministically instead of leaking them. // tears the remaining children down deterministically.
let mut survivors: Vec<(Pid, usize)> = by_pid.iter().map(|(p, i)| (*p, *i)).collect(); let mut survivors: Vec<(Pid, usize)> = live.iter().map(|(p, i)| (*p, *i)).collect();
survivors.sort_unstable_by_key(|x| std::cmp::Reverse(x.1)); stop_set(&mut survivors, &mut live, &mut pending);
let mut awaiting: Vec<Pid> = Vec::with_capacity(survivors.len());
for (pid, _) in &survivors {
crate::scheduler::request_stop(*pid);
awaiting.push(*pid);
} }
while !awaiting.is_empty() { }
let s = match next_signal(&mut pending) {
Some(s) => s, /// The live children of a supervisor, with a drop guard: if the supervisor is
None => break, /// unwound (a hard `request_stop`, or a panic in the loop) its children are
}; /// hard-stopped rather than orphaned. Fire-and-forget by necessity — a guard
if let Some(pos) = awaiting.iter().position(|p| *p == s.pid()) { /// running mid-unwind cannot park to await anything.
awaiting.swap_remove(pos); #[derive(Default)]
struct Live(HashMap<Pid, usize>);
impl std::ops::Deref for Live {
type Target = HashMap<Pid, usize>;
fn deref(&self) -> &Self::Target {
&self.0
}
}
impl std::ops::DerefMut for Live {
fn deref_mut(&mut self) -> &mut Self::Target {
&mut self.0
}
}
impl Drop for Live {
fn drop(&mut self) {
if std::thread::panicking() {
for pid in self.0.keys() {
crate::scheduler::request_stop(*pid);
} }
} }
} }
+24 -2
View File
@@ -46,17 +46,32 @@ mod inner {
#[derive(Clone, Debug)] #[derive(Clone, Debug)]
pub enum Event { pub enum Event {
// Actor lifecycle // Actor lifecycle
Spawn { parent: Pid, child: Pid }, Spawn {
parent: Pid,
child: Pid,
},
Resume(Pid), Resume(Pid),
Yield(Pid), Yield(Pid),
Park(Pid), Park(Pid),
Done(Pid), Done(Pid),
/// Root exit found a live forest root (an actor nobody supervises)
/// and delivered `request_shutdown` to it. `trapping` says whether it
/// got the chance to drain (true) or was stopped outright (false).
/// Every such line is an actor whose lifetime was nobody's business
/// but the runtime's — the way to *see* unsupervised leftovers.
RootSweep {
target: Pid,
trapping: bool,
},
// Wakeup paths // Wakeup paths
UnparkDirect(Pid), // unpark() saw Parked -> re-queued immediately UnparkDirect(Pid), // unpark() saw Parked -> re-queued immediately
UnparkDeferred(Pid), // unpark() saw Runnable -> set pending_unpark flag UnparkDeferred(Pid), // unpark() saw Runnable -> set pending_unpark flag
UnparkFlagConsumed(Pid), // scheduler saw flag on Park -> re-queued instead UnparkFlagConsumed(Pid), // scheduler saw flag on Park -> re-queued instead
// Channel // Channel
Send { sender: Pid, receiver: Option<Pid> }, Send {
sender: Pid,
receiver: Option<Pid>,
},
RecvPark(Pid), RecvPark(Pid),
RecvWake(Pid), RecvWake(Pid),
// Queue // Queue
@@ -253,6 +268,13 @@ mod inner {
Event::Yield(p) => ("yield".into(), p.index()), Event::Yield(p) => ("yield".into(), p.index()),
Event::Park(p) => ("park".into(), p.index()), Event::Park(p) => ("park".into(), p.index()),
Event::Done(p) => ("done".into(), p.index()), Event::Done(p) => ("done".into(), p.index()),
Event::RootSweep { target, trapping } => (
format!(
"root_sweep {}",
if *trapping { "shutdown" } else { "stopped" }
),
target.index(),
),
Event::UnparkDirect(p) => ("unpark_direct".into(), p.index()), Event::UnparkDirect(p) => ("unpark_direct".into(), p.index()),
Event::UnparkDeferred(p) => ("unpark_deferred".into(), p.index()), Event::UnparkDeferred(p) => ("unpark_deferred".into(), p.index()),
Event::UnparkFlagConsumed(p) => ("unpark_flag_consumed".into(), p.index()), Event::UnparkFlagConsumed(p) => ("unpark_flag_consumed".into(), p.index()),
+125
View File
@@ -0,0 +1,125 @@
//! Cross-thread wake: a thread that is *not* a smarm scheduler thread must be
//! able to wake (and stop) a parked actor.
//!
//! The gap this pins down: every off-runtime wake primitive (`unpark`,
//! `unpark_at`, `request_stop`) reaches the runtime through the `RUNTIME`
//! thread-local, which is `None` on any non-scheduler thread — so a wake
//! issued from a foreign OS thread is a silent no-op and the parked actor
//! sleeps forever. Both failure modes below manifest as `Runtime::run` never
//! returning, so each test is wrapped in a watchdog: a timeout is the failure.
//!
//! The fix mirrors RFC 018's IO backend — the waker reaches the runtime
//! through a `Weak<RuntimeInner>` it already holds (the receiver captures one
//! when it parks; `Runtime::handle()` hands one to an app thread).
use std::sync::mpsc;
use std::thread;
use std::time::Duration;
const WATCHDOG: Duration = Duration::from_secs(10);
/// Give the target actor time to actually park before the foreign thread pokes
/// it, so we exercise the *wake* of a parked actor rather than the entry-side
/// stop check.
const SETTLE: Duration = Duration::from_millis(200);
fn assert_send_sync<T: Send + Sync>() {}
/// A cross-thread `send` from a plain OS thread must wake a receiver parked in
/// `recv`. Under the thread-local-only wake path the send enqueues the message
/// but never wakes the receiver, so `recv` — and therefore `run` — hangs.
#[test]
fn foreign_thread_send_wakes_parked_receiver() {
let (done_tx, done_rx) = mpsc::channel();
thread::spawn(move || {
let rt = smarm::init(smarm::Config::exact(2));
rt.run(|| {
let (tx, rx) = smarm::channel::<u32>();
// Receiver actor: parks on recv until the foreign thread sends.
let h = smarm::spawn(move || {
assert_eq!(rx.recv().expect("recv"), 42);
});
// Foreign (non-scheduler) OS thread owns the Sender and sends
// after the receiver has parked.
let sender = thread::spawn(move || {
thread::sleep(SETTLE);
tx.send(42).expect("send");
});
let _ = h.join();
sender.join().expect("sender thread");
});
let _ = done_tx.send(());
});
done_rx
.recv_timeout(WATCHDOG)
.expect("run did not return: a foreign-thread send never woke the parked receiver");
}
/// A cross-thread `request_stop` through a `RuntimeHandle` must wake and stop a
/// parked actor. The actor parks on a long sleep (only a stop can end it); the
/// handle is grabbed before `run` and driven from a foreign thread.
#[test]
fn foreign_thread_request_stop_wakes_parked_actor() {
assert_send_sync::<smarm::RuntimeHandle>();
let rt = smarm::init(smarm::Config::exact(2));
let handle = rt.handle();
// Foreign thread: learn the target pid from inside the run, let it park,
// then stop it through the handle.
let (pid_tx, pid_rx) = mpsc::channel::<smarm::Pid>();
let stopper = thread::spawn(move || {
let pid = pid_rx.recv().expect("pid");
thread::sleep(SETTLE);
handle.request_stop(pid);
});
let (done_tx, done_rx) = mpsc::channel();
thread::spawn(move || {
rt.run(move || {
let h = smarm::spawn(|| {
// Parks indefinitely; only a cooperative stop unwinds it.
smarm::sleep(Duration::from_secs(3600));
});
pid_tx.send(h.pid()).expect("send pid");
let _ = h.join();
});
let _ = done_tx.send(());
});
done_rx
.recv_timeout(WATCHDOG)
.expect("run did not return: a foreign-thread request_stop never woke the parked actor");
stopper.join().expect("stopper thread");
}
/// A `RuntimeHandle` held across (and beyond) a run must not keep the runtime
/// alive or block all-done: `run` still returns, and once the `Runtime` is
/// dropped the handle degrades to a harmless no-op (Weak lifecycle) rather than
/// panicking or touching freed memory.
#[test]
fn lingering_handle_does_not_block_all_done() {
let rt = smarm::init(smarm::Config::exact(1));
let handle = rt.handle(); // outlives the run below
let (pid_tx, pid_rx) = mpsc::channel::<smarm::Pid>();
let (done_tx, done_rx) = mpsc::channel();
let runner = thread::spawn(move || {
rt.run(move || {
let h = smarm::spawn(|| {});
pid_tx.send(h.pid()).expect("send pid");
let _ = h.join();
});
// `rt` is dropped here, at the end of this thread.
let _ = done_tx.send(());
});
done_rx
.recv_timeout(WATCHDOG)
.expect("run did not return while a RuntimeHandle was held live");
runner.join().expect("runner thread");
// Runtime is now dropped. A stop through the lingering handle must be a
// silent no-op, not a panic or use-after-free.
let dead_pid = pid_rx.recv().expect("pid");
handle.request_stop(dead_pid);
}
+5 -5
View File
@@ -88,7 +88,7 @@ impl GenServer for Lifecycle {
} }
} }
// init -> handle_call -> (drop last ref closes inbox) -> terminate. // init -> handle_call -> shutdown -> terminate.
#[test] #[test]
fn init_and_terminate_run() { fn init_and_terminate_run() {
let log = Arc::new(Mutex::new(Vec::new())); let log = Arc::new(Mutex::new(Vec::new()));
@@ -96,9 +96,9 @@ fn init_and_terminate_run() {
run(move || { run(move || {
let server = start(Lifecycle { log: log2 }); let server = start(Lifecycle { log: log2 });
server.call(()).unwrap(); server.call(()).unwrap();
// Dropping the only ref closes the inbox; the server breaks out of its // Refs are addresses: dropping one does not end the server. The
// recv loop and runs terminate. run() will not return until it has. // explicit close does, and waits for terminate.
drop(server); server.shutdown();
}); });
assert_eq!(*log.lock().unwrap(), vec!["init", "call", "terminate"]); assert_eq!(*log.lock().unwrap(), vec!["init", "call", "terminate"]);
} }
@@ -702,7 +702,7 @@ fn no_timer_survives_exit() {
let _ = server.call(()).unwrap(); // sync: periodic armed let _ = server.call(()).unwrap(); // sync: periodic armed
smarm::sleep(Duration::from_millis(45)); // a couple of ticks smarm::sleep(Duration::from_millis(45)); // a couple of ticks
let mon = smarm::monitor(server.pid()); let mon = smarm::monitor(server.pid());
drop(server); // inbox closes → loop exits → guard drains timers server.shutdown(); // loop exits → guard drains timers
// Clean Down ⇒ the loop returned without the no-leak assert aborting. // Clean Down ⇒ the loop returned without the no-leak assert aborting.
assert!(mon.rx.recv().is_ok()); assert!(mon.rx.recv().is_ok());
let at_exit = f_read.lock().unwrap().len(); let at_exit = f_read.lock().unwrap().len();
+243
View File
@@ -0,0 +1,243 @@
//! gen_server lifetime is the actor's, not its refs' (OTP: a pid is an
//! address, a process lives until it stops, is shut down, or is killed).
//!
//! - Dropping the last `GenServerRef` does NOT end the server. It ends via
//! `StopHandle::stop`, `request_shutdown` / `GenServerRef::shutdown`,
//! `request_stop`, or a handler panic.
//! - `GenServerBuilder::named(N).run()` runs the loop inline as the *current*
//! actor, so a server is a direct `ChildSpec` child: the supervisor's
//! shutdown reaches it as `handle_shutdown`, a restart re-binds the name,
//! and by-name `call`/`cast` reach whichever incarnation is live.
use smarm::gen_server::{
self, GenServer, GenServerBuilder, GenServerCtx, GenServerName, ShutdownAction, StopHandle,
};
use smarm::registry::RegisterError;
use smarm::supervisor::{ChildSpec, OneForOne, Restart, Shutdown};
use smarm::{monitor, request_shutdown, run, sleep, spawn, DownReason};
use std::sync::atomic::{AtomicUsize, Ordering};
use std::sync::{Arc, Mutex};
use std::time::Duration;
#[derive(Default, Clone)]
struct Log(Arc<Mutex<Vec<String>>>);
impl Log {
fn push(&self, s: impl Into<String>) {
self.0.lock().unwrap().push(s.into());
}
fn get(&self) -> Vec<String> {
self.0.lock().unwrap().clone()
}
}
struct Counter {
log: Log,
n: u64,
trap: bool,
stop: Option<StopHandle<Counter>>,
}
enum Call {
Get,
}
enum Cast {
Inc,
Stop,
}
impl GenServer for Counter {
type Call = Call;
type Reply = u64;
type Cast = Cast;
type Info = ();
type Timer = ();
fn init(&mut self, ctx: &GenServerCtx<Self>) {
if self.trap {
ctx.trap_exit();
}
self.stop = Some(ctx.stop_handle());
self.log.push("init");
}
fn handle_call(&mut self, Call::Get: Call) -> u64 {
self.n
}
fn handle_cast(&mut self, c: Cast) {
match c {
Cast::Inc => self.n += 1,
Cast::Stop => self.stop.as_ref().unwrap().stop(),
}
}
fn handle_shutdown(&mut self) -> ShutdownAction {
self.log.push("handle_shutdown");
ShutdownAction::Exit
}
fn terminate(&mut self) {
self.log.push("terminate");
}
}
fn counter(log: &Log, trap: bool) -> Counter {
Counter {
log: log.clone(),
n: 0,
trap,
stop: None,
}
}
// ---------------------------------------------------------------------------
// Refs are addresses: dropping the last one does not end the server.
// ---------------------------------------------------------------------------
#[test]
fn dropping_last_ref_does_not_end_server() {
let log = Log::default();
let l = log.clone();
run(move || {
let srv = gen_server::start(counter(&l, true));
let pid = srv.pid();
srv.cast(Cast::Inc).unwrap();
assert_eq!(srv.call(Call::Get).unwrap(), 1);
let mon = monitor(pid);
drop(srv);
sleep(Duration::from_millis(30));
assert!(
mon.rx.try_recv().unwrap().is_none(),
"server must outlive its last ref"
);
assert_eq!(l.get(), vec!["init"], "terminate must not have run");
// Explicit teardown still works, and is what ends it.
request_shutdown(pid);
let down = mon.rx.recv().unwrap();
assert_eq!(down.reason, DownReason::Exit);
});
assert_eq!(log.get(), vec!["init", "handle_shutdown", "terminate"]);
}
#[test]
fn ref_shutdown_is_the_explicit_close() {
let log = Log::default();
let l = log.clone();
run(move || {
let srv = gen_server::start(counter(&l, true));
srv.call(Call::Get).unwrap(); // sync: init (and trap_exit) has run
srv.shutdown(); // graceful, waits
assert_eq!(l.get(), vec!["init", "handle_shutdown", "terminate"]);
});
}
#[test]
fn forgotten_server_is_shut_down_at_root_exit() {
// A ref-less server is not a hung run: root exit shuts it down.
let log = Log::default();
let l = log.clone();
run(move || {
let srv = gen_server::start(counter(&l, false));
drop(srv);
});
assert_eq!(log.get(), vec!["init", "terminate"]);
}
// ---------------------------------------------------------------------------
// Inline run: a gen_server as a direct ChildSpec child.
// ---------------------------------------------------------------------------
const COUNTER: GenServerName<Counter> = GenServerName::new("lifetime-counter");
#[test]
fn named_run_is_a_direct_supervised_child_and_gets_shutdown() {
let log = Log::default();
let l = log.clone();
run(move || {
let l2 = l.clone();
let sup = spawn(move || {
let l3 = l2.clone();
OneForOne::new()
.child(
ChildSpec::new(Restart::Permanent, move || {
GenServerBuilder::new(counter(&l3, true))
.named(COUNTER)
.run()
.expect("name free");
})
.shutdown(Shutdown::Infinity),
)
.run();
});
sleep(Duration::from_millis(10));
gen_server::cast(COUNTER, Cast::Inc).unwrap();
assert_eq!(gen_server::call(COUNTER, Call::Get).unwrap(), 1);
request_shutdown(sup.pid());
sup.join()
.expect("ordered shutdown, supervisor returns normally");
assert_eq!(l.get(), vec!["init", "handle_shutdown", "terminate"]);
assert!(gen_server::whereis_server(COUNTER).is_none());
});
}
#[test]
fn named_run_child_restarts_and_rebinds_name() {
let log = Log::default();
let inits = Arc::new(AtomicUsize::new(0));
let l = log.clone();
let i = inits.clone();
run(move || {
let l2 = l.clone();
let i2 = i.clone();
let sup = spawn(move || {
let l3 = l2.clone();
let i3 = i2.clone();
OneForOne::new()
.child(ChildSpec::new(Restart::Permanent, move || {
i3.fetch_add(1, Ordering::SeqCst);
GenServerBuilder::new(counter(&l3, false))
.named(COUNTER)
.run()
.expect("name free on (re)start");
}))
.run();
});
sleep(Duration::from_millis(10));
gen_server::cast(COUNTER, Cast::Inc).unwrap();
assert_eq!(gen_server::call(COUNTER, Call::Get).unwrap(), 1);
// Normal self-exit → Permanent restarts it, fresh state, same name.
gen_server::cast(COUNTER, Cast::Stop).unwrap();
sleep(Duration::from_millis(30));
assert_eq!(i.load(Ordering::SeqCst), 2, "restarted once");
assert_eq!(gen_server::call(COUNTER, Call::Get).unwrap(), 0);
request_shutdown(sup.pid());
sup.join().unwrap();
});
assert_eq!(log.get(), vec!["init", "terminate", "init", "terminate"]);
}
#[test]
fn named_run_name_clash_fails_before_init() {
let log = Log::default();
let l = log.clone();
run(move || {
let first = GenServerBuilder::new(counter(&l, false))
.named(COUNTER)
.start()
.unwrap();
let l2 = l.clone();
let res = Arc::new(Mutex::new(None));
let r2 = res.clone();
let first_pid = first.pid();
spawn(move || {
let r = GenServerBuilder::new(counter(&l2, false))
.named(COUNTER)
.run();
*r2.lock().unwrap() = Some(r);
})
.join()
.unwrap();
assert_eq!(
*res.lock().unwrap(),
Some(Err(RegisterError::NameTaken { holder: first_pid }))
);
assert_eq!(l.get(), vec!["init"], "clashing server never ran init");
first.shutdown();
});
}
+222
View File
@@ -0,0 +1,222 @@
//! gen_server graceful shutdown.
//!
//! - A server that does not opt in (`ctx.trap_exit()` in `init`) is stopped
//! outright by `request_shutdown`, exactly as by `request_stop`.
//! - A trapping server receives the request as `handle_shutdown`. The default
//! returns `ShutdownAction::Exit`: the loop breaks and `terminate` runs on
//! the normal (non-unwind) path, so it may block. `Continue` keeps the loop
//! dispatching; the state later ends itself with a `StopHandle` — the only
//! way for a gen_server to exit *normally* on its own (`request_stop` on
//! self is an abnormal `Stopped`, which `Transient` restarts).
//! - Other exit signals (linked peers dying) reach a trapping server via
//! `handle_exit`.
use smarm::gen_server::{
start, GenServer, GenServerBuilder, GenServerCtx, GenServerRef, ShutdownAction, StopHandle,
};
use smarm::supervisor::{ChildSpec, OneForOne, Restart};
use smarm::{link, monitor, request_shutdown, run, self_pid, sleep, spawn, DownReason, ExitSignal};
use std::sync::atomic::{AtomicBool, AtomicUsize, Ordering};
use std::sync::{Arc, Mutex};
use std::time::Duration;
#[derive(Default, Clone)]
struct Log {
events: Arc<Mutex<Vec<&'static str>>>,
}
impl Log {
fn push(&self, e: &'static str) {
self.events.lock().unwrap().push(e);
}
fn get(&self) -> Vec<&'static str> {
self.events.lock().unwrap().clone()
}
}
/// A server with configurable shutdown behaviour.
struct Srv {
log: Log,
trap: bool,
action: ShutdownAction,
stop: Option<StopHandle<Srv>>,
exits: Arc<Mutex<Vec<ExitSignal>>>,
}
impl Srv {
fn new(log: &Log, trap: bool, action: ShutdownAction) -> Self {
Srv {
log: log.clone(),
trap,
action,
stop: None,
exits: Default::default(),
}
}
}
enum Cast {
Note(&'static str),
StopNow,
}
impl GenServer for Srv {
type Call = ();
type Reply = ();
type Cast = Cast;
type Info = ();
type Timer = ();
fn init(&mut self, ctx: &GenServerCtx<Self>) {
if self.trap {
ctx.trap_exit();
}
self.stop = Some(ctx.stop_handle());
}
fn handle_call(&mut self, _: ()) {}
fn handle_cast(&mut self, c: Cast) {
match c {
Cast::Note(s) => self.log.push(s),
Cast::StopNow => self.stop.as_ref().unwrap().stop(),
}
}
fn handle_shutdown(&mut self) -> ShutdownAction {
self.log.push("handle_shutdown");
self.action
}
fn handle_exit(&mut self, sig: ExitSignal) {
self.log.push("handle_exit");
self.exits.lock().unwrap().push(sig);
}
fn terminate(&mut self) {
// Allowed to block on the graceful path.
if self.trap {
sleep(Duration::from_millis(10));
}
self.log.push("terminate");
}
}
fn spawn_settled<G: GenServer>(state: G) -> GenServerRef<G> {
let r = start(state);
sleep(Duration::from_millis(20)); // let init (trap_exit) run
r
}
#[test]
fn non_trapping_server_is_stopped_outright() {
let log = Log::default();
let l = log.clone();
run(move || {
let r = spawn_settled(Srv::new(&l, false, ShutdownAction::Exit));
let mon = monitor(r.pid());
request_shutdown(r.pid());
let d = mon.rx.recv().unwrap();
assert_eq!(d.reason, DownReason::Stopped);
});
assert_eq!(log.get(), vec!["terminate"]);
}
#[test]
fn trapping_server_exits_normally_via_handle_shutdown() {
let log = Log::default();
let l = log.clone();
run(move || {
let r = spawn_settled(Srv::new(&l, true, ShutdownAction::Exit));
let mon = monitor(r.pid());
request_shutdown(r.pid());
let d = mon.rx.recv().unwrap();
assert_eq!(d.reason, DownReason::Exit);
});
assert_eq!(log.get(), vec!["handle_shutdown", "terminate"]);
}
#[test]
fn continue_keeps_dispatching_until_stop_handle() {
let log = Log::default();
let l = log.clone();
run(move || {
let r = spawn_settled(Srv::new(&l, true, ShutdownAction::Continue));
let mon = monitor(r.pid());
request_shutdown(r.pid());
sleep(Duration::from_millis(20));
r.cast(Cast::Note("after-shutdown-request")).unwrap();
r.cast(Cast::StopNow).unwrap();
let d = mon.rx.recv().unwrap();
assert_eq!(d.reason, DownReason::Exit);
});
assert_eq!(
log.get(),
vec!["handle_shutdown", "after-shutdown-request", "terminate"]
);
}
#[test]
fn stop_handle_is_a_normal_exit_that_transient_does_not_restart() {
let starts = Arc::new(AtomicUsize::new(0));
let s = starts.clone();
run(move || {
let s2 = s.clone();
let sup = spawn(move || {
let s3 = s2.clone();
OneForOne::new()
.child(ChildSpec::new(Restart::Transient, move || {
s3.fetch_add(1, Ordering::SeqCst);
let log = Log::default();
let r = GenServerBuilder::new(Srv::new(&log, false, ShutdownAction::Exit))
.under(self_pid())
.start();
r.cast(Cast::StopNow).unwrap();
// Block until the server is gone; a bare spawn parent
// returning would not itself end the server.
let mon = monitor(r.pid());
let _ = mon.rx.recv();
}))
.run();
});
sup.join().unwrap(); // returns only if the child was not restarted forever
});
assert_eq!(starts.load(Ordering::SeqCst), 1);
}
#[test]
fn linked_peer_death_reaches_handle_exit() {
let log = Log::default();
let l = log.clone();
let alive = Arc::new(AtomicBool::new(false));
let a = alive.clone();
run(move || {
let r = spawn_settled(Srv::new(&l, true, ShutdownAction::Exit));
let pid = r.pid();
let peer = spawn(move || {
link(pid);
panic!("peer dies");
});
let _ = peer.join();
sleep(Duration::from_millis(20));
r.cast(Cast::Note("still-serving")).unwrap();
sleep(Duration::from_millis(20));
a.store(true, Ordering::SeqCst);
r.shutdown(); // graceful; waits for terminate
});
assert!(alive.load(Ordering::SeqCst));
assert_eq!(
log.get(),
vec![
"handle_exit",
"still-serving",
"handle_shutdown",
"terminate"
]
);
}
#[test]
fn gen_server_ref_shutdown_is_graceful_for_a_trapping_server() {
let log = Log::default();
let l = log.clone();
run(move || {
let r = spawn_settled(Srv::new(&l, true, ShutdownAction::Exit));
r.shutdown();
});
assert_eq!(log.get(), vec!["handle_shutdown", "terminate"]);
}
+128
View File
@@ -0,0 +1,128 @@
//! gen_statem lifetime parity with gen_server: a machine lives until it
//! stops, is shut down, or is killed — its refs are addresses. And
//! `gen_statem::run_named` runs a machine inline as the current actor, so it
//! is a direct `ChildSpec` child addressed by name.
use smarm::gen_statem::{self, GenStatemName, Reply};
use smarm::supervisor::{ChildSpec, OneForOne, Restart, Shutdown};
use smarm::{monitor, request_shutdown, run, sleep, spawn, DownReason};
use std::sync::{Arc, Mutex};
use std::time::Duration;
#[derive(Default, Clone)]
struct Log(Arc<Mutex<Vec<&'static str>>>);
impl Log {
fn push(&self, e: &'static str) {
self.0.lock().unwrap().push(e);
}
fn get(&self) -> Vec<&'static str> {
self.0.lock().unwrap().clone()
}
}
#[derive(Clone, Copy, PartialEq, Eq, Debug)]
enum S {
On,
}
struct D {
log: Log,
trap: bool,
n: u64,
}
enum Cast {
Inc,
StopNow,
}
enum Call {
Get(Reply<u64>),
}
smarm::gen_statem! {
machine: Sm { state: S, data: D };
event: Ev { cast: Cast, call: Call, info: () };
context(data, prev, cx);
enter {
S::On => { data.log.push("enter"); if data.trap { cx.trap_exit() } },
}
on S::On => {
cast Cast::Inc => { data.n += 1; prev },
cast Cast::StopNow => stop,
call Call::Get(r) => { r.reply(data.n); prev },
shutdown => { data.log.push("shutdown"); cx.stop(); prev },
state_timeout => unhandled,
timeout _ => unhandled,
}
terminate { data.log.push("terminate"); }
}
fn d(log: &Log, trap: bool) -> D {
D {
log: log.clone(),
trap,
n: 0,
}
}
#[test]
fn dropping_last_ref_does_not_end_machine() {
let log = Log::default();
let l = log.clone();
run(move || {
let m = Sm::start(S::On, d(&l, true));
let pid = m.pid();
m.send(Ev::Cast(Cast::Inc)).unwrap();
assert_eq!(m.call(|r| Ev::Call(Call::Get(r))).unwrap(), 1);
let mon = monitor(pid);
drop(m);
sleep(Duration::from_millis(30));
assert!(
mon.rx.try_recv().unwrap().is_none(),
"machine must outlive its refs"
);
assert_eq!(l.get(), vec!["enter"]);
request_shutdown(pid);
assert_eq!(mon.rx.recv().unwrap().reason, DownReason::Exit);
});
assert_eq!(log.get(), vec!["enter", "shutdown", "terminate"]);
}
const SM: GenStatemName<Sm> = GenStatemName::new("lifetime-sm");
#[test]
fn run_named_is_a_direct_supervised_child() {
let log = Log::default();
let l = log.clone();
run(move || {
let l2 = l.clone();
let sup = spawn(move || {
let l3 = l2.clone();
OneForOne::new()
.child(
ChildSpec::new(Restart::Permanent, move || {
gen_statem::run_named(SM, Sm::new(S::On, d(&l3, true))).expect("name free");
})
.shutdown(Shutdown::Infinity),
)
.run();
});
sleep(Duration::from_millis(10));
gen_statem::send(SM, Ev::Cast(Cast::Inc)).unwrap();
assert_eq!(gen_statem::call(SM, |r| Ev::Call(Call::Get(r))).unwrap(), 1);
// Normal self-exit → Permanent restart → fresh data, same name.
gen_statem::send(SM, Ev::Cast(Cast::StopNow)).unwrap();
sleep(Duration::from_millis(30));
assert_eq!(gen_statem::call(SM, |r| Ev::Call(Call::Get(r))).unwrap(), 0);
request_shutdown(sup.pid());
sup.join().unwrap();
assert!(gen_statem::whereis_machine(SM).is_none());
});
assert_eq!(
log.get(),
vec!["enter", "terminate", "enter", "shutdown", "terminate"]
);
}
+219
View File
@@ -0,0 +1,219 @@
//! gen_statem graceful shutdown — the gen_server surface, in state-machine
//! clothes. Where gen_server routes a shutdown request to a `handle_shutdown`
//! method, a gen_statem gets it as an **event** so it can be routed by state:
//!
//! - A machine that does not opt in (`cx.trap_exit()` in the initial `enter`)
//! is stopped outright by `request_shutdown`, exactly as by `request_stop`.
//! - A trapping machine sees the request as a `shutdown` row (a unit event
//! like `state_timeout`). The macro's default, when a state writes no
//! `shutdown` row, is `stop` — the loop breaks and `terminate` runs on the
//! normal path. A row may instead transition (e.g. into a Draining state)
//! and `stop` later from any row via the `stop` tail keyword.
//! - Linked-peer deaths reach a trapping machine as `exit <pat>` rows; an
//! unmatched exit is silently dropped, like an unmatched info.
//! - `terminate { … }` is an optional macro block, run on every exit path.
use smarm::gen_statem;
use smarm::gen_statem::{GenStatemRef, Reply};
use smarm::{link, monitor, request_shutdown, run, sleep, spawn, DownReason, ExitSignal};
use std::sync::{Arc, Mutex};
use std::time::Duration;
#[derive(Default, Clone)]
struct Log(Arc<Mutex<Vec<&'static str>>>);
impl Log {
fn push(&self, e: &'static str) {
self.0.lock().unwrap().push(e);
}
fn get(&self) -> Vec<&'static str> {
self.0.lock().unwrap().clone()
}
}
#[derive(Clone, Copy, PartialEq, Eq, Debug)]
enum S {
Idle,
Draining,
}
struct D {
log: Log,
trap: bool,
exits: Vec<ExitSignal>,
}
enum Cast {
Note(&'static str),
StopNow,
}
enum Call {
Exits(Reply<usize>),
}
gen_statem! {
machine: Sm { state: S, data: D };
event: Ev { cast: Cast, call: Call, info: () };
context(data, prev, cx);
enter {
S::Idle => if data.trap { cx.trap_exit() },
S::Draining => { data.log.push("draining"); cx.state_timeout(Duration::from_millis(30)); },
}
on S::Idle => {
// Shutdown in Idle: go drain first, stop later.
shutdown => S::Draining,
cast Cast::StopNow => stop,
state_timeout => unhandled,
}
on S::Draining => {
// Drained: end the machine normally.
state_timeout => { data.log.push("drained"); cx.stop(); prev },
// A second request while draining is ignored.
shutdown => unhandled,
cast Cast::StopNow => stop,
}
on _ => {
cast Cast::Note(s) => { data.log.push(s); prev },
call Call::Exits(r) => { r.reply(data.exits.len()); prev },
exit sig => { data.log.push("exit"); data.exits.push(sig); prev },
timeout _ => unhandled,
}
terminate {
data.log.push("terminate");
}
}
/// A machine with no `shutdown` rows at all: the macro default applies.
#[derive(Clone, Copy, PartialEq, Eq, Debug)]
enum P {
On,
}
struct PD {
log: Log,
}
enum PCast {}
enum PCall {}
gen_statem! {
machine: Plain { state: P, data: PD };
event: PEv { cast: PCast, call: PCall, info: () };
context(data, prev, cx);
enter { P::On => cx.trap_exit(), }
on P::On => {
cast _ => unhandled,
call _ => unhandled,
state_timeout => unhandled,
timeout _ => unhandled,
}
terminate { data.log.push("terminate"); }
}
fn settled(log: &Log, trap: bool) -> GenStatemRef<Sm> {
let r = Sm::start(
S::Idle,
D {
log: log.clone(),
trap,
exits: Vec::new(),
},
);
sleep(Duration::from_millis(20)); // let on_start (trap_exit) run
r
}
#[test]
fn non_trapping_machine_is_stopped_outright() {
let log = Log::default();
let l = log.clone();
run(move || {
let r = settled(&l, false);
let mon = monitor(r.pid());
request_shutdown(r.pid());
assert_eq!(mon.rx.recv().unwrap().reason, DownReason::Stopped);
});
assert_eq!(log.get(), vec!["terminate"]);
}
#[test]
fn shutdown_row_routes_by_state_and_stop_tail_exits_normally() {
let log = Log::default();
let l = log.clone();
run(move || {
let r = settled(&l, true);
let mon = monitor(r.pid());
request_shutdown(r.pid());
// The second request lands in Draining and is `unhandled` (ignored).
sleep(Duration::from_millis(5));
request_shutdown(r.pid());
assert_eq!(mon.rx.recv().unwrap().reason, DownReason::Exit);
});
assert_eq!(log.get(), vec!["draining", "drained", "terminate"]);
}
#[test]
fn default_shutdown_is_stop() {
let log = Log::default();
let l = log.clone();
run(move || {
let r = Plain::start(P::On, PD { log: l });
sleep(Duration::from_millis(20));
let mon = monitor(r.pid());
request_shutdown(r.pid());
assert_eq!(mon.rx.recv().unwrap().reason, DownReason::Exit);
});
assert_eq!(log.get(), vec!["terminate"]);
}
#[test]
fn stop_tail_from_a_cast_is_a_normal_exit() {
let log = Log::default();
let l = log.clone();
run(move || {
let r = settled(&l, false);
let mon = monitor(r.pid());
r.send(Ev::Cast(Cast::Note("a"))).unwrap();
r.send(Ev::Cast(Cast::StopNow)).unwrap();
r.send(Ev::Cast(Cast::Note("after-stop"))).unwrap(); // never dispatched
assert_eq!(mon.rx.recv().unwrap().reason, DownReason::Exit);
});
assert_eq!(log.get(), vec!["a", "terminate"]);
}
#[test]
fn linked_peer_death_reaches_exit_row() {
let log = Log::default();
let l = log.clone();
run(move || {
let r = settled(&l, true);
let pid = r.pid();
let peer = spawn(move || {
link(pid);
panic!("peer dies");
});
let _ = peer.join();
sleep(Duration::from_millis(20));
r.send(Ev::Cast(Cast::Note("still-running"))).unwrap();
assert_eq!(r.call(|r| Ev::Call(Call::Exits(r))).unwrap(), 1);
r.shutdown();
});
assert_eq!(
log.get(),
vec!["exit", "still-running", "draining", "drained", "terminate"]
);
}
#[test]
fn ref_shutdown_is_graceful_and_waits() {
let log = Log::default();
let l = log.clone();
run(move || {
let r = settled(&l, true);
r.shutdown();
// terminate has run by the time shutdown() returns.
assert_eq!(l.get(), vec!["draining", "drained", "terminate"]);
});
}
+288
View File
@@ -0,0 +1,288 @@
//! Root exit — the run's initial actor returning means "the program is done".
//!
//! When the root finalizes, the runtime delivers `request_shutdown` to every
//! **forest root**: each live actor whose parent is the run itself (a plain
//! `spawn` from the root closure) or is already dead. Nothing below a live
//! parent is touched directly — a supervisor gets one request and runs its
//! own ordered shutdown per child `Shutdown` policy.
//!
//! - Non-trapping actors are stopped outright, exactly as by
//! `request_shutdown` — a `spawn(|| { sleep(..); work() })` the root did
//! not `join` does NOT get to finish. Join it, supervise it, or trap.
//! - Trapping actors get `handle_shutdown` / an `ExitSignal{Shutdown}` and
//! may keep running (`Continue`, drain, then stop themselves) — timers and
//! all; the run ends when they do. There is no second, forcing sweep.
//! - A periodic-timer daemon (the classic wedge) never blocks `run()`.
use smarm::gen_server::{start, GenServer, GenServerCtx, ShutdownAction, StopHandle, TimerHandle};
use smarm::supervisor::{ChildSpec, OneForOne, Restart, Shutdown};
use smarm::{run, sleep, spawn, trap_exit, DownReason};
use std::sync::atomic::{AtomicBool, AtomicUsize, Ordering};
use std::sync::{Arc, Mutex};
use std::time::{Duration, Instant};
fn assert_prompt(start: Instant, what: &str) {
assert!(
start.elapsed() < Duration::from_secs(2),
"{what}: run() took {:?}",
start.elapsed()
);
}
// ---------------------------------------------------------------------------
// Bare (non-gen_server) actors
// ---------------------------------------------------------------------------
/// A non-trapping sleeper the root did not join is stopped, not waited for.
#[test]
fn unjoined_non_trapping_sleeper_is_stopped() {
let finished = Arc::new(AtomicBool::new(false));
let f = finished.clone();
let t = Instant::now();
run(move || {
spawn(move || {
sleep(Duration::from_secs(5));
f.store(true, Ordering::SeqCst);
});
});
assert_prompt(t, "sleeper");
assert!(
!finished.load(Ordering::SeqCst),
"sleeper should have been stopped"
);
}
/// A trapping bare actor sees `Shutdown` from the root's exit and may keep
/// working — here it sleeps (a timer!) after the signal, then returns. The
/// run waits for it: no forcing sweep.
#[test]
fn trapping_actor_may_finish_after_shutdown_signal() {
let finished = Arc::new(AtomicBool::new(false));
let f = finished.clone();
run(move || {
spawn(move || {
let inbox = trap_exit();
let sig = inbox.recv().expect("shutdown signal");
assert_eq!(sig.reason, DownReason::Shutdown);
sleep(Duration::from_millis(100));
f.store(true, Ordering::SeqCst);
});
// Trapping is a runtime opt-in: give the actor a chance to run
// `trap_exit()`; a not-yet-run actor is non-trapping and is stopped.
sleep(Duration::from_millis(20));
});
assert!(
finished.load(Ordering::SeqCst),
"trapping actor must be allowed to finish"
);
}
/// The classic wedge: a lazily spawned daemon that never returns on its own.
#[test]
fn parked_forever_daemon_does_not_block_run() {
let t = Instant::now();
run(|| {
let (_tx, rx) = smarm::channel::<()>();
spawn(move || {
let _ = rx.recv(); // parked forever: sender is held by the root, which returns
});
// Leak the sender into the daemon's own scope so nothing else drops it.
std::mem::forget(_tx);
});
assert_prompt(t, "daemon");
}
// ---------------------------------------------------------------------------
// gen_servers
// ---------------------------------------------------------------------------
/// A non-trapping ticker with a periodic timer: the timer wheel is never empty,
/// and root exit must still end the run.
struct Ticker {
ticks: Arc<AtomicUsize>,
}
impl GenServer for Ticker {
type Call = ();
type Reply = ();
type Cast = ();
type Info = ();
type Timer = ();
fn init(&mut self, ctx: &GenServerCtx<Self>) {
ctx.timer().tick_every(Duration::from_millis(5), ());
}
fn handle_call(&mut self, _: ()) {}
fn handle_cast(&mut self, _: ()) {}
fn handle_timer(&mut self, _: ()) {
self.ticks.fetch_add(1, Ordering::SeqCst);
}
}
#[test]
fn periodic_timer_daemon_does_not_block_run() {
let ticks = Arc::new(AtomicUsize::new(0));
let tk = ticks.clone();
let t = Instant::now();
run(move || {
let _r = start(Ticker { ticks: tk });
sleep(Duration::from_millis(50));
});
assert_prompt(t, "ticker");
assert!(
ticks.load(Ordering::SeqCst) >= 3,
"ticker should have ticked"
);
}
/// A trapping server that answers `Continue`, keeps ticking on its own timer
/// (draining), and stops itself later. Root exit must not cut it short.
struct Drainer {
log: Arc<Mutex<Vec<&'static str>>>,
shutdowns: Arc<AtomicUsize>,
ticks_after_shutdown: usize,
stop: Option<StopHandle<Drainer>>,
timer: Option<TimerHandle<Drainer>>,
draining: bool,
}
impl GenServer for Drainer {
type Call = ();
type Reply = ();
type Cast = ();
type Info = ();
type Timer = ();
fn init(&mut self, ctx: &GenServerCtx<Self>) {
ctx.trap_exit();
self.stop = Some(ctx.stop_handle());
self.timer = Some(ctx.timer());
}
fn handle_call(&mut self, _: ()) {}
fn handle_cast(&mut self, _: ()) {}
fn handle_shutdown(&mut self) -> ShutdownAction {
self.shutdowns.fetch_add(1, Ordering::SeqCst);
self.log.lock().unwrap().push("handle_shutdown");
self.draining = true;
self.timer
.as_ref()
.unwrap()
.tick_every(Duration::from_millis(10), ());
ShutdownAction::Continue
}
fn handle_timer(&mut self, _: ()) {
if !self.draining {
return;
}
self.ticks_after_shutdown += 1;
if self.ticks_after_shutdown == 3 {
self.log.lock().unwrap().push("drained");
self.stop.as_ref().unwrap().stop();
}
}
fn terminate(&mut self) {
self.log.lock().unwrap().push("terminate");
}
}
fn drainer(log: &Arc<Mutex<Vec<&'static str>>>, shutdowns: &Arc<AtomicUsize>) -> Drainer {
Drainer {
log: log.clone(),
shutdowns: shutdowns.clone(),
ticks_after_shutdown: 0,
stop: None,
timer: None,
draining: false,
}
}
#[test]
fn trapping_server_drains_with_timers_after_root_exit() {
let log = Arc::new(Mutex::new(Vec::new()));
let shutdowns = Arc::new(AtomicUsize::new(0));
let (l, s) = (log.clone(), shutdowns.clone());
run(move || {
let r = start(drainer(&l, &s));
// A gen_server's lifetime is governed by its refs: dropping the last
// one closes the inbox and ends the loop cleanly, which would cut the
// drain short for a reason unrelated to root exit. Pin it the way a
// registered name would.
std::mem::forget(r);
sleep(Duration::from_millis(20)); // let init (trap_exit) run
});
assert_eq!(
*log.lock().unwrap(),
vec!["handle_shutdown", "drained", "terminate"]
);
assert_eq!(shutdowns.load(Ordering::SeqCst), 1);
}
// ---------------------------------------------------------------------------
// Supervision trees: only forest roots are addressed
// ---------------------------------------------------------------------------
/// The supervisor gets ONE request and runs its ordered shutdown; a trapping
/// child under it sees exactly one `Shutdown` — from the supervisor, not a
/// second one from the runtime — and is allowed to finish its drain (a sleep,
/// i.e. a timer) under `Shutdown::Infinity`.
#[test]
fn supervised_children_are_shut_down_only_via_their_supervisor() {
let signals = Arc::new(AtomicUsize::new(0));
let drained = Arc::new(AtomicBool::new(false));
let sup_returned = Arc::new(AtomicBool::new(false));
let (sg, dr, sr) = (signals.clone(), drained.clone(), sup_returned.clone());
run(move || {
spawn(move || {
OneForOne::new()
.child(
ChildSpec::new(Restart::Permanent, move || {
let inbox = trap_exit();
while let Ok(sig) = inbox.recv() {
if sig.reason == DownReason::Shutdown {
sg.fetch_add(1, Ordering::SeqCst);
sleep(Duration::from_millis(100));
// A late second signal would land here.
while let Ok(Some(sig)) = inbox.try_recv() {
if sig.reason == DownReason::Shutdown {
sg.fetch_add(1, Ordering::SeqCst);
}
}
dr.store(true, Ordering::SeqCst);
return;
}
}
})
.shutdown(Shutdown::Infinity),
)
.run();
sr.store(true, Ordering::SeqCst);
});
sleep(Duration::from_millis(30)); // let the tree settle
});
assert!(
sup_returned.load(Ordering::SeqCst),
"supervisor should return normally"
);
assert!(
drained.load(Ordering::SeqCst),
"child should finish its drain"
);
assert_eq!(signals.load(Ordering::SeqCst), 1);
}
/// A supervised non-trapping child under `Shutdown::Timeout` is stopped by
/// the supervisor's policy, and the run ends promptly.
#[test]
fn supervisor_tree_is_torn_down_promptly_on_root_exit() {
let t = Instant::now();
run(|| {
spawn(|| {
OneForOne::new()
.child(
ChildSpec::new(Restart::Permanent, || loop {
sleep(Duration::from_millis(5));
})
.shutdown(Shutdown::Timeout(Duration::from_millis(50))),
)
.run();
});
sleep(Duration::from_millis(30));
});
assert_prompt(t, "tree");
}
+36
View File
@@ -0,0 +1,36 @@
//! Under `smarm-trace`, every actor the root-exit sweep reaches is recorded
//! as a `root_sweep` event — the way to *see* unsupervised leftovers. One test
//! per binary: the trace file is process-global.
#![cfg(feature = "smarm-trace")]
use smarm::gen_server::{self, GenServer, GenServerCtx};
use smarm::run;
struct Quiet;
impl GenServer for Quiet {
type Call = ();
type Reply = ();
type Cast = ();
type Info = ();
type Timer = ();
fn init(&mut self, _: &GenServerCtx<Self>) {}
fn handle_call(&mut self, _: ()) {}
fn handle_cast(&mut self, _: ()) {}
}
#[test]
fn forgotten_server_shows_up_as_root_sweep() {
let path = std::env::temp_dir().join(format!("smarm_root_sweep_{}.json", std::process::id()));
std::env::set_var("SMARM_TRACE_FILE", &path);
run(|| {
let srv = gen_server::start(Quiet);
srv.call(()).unwrap();
drop(srv); // forgotten: nobody supervises it, nobody holds it
});
let trace = std::fs::read_to_string(&path).expect("trace file written");
let _ = std::fs::remove_file(&path);
assert!(
trace.contains("root_sweep stopped"),
"expected a root_sweep line for the non-trapping leftover; got:\n{trace}"
);
}
+122
View File
@@ -0,0 +1,122 @@
//! Graceful shutdown — `request_shutdown` (OTP `exit(Pid, shutdown)`).
//!
//! `request_stop` is `exit(Pid, kill)`: an uncatchable unwind at the target's
//! next observation point. `request_shutdown` is the polite form:
//! - a target that is NOT trapping exits is stopped exactly as by
//! `request_stop` (OTP's rule: don't trap, you die);
//! - a target that IS trapping receives an `ExitSignal { reason: Shutdown }`
//! on its trap inbox and keeps running — it is expected to wind down and
//! exit normally on its own.
use smarm::{monitor, request_shutdown, run, self_pid, sleep, spawn, trap_exit, DownReason, Pid};
use std::sync::atomic::{AtomicBool, Ordering};
use std::sync::{mpsc, Arc};
use std::thread;
use std::time::Duration;
const WATCHDOG: Duration = Duration::from_secs(10);
#[test]
fn request_shutdown_stops_a_non_trapping_actor() {
run(|| {
let h = spawn(|| sleep(Duration::from_secs(3600)));
let mon = monitor(h.pid());
request_shutdown(h.pid());
let down = mon.rx.recv().expect("down");
assert_eq!(down.reason, DownReason::Stopped);
});
}
#[test]
fn request_shutdown_is_a_message_to_a_trapping_actor() {
let unwound = Arc::new(AtomicBool::new(false));
let u = unwound.clone();
run(move || {
struct Unwound(Arc<AtomicBool>);
impl Drop for Unwound {
fn drop(&mut self) {
if std::thread::panicking() {
self.0.store(true, Ordering::SeqCst);
}
}
}
let (tx, rx) = smarm::channel::<(Pid, DownReason)>();
let (ready_tx, ready_rx) = smarm::channel::<()>();
let h = spawn(move || {
let _g = Unwound(u);
let inbox = trap_exit();
let _ = ready_tx.send(());
let sig = inbox.recv().expect("exit signal");
let _ = tx.send((sig.from, sig.reason));
// Keep doing work after the request: shutdown is advisory.
sleep(Duration::from_millis(20));
});
// Trapping is set by the target itself; a request that beats it is a
// plain stop (same window as OTP's exit-before-process_flag).
ready_rx.recv().expect("ready");
let me = self_pid();
let mon = monitor(h.pid());
request_shutdown(h.pid());
let (from, reason) = rx.recv().expect("relayed");
assert_eq!(from, me);
assert_eq!(reason, DownReason::Shutdown);
let down = mon.rx.recv().expect("down");
assert_eq!(
down.reason,
DownReason::Exit,
"target exited normally, not stopped"
);
});
assert!(
!unwound.load(Ordering::SeqCst),
"trapping target must not be unwound"
);
}
#[test]
fn request_shutdown_on_dead_pid_is_a_no_op() {
run(|| {
let h = spawn(|| {});
let pid = h.pid();
let _ = h.join();
request_shutdown(pid); // must not panic
});
}
#[test]
fn handle_request_shutdown_from_foreign_thread() {
let rt = smarm::init(smarm::Config::exact(2));
let handle = rt.handle();
let (pid_tx, pid_rx) = mpsc::channel::<Pid>();
let requester = thread::spawn(move || {
let pid = pid_rx.recv().expect("pid");
thread::sleep(Duration::from_millis(50));
handle.request_shutdown(pid);
});
let (done_tx, done_rx) = mpsc::channel();
thread::spawn(move || {
rt.run(move || {
let (tx, rx) = smarm::channel::<DownReason>();
let (ready_tx, ready_rx) = smarm::channel::<()>();
let h = spawn(move || {
let inbox = trap_exit();
let _ = ready_tx.send(());
let sig = inbox.recv().expect("exit signal");
let _ = tx.send(sig.reason);
});
ready_rx.recv().expect("ready");
pid_tx.send(h.pid()).expect("send pid");
let reason = rx.recv().expect("relayed");
assert_eq!(reason, DownReason::Shutdown);
let _ = h.join();
});
let _ = done_tx.send(());
});
done_rx
.recv_timeout(WATCHDOG)
.expect("run did not return: foreign-thread request_shutdown never reached the target");
requester.join().expect("requester thread");
}
+303
View File
@@ -0,0 +1,303 @@
//! Supervisor shutdown — the OTP child-spec `shutdown` policy.
//!
//! A supervisor traps exits. A `request_shutdown` reaching it (from its parent
//! supervisor, or from the app via `request_shutdown`/`RuntimeHandle`) runs
//! the ordered shutdown: children are stopped in reverse start order, each
//! per its `Shutdown` policy — `request_shutdown`, wait up to the timeout for
//! its termination signal, `request_stop` if it overstays — and then `run()`
//! returns normally. Every supervisor-initiated child stop (ordered shutdown,
//! OneForAll/RestForOne sibling cycling) goes through the same policy.
//!
//! A *hard* `request_stop` on a supervisor unwinds it; a drop guard then
//! hard-stops its live children so the subtree is never orphaned.
use smarm::supervisor::{ChildSpec, OneForOne, Restart, Shutdown, Strategy};
use smarm::{
monitor, request_shutdown, request_stop, run, sleep, spawn, trap_exit, DownReason, JoinHandle,
};
use std::sync::atomic::{AtomicBool, AtomicUsize, Ordering};
use std::sync::{Arc, Mutex};
use std::time::{Duration, Instant};
/// A child that traps exits, records the order it was shut down in, and exits
/// normally on the request (after `delay`). Ignores the request if `comply`
/// is false — a straggler that must be hard-stopped.
fn polite_child(
tag: usize,
log: &Arc<Mutex<Vec<usize>>>,
delay: Duration,
comply: bool,
) -> impl Fn() + Send + Sync + 'static {
let log = log.clone();
move || {
let inbox = trap_exit();
loop {
let sig = match inbox.recv() {
Ok(s) => s,
Err(_) => return,
};
if sig.reason == DownReason::Shutdown {
log.lock().unwrap().push(tag);
if comply {
sleep(delay);
return;
}
// Not complying: keep running until hard-stopped.
loop {
sleep(Duration::from_millis(5));
}
}
}
}
}
/// Spawn `sup`, let its children reach `trap_exit`, return the handle.
fn spawn_settled(sup: OneForOne) -> JoinHandle {
let h = spawn(move || sup.run());
sleep(Duration::from_millis(30));
h
}
#[test]
fn shutdown_stops_children_in_reverse_order_and_returns_normally() {
let log = Arc::new(Mutex::new(Vec::new()));
let l = log.clone();
run(move || {
let sup = OneForOne::new()
.child(ChildSpec::new(
Restart::Permanent,
polite_child(1, &l, Duration::ZERO, true),
))
.child(ChildSpec::new(
Restart::Permanent,
polite_child(2, &l, Duration::ZERO, true),
))
.child(ChildSpec::new(
Restart::Permanent,
polite_child(3, &l, Duration::ZERO, true),
));
let h = spawn_settled(sup);
let mon = monitor(h.pid());
request_shutdown(h.pid());
let down = mon.rx.recv().expect("down");
assert_eq!(
down.reason,
DownReason::Exit,
"supervisor exits normally after shutdown"
);
});
assert_eq!(*log.lock().unwrap(), vec![3, 2, 1]);
}
#[test]
fn non_trapping_child_is_simply_stopped() {
let dropped = Arc::new(AtomicBool::new(false));
let d = dropped.clone();
run(move || {
struct G(Arc<AtomicBool>);
impl Drop for G {
fn drop(&mut self) {
self.0.store(true, Ordering::SeqCst);
}
}
let sup = OneForOne::new().child(ChildSpec::new(Restart::Permanent, move || {
let _g = G(d.clone());
loop {
sleep(Duration::from_millis(5));
}
}));
let h = spawn_settled(sup);
request_shutdown(h.pid());
h.join().expect("sup");
});
assert!(dropped.load(Ordering::SeqCst));
}
#[test]
fn straggler_is_hard_stopped_after_timeout() {
let log = Arc::new(Mutex::new(Vec::new()));
let l = log.clone();
run(move || {
let sup = OneForOne::new().child(
ChildSpec::new(
Restart::Permanent,
polite_child(1, &l, Duration::ZERO, false),
)
.shutdown(Shutdown::Timeout(Duration::from_millis(50))),
);
let h = spawn_settled(sup);
let t0 = Instant::now();
request_shutdown(h.pid());
h.join().expect("sup");
let took = t0.elapsed();
assert!(
took >= Duration::from_millis(50),
"returned before the grace period: {took:?}"
);
assert!(
took < Duration::from_secs(2),
"did not fall back to a hard stop: {took:?}"
);
});
assert_eq!(
*log.lock().unwrap(),
vec![1],
"the straggler did receive the request"
);
}
#[test]
fn infinity_waits_for_a_slow_but_compliant_child() {
let log = Arc::new(Mutex::new(Vec::new()));
let l = log.clone();
let finished = Arc::new(AtomicBool::new(false));
let f = finished.clone();
run(move || {
let f2 = f.clone();
let l2 = l.clone();
let sup = OneForOne::new().child(
ChildSpec::new(Restart::Permanent, move || {
let inbox = trap_exit();
let _ = inbox.recv();
l2.lock().unwrap().push(1);
sleep(Duration::from_millis(150));
f2.store(true, Ordering::SeqCst); // only reached if not hard-stopped
})
.shutdown(Shutdown::Infinity),
);
let h = spawn_settled(sup);
request_shutdown(h.pid());
h.join().expect("sup");
});
assert!(
finished.load(Ordering::SeqCst),
"Infinity must not hard-stop a compliant child"
);
}
#[test]
fn brutal_kill_skips_the_request() {
let log = Arc::new(Mutex::new(Vec::new()));
let l = log.clone();
run(move || {
let sup = OneForOne::new().child(
ChildSpec::new(
Restart::Permanent,
polite_child(1, &l, Duration::ZERO, true),
)
.shutdown(Shutdown::BrutalKill),
);
let h = spawn_settled(sup);
request_shutdown(h.pid());
h.join().expect("sup");
});
assert!(
log.lock().unwrap().is_empty(),
"a BrutalKill child never sees the request"
);
}
#[test]
fn hard_stop_of_supervisor_does_not_orphan_children() {
let alive = Arc::new(AtomicUsize::new(0));
let a = alive.clone();
run(move || {
struct Alive(Arc<AtomicUsize>);
impl Drop for Alive {
fn drop(&mut self) {
self.0.fetch_sub(1, Ordering::SeqCst);
}
}
let mk = |a: Arc<AtomicUsize>| {
move || {
a.fetch_add(1, Ordering::SeqCst);
let _g = Alive(a.clone());
loop {
sleep(Duration::from_millis(5));
}
}
};
let sup = OneForOne::new()
.child(ChildSpec::new(Restart::Permanent, mk(a.clone())))
.child(ChildSpec::new(Restart::Permanent, mk(a.clone())));
let h = spawn_settled(sup);
assert_eq!(a.load(Ordering::SeqCst), 2);
let mon = monitor(h.pid());
request_stop(h.pid());
let _ = mon.rx.recv();
sleep(Duration::from_millis(50));
assert_eq!(
a.load(Ordering::SeqCst),
0,
"children orphaned by a hard supervisor stop"
);
});
}
#[test]
fn nested_shutdown_reaches_grandchildren() {
let log = Arc::new(Mutex::new(Vec::new()));
let l = log.clone();
run(move || {
let l_inner = l.clone();
let inner = move || {
OneForOne::new()
.child(ChildSpec::new(
Restart::Permanent,
polite_child(10, &l_inner, Duration::ZERO, true),
))
.child(ChildSpec::new(
Restart::Permanent,
polite_child(11, &l_inner, Duration::ZERO, true),
))
.run()
};
let sup = OneForOne::new()
.child(ChildSpec::new(
Restart::Permanent,
polite_child(1, &l, Duration::ZERO, true),
))
.child(ChildSpec::new(Restart::Permanent, inner).shutdown(Shutdown::Infinity));
let h = spawn_settled(sup);
request_shutdown(h.pid());
h.join().expect("sup");
});
assert_eq!(*log.lock().unwrap(), vec![11, 10, 1]);
}
#[test]
fn sibling_cycling_uses_graceful_shutdown() {
// OneForAll: when child A dies, sibling B (trapping) must receive a
// Shutdown request rather than a bare stop.
let log = Arc::new(Mutex::new(Vec::new()));
let l = log.clone();
let a_runs = Arc::new(AtomicUsize::new(0));
let ar = a_runs.clone();
run(move || {
let ar2 = ar.clone();
let sup = OneForOne::new()
.strategy(Strategy::OneForAll)
.intensity(5, Duration::from_secs(60))
.child(ChildSpec::new(Restart::Transient, move || {
let n = ar2.fetch_add(1, Ordering::SeqCst) + 1;
sleep(Duration::from_millis(30));
if n == 1 {
panic!("first run dies");
}
// Second run: park until shut down.
let inbox = trap_exit();
let _ = inbox.recv();
}))
.child(ChildSpec::new(
Restart::Permanent,
polite_child(2, &l, Duration::ZERO, true),
));
let h = spawn(move || sup.run());
sleep(Duration::from_millis(150));
request_shutdown(h.pid());
h.join().expect("sup");
});
// B was shut down once by the cycle and once by the final shutdown.
assert_eq!(*log.lock().unwrap(), vec![2, 2]);
assert_eq!(a_runs.load(Ordering::SeqCst), 2);
}