Compare commits
9
Commits
v0.6.1
..
741c10337b
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
741c10337b | ||
|
|
e570138da5 | ||
|
|
415effb2e9 | ||
|
|
849a424c8e | ||
|
|
6ceb138f5f | ||
|
|
250f31265b | ||
|
|
9c8f59ca53 | ||
|
|
1002777ef3 | ||
|
|
8f2d513940 |
+1
-1
@@ -1,6 +1,6 @@
|
||||
[package]
|
||||
name = "smarm"
|
||||
version = "0.6.1"
|
||||
version = "0.7.0"
|
||||
edition = "2021"
|
||||
rust-version = "1.95"
|
||||
|
||||
|
||||
@@ -1,35 +1,38 @@
|
||||
# smarm
|
||||
|
||||
> SMARM — Smarm, Marks Actor Runtime Machinery. A proof-of-concept green-thread actor runtime for Rust.
|
||||
> SMARM: Smarm, Marks Actor Runtime Machinery. A proof-of-concept green-thread actor runtime for Rust.
|
||||
|
||||
Implements the core ideas in [`Achitecture.md`](.docs/Architecture.md): green-thread actors on a
|
||||
shared heap, scheduled cooperatively, communicating only by `Send` messages.
|
||||
Erlang's isolation model without Erlang's copying GC, Rust's zero-copy
|
||||
ownership transfers without async's function colouring.
|
||||
SMARM is my attempt to implement the erlang/OTP philosophy in the Rust programming language. This has yielded a fault-tolerant, fast, and scalable runtime. This runtime allows the creation of asynchronous applications in Rust without the function coloring associated with the async/await system. It encourages the writing of simple, synchronous code, and largely elides the need for lifetime annotations.
|
||||
|
||||
The scheduler is multi-threaded — one OS thread per available CPU, all drawing
|
||||
from a shared run queue. The single-threaded `run()` entry point is kept as a
|
||||
convenience wrapper around `runtime::init(Config::exact(1)).run(f)`.
|
||||
|
||||
## What's here
|
||||
## Overview
|
||||
|
||||
SMARM implements green-thread actors on a shared heap, communicating only by `Send` messages. By sharing the heap, SMARM avoids the copying overhead of Erlang, which is safe to do due to Rust's borrow checker.
|
||||
|
||||
On top of the core runtime mechanics, SMARM also provides a library of primitives for making applications closely inspired by erlang/OTP. This includes generic servers (gen_servers), generic state machines (gen_statem), and supervision trees.
|
||||
|
||||
Supervision trees are the core primitive to allow your application to survive an unexpected panic. Supervisors are processes dedicated to monitoring other processes, which can restart these should they fail. This means that when set up properly an application may 'self-heal' when encountering unforeseen circumstances.
|
||||
|
||||
SMARM is not cooperatively scheduled; it uses preemption. This means a heavy task will not starve out other lighter tasks. Everything will make steady progress, which translates to very beneficial behaviour under (over)load: average latency goes up, but tail latency does not blow up.
|
||||
|
||||
To help diagnose these unforeseen circumstances, smarm may be compiled with its `tracing` feature, which emits a full trace using [Perfetto](https://perfetto.dev/).
|
||||
|
||||
Should you want to optimize your application, SMARM is unusually well poised to help. As the runtime functionally controls time, SMARM comes with a built in causal profiler, under the `causal` feature.
|
||||
|
||||
I also built a Phoenix-Framework inspired HTTP 1.1 library on top of SMARM called [URUS](https://git.kalsbeek.dev/Markk116/urus), which implements Pub/Sub, Channels, and basic amenities like Websockets and Server Sent Events.
|
||||
|
||||
## Limitations
|
||||
|
||||
This runtime requires naked assembly to function, and has thus far only been implemented for x86-64 assembly. It expects an operating system that supports virtual address space, and is therefore not (yet) suited for embedded targets. The IO implementation is currently based around the Linux kernel's `epoll` mechanism, meaning it requires a (GNU+)Linux distribution to run.
|
||||
|
||||
The preemption mechanism works by wrapping the memory allocator and checking how many CPU cycles you have used compared to your timeslice budget. This allows preemption to fire in most normal code, but tight zero-allocation loops do not get caught and require manual insertion of `check!()` if you want preemption to function.
|
||||
|
||||
This library is still in its early stages, and while I try my best with loom and tests, stable operation cannot be guaranteed. Therefore it is not (yet) recommended for production use.
|
||||
|
||||
At this moment, stack memory for each green thread is capped. Uncapping this may lead to performance benefits for deeply recursive algorithms that in a traditional async runtime might require pointer-chases through the heap. This is as yet unrealised.
|
||||
|
||||
At this stage, the codebase is largely LLM-generated, which is obvious if you start to read through the internals. While I did the design, and I keep the LLM under tight rein, the codebase is not in a state that I am very happy with. This also goes for the documentation.
|
||||
|
||||
| Module | What it does |
|
||||
|--------------|------------------------------------------------------------------------|
|
||||
| `stack` | `mmap`'d growable stack with guard page; SIGSEGV on overflow |
|
||||
| `context` | `#[naked]` x86-64 context-switch shims, callee-saved regs only |
|
||||
| `preempt` | Allocator-driven preemption; `check!()` macro for no-alloc loops |
|
||||
| `pid` | `(index, generation)` PIDs; stale handles are detectable, not silent |
|
||||
| `actor` | Trampoline + `catch_unwind` boundary at the actor entry point |
|
||||
| `scheduler` | Run queue, slot table, spawn/join, parking, idle path |
|
||||
| `channel` | Unbounded MPSC channel; `recv` parks the actor; `recv_timeout` bounds it; `select`/`select_timeout` park on many receivers at once (ready-index, priority order) |
|
||||
| `mutex` | `Mutex<T>` with mandatory timeout; FIFO waiters; parks the green thread |
|
||||
| `timer` | Min-heap of `(deadline, reason)`; `Sleep` and `WaitTimeout` reasons |
|
||||
| `io` | `block_on_io` for blocking work; `wait_readable`/`wait_writable` + `read`/`write` via epoll |
|
||||
| `supervisor` | `Signal::Exit`/`Panic`/`Stopped` funnelled to a parent; `OneForOne`/`OneForAll`/`RestForOne` strategies + restart-intensity cap |
|
||||
| `monitor` | `monitor(pid)` → `Monitor { id, target, rx }`; one-shot `Down` via `rx`; `demonitor(&m)` tears one registration down; unidirectional death notice |
|
||||
| `link` | bidirectional `link`/`unlink`; abnormal death propagates (cooperative stop, or an `ExitSignal` message under `trap_exit`) |
|
||||
| `gen_server` | `call`/`call_timeout` (sync request-reply) / `cast` (async) over one inbox; `handle_info` over static info arms + `handle_down` via `Watcher`-fed monitors, selected ahead of the inbox; `ServerRef`/`ServerBuilder` + `init`/`terminate` hooks; server-down via channel closure |
|
||||
| `registry` | `register`/`whereis`/`name_of`: name ↔ pid bimap; lazy generation-checked cleanup |
|
||||
|
||||
## Quick taste
|
||||
|
||||
@@ -51,6 +54,29 @@ run(|| {
|
||||
});
|
||||
```
|
||||
|
||||
## Stopping actors
|
||||
|
||||
Two strengths, as in OTP. `request_stop(pid)` is `exit(Pid, kill)`: a cooperative
|
||||
hard stop, unwinding at the actor's next observation point. `request_shutdown(pid)`
|
||||
is `exit(Pid, shutdown)`: an actor that traps exits (`trap_exit()`, or
|
||||
`ctx.trap_exit()` in a gen_server / `cx.trap_exit()` in a gen_statem) receives it
|
||||
as a signal — `handle_shutdown` / a `shutdown` row — and may drain before stopping
|
||||
itself; one that does not trap is stopped outright. Supervisors trap:
|
||||
`request_shutdown(sup)` tears the tree down top-down, each child per its
|
||||
`ChildSpec` `Shutdown` policy (`Timeout(d)`, `Infinity`, `BrutalKill`). The run's
|
||||
root actor returning means "the program is done": every top-level actor gets a
|
||||
`request_shutdown`, and `run()` returns when they are gone. From outside the
|
||||
runtime (a signal thread), `Runtime::handle().request_shutdown(pid)` does the
|
||||
same. `examples/graceful_shutdown.rs` shows all of it.
|
||||
|
||||
A gen_server or gen_statem lives until it stops, is shut down, or is killed; its
|
||||
refs are addresses — dropping them never ends it (a forgotten one is swept at
|
||||
root exit; with `--features smarm-trace` each such sweep is a `root_sweep`
|
||||
trace line). The supervised shape is `GenServerBuilder::named(N).run()` /
|
||||
`gen_statem::run_named(N, m)`: the server runs inline as the `ChildSpec` child
|
||||
itself, so the supervisor's shutdown reaches it directly, a restart re-binds the
|
||||
name, and the program addresses it by name.
|
||||
|
||||
## Layout
|
||||
|
||||
```
|
||||
@@ -67,12 +93,9 @@ benches/
|
||||
|
||||
## Building and running
|
||||
|
||||
Standard Cargo. Requires Rust 1.95 or newer (the `#[naked]` attribute went stable
|
||||
in 1.88; we use a few unrelated post-1.88 features). `master` is x86-64 Linux
|
||||
only. An experimental, **untested** aarch64 context-switch backend lives on the
|
||||
`arm-port` branch (extracted into a `target_arch`-gated `src/arch/`); it has not
|
||||
been validated on hardware yet. macOS remains on the deferred list because of the
|
||||
epoll dependency.
|
||||
Standard Cargo. Requires Rust 1.95 or newer (the `#[naked]` attribute went stable in 1.88; we use a few unrelated post-1.88 features). I have worked hard to keep this library as dependency-free as possible. `master` is x86-64 Linux only. An experimental, **untested** aarch64 context-switch backend lives on the `arm-port` branch (extracted into a `target_arch`-gated `src/arch/`); it has not been validated on hardware yet. macOS remains on the deferred list because of the epoll dependency.
|
||||
|
||||
|
||||
|
||||
```sh
|
||||
cargo test # all tests
|
||||
@@ -80,26 +103,29 @@ cargo test --test mutex # one module
|
||||
cargo bench # primes benchmark vs tokio
|
||||
```
|
||||
|
||||
## What's not here
|
||||
|
||||
See the **Defer** section of `Architecture.md`.
|
||||
`join!` for handle groups, stack growth via remap,
|
||||
hierarchical timer wheel, fd-wait timeouts, `Signal::Timeout`. Each is
|
||||
mechanism we know how to add; none belongs in this iteration.
|
||||
|
||||
## Docs
|
||||
|
||||
| Document | What it covers |
|
||||
|---|---|
|
||||
| [`Architecture.md`](./docs/Architecture.md) | Design intent, runtime model, and deferred work |
|
||||
| [`smarm - Deep Dive.html`](./docs/smarm%20-%20Deep%20Dive.html) | Generated walkthrough of the system; good starting point |
|
||||
| [`smarm - Deep Dive.html`](./docs/smarm%20-%20Deep%20Dive.html) | Generated walkthrough of the system; good starting point if you want to learn about the internals |
|
||||
| [`BENCHMARKS_AND_TUNING.md`](./docs/BENCHMARKS_AND_TUNING.md) | Where smarm wins and loses vs tokio, preemption knob recommendations |
|
||||
| [`benchmarks.md`](./docs/benchmarks.md) | Raw benchmark results, methodology, and tuning experiment log |
|
||||
|
||||
## Coming up
|
||||
|
||||
Clustering: clustering multiple SMARM nodes together is in the pipeline.
|
||||
SMARM-BEAM Interop: Running SMARM as a supervised node under the BEAM via a Rustler NIF works, including message passing and supervision trees that span the runtimes. However, this library is still too unstable to release.
|
||||
SMARM is an interesting platform for implementing a 'dataflow' library, but work on this has not yet started.
|
||||
|
||||
|
||||
## Contributing
|
||||
|
||||
This is a personal proof-of-concept. There's no PR workflow. If you fork it and do something interesting, just send me an email. If it's nice, I'll upstream the changes.
|
||||
This started as a personal proof-of-concept, but it is starting to outgrow that name. If you want to contribute, please get in contact to discuss what you want to work on. Code without prior communication is not welcome.
|
||||
|
||||
|
||||
## A note on open source
|
||||
|
||||
An open source project is a gift, and by giving it, it is no longer mine. I highly enourage you to fork it, to make it your own. This repository, however, is still mine.
|
||||
|
||||
---
|
||||
|
||||
|
||||
|
||||
+48
-4
@@ -77,10 +77,10 @@ Delivered surface:
|
||||
`ServerBuilder::start` untouched; free `call` / `cast` / `whereis_server`;
|
||||
`ServerRef::shutdown` + free `shutdown` as the sys-style synchronous stop.
|
||||
- **Root-exit teardown** (final phase): the run's initial actor is the root;
|
||||
when it exits, the scheduler's idle verdict stops the parked-forever remainder
|
||||
(deferred past the queue drain, so actors with in-flight work finish rather
|
||||
than unwinding on the stop). Closes the "app actor blocks AllDone" stall — see
|
||||
Look into, below.
|
||||
when it exits the run winds down. *(Reworked with the graceful-shutdown work:
|
||||
root exit now delivers `request_shutdown` to every forest root — see
|
||||
"Root exit" below and `tests/root_exit.rs`.)* Closes the "app actor blocks
|
||||
AllDone" stall — see Look into, below.
|
||||
|
||||
Extends — does not retire — the "select exists; a unified per-process mailbox
|
||||
still does not" invariant: 014 adds addressable *delivery*, not a unified inbox;
|
||||
@@ -260,8 +260,52 @@ path the atomic-bool workaround stood in for. Re-check the urus crud repro to
|
||||
confirm the workaround can be retired (the teardown is cooperative — an actor in
|
||||
a tight loop with no observation point still can't be stopped).
|
||||
|
||||
**Update (graceful shutdown):** the RFC 014 sweep was a hard `request_stop` of
|
||||
every live slot, deferred until nothing was runnable — which killed a sleeping
|
||||
actor (timer pending) but drained a queued one, for no principled reason. It is
|
||||
now the OTP semantics: root exit = "the program is done" = `request_shutdown`
|
||||
to every **forest root** (live actor whose parent is the run or is dead), run
|
||||
synchronously on the root's finalize path. Supervisors cascade with their child
|
||||
`Shutdown` policies; trapping actors may `Continue`/drain (timers keep working)
|
||||
and end the run when they stop themselves; non-trapping actors are stopped
|
||||
outright — `join` what you need finished. No forcing sweep follows.
|
||||
|
||||
---
|
||||
|
||||
### Open items from the graceful-shutdown work (not scheduled)
|
||||
- ~~gen_server / gen_statem as a direct supervised child.~~ Done: lifetime is
|
||||
the actor's (refs are addresses, the loop holds an inbox sender);
|
||||
`NamedGenServerBuilder::run` / `gen_statem::run_named` run the loop inline as
|
||||
the `ChildSpec` child; the root-exit sweep traces each leftover as
|
||||
`root_sweep` under `smarm-trace`.
|
||||
- Supervisor `Live` drop-guard sweep is `request_stop` (kill propagates as
|
||||
kill); OTP would deliver a trappable `killed`. Chosen for boundedness.
|
||||
- A root-exit shutdown reaches only actors live *at that instant*; a
|
||||
non-trapping forest root that spawns before it unwinds leaves that spawn
|
||||
to itself (Erlang: an unlinked spawn is nobody's child).
|
||||
- **Supervisor start *order* is not start *readiness*.** `start_child`
|
||||
spawns and moves straight on, so an earlier child is merely *scheduled*,
|
||||
not initialised, when a later sibling starts. A later child that resolves
|
||||
an earlier one by name (`whereis_server`) can therefore miss it — the
|
||||
classic "named registry sibling, then its consumers" tree. Ordered
|
||||
`OneForOne`/`RestForOne` shutdown is unaffected (reverse order is honoured
|
||||
and each stop *is* awaited); this is a start-side gap only.
|
||||
Making `spawn` itself block does NOT fix it — it would only shrink the
|
||||
window to "child has begun executing", while the property callers need is
|
||||
"child has bound its name / opened its socket", which only the child can
|
||||
declare. It would also tax the hot path (one round-trip per accepted
|
||||
connection) and turn every spawn into a context-switch point. OTP has the
|
||||
same async `spawn` and puts the synchronisation one level up:
|
||||
`gen_server:start_link` blocks the caller until `init/1` returns.
|
||||
Fix shape when scheduled: a readiness ack in the supervisor's child-start
|
||||
path (`ChildSpec` variant whose factory receives a ready-signal;
|
||||
`NamedGenServerBuilder::run` acks after its name bind, gen_server default
|
||||
acks after `init`; plain closures ack at spawn as today, i.e. opt-in with
|
||||
no cost to existing children). Until then the workaround is structural:
|
||||
have the registrar spawn its own consumers so the ordering is program
|
||||
order inside one actor, not a cross-actor guarantee (urus v0.3 endpoint
|
||||
does exactly this).
|
||||
|
||||
## Invariants & gotchas (respect these across all cycles)
|
||||
|
||||
- **Shared mutex is non-reentrant.** `Sender::send` can call `unpark` →
|
||||
|
||||
+229
@@ -0,0 +1,229 @@
|
||||
# urus / smarm handoff — updated 2026-08-19 (session 3)
|
||||
|
||||
## TL;DR for the next session
|
||||
**smarm is done for now** (5 unpushed commits on local `master`, see below).
|
||||
**Next = urus v0.3 endpoint refactor.** You should NOT need to read smarm
|
||||
scheduler internals; the contract you build on is fully described here and in
|
||||
`smarm_full/examples/graceful_shutdown.rs` (read that file first — it is the
|
||||
exact shape urus's tree will take) plus `smarm_full/tests/root_exit.rs`.
|
||||
|
||||
### The smarm contract urus builds on (all on local master, verified by tests)
|
||||
- `request_stop(pid)` = kill (cooperative hard stop). `request_shutdown(pid)` =
|
||||
polite: trapping target gets `ExitSignal{reason: Shutdown}`, non-trapping is
|
||||
stopped outright. `RuntimeHandle::{request_stop,request_shutdown}` do the same
|
||||
from any OS thread (signal handler); grab `rt.handle()` before `rt.run`.
|
||||
- Supervisor traps; `request_shutdown(sup)` = ordered reverse-start shutdown,
|
||||
per-child `ChildSpec::shutdown(Shutdown::{Timeout(d)|Infinity|BrutalKill})`
|
||||
(default Timeout(5s)); sup then returns normally. `request_stop(sup)`
|
||||
hard-stops children too (no orphans).
|
||||
- gen_server: `ctx.trap_exit()` in init; `handle_shutdown() -> Exit|Continue`;
|
||||
`handle_exit(sig)`; `ctx.stop_handle().stop()` = normal self-exit;
|
||||
`terminate()` may block only on the graceful path (Exit / stop / inbox close).
|
||||
`GenServerRef::shutdown()` is graceful and waits.
|
||||
- gen_statem: same in event clothes — `cx.trap_exit()` in initial enter,
|
||||
`shutdown` rows (default `stop`), `exit sig` rows, `cx.stop()` / `stop` tail,
|
||||
optional `terminate { }` block. `GenStatemRef::shutdown()`.
|
||||
- **Root exit = program done**: when the root actor returns, the runtime
|
||||
`request_shutdown`s every *forest root* (live actor whose parent is the run
|
||||
or dead). Supervisors cascade; trapping actors may drain (timers keep
|
||||
working) and end the run when they stop; non-trapping are stopped; **no
|
||||
forcing sweep** (`join` what must finish). The old "wait until nothing
|
||||
runnable then kill all" deferral is gone.
|
||||
- **Gotcha for urus:** a gen_server's lifetime is governed by its refs — drop
|
||||
the last `GenServerRef` and the inbox closes → clean exit *even mid-drain*.
|
||||
The endpoint must be pinned (named, or its ref held by the supervisor
|
||||
wrapper) or it will terminate the moment the root drops its ref.
|
||||
- **Known gap (ROADMAP open item):** a gen_server can't be a direct `ChildSpec`
|
||||
child; use the trapping wrapper pattern in `examples/graceful_shutdown.rs::
|
||||
drainer_child` (starts `under(self_pid())`, forwards shutdown, waits). Doing
|
||||
an inline `GenServerBuilder::run()` first may be worth a short smarm detour —
|
||||
decide with Markk.
|
||||
|
||||
### smarm commits this session (local master, NOT pushed, NOT tagged)
|
||||
`250f312` root-exit = graceful shutdown of forest roots (tests/root_exit.rs)
|
||||
`6ceb138` gen_statem shutdown parity (tests/gen_statem_shutdown.rs)
|
||||
`849a424` docs + examples/graceful_shutdown.rs + README "Stopping actors"
|
||||
On top of `1002777` (cross-thread wake) and `9c8f59c` (graceful shutdown).
|
||||
Cargo.toml still `0.6.1`. Release cut (push, tag — v0.7 is justified by the
|
||||
API surface — version bump) is Markk's. Full suite, doc tests, examples,
|
||||
`cargo fmt`, `cargo clippy --lib` all clean. (`clippy --tests` has pre-existing
|
||||
unwrap lints in tests/fd_select.rs, untouched.)
|
||||
|
||||
### Decisions taken this session (Markk)
|
||||
- Root exit means "program done" (Go/tokio/OTP), not "wait for pending work";
|
||||
the previously agreed "sleep(50ms) must finish" test was dropped as encoding
|
||||
the wrong contract (a timer-wheel gate would re-wedge periodic-timer daemons).
|
||||
- No behaviour-preserving deferral, no forcing second sweep.
|
||||
- Examples/docs done in the same session; urus next session.
|
||||
|
||||
---
|
||||
# Previous handoff (still accurate where not superseded above)
|
||||
|
||||
|
||||
## Next-session goal
|
||||
Phase 1 is **done and committed**; Phase 2 is next:
|
||||
1. **smarm v0.6.2** — cross-thread wake root fix. **DONE**, committed on `master`
|
||||
as `1002777`. Not yet tagged, not yet version-bumped (Cargo.toml still reads
|
||||
`0.6.1`), and **not yet pushed to origin** — it exists only in the delivered
|
||||
snapshot zip and the local sandbox clone. Cutting the release (push + tag
|
||||
`v0.6.2` + bump `0.6.1`→`0.6.2`) is Markk's step.
|
||||
2. **urus v0.3** — endpoint refactor. Working against `smarm = { path = "../smarm_full" }`
|
||||
with `git update-index --skip-worktree Cargo.toml` (Markk approved); release commit
|
||||
swaps back to the tag once Markk cuts it. Phase 2 plan below is STALE where it says
|
||||
drain-in-terminate; the endpoint is a trapping GenServer: `handle_shutdown` →
|
||||
`Continue`, enter Draining, `StopHandle::stop()` when the conn set empties.
|
||||
|
||||
Decisions below are locked unless marked *(confirm)*.
|
||||
|
||||
## Reconstruction (the sandbox resets between sessions)
|
||||
A fresh sandbox has an empty home and **no Rust toolchain**. To restore:
|
||||
- Install rustup/cargo. smarm reformats under **rustc 1.97.1**; urus `rust-version`
|
||||
is 1.95. Use 1.97.1.
|
||||
- urus: `git clone https://git.kalsbeek.dev/Markk116/urus` — `origin` is registered
|
||||
and public-read. master `8bdec97` = the v0.2.x line. (Zips in outputs are stale;
|
||||
prefer the remote now.)
|
||||
- smarm: `git clone https://git.kalsbeek.dev/Markk116/smarm`. Latest tag **v0.6.1**
|
||||
(`ca1c983`). The cross-thread wake fix is committed as `1002777` on top of the
|
||||
post-v0.6.1 README commit `8f2d513` (= origin/master). **It is NOT on origin
|
||||
yet** — a fresh clone won't have it until Markk pushes. Restore it from the
|
||||
snapshot zip if working before the push. **v0.6.2 is not yet tagged.**
|
||||
- urus pins smarm by git **tag** in `Cargo.toml` (currently `v0.6.0`). A trivial
|
||||
first commit bumps it to `v0.6.1` (also picks up `try_spawn` + monitor
|
||||
terminal-outcome fixes).
|
||||
|
||||
## Why (context — the finding that drives the plan)
|
||||
urus's shutdown machinery (the `AtomicBool` listener flag + the `SHUTDOWN_POLL`
|
||||
loop in `serve.rs`) is scaffolding around two smarm properties. Their statuses
|
||||
differ, which is the whole point:
|
||||
|
||||
- **Issue A — lossy stop vs a QUEUED actor: ALREADY FIXED in smarm.** Commit
|
||||
`7bab4d2` added an entry-side `check_cancelled()` in `park_current`. A
|
||||
`request_stop` against a listener parked in `wait_readable_timeout` now unwinds
|
||||
cleanly (it parks via `try_select_timeout → park_current`). urus's flag + its
|
||||
stale "smarm's lossy stop-while-QUEUED window" comment can be deleted.
|
||||
- **Issue B — foreign-thread wake is a no-op: FIXED in `1002777` (was present
|
||||
through v0.6.1).** The gap: `unpark`/`unpark_at`/`request_stop` all route through
|
||||
`try_with_runtime`, which reads a thread-local that is `None` on any non-scheduler
|
||||
thread, so no cross-thread wake worked — a signal handler / OS thread could not
|
||||
wake *or* stop a parked actor, which is why `serve.rs` polls the shutdown signal
|
||||
instead of parking on it. Now closed (see Phase 1 below): urus's `SHUTDOWN_POLL`
|
||||
loop can be deleted and its `Handle::shutdown` can park on a handle-driven stop.
|
||||
|
||||
## Phase 1 — smarm cross-thread wake (root fix) — DONE (`1002777`)
|
||||
Shipped as one commit generalizing RFC 018 (a producer reaches the runtime through
|
||||
a `Weak` it holds) from the IO backend to channel senders and a new handle:
|
||||
- **`Runtime::handle() -> RuntimeHandle`** (`Send + Sync`), holding a
|
||||
`Weak<RuntimeInner>`. Grab it before `rt.run` and hand it to the signal thread.
|
||||
- **`RuntimeHandle::request_stop<A>(Pid<A>)`** — upgrades the Weak and calls
|
||||
`request_stop_inner` on the inner; no-op if the runtime is gone. This is the
|
||||
signal-handler-drives-shutdown path; it cascades the ordered stop down the tree
|
||||
exactly like an in-runtime `request_stop`.
|
||||
- **Send-wake:** the receiver captures `scheduler::runtime_weak()` into its
|
||||
`parked_receiver` tuple **at park time** (not at channel creation — the resolved
|
||||
sub-decision; a parked receiver is a live actor so the Weak is provably upgradable,
|
||||
and it scopes the capture to when a wake is possible). `send()` and last-sender
|
||||
`drop` wake via `scheduler::unpark_at_via(pid, epoch, &weak)`: thread-local path
|
||||
when on a scheduler thread (preempt-gated, slot-eligible), captured Weak otherwise.
|
||||
In-runtime timer wakes (recv/select) were left on `scheduler::unpark_at`.
|
||||
|
||||
**API scope decision (signed off):** `RuntimeHandle` exposes **`request_stop` only**.
|
||||
No public `unpark`/`unpark_at` on the handle — send-wake needs no user-facing handle,
|
||||
and "unpark off-runtime" is covered because `request_stop` drives `unpark` on the
|
||||
upgraded inner. No `is_alive()`. Both are one-line additions if a consumer appears.
|
||||
|
||||
**No RFC written** — pattern was already established (RFC 018), agreed not needed.
|
||||
|
||||
Tests: `tests/cross_thread_wake.rs` (foreign-thread send wakes a parked receiver;
|
||||
foreign-thread `request_stop` wakes+stops a parked actor; a lingering handle never
|
||||
blocks all-done and degrades to a no-op once the runtime drops). Full suite green;
|
||||
`cargo fmt` + `cargo clippy --lib` clean.
|
||||
|
||||
**Remaining release step (Markk):** push `master`, tag `v0.6.2`, bump Cargo.toml
|
||||
`0.6.1`→`0.6.2`. Left paired with the tag as the release cut, not done in `1002777`.
|
||||
|
||||
## Phase 1b — smarm graceful shutdown (OTP lift) — DONE (`9c8f59c`, on top of `1002777`)
|
||||
Decided this session (Markk): B — fix at the smarm level rather than a two-stop
|
||||
split in urus. No RFC (Markk: "just implement it"). Shipped, tested, committed on
|
||||
the local `master`, **not pushed, not tagged**. It should ship as the same
|
||||
release as 1002777 (v0.6.2, or v0.7 given the API surface — Markk's call).
|
||||
- `request_shutdown(pid)` / `RuntimeHandle::request_shutdown` = `exit(Pid, shutdown)`;
|
||||
`request_stop` = `exit(Pid, kill)`. Trapping target gets `ExitSignal{reason:
|
||||
DownReason::Shutdown}`; non-trapping is stopped outright.
|
||||
- `ChildSpec::shutdown(Shutdown::{BrutalKill, Timeout(d), Infinity})`, default 5s.
|
||||
Supervisor traps exits; `request_shutdown(sup)` = ordered top-down shutdown,
|
||||
returns normally. **Also fixed**: `request_stop(sup)` used to ORPHAN children
|
||||
(probe-verified; the handoff's "cascade" claim was wrong) — `Live` drop guard now
|
||||
hard-stops them.
|
||||
- gen_server: `ctx.trap_exit()`, `handle_shutdown() -> ShutdownAction::{Exit,
|
||||
Continue}`, `handle_exit(ExitSignal)`, `ctx.stop_handle().stop()` = normal
|
||||
self-exit (`{stop, normal}`; previously impossible — only abnormal `Stopped`).
|
||||
`GenServerRef::shutdown()` is graceful now.
|
||||
- Root finding that forced this: gen_server `terminate()` runs from a Drop guard,
|
||||
mid-unwind on the stop path; any park in it = double panic = abort. So
|
||||
"drain-in-terminate()" (the old Phase 2 plan) was never viable.
|
||||
|
||||
### ~~Next-session smarm work~~ DONE this session (see TL;DR)
|
||||
1. **Root-exit sweep**: make `Pop::RootDrain` also require an empty timer wheel
|
||||
(and it already requires nothing runnable; io_out is only checked for AllDone —
|
||||
check whether it should gate RootDrain too). TDD: an actor in `sleep(50ms)` when
|
||||
the root returns must finish, not be swept. Then a `Reservoir`-style test that a
|
||||
*truly* parked-forever daemon still gets swept.
|
||||
2. **gen_statem parity**: `ctx.trap_exit()`, `handle_shutdown -> ShutdownAction`,
|
||||
`handle_exit`, stop handle. Mirror gen_server; mechanical.
|
||||
3. **Examples review**: `examples/*.rs` predate all of this. Rework where they show
|
||||
shutdown/teardown to use `request_shutdown`, `Shutdown` policies, and
|
||||
`StopHandle`; `named_genserver.rs` first (uses `shutdown`). Also
|
||||
`docs/smarm - Deep Dive.html` says terminate() must be non-blocking — now only
|
||||
true on the unwind paths; and README could use a "Stopping actors" paragraph
|
||||
(request_stop = kill, request_shutdown = shutdown, Shutdown policy).
|
||||
|
||||
### Open smarm items found on the way (noted, not scheduled)
|
||||
- (root-exit sweep and gen_statem parity moved up to the scheduled list.)
|
||||
- Sweep in the supervisor `Live` drop guard is `request_stop` (kill propagates as
|
||||
kill); OTP would deliver a trappable `killed`. Chosen for boundedness.
|
||||
|
||||
## Phase 2 — urus v0.3: endpoint refactor (after v0.6.2 is tagged)
|
||||
Target = the spec's original shape (`urus-spec.md` §2.1/§6: `listener_sup` under the
|
||||
**user's** root supervisor). Deviation to unwind: `serve` owning `rt.run`.
|
||||
- App owns the runtime: `smarm::init(cfg).run(|| root_sup.run())`, root e.g.
|
||||
`RestForOne[ app actors…, urus::endpoint(config, pipeline) ]`. This is what kills
|
||||
the `Arc<OnceLock>` idiom for the right reason (app state born in-runtime as a
|
||||
supervised, ordered child).
|
||||
- `urus::endpoint` = one GenServer child owning the registry + an **internal**
|
||||
listener sub-supervisor + drain-in-`terminate()`. Listeners stay internal, not
|
||||
app-visible peers.
|
||||
- Shutdown = `request_stop` the root supervisor (or via the runtime handle from a
|
||||
signal thread) → cascades down → `endpoint.terminate()` runs the drain
|
||||
(`drain_timeout`, force-stop sweep).
|
||||
- **DELETE:** the `AtomicBool` listener flag (A fixed) and the `SHUTDOWN_POLL` loop
|
||||
+ its apologetic comment (B fixed → park, don't poll).
|
||||
- Nuance: `request_stop` → `Signal::Stopped` is *abnormal* → `Transient` restarts.
|
||||
Stop-without-restart = stop the **supervisor**, not the children.
|
||||
- *(confirm)* Keep `serve`/`serve_with`/`serve_with_shutdown` as thin wrappers that
|
||||
build the one-child tree internally, so the simple case stays one line.
|
||||
- *(confirm)* Keep `Handle`/`ShutdownSignal`? Now that cross-thread wake works,
|
||||
`Handle::shutdown` can map to handle-driven `request_stop` on the endpoint.
|
||||
- Breaking → cut **urus v0.3**; bump the smarm pin to `v0.6.2` here.
|
||||
|
||||
## Working norms
|
||||
- Every bash call: `export PATH=$HOME/.cargo/bin:$PATH` (once the toolchain's in).
|
||||
- **TDD**: failing test first, then implement; keep suites green.
|
||||
- **Hammer ritual** for ANY connection-lifecycle change (the urus #2 shutdown work
|
||||
qualifies): 35× subset (`shutdown timeout reaped slowloris streaming chunked sse
|
||||
stalled ws_ channels session`) + 3× full + 1× trace. `scripts/hammer.sh` does NOT
|
||||
pass feature flags — loop manually with `--features phoenix`. Subset filter must
|
||||
NOT use `--test integration` (session tests live in lib).
|
||||
- Example smoke tests: hold the server's stdin open (`mkfifo` + `sleep > fifo`) or
|
||||
the Enter-to-shutdown thread fires on EOF instantly.
|
||||
- Background procs are reaped BETWEEN bash calls; `pkill -f` matches your own shell.
|
||||
- Artefact store (specs): `curl -H "Authorization: Bearer sk-llmingest-2e45d80c63db24c6781e761eb2a9a58e83d9f48ef77a42185bad311d07c80e68" https://artefacts.kalsbeek.dev/artifacts/<name>`
|
||||
— `urus-spec.md`, `urus-bench-spec.md`, `rfc_008-implementation-notes.md`, …
|
||||
- smarm feature flags: `smarm-trace`, `smarm-causal` (urus re-exports both).
|
||||
|
||||
## Local cross-repo testing — KEEP OUT OF COMMITS
|
||||
To test urus #2 against un-tagged smarm 0.6.2, point urus's `Cargo.toml` smarm dep
|
||||
at a local path (`smarm = { path = "../smarm" }`) instead of the git tag.
|
||||
- Must NOT land in commits. Guard: `git update-index --skip-worktree Cargo.toml`
|
||||
after editing (undo with `--no-skip-worktree`), or stash before committing.
|
||||
- The committed `Cargo.toml` stays pinned to the git tag; restore the tag (bumped to
|
||||
`v0.6.2`) for the release commit.
|
||||
@@ -1620,8 +1620,9 @@
|
||||
wait: <code>select</code> priority is <strong>Down arms › Watcher arm › info channels (declaration order) ›
|
||||
inbox</strong>, rebuilt each turn. A hot inbox can't starve a death notice or a system message;
|
||||
conversely a hot info channel <em>can</em> starve the inbox — deliberately. A closed info arm is
|
||||
silently dropped from the set; a closed <em>inbox</em> (every <code>ServerRef</code> gone) is graceful
|
||||
shutdown.</p>
|
||||
silently dropped from the set. The inbox never closes — the loop holds one sender for its whole
|
||||
life, so a <code>GenServerRef</code> is an address, not an owner: the server ends only by
|
||||
<code>StopHandle::stop</code>, a shutdown, a hard stop, or a panic.</p>
|
||||
|
||||
<h3>Death needs no monitor</h3>
|
||||
<p>Server death detection falls out of channel closure. Already dead → the inbox is closed and
|
||||
@@ -1700,7 +1701,7 @@
|
||||
</div>
|
||||
<div class="module-card">
|
||||
<div class="module-name" style="color:var(--red)">Panics in <code>terminate()</code></div>
|
||||
<p>gen_server's <code>terminate()</code> runs from a drop guard, possibly mid-unwind. A panic inside it during an unwind is a double panic → process abort, no supervision tree to save you. Keep it cheap, non-blocking, non-panicking.</p>
|
||||
<p>gen_server's <code>terminate()</code> runs from a drop guard, possibly mid-unwind. A panic inside it during an unwind is a double panic → process abort, no supervision tree to save you. On the panic and hard-stop paths keep it cheap, non-blocking, non-panicking. Only the graceful path (<code>handle_shutdown → Exit</code>, <code>StopHandle::stop</code>) runs it outside an unwind, where it may do real work.</p>
|
||||
</div>
|
||||
<div class="module-card">
|
||||
<div class="module-name" style="color:var(--yellow)">Cold locks are leaf locks</div>
|
||||
|
||||
@@ -0,0 +1,139 @@
|
||||
//! Graceful shutdown, end to end: a supervised app tree, a server that
|
||||
//! drains before it exits, and the two ways the whole thing winds down.
|
||||
//!
|
||||
//! Stopping an actor comes in two strengths, as in OTP:
|
||||
//! - `request_stop(pid)` = `exit(Pid, kill)`: cooperative hard stop,
|
||||
//! unwinds at the next observation point.
|
||||
//! - `request_shutdown(pid)` = `exit(Pid, shutdown)`: a trapping target gets
|
||||
//! an `ExitSignal { reason: Shutdown }` and winds
|
||||
//! down on its own terms; a non-trapping one is
|
||||
//! stopped outright.
|
||||
//!
|
||||
//! A supervisor traps exits. `request_shutdown(sup)` runs its ordered
|
||||
//! shutdown — children in reverse start order, each per its `ChildSpec`
|
||||
//! `Shutdown` policy (`Timeout(d)` default 5s, `Infinity`, `BrutalKill`) —
|
||||
//! and the supervisor then returns normally.
|
||||
//!
|
||||
//! Two triggers are shown:
|
||||
//! 1. **Root exit.** The run's root actor returning means "the program is
|
||||
//! done": the runtime delivers `request_shutdown` to every top-level actor
|
||||
//! (here: the supervisor). Trapping actors may keep running to drain and
|
||||
//! end the run when they stop themselves; non-trapping ones are stopped.
|
||||
//! 2. **An outside thread** (e.g. a signal handler) driving it via
|
||||
//! `RuntimeHandle::request_shutdown` on the supervisor — the root then
|
||||
//! just waits for the tree to come down.
|
||||
|
||||
use smarm::gen_server::{
|
||||
GenServer, GenServerBuilder, GenServerCtx, GenServerName, ShutdownAction, StopHandle,
|
||||
TimerHandle,
|
||||
};
|
||||
use smarm::supervisor::{ChildSpec, OneForOne, Restart, Shutdown};
|
||||
use smarm::{sleep, spawn};
|
||||
use std::thread;
|
||||
use std::time::Duration;
|
||||
|
||||
/// A server with in-flight work: on shutdown it stops accepting, finishes what
|
||||
/// it has (simulated with a ticking timer), then ends itself.
|
||||
struct Drainer {
|
||||
pending: u32,
|
||||
stop: Option<StopHandle<Drainer>>,
|
||||
timer: Option<TimerHandle<Drainer>>,
|
||||
}
|
||||
|
||||
impl GenServer for Drainer {
|
||||
type Call = ();
|
||||
type Reply = ();
|
||||
type Cast = ();
|
||||
type Info = ();
|
||||
type Timer = ();
|
||||
|
||||
fn init(&mut self, ctx: &GenServerCtx<Self>) {
|
||||
ctx.trap_exit(); // opt in: shutdown arrives as handle_shutdown
|
||||
self.stop = Some(ctx.stop_handle());
|
||||
self.timer = Some(ctx.timer());
|
||||
}
|
||||
fn handle_call(&mut self, _: ()) {}
|
||||
fn handle_cast(&mut self, _: ()) {}
|
||||
fn handle_shutdown(&mut self) -> ShutdownAction {
|
||||
println!(
|
||||
"drainer: shutdown requested, {} items pending",
|
||||
self.pending
|
||||
);
|
||||
self.timer
|
||||
.as_ref()
|
||||
.unwrap()
|
||||
.tick_every(Duration::from_millis(20), ());
|
||||
ShutdownAction::Continue // keep serving until drained
|
||||
}
|
||||
fn handle_timer(&mut self, _: ()) {
|
||||
self.pending -= 1;
|
||||
if self.pending == 0 {
|
||||
println!("drainer: drained, stopping");
|
||||
self.stop.as_ref().unwrap().stop(); // normal exit
|
||||
}
|
||||
}
|
||||
fn terminate(&mut self) {
|
||||
// Graceful path: this runs on the normal path and may block.
|
||||
println!("drainer: terminate");
|
||||
}
|
||||
}
|
||||
|
||||
/// The server's name: how the rest of the app reaches it (and the only handle
|
||||
/// that survives a restart).
|
||||
const DRAINER: GenServerName<Drainer> = GenServerName::new("drainer");
|
||||
|
||||
fn app_tree() -> OneForOne {
|
||||
OneForOne::new()
|
||||
.child(
|
||||
ChildSpec::new(Restart::Permanent, || {
|
||||
// A plain worker that does not trap: stopped outright on shutdown.
|
||||
loop {
|
||||
sleep(Duration::from_millis(10));
|
||||
}
|
||||
})
|
||||
.shutdown(Shutdown::Timeout(Duration::from_millis(100))),
|
||||
)
|
||||
// A gen_server is a direct child: `named(N).run()` runs the loop as
|
||||
// the child actor itself, so the supervisor's shutdown arrives as
|
||||
// `handle_shutdown` and a restart re-binds the name.
|
||||
.child(
|
||||
ChildSpec::new(Restart::Permanent, || {
|
||||
GenServerBuilder::new(Drainer {
|
||||
pending: 3,
|
||||
stop: None,
|
||||
timer: None,
|
||||
})
|
||||
.named(DRAINER)
|
||||
.run()
|
||||
.expect("drainer name is free");
|
||||
})
|
||||
.shutdown(Shutdown::Infinity),
|
||||
)
|
||||
}
|
||||
|
||||
fn main() {
|
||||
println!("--- 1. root exit drives the shutdown ---");
|
||||
smarm::run(|| {
|
||||
spawn(|| app_tree().run());
|
||||
sleep(Duration::from_millis(50)); // the app "runs" for a while
|
||||
// Returning here asks the supervisor to shut down; the run ends when
|
||||
// the tree — drainer included — is gone.
|
||||
});
|
||||
|
||||
println!("--- 2. an outside thread drives the shutdown ---");
|
||||
let rt = smarm::init(smarm::Config::default());
|
||||
let handle = rt.handle(); // Send + Sync; grab it before run
|
||||
rt.run(move || {
|
||||
let sup = spawn(|| app_tree().run());
|
||||
let sup_pid = sup.pid();
|
||||
// Stand-in for a SIGTERM handler thread.
|
||||
thread::spawn(move || {
|
||||
thread::sleep(Duration::from_millis(50));
|
||||
println!("signal thread: requesting shutdown");
|
||||
handle.request_shutdown(sup_pid);
|
||||
});
|
||||
sup.join()
|
||||
.expect("supervisor returns normally after ordered shutdown");
|
||||
println!("supervisor down; root returns");
|
||||
});
|
||||
}
|
||||
@@ -66,6 +66,12 @@ fn main() {
|
||||
let svc: Option<GenServerRef<Counter>> = whereis_server(COUNTER);
|
||||
if let Some(svc) = svc {
|
||||
let _ = svc.call(Query::Get);
|
||||
// A named server is pinned alive by the registry, so dropping refs
|
||||
// does not end it. Stop it explicitly: `shutdown()` asks politely
|
||||
// (a trapping server drains first; this one is stopped outright)
|
||||
// and waits until it is gone. Left running, the root's return
|
||||
// would shut it down the same way — see examples/graceful_shutdown.rs.
|
||||
svc.shutdown();
|
||||
}
|
||||
});
|
||||
}
|
||||
|
||||
+34
-21
@@ -90,8 +90,9 @@
|
||||
|
||||
use crate::pid::Pid;
|
||||
use crate::raw_mutex::RawMutex;
|
||||
use crate::runtime::RuntimeInner;
|
||||
use std::collections::VecDeque;
|
||||
use std::sync::Arc;
|
||||
use std::sync::{Arc, Weak};
|
||||
|
||||
/// Create a new channel and return its `(Sender, Receiver)` halves.
|
||||
///
|
||||
@@ -114,12 +115,15 @@ pub fn channel<T>() -> (Sender<T>, Receiver<T>) {
|
||||
|
||||
struct Inner<T> {
|
||||
queue: VecDeque<T>,
|
||||
/// The parked receiver's `(pid, park-epoch)`, if one is currently
|
||||
/// The parked receiver's `(pid, park-epoch, runtime)`, if one is currently
|
||||
/// waiting. The epoch identifies exactly which wait this is, so a waker
|
||||
/// left over from a wait that already ended (a losing `select` arm, a
|
||||
/// `recv_timeout` whose timer fired after it was already satisfied) is
|
||||
/// inert and does nothing when it fires.
|
||||
parked_receiver: Option<(Pid, u32)>,
|
||||
/// inert and does nothing when it fires. The `Weak<RuntimeInner>` is the
|
||||
/// receiver's runtime, captured while it parked (so provably alive then);
|
||||
/// it lets a sender on a foreign OS thread wake the receiver without the
|
||||
/// `RUNTIME` thread-local, which is unset off a scheduler thread.
|
||||
parked_receiver: Option<(Pid, u32, Weak<RuntimeInner>)>,
|
||||
senders: usize,
|
||||
receiver_alive: bool,
|
||||
}
|
||||
@@ -206,8 +210,8 @@ impl<T> Drop for Sender<T> {
|
||||
None
|
||||
}
|
||||
};
|
||||
if let Some((pid, epoch)) = unpark {
|
||||
crate::scheduler::unpark_at(pid, epoch);
|
||||
if let Some((pid, epoch, rt)) = unpark {
|
||||
crate::scheduler::unpark_at_via(pid, epoch, &rt);
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -254,13 +258,13 @@ impl<T> Sender<T> {
|
||||
g.queue.push_back(value);
|
||||
g.parked_receiver.take()
|
||||
};
|
||||
if let Some((pid, epoch)) = unpark {
|
||||
if let Some((pid, epoch, rt)) = unpark {
|
||||
crate::te!(crate::trace::Event::Send {
|
||||
sender: crate::actor::current_pid()
|
||||
.unwrap_or(crate::pid::Pid::new(u32::MAX, u32::MAX)),
|
||||
receiver: Some(pid)
|
||||
});
|
||||
crate::scheduler::unpark_at(pid, epoch);
|
||||
crate::scheduler::unpark_at_via(pid, epoch, &rt);
|
||||
} else {
|
||||
crate::te!(crate::trace::Event::Send {
|
||||
sender: crate::actor::current_pid()
|
||||
@@ -293,13 +297,17 @@ impl<T> Receiver<T> {
|
||||
None => panic!("smarm: recv() called outside an actor"),
|
||||
};
|
||||
debug_assert!(
|
||||
g.parked_receiver.is_none_or(|(p, _)| p == me),
|
||||
g.parked_receiver.as_ref().is_none_or(|(p, _, _)| *p == me),
|
||||
"channel has more than one receiver"
|
||||
);
|
||||
// begin_wait is lock-free, so it's legal under the Channel lock;
|
||||
// registering in the same critical section makes the epoch
|
||||
// atomic with the senders' view of the registration.
|
||||
g.parked_receiver = Some((me, crate::scheduler::begin_wait()));
|
||||
g.parked_receiver = Some((
|
||||
me,
|
||||
crate::scheduler::begin_wait(),
|
||||
crate::scheduler::runtime_weak(),
|
||||
));
|
||||
crate::te!(crate::trace::Event::RecvPark(me));
|
||||
}
|
||||
// Release the lock before parking: the unparker will need it.
|
||||
@@ -347,11 +355,11 @@ impl<T> Receiver<T> {
|
||||
return Err(RecvTimeoutError::Disconnected);
|
||||
}
|
||||
debug_assert!(
|
||||
g.parked_receiver.is_none_or(|(p, _)| p == me),
|
||||
g.parked_receiver.as_ref().is_none_or(|(p, _, _)| *p == me),
|
||||
"channel has more than one receiver"
|
||||
);
|
||||
epoch = crate::scheduler::begin_wait();
|
||||
g.parked_receiver = Some((me, epoch));
|
||||
g.parked_receiver = Some((me, epoch, crate::scheduler::runtime_weak()));
|
||||
crate::te!(crate::trace::Event::RecvPark(me));
|
||||
}
|
||||
|
||||
@@ -423,10 +431,14 @@ impl<T> Receiver<T> {
|
||||
None => panic!("smarm: recv_match() called outside an actor"),
|
||||
};
|
||||
debug_assert!(
|
||||
g.parked_receiver.is_none_or(|(p, _)| p == me),
|
||||
g.parked_receiver.as_ref().is_none_or(|(p, _, _)| *p == me),
|
||||
"channel has more than one receiver"
|
||||
);
|
||||
g.parked_receiver = Some((me, crate::scheduler::begin_wait()));
|
||||
g.parked_receiver = Some((
|
||||
me,
|
||||
crate::scheduler::begin_wait(),
|
||||
crate::scheduler::runtime_weak(),
|
||||
));
|
||||
crate::te!(crate::trace::Event::RecvPark(me));
|
||||
}
|
||||
// Release the lock before parking: the unparker will need it.
|
||||
@@ -497,11 +509,12 @@ impl<T: Send + 'static> crate::timer::TimerTarget for RawMutex<Inner<T>> {
|
||||
// keeps the registration bookkeeping exact.)
|
||||
let unpark = {
|
||||
let mut g = self.lock();
|
||||
if g.parked_receiver == Some((pid, epoch)) {
|
||||
g.parked_receiver = None;
|
||||
true
|
||||
} else {
|
||||
false
|
||||
match g.parked_receiver {
|
||||
Some((p, e, _)) if p == pid && e == epoch => {
|
||||
g.parked_receiver = None;
|
||||
true
|
||||
}
|
||||
_ => false,
|
||||
}
|
||||
};
|
||||
// Unpark outside the channel lock: it may take the run-queue lock;
|
||||
@@ -562,10 +575,10 @@ impl<T> Selectable for Receiver<T> {
|
||||
return Ok(false);
|
||||
}
|
||||
debug_assert!(
|
||||
g.parked_receiver.is_none_or(|(p, _)| p == pid),
|
||||
g.parked_receiver.as_ref().is_none_or(|(p, _, _)| *p == pid),
|
||||
"channel has more than one receiver"
|
||||
);
|
||||
g.parked_receiver = Some((pid, epoch));
|
||||
g.parked_receiver = Some((pid, epoch, crate::scheduler::runtime_weak()));
|
||||
Ok(true)
|
||||
}
|
||||
|
||||
|
||||
+222
-42
@@ -127,16 +127,51 @@
|
||||
//! - [`GenServer::init`] runs once before the first message. Use it to start
|
||||
//! timers or set up monitors; see the [`GenServerCtx`] it receives.
|
||||
//! - [`GenServer::terminate`] runs when the server is about to exit. It fires
|
||||
//! on every exit path (all `GenServerRef`s dropped, a handler panic, or an
|
||||
//! explicit [`GenServerRef::shutdown`]), not only on clean shutdown. Keep it
|
||||
//! short and non-blocking: if `terminate` panics while the server is already
|
||||
//! unwinding from a handler panic, the process aborts.
|
||||
//! on every exit path (a self-stop, a graceful shutdown, a cooperative hard
|
||||
//! stop, or a handler panic), not only on clean shutdown.
|
||||
//! Keep it non-panicking: on the panic and hard-stop paths it runs
|
||||
//! mid-unwind, where a second panic aborts the process and where it must
|
||||
//! not block (any park re-observes the stop). Only on the graceful path
|
||||
//! (see below) may it do real work.
|
||||
//!
|
||||
//! ## When the server stops
|
||||
//!
|
||||
//! The server runs as long as at least one [`GenServerRef`] exists. When the last
|
||||
//! one is dropped, the inbox closes and the loop exits gracefully. To stop a
|
||||
//! server explicitly and wait for it to finish, call [`GenServerRef::shutdown`].
|
||||
//! A server lives until it stops, is shut down, or is killed — as an OTP
|
||||
//! process does. A [`GenServerRef`] is an *address*: cloning and dropping it
|
||||
//! never changes the server's lifetime, and a ref that nobody holds is not a
|
||||
//! leak — a forgotten server idles until the run ends, when the root-exit
|
||||
//! shutdown (see [`Runtime::run`](crate::Runtime::run)) takes it down with
|
||||
//! every other unsupervised actor. Anything meant to live long should be
|
||||
//! supervised (see *Supervised servers* below); [`start`] / [`start_under`]
|
||||
//! are for scripts, tests and short-lived helpers, and the explicit close is
|
||||
//! [`GenServerRef::shutdown`].
|
||||
//!
|
||||
//! A server can end itself: clone a [`StopHandle`] from
|
||||
//! [`GenServerCtx::stop_handle`] in `init` and call [`StopHandle::stop`] from
|
||||
//! any handler — the loop breaks after the current message and exits
|
||||
//! *normally* (OTP's `{stop, normal}`). This is distinct from
|
||||
//! `request_stop(self_pid())`, which is an abnormal `Stopped` and gets a
|
||||
//! `Transient` child restarted.
|
||||
//!
|
||||
//! ## Graceful shutdown
|
||||
//!
|
||||
//! From outside, [`GenServerRef::shutdown`] (or a plain
|
||||
//! [`request_shutdown`](crate::request_shutdown), which is what a supervisor
|
||||
//! sends) asks the server to stop. What happens next is the server's choice:
|
||||
//!
|
||||
//! - By default a server does not trap exits, and the request stops it
|
||||
//! outright at its next observation point — `terminate` runs mid-unwind.
|
||||
//! - A server that calls [`GenServerCtx::trap_exit`] in `init` receives the
|
||||
//! request as [`GenServer::handle_shutdown`]. Return
|
||||
//! [`ShutdownAction::Exit`] (the default) to have the loop break and
|
||||
//! `terminate` run on the normal path, where it may block; return
|
||||
//! [`ShutdownAction::Continue`] to keep serving — e.g. to drain in-flight
|
||||
//! work — and end the server later with a [`StopHandle`]. The supervisor's
|
||||
//! [`Shutdown`](crate::supervisor::Shutdown) policy bounds how long that
|
||||
//! may take before it falls back to a hard stop.
|
||||
//!
|
||||
//! A trapping server also receives the deaths of its linked peers as
|
||||
//! [`GenServer::handle_exit`] messages instead of dying with them.
|
||||
//!
|
||||
//! If the server panics inside a handler, the panic unwinds the server thread.
|
||||
//! Any caller currently waiting in `call` sees `Err(ServerDown)`: the reply
|
||||
@@ -167,9 +202,23 @@
|
||||
//! Names and registration: a server can be given a static name so other
|
||||
//! actors can reach it without holding a `GenServerRef`. Use
|
||||
//! [`GenServerBuilder::named`] to register on start, and the free functions
|
||||
//! [`call`], [`cast`], and [`whereis_server`] to address it by name. Registered
|
||||
//! servers are a natural fit for supervision; see `supervisor` for how to
|
||||
//! build a tree that restarts servers on failure.
|
||||
//! [`call`], [`cast`], and [`whereis_server`] to address it by name.
|
||||
//!
|
||||
//! ## Supervised servers
|
||||
//!
|
||||
//! The supervised shape is [`NamedGenServerBuilder::run`]: it runs the loop
|
||||
//! **inline, as the current actor**, so the closure of a
|
||||
//! [`ChildSpec`](crate::supervisor::ChildSpec) *is* the server — the
|
||||
//! supervisor's shutdown arrives as [`GenServer::handle_shutdown`], a restart
|
||||
//! runs the factory again and re-binds the name, and the rest of the program
|
||||
//! addresses it by name (a held ref would go stale on restart anyway).
|
||||
//!
|
||||
//! ```ignore
|
||||
//! const COUNTER: GenServerName<Counter> = GenServerName::new("counter");
|
||||
//! OneForOne::new().child(ChildSpec::new(Restart::Permanent, || {
|
||||
//! GenServerBuilder::new(Counter::default()).named(COUNTER).run().unwrap();
|
||||
//! }));
|
||||
//! ```
|
||||
//!
|
||||
//! ## Limitations
|
||||
//!
|
||||
@@ -181,10 +230,12 @@
|
||||
use crate::channel::{
|
||||
channel, select, select_timeout, Receiver, RecvTimeoutError, Selectable, Sender,
|
||||
};
|
||||
use crate::link::ExitSignal;
|
||||
use crate::monitor::DownReason;
|
||||
use crate::monitor::{demonitor, monitor, Down, Monitor};
|
||||
use crate::pid::Pid;
|
||||
use crate::registry::{register_with, resolve_named_sender, RegisterError};
|
||||
use crate::scheduler::{cancel_timer, request_stop, send_after_to};
|
||||
use crate::scheduler::{cancel_timer, request_shutdown, send_after_to};
|
||||
use crate::timer::TimerId;
|
||||
use std::cell::Cell;
|
||||
use std::collections::HashMap;
|
||||
@@ -254,10 +305,34 @@ pub trait GenServer: Send + 'static {
|
||||
/// Default: no-op.
|
||||
fn handle_idle(&mut self) {}
|
||||
|
||||
/// A graceful shutdown request (a [`request_shutdown`](crate::request_shutdown)
|
||||
/// reaching this server), delivered only if `init` called
|
||||
/// [`GenServerCtx::trap_exit`]. Return [`ShutdownAction::Exit`] to stop
|
||||
/// now (the default), or [`ShutdownAction::Continue`] to keep serving and
|
||||
/// end the server later with a [`StopHandle`].
|
||||
fn handle_shutdown(&mut self) -> ShutdownAction {
|
||||
ShutdownAction::Exit
|
||||
}
|
||||
|
||||
/// A linked peer's abnormal death (an [`ExitSignal`] that is not a
|
||||
/// shutdown request), delivered only if `init` called
|
||||
/// [`GenServerCtx::trap_exit`]. Default: drop it.
|
||||
fn handle_exit(&mut self, _sig: ExitSignal) {}
|
||||
|
||||
/// Runs as the server actor exits, on any exit path (see module docs).
|
||||
fn terminate(&mut self) {}
|
||||
}
|
||||
|
||||
/// What a server does with a shutdown request; see [`GenServer::handle_shutdown`].
|
||||
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
|
||||
pub enum ShutdownAction {
|
||||
/// Break the loop now. `terminate` runs on the normal path and may block.
|
||||
Exit,
|
||||
/// Keep dispatching. The server is expected to end itself with a
|
||||
/// [`StopHandle`] once it is done winding down.
|
||||
Continue,
|
||||
}
|
||||
|
||||
/// What travels the server's single inbox channel: a synchronous call (with a
|
||||
/// reply sender) or an asynchronous cast. Private — callers use [`GenServerRef`].
|
||||
enum Envelope<G: GenServer> {
|
||||
@@ -265,9 +340,9 @@ enum Envelope<G: GenServer> {
|
||||
Cast(G::Cast),
|
||||
}
|
||||
|
||||
/// A clonable handle to a running server. Cloning yields another sender to the
|
||||
/// same inbox; the server lives until the last `GenServerRef` is dropped, at which
|
||||
/// point its inbox closes and the loop exits normally.
|
||||
/// A clonable handle to a running server: an *address*, not an owner. Cloning
|
||||
/// yields another sender to the same inbox; dropping refs never ends the
|
||||
/// server (see the module docs, *When the server stops*).
|
||||
pub struct GenServerRef<G: GenServer> {
|
||||
tx: Sender<Envelope<G>>,
|
||||
pid: Pid,
|
||||
@@ -362,18 +437,22 @@ impl<G: GenServer> GenServerRef<G> {
|
||||
|
||||
/// Stop the server and block until it has fully exited.
|
||||
///
|
||||
/// Sends a cooperative stop signal to the server actor and waits for it to
|
||||
/// exit, so [`GenServer::terminate`] has run by the time this returns.
|
||||
/// Returns immediately if the server is already gone.
|
||||
/// Asks the server to shut down (a [`request_shutdown`](crate::request_shutdown))
|
||||
/// and waits for it to exit, so [`GenServer::terminate`] has run by the
|
||||
/// time this returns. A trapping server gets to wind down via
|
||||
/// [`GenServer::handle_shutdown`]; any other is stopped outright. Returns
|
||||
/// immediately if the server is already gone. Waits as long as the server
|
||||
/// takes — the caller, not the server, decides whether that is acceptable;
|
||||
/// a supervisor uses its child's [`Shutdown`](crate::supervisor::Shutdown)
|
||||
/// policy to bound it.
|
||||
///
|
||||
/// This is the right teardown for a server kept alive by a registered
|
||||
/// [`GenServerName`], where dropping every external `GenServerRef` is not enough
|
||||
/// to close the inbox. Like all cooperative cancellation, it is best-effort:
|
||||
/// This is the explicit close: dropping refs never ends a server. Like
|
||||
/// all cooperative cancellation, it is best-effort:
|
||||
/// a server wedged in a tight loop with no observation point cannot be
|
||||
/// stopped this way. Panics if called outside `Runtime::run()`.
|
||||
pub fn shutdown(&self) {
|
||||
let mon = monitor(self.pid);
|
||||
request_stop(self.pid);
|
||||
request_shutdown(self.pid);
|
||||
// The Down lands when the server finalizes; an already-dead target makes
|
||||
// `monitor` deliver NoProc immediately, so this never blocks forever.
|
||||
let _ = mon.rx.recv();
|
||||
@@ -396,6 +475,8 @@ enum Sys<G: GenServer> {
|
||||
/// the payload factory, dispatches it to [`GenServer::handle_timer`], and
|
||||
/// re-arms the next tick before returning.
|
||||
Tick(crate::timer::TimerId),
|
||||
/// The state asked to end the server (via [`StopHandle::stop`]).
|
||||
Stop,
|
||||
}
|
||||
|
||||
/// The server loop's runtime hook, passed to [`GenServer::init`]. Hands out the
|
||||
@@ -411,9 +492,29 @@ pub struct GenServerCtx<G: GenServer> {
|
||||
/// because `init` holds only `&ctx`; not `Send`, but `GenServerCtx` is only ever
|
||||
/// borrowed on the actor's own stack during `init`, never sent.
|
||||
idle: Cell<Option<Duration>>,
|
||||
/// Whether the loop should trap exits (set via [`trap_exit`](Self::trap_exit)
|
||||
/// during `init`, read by the loop after).
|
||||
trap: Cell<bool>,
|
||||
}
|
||||
|
||||
impl<G: GenServer> GenServerCtx<G> {
|
||||
/// Trap exits for the server's lifetime: shutdown requests then arrive as
|
||||
/// [`GenServer::handle_shutdown`] and linked-peer deaths as
|
||||
/// [`GenServer::handle_exit`], instead of stopping the server outright.
|
||||
/// Call this once during `init`.
|
||||
pub fn trap_exit(&self) {
|
||||
self.trap.set(true);
|
||||
}
|
||||
|
||||
/// A clonable handle that lets the state end the server from any handler
|
||||
/// (a normal exit; see the module docs). Store it on the state during
|
||||
/// `init`.
|
||||
pub fn stop_handle(&self) -> StopHandle<G> {
|
||||
StopHandle {
|
||||
sys_tx: self.sys_tx.clone(),
|
||||
}
|
||||
}
|
||||
|
||||
/// A clonable handle to the loop's monitor intake. Store it in the state
|
||||
/// during `init` to watch monitors from later handlers.
|
||||
pub fn watcher(&self) -> Watcher<G> {
|
||||
@@ -452,6 +553,30 @@ impl<G: GenServer> GenServerCtx<G> {
|
||||
}
|
||||
}
|
||||
|
||||
/// Lets a server's state end the server, cloned from
|
||||
/// [`GenServerCtx::stop_handle`] during `init`. [`stop`](Self::stop) makes the
|
||||
/// loop break after the current message and exit normally; `terminate` runs on
|
||||
/// the normal path.
|
||||
pub struct StopHandle<G: GenServer> {
|
||||
sys_tx: Sender<Sys<G>>,
|
||||
}
|
||||
|
||||
impl<G: GenServer> Clone for StopHandle<G> {
|
||||
fn clone(&self) -> Self {
|
||||
StopHandle {
|
||||
sys_tx: self.sys_tx.clone(),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
impl<G: GenServer> StopHandle<G> {
|
||||
/// End the server after the current message. Idempotent; a no-op once the
|
||||
/// server is gone.
|
||||
pub fn stop(&self) {
|
||||
let _ = self.sys_tx.send(Sys::Stop);
|
||||
}
|
||||
}
|
||||
|
||||
/// Per-server timer bookkeeping, shared between the loop and every
|
||||
/// [`TimerHandle`] clone. A gen_server actor is single-threaded — handlers
|
||||
/// and the loop never run concurrently — so this `Mutex` is always
|
||||
@@ -704,9 +829,10 @@ impl<G: GenServer> GenServerBuilder<G> {
|
||||
self
|
||||
}
|
||||
|
||||
/// Spawn the server actor and hand back its [`GenServerRef`]. The server's
|
||||
/// lifetime is governed by its refs, not by joining, so the backing join
|
||||
/// handle is dropped.
|
||||
/// Spawn the server actor and hand back its [`GenServerRef`] (an address;
|
||||
/// the server's lifetime is its own, see the module docs). The backing
|
||||
/// join handle is dropped. For a supervised server use
|
||||
/// [`named`](Self::named) + [`NamedGenServerBuilder::run`] instead.
|
||||
pub fn start(self) -> GenServerRef<G> {
|
||||
self.spawn_server()
|
||||
}
|
||||
@@ -734,13 +860,14 @@ impl<G: GenServer> GenServerBuilder<G> {
|
||||
supervisor,
|
||||
stack_opts,
|
||||
} = self;
|
||||
let keep = tx.clone();
|
||||
let handle = match supervisor {
|
||||
Some(sup) => crate::scheduler::spawn_under_with(sup, stack_opts, move || {
|
||||
server_loop::<G>(rx, state, infos)
|
||||
server_loop::<G>(keep, rx, state, infos)
|
||||
}),
|
||||
None => crate::scheduler::spawn_with(stack_opts, move || {
|
||||
server_loop::<G>(keep, rx, state, infos)
|
||||
}),
|
||||
None => {
|
||||
crate::scheduler::spawn_with(stack_opts, move || server_loop::<G>(rx, state, infos))
|
||||
}
|
||||
};
|
||||
GenServerRef {
|
||||
tx,
|
||||
@@ -828,19 +955,38 @@ impl<G: GenServer> NamedGenServerBuilder<G> {
|
||||
/// The inbox sender is published under the name **from the parent side,
|
||||
/// before this returns**, so a by-name `call` / `cast` resolves the instant
|
||||
/// `start()` returns — no race with the server body. On a name clash the
|
||||
/// just-spawned server is wound down (its only ref is dropped, closing the
|
||||
/// inbox), so a failed bind leaks no actor.
|
||||
/// just-spawned server is stopped, so a failed bind leaks no actor.
|
||||
pub fn start(self) -> Result<GenServerRef<G>, RegisterError> {
|
||||
let NamedGenServerBuilder { builder, name } = self;
|
||||
let server = builder.spawn_server();
|
||||
match register_with::<Envelope<G>>(server.pid, name, server.tx.clone()) {
|
||||
Ok(()) => Ok(server),
|
||||
Err(e) => {
|
||||
drop(server); // inbox closes → loop exits gracefully
|
||||
crate::scheduler::request_stop(server.pid); // never ran init
|
||||
Err(e)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Run the server **inline, as the current actor**, bound to its name.
|
||||
/// This is the supervised shape: the closure of a
|
||||
/// [`ChildSpec`](crate::supervisor::ChildSpec) *is* the server, so the
|
||||
/// supervisor's shutdown reaches it as [`GenServer::handle_shutdown`], a
|
||||
/// restart runs the factory again and re-binds the name, and clients
|
||||
/// address it by name ([`call`], [`cast`], [`whereis_server`]). Returns
|
||||
/// when the server exits; [`RegisterError::NameTaken`] (before `init`) if
|
||||
/// the name is held by a different live actor.
|
||||
///
|
||||
/// `under` / `stack_opts` are spawn options and do not apply here — the
|
||||
/// actor already exists.
|
||||
pub fn run(self) -> Result<(), RegisterError> {
|
||||
let NamedGenServerBuilder { builder, name } = self;
|
||||
let GenServerBuilder { state, infos, .. } = builder;
|
||||
let (tx, rx) = channel::<Envelope<G>>();
|
||||
register_with::<Envelope<G>>(crate::scheduler::self_pid(), name, tx.clone())?;
|
||||
server_loop::<G>(tx, rx, state, infos);
|
||||
Ok(())
|
||||
}
|
||||
}
|
||||
|
||||
/// Resolve a [`GenServerName`] to a [`GenServerRef`] when you want a handle to hold or
|
||||
@@ -899,10 +1045,17 @@ pub fn start_under<G: GenServer>(supervisor: Pid, state: G) -> GenServerRef<G> {
|
||||
}
|
||||
|
||||
fn server_loop<G: GenServer>(
|
||||
keep: Sender<Envelope<G>>,
|
||||
rx: Receiver<Envelope<G>>,
|
||||
state: G,
|
||||
mut infos: Vec<Receiver<G::Info>>,
|
||||
) {
|
||||
// The loop holds one inbox sender for its whole life: the inbox never
|
||||
// closes, so refs are addresses and the server's lifetime is the actor's
|
||||
// (stop handle, shutdown, stop, panic). The `Disconnected` arms below are
|
||||
// defensive only.
|
||||
let _keep = keep;
|
||||
|
||||
// Drop guard — owns the server state and the timer registry.
|
||||
//
|
||||
// Why a guard rather than code after the loop:
|
||||
@@ -974,9 +1127,14 @@ fn server_loop<G: GenServer>(
|
||||
sys_tx,
|
||||
reg: reg.clone(),
|
||||
idle: Cell::new(None),
|
||||
trap: Cell::new(false),
|
||||
};
|
||||
guard.0.init(&ctx);
|
||||
let idle = ctx.idle.get();
|
||||
// Trapping is opted into during init and fixed for the loop's life. The
|
||||
// inbox is armed only when set: an untrapped server keeps the fast path,
|
||||
// and a shutdown request simply stops it as `request_stop` would.
|
||||
let exits: Option<Receiver<ExitSignal>> = ctx.trap.get().then(crate::link::trap_exit);
|
||||
drop(ctx);
|
||||
|
||||
let mut monitors: Vec<Monitor> = Vec::new();
|
||||
@@ -992,7 +1150,7 @@ fn server_loop<G: GenServer>(
|
||||
};
|
||||
|
||||
loop {
|
||||
if monitors.is_empty() && !sys_open && infos.is_empty() {
|
||||
if exits.is_none() && monitors.is_empty() && !sys_open && infos.is_empty() {
|
||||
// Fast path: no extra arms, no select overhead — park directly on
|
||||
// the inbox. Mirrors the inbox arm of the select path below; any
|
||||
// change there must be applied here too.
|
||||
@@ -1009,7 +1167,7 @@ fn server_loop<G: GenServer>(
|
||||
guard.0.handle_idle();
|
||||
reset_idle(&mut idle_deadline);
|
||||
}
|
||||
// All ServerRefs dropped → inbox closed → shutdown.
|
||||
// Defensive: the loop holds a sender, so unreachable.
|
||||
Err(RecvTimeoutError::Disconnected) => break,
|
||||
}
|
||||
}
|
||||
@@ -1020,16 +1178,21 @@ fn server_loop<G: GenServer>(
|
||||
}
|
||||
} else {
|
||||
// Slow path: one or more extra arms live — build the arm slice and
|
||||
// select. Arm order encodes priority: downs → system → infos →
|
||||
// inbox. The slice is rebuilt each iteration because the monitor
|
||||
// and info sets shrink/grow. Mirrors the fast-path inbox park
|
||||
// above; keep them in sync.
|
||||
let nd = monitors.len(); // monitor band: [0, nd)
|
||||
let nw = sys_open as usize; // system arm: [nd, nd+nw)
|
||||
// info band: [nd+nw, nd+nw+ni)
|
||||
// select. Arm order encodes priority: exits → downs → system →
|
||||
// infos → inbox (a shutdown request is noticed under any load).
|
||||
// The slice is rebuilt each iteration because the monitor and info
|
||||
// sets shrink/grow. Mirrors the fast-path inbox park above; keep
|
||||
// them in sync.
|
||||
let ne = exits.is_some() as usize; // exit arm: [0, ne)
|
||||
let nd = ne + monitors.len(); // monitor band: [ne, nd)
|
||||
let nw = sys_open as usize; // system arm: [nd, nd+nw)
|
||||
// info band: [nd+nw, nd+nw+ni)
|
||||
// inbox arm: [nd+nw+ni]
|
||||
let sel = {
|
||||
let mut arms: Vec<&dyn Selectable> = Vec::with_capacity(nd + nw + infos.len() + 1);
|
||||
if let Some(e) = &exits {
|
||||
arms.push(e);
|
||||
}
|
||||
for m in &monitors {
|
||||
arms.push(&m.rx);
|
||||
}
|
||||
@@ -1055,10 +1218,25 @@ fn server_loop<G: GenServer>(
|
||||
continue;
|
||||
}
|
||||
};
|
||||
if i < nd {
|
||||
if i < ne {
|
||||
// Exit arm: a shutdown request or a linked peer's death.
|
||||
// The inbox lives for the loop's life, so it never closes.
|
||||
let sig = exits.as_ref().and_then(|e| e.try_recv().ok().flatten());
|
||||
if let Some(sig) = sig {
|
||||
if sig.reason == DownReason::Shutdown {
|
||||
match guard.0.handle_shutdown() {
|
||||
ShutdownAction::Exit => break,
|
||||
ShutdownAction::Continue => {}
|
||||
}
|
||||
} else {
|
||||
guard.0.handle_exit(sig);
|
||||
}
|
||||
reset_idle(&mut idle_deadline);
|
||||
}
|
||||
} else if i < nd {
|
||||
// Monitor band: a Down retires its arm either way (one-shot)
|
||||
// or closes without delivering (defensive; shouldn't happen).
|
||||
let m = monitors.remove(i);
|
||||
let m = monitors.remove(i - ne);
|
||||
if let Ok(Some(down)) = m.rx.try_recv() {
|
||||
guard.0.handle_down(down);
|
||||
reset_idle(&mut idle_deadline);
|
||||
@@ -1067,6 +1245,8 @@ fn server_loop<G: GenServer>(
|
||||
match sys_rx.try_recv() {
|
||||
// Control intake, not a dispatched message: no idle reset.
|
||||
Ok(Some(Sys::Watch(m))) => monitors.push(m),
|
||||
// The state ended the server: a normal exit.
|
||||
Ok(Some(Sys::Stop)) => break,
|
||||
Ok(Some(Sys::Timer(id, msg))) => {
|
||||
// The one-shot fired: retire its registry entry so the
|
||||
// live set tracks only still-pending timers, then
|
||||
|
||||
+469
-66
@@ -68,10 +68,38 @@
|
||||
//! inbox or timer event; a replayed event may postpone again (it re-queues for
|
||||
//! the next transition). See the macro docs for the row surface and [`Step`]
|
||||
//! for how a postpone surfaces to the loop.
|
||||
//!
|
||||
//! ## Stopping, shutdown, and exits
|
||||
//!
|
||||
//! A machine lives until it stops, is shut down, or is killed; a
|
||||
//! [`GenStatemRef`] is an address, and dropping refs never ends it (the
|
||||
//! gen_server rule — see its *When the server stops*). The supervised shape
|
||||
//! is [`run_named`], which runs the machine inline as the current actor so it
|
||||
//! is a direct `ChildSpec` child addressed by [`GenStatemName`]. A machine
|
||||
//! can end itself: any row
|
||||
//! body may call [`cx.stop()`](Cx::stop) (or use the `stop` tail keyword,
|
||||
//! sugar for `{ cx.stop(); prev }`) — the loop breaks after that event and the
|
||||
//! actor exits *normally* (OTP's `{stop, normal}`; a `Transient` child is not
|
||||
//! restarted). [`terminate`](Machine::terminate) — the optional `terminate
|
||||
//! { … }` macro block — runs on every exit path.
|
||||
//!
|
||||
//! From outside, [`GenStatemRef::shutdown`] (or a plain
|
||||
//! [`request_shutdown`](crate::request_shutdown), which is what a supervisor
|
||||
//! sends) asks the machine to stop. By default a machine does not trap exits
|
||||
//! and the request stops it outright. A machine that calls
|
||||
//! [`cx.trap_exit()`](Cx::trap_exit) in its initial `enter` instead receives
|
||||
//! it as the **`shutdown` event**, routed by state like any other — a
|
||||
//! `Connected` state may transition into `Draining` and stop later from a
|
||||
//! timeout row, while a state with no `shutdown` row takes the macro's
|
||||
//! default, `stop`. Linked-peer deaths reach a trapping machine as
|
||||
//! `exit <pat>` events; an unmatched one is dropped like an unmatched info.
|
||||
|
||||
use crate::channel::{channel, select, Receiver, Sender};
|
||||
use crate::channel::{channel, select, Receiver, Selectable, Sender};
|
||||
use crate::link::ExitSignal;
|
||||
use crate::monitor::{monitor, DownReason};
|
||||
use crate::pid::Pid;
|
||||
use crate::scheduler::{cancel_timer, send_after_to};
|
||||
use crate::registry::{register_with, resolve_named_sender, RegisterError};
|
||||
use crate::scheduler::{cancel_timer, request_shutdown, send_after_to};
|
||||
use crate::timer::TimerId;
|
||||
use std::collections::{HashMap, VecDeque};
|
||||
use std::marker::PhantomData;
|
||||
@@ -119,6 +147,33 @@ pub trait Machine: Send + 'static {
|
||||
/// state's `enter`, and returns [`Step::Transitioned`]; a stay or unmatched
|
||||
/// event returns [`Step::Stayed`].
|
||||
fn handle(&mut self, ev: Self::Ev, cx: &mut Cx<Self::Ev>) -> Step<Self::Ev>;
|
||||
|
||||
/// Wrap a graceful shutdown request into this machine's event, so a
|
||||
/// trapping machine (see [`Cx::trap_exit`]) can route it **by state**. The
|
||||
/// macro generates it as `Ev::Shutdown` and matches it in `shutdown` rows;
|
||||
/// its default for a state that writes no such row is `stop`. A
|
||||
/// hand-written machine that returns `None` (the default) is simply
|
||||
/// stopped — the loop breaks and `terminate` runs on the normal path.
|
||||
fn shutdown_ev() -> Option<Self::Ev> {
|
||||
None
|
||||
}
|
||||
|
||||
/// Wrap a linked peer's death (an [`ExitSignal`] that is not a shutdown
|
||||
/// request, delivered only when trapping) into this machine's event. The
|
||||
/// macro generates it as `Ev::Exit(sig)` and matches it in `exit <pat>`
|
||||
/// rows; an unmatched exit is silently dropped, like an unmatched info.
|
||||
/// A hand-written machine that returns `None` (the default) drops it.
|
||||
fn exit_ev(_sig: ExitSignal) -> Option<Self::Ev> {
|
||||
None
|
||||
}
|
||||
|
||||
/// Runs as the machine actor exits, on any exit path (a `stop`, a
|
||||
/// graceful shutdown, a handler panic, a hard stop). Like
|
||||
/// `gen_server::terminate`: on the panic and hard-stop paths it runs
|
||||
/// mid-unwind — do not panic or park there; only on the normal path
|
||||
/// (`stop`, shutdown rows) may it do real work. The macro's optional
|
||||
/// `terminate { … }` block generates it.
|
||||
fn terminate(&mut self) {}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
@@ -247,6 +302,11 @@ impl Timers {
|
||||
pub struct Cx<Ev> {
|
||||
sys_tx: Sender<Sys>,
|
||||
reg: Arc<Mutex<Timers>>,
|
||||
/// Set by [`trap_exit`](Self::trap_exit) during `on_start`; read once by
|
||||
/// the loop right after, fixed for the machine's life.
|
||||
trap: bool,
|
||||
/// Set by [`stop`](Self::stop); the loop breaks after the current event.
|
||||
stop: bool,
|
||||
_ev: PhantomData<fn() -> Ev>,
|
||||
}
|
||||
|
||||
@@ -255,10 +315,30 @@ impl<Ev> Cx<Ev> {
|
||||
Cx {
|
||||
sys_tx,
|
||||
reg,
|
||||
trap: false,
|
||||
stop: false,
|
||||
_ev: PhantomData,
|
||||
}
|
||||
}
|
||||
|
||||
/// Trap exits for the machine's lifetime: a shutdown request then arrives
|
||||
/// as the `shutdown` event (routed by state) and linked-peer deaths as
|
||||
/// `exit` events, instead of stopping the machine outright. Call it in the
|
||||
/// initial state's `enter` (i.e. during `on_start`); later calls have no
|
||||
/// effect.
|
||||
pub fn trap_exit(&mut self) {
|
||||
self.trap = true;
|
||||
}
|
||||
|
||||
/// End the machine after the current event: the loop breaks and the actor
|
||||
/// exits *normally* (OTP's `{stop, normal}`); `terminate` runs on the
|
||||
/// normal path. Anything still queued or postponed is dropped. The `stop`
|
||||
/// tail keyword in a macro row is sugar for `{ cx.stop(); prev }`.
|
||||
/// Mirrors gen_server's [`StopHandle`](crate::gen_server::StopHandle).
|
||||
pub fn stop(&mut self) {
|
||||
self.stop = true;
|
||||
}
|
||||
|
||||
/// Arm the **state timeout**: fire a `state_timeout` event after `after` in
|
||||
/// the current state. Auto-reset on any state change (the loop cancels and
|
||||
/// clears it on every real transition), so it measures quiet time *within* a
|
||||
@@ -386,8 +466,7 @@ pub enum SendError {
|
||||
}
|
||||
|
||||
/// A clonable handle to a running machine. Cloning yields another sender to the
|
||||
/// same inbox; the machine lives until the last `GenStatemRef` is dropped, at which
|
||||
/// point its inbox closes and the loop exits.
|
||||
/// same inbox. An address, not an owner: dropping refs never ends the machine.
|
||||
pub struct GenStatemRef<M: Machine> {
|
||||
tx: Sender<M::Ev>,
|
||||
pid: Pid,
|
||||
@@ -433,6 +512,19 @@ impl<M: Machine> GenStatemRef<M> {
|
||||
self.send(ev).map_err(|_| CallError::Down)?;
|
||||
rx.recv().map_err(|_| CallError::Down)
|
||||
}
|
||||
|
||||
/// Ask the machine to shut down and block until it has fully exited, so
|
||||
/// [`Machine::terminate`] has run by the time this returns. A trapping
|
||||
/// machine winds down through its `shutdown` rows; any other is stopped
|
||||
/// outright. Returns immediately if the machine is already gone. Waits as
|
||||
/// long as the machine takes — a supervisor bounds that with its child's
|
||||
/// [`Shutdown`](crate::supervisor::Shutdown) policy. Mirrors
|
||||
/// [`GenServerRef::shutdown`](crate::gen_server::GenServerRef::shutdown).
|
||||
pub fn shutdown(&self) {
|
||||
let mon = monitor(self.pid);
|
||||
request_shutdown(self.pid);
|
||||
let _ = mon.rx.recv();
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
@@ -441,7 +533,8 @@ impl<M: Machine> GenStatemRef<M> {
|
||||
|
||||
/// Spawn `machine` as an actor and hand back its [`GenStatemRef`]. Shape mirrors
|
||||
/// `gen_server::start`: make the inbox, spawn the loop, return the ref; the
|
||||
/// backing join handle is dropped (lifetime is governed by refs, not joining).
|
||||
/// backing join handle is dropped (the machine's lifetime is its own). For a
|
||||
/// supervised machine use [`run_named`].
|
||||
///
|
||||
/// Panics if called outside `Runtime::run()`.
|
||||
pub fn spawn<M: Machine>(machine: M) -> GenStatemRef<M> {
|
||||
@@ -456,40 +549,188 @@ pub fn spawn<M: Machine>(machine: M) -> GenStatemRef<M> {
|
||||
/// Panics if called outside `Runtime::run()`.
|
||||
pub fn spawn_with<M: Machine>(opts: crate::scheduler::SpawnOpts, machine: M) -> GenStatemRef<M> {
|
||||
let (tx, rx) = channel::<M::Ev>();
|
||||
let handle = crate::scheduler::spawn_with(opts, move || statem_loop(rx, machine));
|
||||
let keep = tx.clone();
|
||||
let handle = crate::scheduler::spawn_with(opts, move || statem_loop(keep, rx, machine));
|
||||
GenStatemRef {
|
||||
tx,
|
||||
pid: handle.pid(),
|
||||
}
|
||||
}
|
||||
|
||||
/// A typed, static name for a gen_statem, used to address a machine through
|
||||
/// the registry without holding a [`GenStatemRef`]. Mirrors
|
||||
/// [`GenServerName`](crate::gen_server::GenServerName): declare it as a
|
||||
/// constant and bind it with [`run_named`].
|
||||
pub struct GenStatemName<M> {
|
||||
name: &'static str,
|
||||
_marker: PhantomData<fn() -> M>,
|
||||
}
|
||||
|
||||
impl<M> GenStatemName<M> {
|
||||
/// Bind a static string as a machine name.
|
||||
#[inline]
|
||||
pub const fn new(name: &'static str) -> Self {
|
||||
Self {
|
||||
name,
|
||||
_marker: PhantomData,
|
||||
}
|
||||
}
|
||||
|
||||
/// The underlying registry key.
|
||||
#[inline]
|
||||
pub const fn as_str(self) -> &'static str {
|
||||
self.name
|
||||
}
|
||||
}
|
||||
|
||||
impl<M> Copy for GenStatemName<M> {}
|
||||
impl<M> Clone for GenStatemName<M> {
|
||||
fn clone(&self) -> Self {
|
||||
*self
|
||||
}
|
||||
}
|
||||
|
||||
/// Run `machine` **inline, as the current actor**, bound to `name`. The
|
||||
/// supervised shape: the closure of a
|
||||
/// [`ChildSpec`](crate::supervisor::ChildSpec) *is* the machine, so the
|
||||
/// supervisor's shutdown reaches it as a `shutdown` row, a restart runs the
|
||||
/// factory again and re-binds the name, and clients address it by name
|
||||
/// ([`send`], [`call`], [`whereis_machine`]). Returns when the machine exits;
|
||||
/// [`RegisterError::NameTaken`] (before `on_start`) if the name is held by a
|
||||
/// different live actor. Mirrors
|
||||
/// [`NamedGenServerBuilder::run`](crate::gen_server::NamedGenServerBuilder::run).
|
||||
pub fn run_named<M: Machine>(name: GenStatemName<M>, machine: M) -> Result<(), RegisterError> {
|
||||
let (tx, rx) = channel::<M::Ev>();
|
||||
register_with::<M::Ev>(crate::scheduler::self_pid(), name.as_str(), tx.clone())?;
|
||||
statem_loop(tx, rx, machine);
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Resolve a [`GenStatemName`] to a [`GenStatemRef`]; `None` if no live
|
||||
/// machine holds the name. Panics if called outside `Runtime::run()`.
|
||||
pub fn whereis_machine<M: Machine>(name: GenStatemName<M>) -> Option<GenStatemRef<M>> {
|
||||
resolve_named_sender::<M::Ev>(name.as_str()).map(|(pid, tx)| GenStatemRef { tx, pid })
|
||||
}
|
||||
|
||||
/// Push an event to the machine registered under `name`, resolving per send.
|
||||
/// [`SendError::Down`] if no live machine holds the name.
|
||||
pub fn send<M: Machine>(name: GenStatemName<M>, ev: M::Ev) -> Result<(), SendError> {
|
||||
match whereis_machine(name) {
|
||||
Some(m) => m.send(ev),
|
||||
None => Err(SendError::Down),
|
||||
}
|
||||
}
|
||||
|
||||
/// Synchronous request-reply to the machine registered under `name`,
|
||||
/// resolving per call (a machine restarted under the same name is reached
|
||||
/// transparently). [`CallError::Down`] if no live machine holds the name.
|
||||
pub fn call<M, T, F>(name: GenStatemName<M>, make: F) -> Result<T, CallError>
|
||||
where
|
||||
M: Machine,
|
||||
T: Send + 'static,
|
||||
F: FnOnce(Reply<T>) -> M::Ev,
|
||||
{
|
||||
match whereis_machine(name) {
|
||||
Some(m) => m.call(make),
|
||||
None => Err(CallError::Down),
|
||||
}
|
||||
}
|
||||
|
||||
/// Shut down the machine registered under `name` and wait for it (see
|
||||
/// [`GenStatemRef::shutdown`]). A no-op if no live machine holds the name.
|
||||
pub fn shutdown<M: Machine>(name: GenStatemName<M>) {
|
||||
if let Some(m) = whereis_machine(name) {
|
||||
m.shutdown();
|
||||
}
|
||||
}
|
||||
|
||||
/// The machine actor body: `on_start`, then one `handle` per event until the
|
||||
/// inbox closes (all refs dropped → graceful shutdown).
|
||||
/// row resolves to `stop`, a shutdown row stops it, or the actor is stopped
|
||||
/// from outside.
|
||||
///
|
||||
/// Two intake sources are selected each iteration with the **timer arm above
|
||||
/// the inbox**, so a timeout fire is never starved by inbox traffic: `sys_rx`
|
||||
/// carries timer fires armed through `cx`, `rx` is the user inbox. A fire is
|
||||
/// turned into the matching internal event (`state_timeout` / `timeout(name)`)
|
||||
/// and run through the same `handle` dispatch as an inbox event — the
|
||||
/// gen_statem model, where timeouts surface as ordinary events.
|
||||
/// Intake arms are selected each iteration in priority order — **exits**
|
||||
/// (only when trapping) above **timers** above the **inbox** — so a shutdown
|
||||
/// request or a timeout fire is never starved by inbox traffic. `sys_rx`
|
||||
/// carries timer fires armed through `cx`; a fire is turned into the matching
|
||||
/// internal event (`state_timeout` / `timeout(name)`) and run through the same
|
||||
/// `handle` dispatch as an inbox event — the gen_statem model, where timeouts
|
||||
/// (and, when trapping, shutdown and exits) surface as ordinary events.
|
||||
///
|
||||
/// The loop owns the **postpone queue**: a `handle` that defers its event hands
|
||||
/// it back ([`Step::Postponed`]) for the queue; a `handle` that transitions
|
||||
/// ([`Step::Transitioned`]) triggers a [`replay`] of the queue in the new
|
||||
/// state, ahead of the next intake.
|
||||
fn statem_loop<M: Machine>(rx: Receiver<M::Ev>, mut machine: M) {
|
||||
fn statem_loop<M: Machine>(keep: Sender<M::Ev>, rx: Receiver<M::Ev>, machine: M) {
|
||||
// One inbox sender lives with the loop: the inbox never closes, refs are
|
||||
// addresses, the machine's lifetime is the actor's (stop row, shutdown,
|
||||
// stop, panic). The `Disconnected` inbox arm below is defensive only.
|
||||
let _keep = keep;
|
||||
// Drop guard — owns the machine and the timer registry, so `terminate`
|
||||
// fires on every exit path (clean close, `stop`, a handler panic, a hard
|
||||
// stop) and the timer drain is sequenced before it. Same shape as
|
||||
// gen_server's guard; see the rationale there.
|
||||
struct Terminate<M: Machine>(M, Arc<Mutex<Timers>>);
|
||||
impl<M: Machine> Drop for Terminate<M> {
|
||||
fn drop(&mut self) {
|
||||
{
|
||||
let mut reg = match self.1.lock() {
|
||||
Ok(g) => g,
|
||||
Err(e) => panic!("smarm: gen_statem reg lock poisoned (core corrupt): {e}"),
|
||||
};
|
||||
if let Some((_, sub)) = reg.state.take() {
|
||||
cancel_timer(sub);
|
||||
}
|
||||
for (_, (_, sub)) in reg.named.drain() {
|
||||
cancel_timer(sub);
|
||||
}
|
||||
}
|
||||
self.0.terminate();
|
||||
}
|
||||
}
|
||||
|
||||
let (sys_tx, sys_rx) = channel::<Sys>();
|
||||
let reg = Arc::new(Mutex::new(Timers::new()));
|
||||
let mut guard = Terminate(machine, reg.clone());
|
||||
// The loop owns `cx` (and through it a `sys_tx` clone) for its whole life,
|
||||
// so the sys arm never closes from under us — no auto-close dance needed.
|
||||
let mut cx = Cx::new(sys_tx, reg.clone());
|
||||
// Events deferred by `postpone` rows, replayed FIFO on the next transition.
|
||||
let mut postpone: VecDeque<M::Ev> = VecDeque::new();
|
||||
machine.on_start(&mut cx);
|
||||
guard.0.on_start(&mut cx);
|
||||
// Trapping is opted into during on_start and fixed for the loop's life.
|
||||
// The inbox is armed only when set: an untrapped machine keeps the
|
||||
// two-arm select, and a shutdown request simply stops it as
|
||||
// `request_stop` would.
|
||||
let exits: Option<Receiver<ExitSignal>> = cx.trap.then(crate::link::trap_exit);
|
||||
loop {
|
||||
// Timer arm first: a ready fire is taken in preference to the inbox.
|
||||
let i = select(&[&sys_rx, &rx]);
|
||||
if i == 0 {
|
||||
// Arm order encodes priority: exits → timers → inbox.
|
||||
let ne = exits.is_some() as usize;
|
||||
let i = {
|
||||
let mut arms: Vec<&dyn Selectable> = Vec::with_capacity(3);
|
||||
if let Some(e) = &exits {
|
||||
arms.push(e);
|
||||
}
|
||||
arms.push(&sys_rx);
|
||||
arms.push(&rx);
|
||||
select(&arms)
|
||||
};
|
||||
if i < ne {
|
||||
// Exit arm: a shutdown request or a linked peer's death. The trap
|
||||
// inbox lives for the loop's life, so it never closes.
|
||||
let sig = exits.as_ref().and_then(|e| e.try_recv().ok().flatten());
|
||||
match sig {
|
||||
Some(sig) if sig.reason == DownReason::Shutdown => match M::shutdown_ev() {
|
||||
Some(ev) => dispatch(&mut guard.0, &mut cx, &mut postpone, ev),
|
||||
None => cx.stop(),
|
||||
},
|
||||
Some(sig) => {
|
||||
if let Some(ev) = M::exit_ev(sig) {
|
||||
dispatch(&mut guard.0, &mut cx, &mut postpone, ev);
|
||||
}
|
||||
}
|
||||
None => {}
|
||||
}
|
||||
} else if i == ne {
|
||||
match sys_rx.try_recv() {
|
||||
Ok(Some(fire)) => {
|
||||
// Confirm the fire is still the live one before dispatching:
|
||||
@@ -528,7 +769,7 @@ fn statem_loop<M: Machine>(rx: Receiver<M::Ev>, mut machine: M) {
|
||||
}
|
||||
};
|
||||
if let Some(ev) = ev {
|
||||
dispatch(&mut machine, &mut cx, &mut postpone, ev);
|
||||
dispatch(&mut guard.0, &mut cx, &mut postpone, ev);
|
||||
}
|
||||
}
|
||||
// Single-receiver: nothing can drain the arm between select's
|
||||
@@ -540,12 +781,16 @@ fn statem_loop<M: Machine>(rx: Receiver<M::Ev>, mut machine: M) {
|
||||
}
|
||||
} else {
|
||||
match rx.try_recv() {
|
||||
Ok(Some(ev)) => dispatch(&mut machine, &mut cx, &mut postpone, ev),
|
||||
Ok(Some(ev)) => dispatch(&mut guard.0, &mut cx, &mut postpone, ev),
|
||||
Ok(None) => debug_assert!(false, "ready inbox was empty"),
|
||||
// All GenStatemRefs dropped → inbox closed → shutdown.
|
||||
// Defensive: the loop holds a sender, so unreachable.
|
||||
Err(_) => break,
|
||||
}
|
||||
}
|
||||
// A handler (or a replay) asked to stop: a normal exit.
|
||||
if cx.stop {
|
||||
break;
|
||||
}
|
||||
// Observation point so a machine fed a hot inbox stays preemptible and
|
||||
// cancellable.
|
||||
crate::check!();
|
||||
@@ -554,7 +799,8 @@ fn statem_loop<M: Machine>(rx: Receiver<M::Ev>, mut machine: M) {
|
||||
|
||||
/// Run one event through `handle` and act on its [`Step`]: stash a deferred
|
||||
/// event on the postpone queue, or — on a real transition — [`replay`] the
|
||||
/// queue in the new state. A stay/unmatched event needs nothing further.
|
||||
/// queue in the new state. A stay/unmatched event needs nothing further. A
|
||||
/// [`Cx::stop`] raised by the handler skips the replay; the loop breaks next.
|
||||
fn dispatch<M: Machine>(
|
||||
machine: &mut M,
|
||||
cx: &mut Cx<M::Ev>,
|
||||
@@ -564,14 +810,19 @@ fn dispatch<M: Machine>(
|
||||
match machine.handle(ev, cx) {
|
||||
Step::Postponed(ev) => postpone.push_back(ev),
|
||||
Step::Stayed => {}
|
||||
Step::Transitioned => replay(machine, cx, postpone),
|
||||
Step::Transitioned => {
|
||||
if !cx.stop {
|
||||
replay(machine, cx, postpone)
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Replay deferred events after a real transition: each goes back through
|
||||
/// `handle` in FIFO order, in the now-current state. An event that postpones
|
||||
/// again re-queues (to wait for the *next* transition); one that transitions
|
||||
/// re-arms the replay, so a later state can in turn drain what is still pending.
|
||||
/// re-arms the replay, so a later state can in turn drain what is still pending;
|
||||
/// one that raises [`Cx::stop`] ends the replay (and the machine) at once.
|
||||
/// Subsequent events in a batch already see the post-transition state, since
|
||||
/// `handle` reads the live state cell — the outer loop only re-runs to give
|
||||
/// re-queued events another pass once a transition has occurred within a batch.
|
||||
@@ -590,6 +841,9 @@ fn replay<M: Machine>(machine: &mut M, cx: &mut Cx<M::Ev>, postpone: &mut VecDeq
|
||||
Step::Stayed => {}
|
||||
Step::Transitioned => transitioned = true,
|
||||
}
|
||||
if cx.stop {
|
||||
return;
|
||||
}
|
||||
}
|
||||
if !transitioned {
|
||||
return;
|
||||
@@ -691,8 +945,9 @@ fn replay<M: Machine>(machine: &mut M, cx: &mut Cx<M::Ev>, postpone: &mut VecDeq
|
||||
/// // The transition table. Group rows by current state with `on <pat>`.
|
||||
/// // A row is: <kind> <event-pattern> [if <guard>] => <tail> ,
|
||||
/// // where <kind> is one of `cast`, `call`, `info`, `state_timeout`
|
||||
/// // (no pattern — it is a unit event), or `timeout <name-pattern>`,
|
||||
/// // and the tail is one of:
|
||||
/// // (no pattern — it is a unit event), `timeout <name-pattern>`,
|
||||
/// // `shutdown` (unit; trapping machines only) or `exit <sig-pattern>`
|
||||
/// // (trapping only), and the tail is one of:
|
||||
/// // * a state tag `Door::Closed` (transition, or "stay"
|
||||
/// // if it equals current)
|
||||
/// // * a block ending in one `{ data.enters += 1; Door::Closed }`
|
||||
@@ -701,6 +956,9 @@ fn replay<M: Machine>(machine: &mut M, cx: &mut Cx<M::Ev>, postpone: &mut VecDeq
|
||||
/// // * the keyword `postpone` (defer until next
|
||||
/// // transition; cast/call/
|
||||
/// // info only)
|
||||
/// // * the keyword `stop` (end the machine
|
||||
/// // normally; sugar for
|
||||
/// // `{ cx.stop(); prev }`)
|
||||
/// on Door::Open => {
|
||||
/// cast DoorCast::Push => Door::Closed,
|
||||
/// // An armed state-timeout surfaces as an ordinary event:
|
||||
@@ -724,6 +982,9 @@ fn replay<M: Machine>(machine: &mut M, cx: &mut Cx<M::Ev>, postpone: &mut VecDeq
|
||||
/// on _ => {
|
||||
/// call DoorCall::GetState(r) => { r.reply(prev); prev },
|
||||
/// }
|
||||
///
|
||||
/// // Optional: runs as the machine exits, on every exit path.
|
||||
/// terminate { data.enters = 0; }
|
||||
/// }
|
||||
/// ```
|
||||
///
|
||||
@@ -748,6 +1009,12 @@ fn replay<M: Machine>(machine: &mut M, cx: &mut Cx<M::Ev>, postpone: &mut VecDeq
|
||||
/// ignores info writes no `info` rows at all. State-timeouts and named
|
||||
/// timeouts have **no** such default: a state that can see one must handle it
|
||||
/// (or `unhandled` it) or the match is non-exhaustive.
|
||||
/// * **`shutdown` defaults to `stop`, `exit` to a silent drop.** Both reach a
|
||||
/// machine only if its initial `enter` called `cx.trap_exit()`; a
|
||||
/// non-trapping machine is stopped outright by a shutdown request. Write
|
||||
/// `shutdown => …` rows only in the states that want to wind down first
|
||||
/// (transition into a draining state, `stop` later); write `exit sig => …`
|
||||
/// rows to react to linked-peer deaths. Neither is postponable.
|
||||
/// * **Stay** = return the current tag. The `prev` you named in `context` is
|
||||
/// bound to the pre-handler state for exactly this — handy in any-state
|
||||
/// (`on _`) rows where there is no single literal tag to write.
|
||||
@@ -773,12 +1040,12 @@ fn replay<M: Machine>(machine: &mut M, cx: &mut Cx<M::Ev>, postpone: &mut VecDeq
|
||||
/// # What it emits
|
||||
///
|
||||
/// The unified `enum $Ev` (the `Cast`/`Call`/`Info` wrappers plus the internal
|
||||
/// `StateTimeout` / `Timeout` events), `struct $Sm { state, data }`,
|
||||
/// `$Sm::start(init, data) -> GenStatemRef<$Sm>`, and the `Machine` impl:
|
||||
/// `on_start` runs the initial `enter`; `handle` is the dispatch match plus the
|
||||
/// stay/transition/unhandled apply-tail (the cell's sole writer, which also
|
||||
/// auto-resets the state-timeout on every real transition); and the `enter`
|
||||
/// dispatch.
|
||||
/// `StateTimeout` / `Timeout` / `Shutdown` / `Exit` events), `struct $Sm {
|
||||
/// state, data }`, `$Sm::start(init, data) -> GenStatemRef<$Sm>`, and the
|
||||
/// `Machine` impl: `on_start` runs the initial `enter`; `handle` is the
|
||||
/// dispatch match plus the stay/transition/unhandled apply-tail (the cell's
|
||||
/// sole writer, which also auto-resets the state-timeout on every real
|
||||
/// transition); the `enter` dispatch; and `terminate` when the block is given.
|
||||
///
|
||||
/// # Limitation
|
||||
///
|
||||
@@ -794,9 +1061,10 @@ macro_rules! gen_statem {
|
||||
context ( $data:ident , $cur:ident , $cx:ident ) ;
|
||||
enter { $( $est:pat => $ebody:expr ),+ $(,)? }
|
||||
$( on $st:pat => { $($rows:tt)* } )+
|
||||
$( terminate { $($tbody:tt)* } )?
|
||||
) => {
|
||||
/// Unified inbox payload: the user's `cast`/`call`/`info` enums folded
|
||||
/// together with the runtime's internal timeout events.
|
||||
/// together with the runtime's internal events.
|
||||
enum $Ev {
|
||||
Cast($Cast),
|
||||
Call($Call),
|
||||
@@ -807,6 +1075,12 @@ macro_rules! gen_statem {
|
||||
StateTimeout,
|
||||
/// A named timeout fired (matched in `timeout <pat>` rows).
|
||||
Timeout(&'static str),
|
||||
/// A graceful shutdown request reached this (trapping) machine
|
||||
/// (matched in `shutdown` rows). A state with no such row `stop`s.
|
||||
Shutdown,
|
||||
/// A linked peer died (trapping only; matched in `exit <pat>`
|
||||
/// rows). An unmatched exit is silently dropped.
|
||||
Exit($crate::ExitSignal),
|
||||
}
|
||||
|
||||
struct $sm {
|
||||
@@ -815,8 +1089,14 @@ macro_rules! gen_statem {
|
||||
}
|
||||
|
||||
impl $sm {
|
||||
/// The machine value, for [`gen_statem::run_named`]
|
||||
/// (`$crate::gen_statem::run_named`) or `spawn`.
|
||||
fn new(init: $State, data: $Data) -> $sm {
|
||||
$sm { state: init, data }
|
||||
}
|
||||
|
||||
fn start(init: $State, data: $Data) -> $crate::gen_statem::GenStatemRef<$sm> {
|
||||
$crate::gen_statem::spawn($sm { state: init, data })
|
||||
$crate::gen_statem::spawn($sm::new(init, data))
|
||||
}
|
||||
|
||||
#[allow(unused_variables)]
|
||||
@@ -840,11 +1120,27 @@ macro_rules! gen_statem {
|
||||
$Ev::Timeout(name)
|
||||
}
|
||||
|
||||
fn shutdown_ev() -> Option<$Ev> {
|
||||
Some($Ev::Shutdown)
|
||||
}
|
||||
|
||||
fn exit_ev(sig: $crate::ExitSignal) -> Option<$Ev> {
|
||||
Some($Ev::Exit(sig))
|
||||
}
|
||||
|
||||
fn on_start(&mut self, $cx: &mut $crate::gen_statem::Cx<$Ev>) {
|
||||
let s = self.state;
|
||||
self.enter(s, $cx);
|
||||
}
|
||||
|
||||
$(
|
||||
#[allow(unused_variables)]
|
||||
fn terminate(&mut self) {
|
||||
let $data = &mut self.data;
|
||||
$($tbody)*
|
||||
}
|
||||
)?
|
||||
|
||||
#[allow(unused_variables)]
|
||||
#[deny(unreachable_patterns)] // conflicting rows must fail even though
|
||||
// this match is external-macro-expanded
|
||||
@@ -860,7 +1156,7 @@ macro_rules! gen_statem {
|
||||
// untouched), then the consuming `match (state, event)` whose
|
||||
// value is this `Resolution`.
|
||||
let next: $crate::gen_statem::Resolution<$State> =
|
||||
$crate::gen_statem!(@arms ($Ev) ($cur, ev) [ ] [ ]
|
||||
$crate::gen_statem!(@arms ($Ev) ($cur, ev, $cx) [ ] [ ]
|
||||
$( on $st => { $($rows)* } )+);
|
||||
match next {
|
||||
$crate::gen_statem::Resolution::To(s) if s == $cur => {
|
||||
@@ -896,7 +1192,7 @@ macro_rules! gen_statem {
|
||||
// is the `Info` silent-drop — last and broadest, so per-state `info` rows
|
||||
// stay reachable; cast/call/timeouts get no fallback, so a forgotten pair is
|
||||
// still E0004).
|
||||
(@arms ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ]) => {
|
||||
(@arms ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ]) => {
|
||||
{
|
||||
// Phase 1 — postpone routing (borrow-only). A guard on a postpone
|
||||
// row runs here, against by-ref bindings, so it must not depend on
|
||||
@@ -913,140 +1209,247 @@ macro_rules! gen_statem {
|
||||
// returned for it.
|
||||
match ($ss, $se) {
|
||||
$($arms)*
|
||||
// Macro-injected defaults, last and broadest so per-state rows
|
||||
// stay reachable; a user's own catch-all row may shadow them
|
||||
// entirely, hence the allow.
|
||||
#[allow(unreachable_patterns)]
|
||||
(_, $Ev::Info(_)) => $crate::gen_statem::Resolution::Unhandled,
|
||||
#[allow(unreachable_patterns)]
|
||||
(_, $Ev::Exit(_)) => $crate::gen_statem::Resolution::Unhandled,
|
||||
#[allow(unreachable_patterns)]
|
||||
(_, $Ev::Shutdown) => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) },
|
||||
}
|
||||
}
|
||||
};
|
||||
// Open an on-block: remember its state pat, drain its rows, then continue.
|
||||
(@arms ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ]
|
||||
(@arms ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ]
|
||||
on $st:pat => { $($rows:tt)* } $($more:tt)*
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se) [ $($arms)* ] [ $($post)* ] ($st)
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx) [ $($arms)* ] [ $($post)* ] ($st)
|
||||
{ $($rows)* } { $($more)* })
|
||||
};
|
||||
|
||||
// ===== @rows: drain one on-block's rows, threading both accs =============
|
||||
// cast, explicit refusal
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ cast $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se)
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Cast($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// cast, stop (end the machine normally after this event)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ cast $ev:pat $(if $g:expr)? => stop , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Cast($ev)) $(if $g)? => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) }, ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// cast, postpone (defer the event; the replay in a later state handles it)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ cast $ev:pat $(if $g:expr)? => postpone , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se)
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Cast($ev)) $(if $g)? => unreachable!("postponed event is replayed, not dispatched here"), ]
|
||||
[ $($post)* ($st, $Ev::Cast($ev)) $(if $g)? => true, ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// cast, transition / stay / branch
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ cast $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se)
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Cast($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// call, explicit refusal
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ call $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se)
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Call($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// call, stop (end the machine normally after this event)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ call $ev:pat $(if $g:expr)? => stop , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Call($ev)) $(if $g)? => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) }, ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// call, postpone (the Reply rides inside the event onto the queue)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ call $ev:pat $(if $g:expr)? => postpone , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se)
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Call($ev)) $(if $g)? => unreachable!("postponed event is replayed, not dispatched here"), ]
|
||||
[ $($post)* ($st, $Ev::Call($ev)) $(if $g)? => true, ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// call, transition / stay / branch
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ call $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se)
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Call($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// info, explicit refusal
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ info $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se)
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Info($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// info, stop (end the machine normally after this event)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ info $ev:pat $(if $g:expr)? => stop , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Info($ev)) $(if $g)? => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) }, ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// info, postpone
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ info $ev:pat $(if $g:expr)? => postpone , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se)
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Info($ev)) $(if $g)? => unreachable!("postponed event is replayed, not dispatched here"), ]
|
||||
[ $($post)* ($st, $Ev::Info($ev)) $(if $g)? => true, ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// info, transition / stay / branch
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ info $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se)
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Info($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// state_timeout, explicit refusal (unit event — no pattern; not postponable)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ state_timeout $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se)
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::StateTimeout) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// state_timeout, stop (end the machine normally after this event)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ state_timeout $(if $g:expr)? => stop , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::StateTimeout) $(if $g)? => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) }, ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// state_timeout, transition / stay / branch
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ state_timeout $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se)
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::StateTimeout) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// timeout, explicit refusal (pattern matches the name; not postponable)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ timeout $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se)
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Timeout($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// timeout, stop (end the machine normally after this event)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ timeout $ev:pat $(if $g:expr)? => stop , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Timeout($ev)) $(if $g)? => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) }, ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// timeout, transition / stay / branch
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ timeout $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se)
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Timeout($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// shutdown, explicit refusal (unit event — no pattern; not postponable; a state with no shutdown row defaults to `stop`)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ shutdown $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Shutdown) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// shutdown, stop (end the machine normally after this event)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ shutdown $(if $g:expr)? => stop , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Shutdown) $(if $g)? => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) }, ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// shutdown, transition / stay / branch
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ shutdown $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Shutdown) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// exit, explicit refusal (pattern matches the ExitSignal; not postponable)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ exit $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Exit($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// exit, stop (end the machine normally after this event)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ exit $ev:pat $(if $g:expr)? => stop , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Exit($ev)) $(if $g)? => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) }, ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// exit, transition / stay / branch
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ exit $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Exit($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// this block is drained: hand the remaining on-blocks back to @arms
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@arms ($Ev) ($ss, $se) [ $($arms)* ] [ $($post)* ] $($more)*)
|
||||
$crate::gen_statem!(@arms ($Ev) ($ss, $se, $cx) [ $($arms)* ] [ $($post)* ] $($more)*)
|
||||
};
|
||||
}
|
||||
|
||||
+9
-9
@@ -60,10 +60,10 @@ pub use channel::{
|
||||
pub use gen_server::{
|
||||
call, cast, shutdown, whereis_server, CallError, CallTimeoutError, CastError, GenServer,
|
||||
GenServerBuilder, GenServerCtx, GenServerName, GenServerRef, NamedGenServerBuilder,
|
||||
TimerHandle, Watcher,
|
||||
ShutdownAction, StopHandle, TimerHandle, Watcher,
|
||||
};
|
||||
pub use gen_statem::{
|
||||
CallError as GenStatemCallError, Cx, GenStatemRef, Machine, Reply, Resolution,
|
||||
CallError as GenStatemCallError, Cx, GenStatemName, GenStatemRef, Machine, Reply, Resolution,
|
||||
SendError as GenStatemSendError,
|
||||
};
|
||||
pub use introspect::{
|
||||
@@ -85,15 +85,15 @@ pub use registry::{
|
||||
install, lookup_as, register, resolve_name, send, send_dyn, send_to, unregister, whereis,
|
||||
NameResolution, RegisterError, SendError,
|
||||
};
|
||||
pub use runtime::{init, Config, Runtime};
|
||||
pub use runtime::{init, Config, Runtime, RuntimeHandle};
|
||||
pub use scheduler::{
|
||||
block_on_io, cancel_timer, request_stop, run, self_pid, send_after, send_after_named,
|
||||
send_after_named_wall, send_after_wall, sleep, sleep_wall, spawn, spawn_addr, spawn_addr_with,
|
||||
spawn_under, spawn_under_with, spawn_with, try_spawn, try_spawn_under_with, wait_readable,
|
||||
wait_readable_timeout, wait_writable, wait_writable_timeout, yield_now, FdArm, JoinError,
|
||||
JoinHandle, SpawnError, SpawnOpts,
|
||||
block_on_io, cancel_timer, request_shutdown, request_stop, run, self_pid, send_after,
|
||||
send_after_named, send_after_named_wall, send_after_wall, sleep, sleep_wall, spawn, spawn_addr,
|
||||
spawn_addr_with, spawn_under, spawn_under_with, spawn_with, try_spawn, try_spawn_under_with,
|
||||
wait_readable, wait_readable_timeout, wait_writable, wait_writable_timeout, yield_now, FdArm,
|
||||
JoinError, JoinHandle, SpawnError, SpawnOpts,
|
||||
};
|
||||
pub use supervisor::{ChildSpec, OneForOne, Restart, Signal, Strategy};
|
||||
pub use supervisor::{ChildSpec, OneForOne, Restart, Shutdown, Signal, Strategy};
|
||||
pub use timer::TimerId;
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
@@ -98,6 +98,11 @@ pub enum DownReason {
|
||||
Panic,
|
||||
/// The target was cooperatively cancelled via `request_stop`.
|
||||
Stopped,
|
||||
/// A graceful shutdown was requested via `request_shutdown`. Only ever
|
||||
/// appears in an [`ExitSignal`](crate::link::ExitSignal) delivered to a
|
||||
/// trapping actor — never in a [`Down`]: a target that honours the request
|
||||
/// exits *normally*, one that does not trap is `Stopped`.
|
||||
Shutdown,
|
||||
/// The target was already gone (finished and reclaimed, or never alive)
|
||||
/// at the moment `monitor()` was called.
|
||||
NoProc,
|
||||
|
||||
+132
-54
@@ -127,7 +127,7 @@ use crate::supervisor::Signal;
|
||||
use crate::timer::Timers;
|
||||
|
||||
use std::sync::atomic::{AtomicBool, AtomicPtr, AtomicU32, AtomicU64, AtomicUsize, Ordering};
|
||||
use std::sync::{Arc, Mutex};
|
||||
use std::sync::{Arc, Mutex, Weak};
|
||||
use std::thread;
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
@@ -904,17 +904,10 @@ pub(crate) struct RuntimeInner {
|
||||
pub(crate) live_actors: AtomicU32,
|
||||
/// Packed `(index << 32 | generation)` of the run's root (initial) actor,
|
||||
/// or `u64::MAX` (the ROOT_PID sentinel) before one is set. When this actor
|
||||
/// finalizes it flags `root_exited`; the scheduler's idle verdict then
|
||||
/// stops the remaining (parked-forever) actors. Set once per `run()`, right
|
||||
/// after the initial spawn.
|
||||
/// finalizes, `finalize_actor` runs the root-exit shutdown (see
|
||||
/// `shutdown_forest_roots`). Set once per `run()`, right after the initial
|
||||
/// spawn.
|
||||
pub(crate) root_bits: AtomicU64,
|
||||
/// Set when the root actor finalizes; read by the scheduler's idle verdict
|
||||
/// to trigger the one-shot teardown sweep. Reset per `run()`.
|
||||
pub(crate) root_exited: AtomicBool,
|
||||
/// Guards the teardown sweep to fire at most once per run (a parked-forever
|
||||
/// remainder that survives the sweep falls through to the normal idle wait
|
||||
/// rather than busy-spinning). Reset per `run()`.
|
||||
pub(crate) root_swept: AtomicBool,
|
||||
/// Timer heap. Independent lock: never nested with any other.
|
||||
pub(crate) timers: Mutex<Timers>,
|
||||
/// IO subsystem. `None` between runs. Lock order: io before everything.
|
||||
@@ -1004,8 +997,6 @@ impl RuntimeInner {
|
||||
free: RawMutex::new(free),
|
||||
live_actors: AtomicU32::new(0),
|
||||
root_bits: AtomicU64::new(u64::MAX),
|
||||
root_exited: AtomicBool::new(false),
|
||||
root_swept: AtomicBool::new(false),
|
||||
timers: Mutex::new(timers),
|
||||
io: Mutex::new(None),
|
||||
next_monitor_id: AtomicU64::new(0),
|
||||
@@ -1338,10 +1329,9 @@ impl Runtime {
|
||||
// requires a running runtime in the thread-local).
|
||||
RUNTIME.with(|r| *r.borrow_mut() = Some(self.inner.clone()));
|
||||
let initial_handle = crate::scheduler::spawn(f);
|
||||
// The initial actor is the run's root: when it exits, remaining actors
|
||||
// are stopped so the run winds down (see finalize_actor / schedule_loop).
|
||||
self.inner.root_exited.store(false, Ordering::Relaxed);
|
||||
self.inner.root_swept.store(false, Ordering::Relaxed);
|
||||
// The initial actor is the run's root: its exit means "the program is
|
||||
// done" — every remaining top-level actor is asked to shut down (see
|
||||
// `finalize_actor` / `shutdown_forest_roots`).
|
||||
self.inner.set_root(initial_handle.pid());
|
||||
|
||||
// Launch N-1 extra scheduler threads, named `smarm-sched-{slot}` so
|
||||
@@ -1457,6 +1447,71 @@ impl Runtime {
|
||||
inner: self.inner.clone(),
|
||||
}
|
||||
}
|
||||
|
||||
/// A `Send + Sync` handle to this runtime, usable from any thread —
|
||||
/// including threads that are not smarm schedulers (an OS-signal handler
|
||||
/// thread, an external event source). Grab it *before* [`run`](Self::run)
|
||||
/// and hand it to e.g. a signal thread; that thread can then
|
||||
/// [`request_stop`](RuntimeHandle::request_stop) the runtime's top
|
||||
/// supervisor to drive an ordered shutdown from outside the runtime.
|
||||
///
|
||||
/// The in-runtime primitives ([`scheduler::request_stop`](crate::request_stop)
|
||||
/// and friends) reach the runtime through a thread-local that is unset on
|
||||
/// any non-scheduler thread, so they are silent no-ops off-runtime; this
|
||||
/// handle carries its own reference and closes that gap.
|
||||
pub fn handle(&self) -> RuntimeHandle {
|
||||
RuntimeHandle {
|
||||
inner: Arc::downgrade(&self.inner),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// RuntimeHandle — off-runtime wake/stop
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/// A `Send + Sync` handle to a [`Runtime`], obtained from
|
||||
/// [`Runtime::handle`]. Lets a thread that is *not* a smarm scheduler thread
|
||||
/// drive a cooperative stop into the runtime — the off-runtime counterpart to
|
||||
/// [`scheduler::request_stop`](crate::request_stop).
|
||||
///
|
||||
/// Holds a [`Weak`] to the runtime, for the same reason the IO backend does
|
||||
/// (RFC 018): a lingering handle can never keep the runtime's slot table alive
|
||||
/// and can never block [`Runtime::run`] from finishing. Once the `Runtime` is
|
||||
/// dropped every method is a harmless no-op — the same end state as calling
|
||||
/// `request_stop` on an actor that has already exited.
|
||||
#[derive(Clone)]
|
||||
pub struct RuntimeHandle {
|
||||
inner: Weak<RuntimeInner>,
|
||||
}
|
||||
|
||||
impl RuntimeHandle {
|
||||
/// Ask `pid` to stop cooperatively, from any thread. The off-runtime
|
||||
/// equivalent of [`scheduler::request_stop`](crate::request_stop): it sets
|
||||
/// the target's stop flag and wakes it, so a parked actor unwinds at its
|
||||
/// next checkpoint exactly as it would for an in-runtime stop. A no-op if
|
||||
/// the runtime has been dropped, or if the actor has already exited.
|
||||
pub fn request_stop<A>(&self, pid: Pid<A>) {
|
||||
let pid = pid.erase();
|
||||
// Upgrade the Weak per call, like the IO backend does (io.rs): a live
|
||||
// runtime yields the inner and we drive the same stop the in-runtime
|
||||
// path would; a dropped runtime makes this a no-op.
|
||||
if let Some(inner) = self.inner.upgrade() {
|
||||
crate::scheduler::request_stop_inner(&inner, pid);
|
||||
}
|
||||
}
|
||||
|
||||
/// Ask `pid` to shut down gracefully, from any thread. The off-runtime
|
||||
/// equivalent of [`scheduler::request_shutdown`](crate::request_shutdown);
|
||||
/// the delivered [`ExitSignal`](crate::ExitSignal) carries `from ==
|
||||
/// ROOT_PID`, since no actor made the request. A no-op if the runtime has
|
||||
/// been dropped, or if the actor has already exited.
|
||||
pub fn request_shutdown<A>(&self, pid: Pid<A>) {
|
||||
let pid = pid.erase();
|
||||
if let Some(inner) = self.inner.upgrade() {
|
||||
crate::scheduler::request_shutdown_inner(&inner, pid, ROOT_PID);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
@@ -1851,14 +1906,12 @@ fn finalize_actor(inner: &Arc<RuntimeInner>, pid: Pid, outcome: Outcome) {
|
||||
// Reclaim if no outstanding handles (re-verified inside).
|
||||
reclaim_slot(inner, pid);
|
||||
|
||||
// Root-exit teardown is DEFERRED to the scheduler's idle verdict, not done
|
||||
// here: stopping eagerly would cut off actors that still have queued work
|
||||
// (they'd unwind on the stop before draining their mailbox). Flagging it
|
||||
// instead lets the run queue drain naturally first; only the parked-forever
|
||||
// remainder (e.g. a server pinned alive by a registered name) is then
|
||||
// stopped, once nothing runnable is left. See `schedule_loop`.
|
||||
// Root exit = the program is done. Ask every top-level survivor to shut
|
||||
// down, right here, before the live-count decrement below: any wake this
|
||||
// produces is then ordered before `live_actors` can be observed at its
|
||||
// decremented value, same as every other wakeup finalize issues.
|
||||
if inner.is_root(pid) {
|
||||
inner.root_exited.store(true, Ordering::Release);
|
||||
shutdown_forest_roots(inner, pid);
|
||||
}
|
||||
|
||||
// The decrement is LAST: every wakeup this finalize produced (joiners,
|
||||
@@ -1869,16 +1922,62 @@ fn finalize_actor(inner: &Arc<RuntimeInner>, pid: Pid, outcome: Outcome) {
|
||||
debug_assert!(prev >= 1, "live_actors underflow — double finalize");
|
||||
}
|
||||
|
||||
/// Cooperatively stop every live actor — the root-exit teardown sweep, run from
|
||||
/// `schedule_loop` once the run queue is empty after the root has exited. Each
|
||||
/// [`request_stop_inner`](crate::scheduler::request_stop_inner) re-verifies the
|
||||
/// target under its cold lock, so the racy per-slot generation read is safe: a
|
||||
/// vacant, dead, or reused slot no-ops. The swept actors unpark, unwind at their
|
||||
/// next observation point, and finalize, dropping `live_actors` to zero.
|
||||
fn stop_live_actors(inner: &Arc<RuntimeInner>) {
|
||||
/// The root-exit shutdown. Delivers [`request_shutdown`](crate::request_shutdown)
|
||||
/// to every **forest root**: each live actor whose recorded parent
|
||||
/// (`Actor::supervisor` — the spawner for a plain `spawn`, the supervisor for
|
||||
/// `spawn_under`) is the run itself (`ROOT_PID`) or is no longer live. Actors
|
||||
/// under a live parent are not addressed — that parent is responsible for
|
||||
/// them: a supervisor traps and runs its ordered, policy-driven shutdown; a
|
||||
/// bare parent that dies takes non-trapping children with it via the next
|
||||
/// pass of this same rule only if it dies *now*, so a parent that outlives
|
||||
/// this scan and later dies leaves its subtree to itself (Erlang semantics: an
|
||||
/// unlinked spawn is nobody's child).
|
||||
///
|
||||
/// Semantics per target follow `request_shutdown`: a trapping actor receives
|
||||
/// `ExitSignal { from: root, reason: Shutdown }` and may finish work — drain,
|
||||
/// keep its timers ticking, then stop itself; a non-trapping one is stopped
|
||||
/// outright. There is no second, forcing sweep: an actor that traps and never
|
||||
/// stops keeps the run alive by design (put it under a supervisor with a
|
||||
/// `Shutdown::Timeout` policy if that is not wanted). Runs once, on the root's
|
||||
/// finalize path, so it races only against actors that are still running —
|
||||
/// each `request_shutdown_inner` re-verifies its target under the cold lock,
|
||||
/// so a slot that dies or is reused mid-scan is a no-op.
|
||||
fn shutdown_forest_roots(inner: &Arc<RuntimeInner>, root: Pid) {
|
||||
for idx in 0..inner.slots.len() as u32 {
|
||||
let pid = Pid::new(idx, inner.slots[idx as usize].generation());
|
||||
crate::scheduler::request_stop_inner(inner, pid);
|
||||
let slot = &inner.slots[idx as usize];
|
||||
let pid = Pid::new(idx, slot.generation());
|
||||
if pid == root {
|
||||
continue;
|
||||
}
|
||||
// Read the parent under the cold lock (generation-verified); act
|
||||
// outside it — `request_shutdown_inner` sends and may unpark.
|
||||
let parent = {
|
||||
let cold = slot.cold.lock();
|
||||
if slot.generation() != pid.generation() {
|
||||
continue;
|
||||
}
|
||||
match cold.actor.as_ref() {
|
||||
Some(a) => a.supervisor,
|
||||
None => continue,
|
||||
}
|
||||
};
|
||||
let parent_live = inner
|
||||
.slot_at(parent)
|
||||
.is_some_and(|ps| ps.is_live_for(parent));
|
||||
if !parent_live {
|
||||
// `_probe` so the trace can say what each leftover was: under
|
||||
// `smarm-trace` every swept actor is a `root_sweep` line — the
|
||||
// visibility that makes "a forgotten actor costs only a slot, never
|
||||
// a hung run" a checkable claim rather than a hope.
|
||||
let found = crate::scheduler::request_shutdown_inner_probe(inner, pid, root);
|
||||
#[cfg_attr(not(feature = "smarm-trace"), allow(unused_variables))]
|
||||
if let Some(trapping) = found {
|
||||
crate::te!(crate::trace::Event::RootSweep {
|
||||
target: pid,
|
||||
trapping
|
||||
});
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1962,9 +2061,6 @@ fn schedule_loop(inner: &Arc<RuntimeInner>, slot_idx: usize) {
|
||||
Got(Pid),
|
||||
Idle,
|
||||
AllDone,
|
||||
/// Root has exited and nothing is runnable: stop the parked-forever
|
||||
/// remainder, then re-pop. Fires at most once per run.
|
||||
RootDrain,
|
||||
}
|
||||
|
||||
// 2a. RFC 005: drain this thread's wake slot before touching the
|
||||
@@ -2014,15 +2110,6 @@ fn schedule_loop(inner: &Arc<RuntimeInner>, slot_idx: usize) {
|
||||
let live = inner.live_actors.load(Ordering::Acquire);
|
||||
if live == 0 && io_out == 0 {
|
||||
Pop::AllDone
|
||||
} else if inner.root_exited.load(Ordering::Acquire)
|
||||
&& !inner.root_swept.swap(true, Ordering::AcqRel)
|
||||
{
|
||||
// Root gone and nothing runnable — the live remainder
|
||||
// are parked-forever daemons (Queued actors with pending
|
||||
// work drained before the queue emptied). Stop them so
|
||||
// the run can end. One-shot: a survivor falls through to
|
||||
// the idle wait below on the next pass.
|
||||
Pop::RootDrain
|
||||
} else {
|
||||
Pop::Idle
|
||||
}
|
||||
@@ -2050,13 +2137,6 @@ fn schedule_loop(inner: &Arc<RuntimeInner>, slot_idx: usize) {
|
||||
inner.coord.wake_all();
|
||||
return;
|
||||
}
|
||||
Pop::RootDrain => {
|
||||
// Root has exited and nothing is runnable: stop the
|
||||
// parked-forever remainder, then loop back to re-pop the
|
||||
// now-runnable (stopping) actors.
|
||||
stop_live_actors(inner);
|
||||
continue;
|
||||
}
|
||||
Pop::Idle => {
|
||||
// Something is still in flight. Park on our own futex
|
||||
// until a producer wakes us (enqueue tail), a deadline
|
||||
@@ -2087,8 +2167,6 @@ fn schedule_loop(inner: &Arc<RuntimeInner>, slot_idx: usize) {
|
||||
|| (inner.live_actors.load(Ordering::Acquire) == 0
|
||||
&& inner.io_outstanding.load(Ordering::Acquire) == 0
|
||||
&& inner.io_fd_waiters.load(Ordering::Acquire) == 0)
|
||||
|| (inner.root_exited.load(Ordering::Acquire)
|
||||
&& !inner.root_swept.load(Ordering::Acquire))
|
||||
|| inner.coord.deadline_due()
|
||||
});
|
||||
if tk_deadline.is_some() {
|
||||
|
||||
+94
-6
@@ -71,7 +71,7 @@ use crate::pid::{Name, Pid};
|
||||
use crate::runtime::{self, RuntimeInner, YieldIntent, RUNTIME};
|
||||
use crate::supervisor::Signal;
|
||||
use std::sync::atomic::Ordering;
|
||||
use std::sync::Arc;
|
||||
use std::sync::{Arc, Weak};
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// with_runtime / try_with_runtime
|
||||
@@ -544,6 +544,29 @@ pub(crate) fn unpark_at(pid: Pid, epoch: u32) {
|
||||
let _ = try_with_runtime(|inner| inner.unpark_at(pid, epoch));
|
||||
}
|
||||
|
||||
// The current actor's runtime as a `Weak`, for a waker that must reach the
|
||||
// runtime from a foreign thread later. A channel captures this when its
|
||||
// receiver parks, so a cross-thread `send` can wake without the `RUNTIME`
|
||||
// thread-local (unset off a scheduler thread). Panics outside `Runtime::run()`,
|
||||
// the same contract as `begin_wait`.
|
||||
pub(crate) fn runtime_weak() -> Weak<RuntimeInner> {
|
||||
with_runtime(Arc::downgrade)
|
||||
}
|
||||
|
||||
// Epoch-matched wake of `pid` from a waker that may or may not be on a
|
||||
// scheduler thread. On a scheduler thread we take the thread-local path
|
||||
// (preemption-gated, slot-eligible); off one that path is a silent no-op, so
|
||||
// we reach the runtime through `rt` — the `Weak` the waker captured while it
|
||||
// was in-runtime. Mirrors the IO backend's cross-context wake (io.rs, RFC 018).
|
||||
pub(crate) fn unpark_at_via(pid: Pid, epoch: u32, rt: &Weak<RuntimeInner>) {
|
||||
if try_with_runtime(|inner| inner.unpark_at(pid, epoch)).is_some() {
|
||||
return;
|
||||
}
|
||||
if let Some(inner) = rt.upgrade() {
|
||||
inner.unpark_at(pid, epoch);
|
||||
}
|
||||
}
|
||||
|
||||
// Open a new wait for the current actor and return its wait identity
|
||||
// ("epoch"). Call once per wait, before registering with any waker. Lock-free,
|
||||
// so it's legal to call while already holding another internal lock.
|
||||
@@ -577,10 +600,12 @@ pub(crate) fn retire_wait() {
|
||||
/// [`JoinHandle::join`] reports it as a normal, non-error exit: cooperative
|
||||
/// stop is a controlled shutdown, not a failure.
|
||||
///
|
||||
/// This is exactly the mechanism `gen_server` shutdown, supervisor restarts,
|
||||
/// and structured teardown are built from: reach for [`GenServerRef::shutdown`](crate::GenServerRef::shutdown)
|
||||
/// or a [`supervisor`](crate::supervisor) instead of calling this directly
|
||||
/// where those apply.
|
||||
/// This is the *hard* stop — OTP's `exit(Pid, kill)`. It is what a supervisor
|
||||
/// falls back to when a child overstays its [`Shutdown`](crate::supervisor::Shutdown)
|
||||
/// grace period. For a stop the target gets to prepare for, use
|
||||
/// [`request_shutdown`]; for structured teardown, reach for
|
||||
/// [`GenServerRef::shutdown`](crate::GenServerRef::shutdown) or a
|
||||
/// [`supervisor`](crate::supervisor) instead of calling this directly.
|
||||
///
|
||||
/// Because it's cooperative, an actor stuck in a tight loop with no
|
||||
/// blocking call, no [`check!`](crate::check), and no allocation cannot be
|
||||
@@ -592,7 +617,8 @@ pub fn request_stop<A>(pid: Pid<A>) {
|
||||
}
|
||||
|
||||
// The core of `request_stop`, taking the runtime directly so it can also be
|
||||
// driven from inside the runtime itself (the root-exit sweep) without
|
||||
// driven from inside the runtime itself (the RuntimeHandle path, supervisor
|
||||
// sweeps) without
|
||||
// re-borrowing the thread-local. Sets the stop flag under the target's lock
|
||||
// (a generation mismatch, or no live actor there, makes it a no-op) and
|
||||
// wakes the target.
|
||||
@@ -612,6 +638,68 @@ pub(crate) fn request_stop_inner(inner: &RuntimeInner, pid: Pid) {
|
||||
}
|
||||
}
|
||||
|
||||
/// Ask an actor to shut down gracefully — OTP's `exit(Pid, shutdown)`, where
|
||||
/// [`request_stop`] is `exit(Pid, kill)`.
|
||||
///
|
||||
/// If the target has called [`trap_exit`](crate::trap_exit), it receives an
|
||||
/// [`ExitSignal`](crate::ExitSignal) with reason
|
||||
/// [`DownReason::Shutdown`](crate::DownReason::Shutdown) on its trap inbox and
|
||||
/// keeps running: the request is advisory, and the target is expected to wind
|
||||
/// down and exit normally in its own time (a supervisor bounds that time with
|
||||
/// its child's [`Shutdown`](crate::supervisor::Shutdown) policy and falls back
|
||||
/// to `request_stop`). A target that is not trapping is stopped exactly as by
|
||||
/// `request_stop`. A dead pid is a no-op.
|
||||
///
|
||||
/// The signal's `from` is the calling actor, or `ROOT_PID` when driven from
|
||||
/// outside the runtime (see [`RuntimeHandle::request_shutdown`](crate::RuntimeHandle::request_shutdown)).
|
||||
pub fn request_shutdown<A>(pid: Pid<A>) {
|
||||
let pid = pid.erase();
|
||||
let from = current_pid().unwrap_or(crate::runtime::ROOT_PID);
|
||||
let _ = try_with_runtime(|inner| request_shutdown_inner(inner, pid, from));
|
||||
}
|
||||
|
||||
// The core of `request_shutdown`. Reads the target's trap sender under its
|
||||
// cold lock (generation-verified), then acts outside the lock: a trap send
|
||||
// may unpark the receiver, and `request_stop_inner` re-takes the lock.
|
||||
pub(crate) fn request_shutdown_inner(inner: &RuntimeInner, pid: Pid, from: Pid) {
|
||||
request_shutdown_inner_probe(inner, pid, from);
|
||||
}
|
||||
|
||||
/// [`request_shutdown_inner`], reporting what it found: `Some(true)` if the
|
||||
/// target was trapping (got the signal), `Some(false)` if it was stopped
|
||||
/// outright, `None` if there was nothing live at `pid`.
|
||||
pub(crate) fn request_shutdown_inner_probe(
|
||||
inner: &RuntimeInner,
|
||||
pid: Pid,
|
||||
from: Pid,
|
||||
) -> Option<bool> {
|
||||
let trap = match inner.slot_at(pid) {
|
||||
Some(slot) => {
|
||||
let cold = slot.cold.lock();
|
||||
if slot.generation() == pid.generation() {
|
||||
cold.actor.as_ref().map(|a| a.trap.clone())
|
||||
} else {
|
||||
None // stale pid: nothing there to shut down
|
||||
}
|
||||
}
|
||||
None => None,
|
||||
};
|
||||
match trap {
|
||||
Some(Some(tx)) => {
|
||||
let _ = tx.send(crate::link::ExitSignal {
|
||||
from,
|
||||
reason: crate::monitor::DownReason::Shutdown,
|
||||
});
|
||||
Some(true)
|
||||
}
|
||||
Some(None) => {
|
||||
request_stop_inner(inner, pid);
|
||||
Some(false)
|
||||
}
|
||||
None => None,
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// NoPreempt
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
+189
-68
@@ -140,7 +140,8 @@ impl Signal {
|
||||
}
|
||||
}
|
||||
|
||||
use crate::channel::channel;
|
||||
use crate::channel::{channel, RecvTimeoutError};
|
||||
use crate::monitor::DownReason;
|
||||
use std::collections::{HashMap, VecDeque};
|
||||
use std::sync::Arc;
|
||||
use std::time::{Duration, Instant};
|
||||
@@ -165,15 +166,53 @@ pub enum Restart {
|
||||
pub struct ChildSpec {
|
||||
start: Arc<dyn Fn() + Send + Sync + 'static>,
|
||||
restart: Restart,
|
||||
shutdown: Shutdown,
|
||||
}
|
||||
|
||||
impl ChildSpec {
|
||||
/// A child with the given restart policy and the default
|
||||
/// [`Shutdown::Timeout`] of 5 seconds.
|
||||
pub fn new(restart: Restart, start: impl Fn() + Send + Sync + 'static) -> Self {
|
||||
Self {
|
||||
start: Arc::new(start),
|
||||
restart,
|
||||
shutdown: Shutdown::default(),
|
||||
}
|
||||
}
|
||||
|
||||
/// Set how the supervisor stops this child (see [`Shutdown`]). A child
|
||||
/// that is itself a supervisor should use [`Shutdown::Infinity`] so its
|
||||
/// own subtree gets its full grace periods.
|
||||
pub fn shutdown(mut self, shutdown: Shutdown) -> Self {
|
||||
self.shutdown = shutdown;
|
||||
self
|
||||
}
|
||||
}
|
||||
|
||||
/// How a supervisor stops a child it is taking down — the OTP child-spec
|
||||
/// `shutdown` value. Applies to every supervisor-initiated stop: the ordered
|
||||
/// shutdown of the whole set and the sibling cycling of
|
||||
/// [`Strategy::OneForAll`] / [`Strategy::RestForOne`].
|
||||
///
|
||||
/// A graceful stop is a [`request_shutdown`](crate::request_shutdown): a child
|
||||
/// that traps exits receives the request as a message and winds down in its
|
||||
/// own time; one that does not is stopped outright.
|
||||
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
|
||||
pub enum Shutdown {
|
||||
/// `request_stop` immediately; no request, no grace period.
|
||||
BrutalKill,
|
||||
/// `request_shutdown`, wait up to the duration for the child to exit, then
|
||||
/// `request_stop` it. The default, at 5 seconds.
|
||||
Timeout(Duration),
|
||||
/// `request_shutdown` and wait however long the child takes. Use for a
|
||||
/// child supervisor, whose subtree has its own timeouts.
|
||||
Infinity,
|
||||
}
|
||||
|
||||
impl Default for Shutdown {
|
||||
fn default() -> Self {
|
||||
Shutdown::Timeout(Duration::from_secs(5))
|
||||
}
|
||||
}
|
||||
|
||||
/// How a supervisor reacts when one child terminates and a restart is due.
|
||||
@@ -244,23 +283,35 @@ impl OneForOne {
|
||||
}
|
||||
|
||||
/// Run the supervision loop on the current actor. Returns when every child
|
||||
/// has reached a terminal, non-restartable state, or when the restart
|
||||
/// intensity cap is tripped.
|
||||
/// has reached a terminal, non-restartable state, when the restart
|
||||
/// intensity cap is tripped, or when the supervisor is asked to shut down
|
||||
/// (a [`request_shutdown`](crate::request_shutdown) — from its own
|
||||
/// supervisor, or from the app). On every one of those exits the survivors
|
||||
/// are stopped in reverse start order, each per its
|
||||
/// [`Shutdown`] policy, before this returns.
|
||||
///
|
||||
/// The supervisor traps exits for the length of the loop (that is how the
|
||||
/// shutdown request reaches it as a message). Should the supervisor itself
|
||||
/// be hard-stopped with [`request_stop`](crate::request_stop), it unwinds
|
||||
/// without waiting for anything — but a drop guard hard-stops its live
|
||||
/// children on the way out, so the subtree is not orphaned (a child
|
||||
/// supervisor unwinds the same way, recursively).
|
||||
pub fn run(self) {
|
||||
let me = crate::scheduler::self_pid();
|
||||
let (tx, rx) = channel::<Signal>();
|
||||
crate::scheduler::register_supervisor_channel(me, tx);
|
||||
let exits = crate::link::trap_exit();
|
||||
|
||||
// pid -> index into `self.children`, for the children currently alive.
|
||||
let mut by_pid: HashMap<Pid, usize> = HashMap::new();
|
||||
let mut live = Live::default();
|
||||
let mut active: usize = 0;
|
||||
// Sliding window of recent restart instants, for the intensity cap.
|
||||
let mut restarts: Vec<Instant> = Vec::new();
|
||||
|
||||
let start_child = |idx: usize, by_pid: &mut HashMap<Pid, usize>| {
|
||||
let start_child = |idx: usize, live: &mut Live| {
|
||||
let start = self.children[idx].start.clone();
|
||||
let h = crate::scheduler::spawn_under(me, move || (start)());
|
||||
by_pid.insert(h.pid(), idx);
|
||||
live.insert(h.pid(), idx);
|
||||
// We supervise via the signal funnel, not by joining; drop the
|
||||
// handle so the child's slot is reclaimed promptly on death (the
|
||||
// termination Signal is delivered before reclamation regardless).
|
||||
@@ -268,28 +319,105 @@ impl OneForOne {
|
||||
};
|
||||
|
||||
for idx in 0..self.children.len() {
|
||||
start_child(idx, &mut by_pid);
|
||||
start_child(idx, &mut live);
|
||||
active += 1;
|
||||
}
|
||||
|
||||
// A signal that arrives while we are awaiting stop-confirmations (for a
|
||||
// child we are *not* currently stopping) is stashed here and processed
|
||||
// by the main loop before it blocks on `recv` again.
|
||||
// by the main loop before it blocks again.
|
||||
let mut pending: VecDeque<Signal> = VecDeque::new();
|
||||
let next_signal = |pending: &mut VecDeque<Signal>| -> Option<Signal> {
|
||||
if let Some(s) = pending.pop_front() {
|
||||
Some(s)
|
||||
} else {
|
||||
rx.recv().ok()
|
||||
|
||||
// Stop one child per its policy and wait for its termination signal.
|
||||
// Signals for other pids that arrive meanwhile are stashed. Bounded by
|
||||
// construction: `request_stop` (used directly, or as the fallback once
|
||||
// the grace period lapses) always produces a signal.
|
||||
let stop_child = |pid: Pid, idx: usize, pending: &mut VecDeque<Signal>| {
|
||||
let await_one = |deadline: Option<Instant>, pending: &mut VecDeque<Signal>| -> bool {
|
||||
loop {
|
||||
let sig = match pending.iter().position(|s| s.pid() == pid) {
|
||||
Some(i) => pending.remove(i),
|
||||
None => match deadline {
|
||||
None => rx.recv().ok(),
|
||||
Some(dl) => {
|
||||
match rx.recv_timeout(dl.saturating_duration_since(Instant::now()))
|
||||
{
|
||||
Ok(s) => Some(s),
|
||||
Err(RecvTimeoutError::Timeout) => return false,
|
||||
Err(RecvTimeoutError::Disconnected) => None,
|
||||
}
|
||||
}
|
||||
},
|
||||
};
|
||||
match sig {
|
||||
Some(s) if s.pid() == pid => return true,
|
||||
Some(s) => pending.push_back(s),
|
||||
None => return true, // funnel closed: nothing more can arrive
|
||||
}
|
||||
}
|
||||
};
|
||||
match self.children[idx].shutdown {
|
||||
Shutdown::BrutalKill => {
|
||||
crate::scheduler::request_stop(pid);
|
||||
await_one(None, pending);
|
||||
}
|
||||
Shutdown::Timeout(grace) => {
|
||||
crate::scheduler::request_shutdown(pid);
|
||||
if !await_one(Some(Instant::now() + grace), pending) {
|
||||
crate::scheduler::request_stop(pid);
|
||||
await_one(None, pending);
|
||||
}
|
||||
}
|
||||
Shutdown::Infinity => {
|
||||
crate::scheduler::request_shutdown(pid);
|
||||
await_one(None, pending);
|
||||
}
|
||||
}
|
||||
};
|
||||
|
||||
// Stop a set of children in reverse start order, one at a time.
|
||||
let stop_set =
|
||||
|set: &mut Vec<(Pid, usize)>, live: &mut Live, pending: &mut VecDeque<Signal>| {
|
||||
set.sort_unstable_by_key(|x| std::cmp::Reverse(x.1));
|
||||
for (pid, idx) in set.iter() {
|
||||
live.remove(pid);
|
||||
stop_child(*pid, *idx, pending);
|
||||
}
|
||||
};
|
||||
|
||||
// Wait for the next event: a stashed signal, a child signal, or a
|
||||
// shutdown request. `Ok(sig)`, or `Err(())` when we must wind down.
|
||||
let next_event = |pending: &mut VecDeque<Signal>| -> Result<Signal, ()> {
|
||||
loop {
|
||||
if let Some(s) = pending.pop_front() {
|
||||
return Ok(s);
|
||||
}
|
||||
// The trap inbox is arm 0: a shutdown request is noticed even
|
||||
// under a flood of child signals.
|
||||
match crate::channel::select(&[&exits, &rx]) {
|
||||
0 => match exits.try_recv() {
|
||||
Ok(Some(sig)) if sig.reason == DownReason::Shutdown => return Err(()),
|
||||
// Any other exit signal (a linked peer's death — a
|
||||
// supervisor links nothing itself, but may be linked
|
||||
// to) is not ours to act on; a closed trap inbox is
|
||||
// impossible while `exits` is held here.
|
||||
_ => {}
|
||||
},
|
||||
_ => match rx.try_recv() {
|
||||
Ok(Some(s)) => return Ok(s),
|
||||
Ok(None) => {}
|
||||
Err(_) => return Err(()), // funnel closed: nothing left to supervise
|
||||
},
|
||||
}
|
||||
}
|
||||
};
|
||||
|
||||
while active > 0 {
|
||||
let sig = match next_signal(&mut pending) {
|
||||
Some(s) => s,
|
||||
None => break, // mailbox closed: nothing left to supervise
|
||||
let sig = match next_event(&mut pending) {
|
||||
Ok(s) => s,
|
||||
Err(()) => break,
|
||||
};
|
||||
let idx = match by_pid.remove(&sig.pid()) {
|
||||
let idx = match live.remove(&sig.pid()) {
|
||||
Some(i) => i,
|
||||
None => continue, // stray/duplicate signal
|
||||
};
|
||||
@@ -321,76 +449,69 @@ impl OneForOne {
|
||||
restarts.push(now);
|
||||
|
||||
// Which *live* siblings get cycled along with the failed child.
|
||||
// (The failed child is already gone — removed from `by_pid` above.)
|
||||
// (The failed child is already gone — removed from `live` above.)
|
||||
let mut to_stop: Vec<(Pid, usize)> = match self.strategy {
|
||||
Strategy::OneForOne => Vec::new(),
|
||||
Strategy::OneForAll => by_pid.iter().map(|(p, i)| (*p, *i)).collect(),
|
||||
Strategy::RestForOne => by_pid
|
||||
Strategy::OneForAll => live.iter().map(|(p, i)| (*p, *i)).collect(),
|
||||
Strategy::RestForOne => live
|
||||
.iter()
|
||||
.filter(|(_, i)| **i > idx)
|
||||
.map(|(p, i)| (*p, *i))
|
||||
.collect(),
|
||||
};
|
||||
// Stop survivors in reverse start order (highest child index first).
|
||||
to_stop.sort_unstable_by_key(|x| std::cmp::Reverse(x.1));
|
||||
|
||||
// The set we will restart: the failed child plus every sibling we
|
||||
// are about to stop, restarted in start (ascending index) order.
|
||||
let mut restart_set: Vec<usize> = Vec::with_capacity(to_stop.len() + 1);
|
||||
restart_set.push(idx);
|
||||
restart_set.extend(to_stop.iter().map(|(_, i)| *i));
|
||||
|
||||
// Request stops, then await each survivor's termination signal
|
||||
// before restarting. `request_stop` on an already-dead pid is a
|
||||
// no-op; in that case its (already-sent) Exit signal serves as the
|
||||
// confirmation. Any signal for a pid we are *not* awaiting is
|
||||
// stashed for the main loop.
|
||||
let mut awaiting: Vec<Pid> = Vec::with_capacity(to_stop.len());
|
||||
for (pid, cidx) in &to_stop {
|
||||
by_pid.remove(pid);
|
||||
restart_set.push(*cidx);
|
||||
crate::scheduler::request_stop(*pid);
|
||||
awaiting.push(*pid);
|
||||
}
|
||||
while !awaiting.is_empty() {
|
||||
let s = match next_signal(&mut pending) {
|
||||
Some(s) => s,
|
||||
None => break, // mailbox closed mid-await; stop waiting
|
||||
};
|
||||
if let Some(pos) = awaiting.iter().position(|p| *p == s.pid()) {
|
||||
awaiting.swap_remove(pos);
|
||||
} else {
|
||||
pending.push_back(s);
|
||||
}
|
||||
}
|
||||
|
||||
// Restart the whole set in start order. Net effect on `active`:
|
||||
// one child died (idx), `to_stop.len()` were stopped, and
|
||||
// `restart_set.len() == 1 + to_stop.len()` are started — so
|
||||
// Stop the survivors (each per its policy, reverse start order),
|
||||
// then restart the whole set in start order. Net effect on
|
||||
// `active`: one child died (idx), `to_stop.len()` were stopped,
|
||||
// and `restart_set.len() == 1 + to_stop.len()` are started — so
|
||||
// `active` is unchanged and needs no adjustment here.
|
||||
stop_set(&mut to_stop, &mut live, &mut pending);
|
||||
restart_set.sort_unstable();
|
||||
for cidx in restart_set {
|
||||
start_child(cidx, &mut by_pid);
|
||||
start_child(cidx, &mut live);
|
||||
}
|
||||
}
|
||||
|
||||
// Ordered shutdown: stop any survivors in reverse start order and await
|
||||
// their termination. On the normal `active == 0` exit `by_pid` is empty
|
||||
// and this is a no-op; on a cap-trip or mailbox-closed break it tears
|
||||
// the remaining children down deterministically instead of leaking them.
|
||||
let mut survivors: Vec<(Pid, usize)> = by_pid.iter().map(|(p, i)| (*p, *i)).collect();
|
||||
survivors.sort_unstable_by_key(|x| std::cmp::Reverse(x.1));
|
||||
let mut awaiting: Vec<Pid> = Vec::with_capacity(survivors.len());
|
||||
for (pid, _) in &survivors {
|
||||
crate::scheduler::request_stop(*pid);
|
||||
awaiting.push(*pid);
|
||||
}
|
||||
while !awaiting.is_empty() {
|
||||
let s = match next_signal(&mut pending) {
|
||||
Some(s) => s,
|
||||
None => break,
|
||||
};
|
||||
if let Some(pos) = awaiting.iter().position(|p| *p == s.pid()) {
|
||||
awaiting.swap_remove(pos);
|
||||
// Ordered shutdown: stop any survivors in reverse start order, each per
|
||||
// its policy. On the normal `active == 0` exit `live` is empty and this
|
||||
// is a no-op; on a shutdown request, a cap-trip, or a closed funnel it
|
||||
// tears the remaining children down deterministically.
|
||||
let mut survivors: Vec<(Pid, usize)> = live.iter().map(|(p, i)| (*p, *i)).collect();
|
||||
stop_set(&mut survivors, &mut live, &mut pending);
|
||||
}
|
||||
}
|
||||
|
||||
/// The live children of a supervisor, with a drop guard: if the supervisor is
|
||||
/// unwound (a hard `request_stop`, or a panic in the loop) its children are
|
||||
/// hard-stopped rather than orphaned. Fire-and-forget by necessity — a guard
|
||||
/// running mid-unwind cannot park to await anything.
|
||||
#[derive(Default)]
|
||||
struct Live(HashMap<Pid, usize>);
|
||||
|
||||
impl std::ops::Deref for Live {
|
||||
type Target = HashMap<Pid, usize>;
|
||||
fn deref(&self) -> &Self::Target {
|
||||
&self.0
|
||||
}
|
||||
}
|
||||
|
||||
impl std::ops::DerefMut for Live {
|
||||
fn deref_mut(&mut self) -> &mut Self::Target {
|
||||
&mut self.0
|
||||
}
|
||||
}
|
||||
|
||||
impl Drop for Live {
|
||||
fn drop(&mut self) {
|
||||
if std::thread::panicking() {
|
||||
for pid in self.0.keys() {
|
||||
crate::scheduler::request_stop(*pid);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
+24
-2
@@ -46,17 +46,32 @@ mod inner {
|
||||
#[derive(Clone, Debug)]
|
||||
pub enum Event {
|
||||
// Actor lifecycle
|
||||
Spawn { parent: Pid, child: Pid },
|
||||
Spawn {
|
||||
parent: Pid,
|
||||
child: Pid,
|
||||
},
|
||||
Resume(Pid),
|
||||
Yield(Pid),
|
||||
Park(Pid),
|
||||
Done(Pid),
|
||||
/// Root exit found a live forest root (an actor nobody supervises)
|
||||
/// and delivered `request_shutdown` to it. `trapping` says whether it
|
||||
/// got the chance to drain (true) or was stopped outright (false).
|
||||
/// Every such line is an actor whose lifetime was nobody's business
|
||||
/// but the runtime's — the way to *see* unsupervised leftovers.
|
||||
RootSweep {
|
||||
target: Pid,
|
||||
trapping: bool,
|
||||
},
|
||||
// Wakeup paths
|
||||
UnparkDirect(Pid), // unpark() saw Parked -> re-queued immediately
|
||||
UnparkDeferred(Pid), // unpark() saw Runnable -> set pending_unpark flag
|
||||
UnparkFlagConsumed(Pid), // scheduler saw flag on Park -> re-queued instead
|
||||
// Channel
|
||||
Send { sender: Pid, receiver: Option<Pid> },
|
||||
Send {
|
||||
sender: Pid,
|
||||
receiver: Option<Pid>,
|
||||
},
|
||||
RecvPark(Pid),
|
||||
RecvWake(Pid),
|
||||
// Queue
|
||||
@@ -253,6 +268,13 @@ mod inner {
|
||||
Event::Yield(p) => ("yield".into(), p.index()),
|
||||
Event::Park(p) => ("park".into(), p.index()),
|
||||
Event::Done(p) => ("done".into(), p.index()),
|
||||
Event::RootSweep { target, trapping } => (
|
||||
format!(
|
||||
"root_sweep {}",
|
||||
if *trapping { "shutdown" } else { "stopped" }
|
||||
),
|
||||
target.index(),
|
||||
),
|
||||
Event::UnparkDirect(p) => ("unpark_direct".into(), p.index()),
|
||||
Event::UnparkDeferred(p) => ("unpark_deferred".into(), p.index()),
|
||||
Event::UnparkFlagConsumed(p) => ("unpark_flag_consumed".into(), p.index()),
|
||||
|
||||
@@ -0,0 +1,125 @@
|
||||
//! Cross-thread wake: a thread that is *not* a smarm scheduler thread must be
|
||||
//! able to wake (and stop) a parked actor.
|
||||
//!
|
||||
//! The gap this pins down: every off-runtime wake primitive (`unpark`,
|
||||
//! `unpark_at`, `request_stop`) reaches the runtime through the `RUNTIME`
|
||||
//! thread-local, which is `None` on any non-scheduler thread — so a wake
|
||||
//! issued from a foreign OS thread is a silent no-op and the parked actor
|
||||
//! sleeps forever. Both failure modes below manifest as `Runtime::run` never
|
||||
//! returning, so each test is wrapped in a watchdog: a timeout is the failure.
|
||||
//!
|
||||
//! The fix mirrors RFC 018's IO backend — the waker reaches the runtime
|
||||
//! through a `Weak<RuntimeInner>` it already holds (the receiver captures one
|
||||
//! when it parks; `Runtime::handle()` hands one to an app thread).
|
||||
|
||||
use std::sync::mpsc;
|
||||
use std::thread;
|
||||
use std::time::Duration;
|
||||
|
||||
const WATCHDOG: Duration = Duration::from_secs(10);
|
||||
/// Give the target actor time to actually park before the foreign thread pokes
|
||||
/// it, so we exercise the *wake* of a parked actor rather than the entry-side
|
||||
/// stop check.
|
||||
const SETTLE: Duration = Duration::from_millis(200);
|
||||
|
||||
fn assert_send_sync<T: Send + Sync>() {}
|
||||
|
||||
/// A cross-thread `send` from a plain OS thread must wake a receiver parked in
|
||||
/// `recv`. Under the thread-local-only wake path the send enqueues the message
|
||||
/// but never wakes the receiver, so `recv` — and therefore `run` — hangs.
|
||||
#[test]
|
||||
fn foreign_thread_send_wakes_parked_receiver() {
|
||||
let (done_tx, done_rx) = mpsc::channel();
|
||||
thread::spawn(move || {
|
||||
let rt = smarm::init(smarm::Config::exact(2));
|
||||
rt.run(|| {
|
||||
let (tx, rx) = smarm::channel::<u32>();
|
||||
// Receiver actor: parks on recv until the foreign thread sends.
|
||||
let h = smarm::spawn(move || {
|
||||
assert_eq!(rx.recv().expect("recv"), 42);
|
||||
});
|
||||
// Foreign (non-scheduler) OS thread owns the Sender and sends
|
||||
// after the receiver has parked.
|
||||
let sender = thread::spawn(move || {
|
||||
thread::sleep(SETTLE);
|
||||
tx.send(42).expect("send");
|
||||
});
|
||||
let _ = h.join();
|
||||
sender.join().expect("sender thread");
|
||||
});
|
||||
let _ = done_tx.send(());
|
||||
});
|
||||
done_rx
|
||||
.recv_timeout(WATCHDOG)
|
||||
.expect("run did not return: a foreign-thread send never woke the parked receiver");
|
||||
}
|
||||
|
||||
/// A cross-thread `request_stop` through a `RuntimeHandle` must wake and stop a
|
||||
/// parked actor. The actor parks on a long sleep (only a stop can end it); the
|
||||
/// handle is grabbed before `run` and driven from a foreign thread.
|
||||
#[test]
|
||||
fn foreign_thread_request_stop_wakes_parked_actor() {
|
||||
assert_send_sync::<smarm::RuntimeHandle>();
|
||||
|
||||
let rt = smarm::init(smarm::Config::exact(2));
|
||||
let handle = rt.handle();
|
||||
|
||||
// Foreign thread: learn the target pid from inside the run, let it park,
|
||||
// then stop it through the handle.
|
||||
let (pid_tx, pid_rx) = mpsc::channel::<smarm::Pid>();
|
||||
let stopper = thread::spawn(move || {
|
||||
let pid = pid_rx.recv().expect("pid");
|
||||
thread::sleep(SETTLE);
|
||||
handle.request_stop(pid);
|
||||
});
|
||||
|
||||
let (done_tx, done_rx) = mpsc::channel();
|
||||
thread::spawn(move || {
|
||||
rt.run(move || {
|
||||
let h = smarm::spawn(|| {
|
||||
// Parks indefinitely; only a cooperative stop unwinds it.
|
||||
smarm::sleep(Duration::from_secs(3600));
|
||||
});
|
||||
pid_tx.send(h.pid()).expect("send pid");
|
||||
let _ = h.join();
|
||||
});
|
||||
let _ = done_tx.send(());
|
||||
});
|
||||
|
||||
done_rx
|
||||
.recv_timeout(WATCHDOG)
|
||||
.expect("run did not return: a foreign-thread request_stop never woke the parked actor");
|
||||
stopper.join().expect("stopper thread");
|
||||
}
|
||||
|
||||
/// A `RuntimeHandle` held across (and beyond) a run must not keep the runtime
|
||||
/// alive or block all-done: `run` still returns, and once the `Runtime` is
|
||||
/// dropped the handle degrades to a harmless no-op (Weak lifecycle) rather than
|
||||
/// panicking or touching freed memory.
|
||||
#[test]
|
||||
fn lingering_handle_does_not_block_all_done() {
|
||||
let rt = smarm::init(smarm::Config::exact(1));
|
||||
let handle = rt.handle(); // outlives the run below
|
||||
|
||||
let (pid_tx, pid_rx) = mpsc::channel::<smarm::Pid>();
|
||||
let (done_tx, done_rx) = mpsc::channel();
|
||||
let runner = thread::spawn(move || {
|
||||
rt.run(move || {
|
||||
let h = smarm::spawn(|| {});
|
||||
pid_tx.send(h.pid()).expect("send pid");
|
||||
let _ = h.join();
|
||||
});
|
||||
// `rt` is dropped here, at the end of this thread.
|
||||
let _ = done_tx.send(());
|
||||
});
|
||||
|
||||
done_rx
|
||||
.recv_timeout(WATCHDOG)
|
||||
.expect("run did not return while a RuntimeHandle was held live");
|
||||
runner.join().expect("runner thread");
|
||||
|
||||
// Runtime is now dropped. A stop through the lingering handle must be a
|
||||
// silent no-op, not a panic or use-after-free.
|
||||
let dead_pid = pid_rx.recv().expect("pid");
|
||||
handle.request_stop(dead_pid);
|
||||
}
|
||||
+6
-6
@@ -88,7 +88,7 @@ impl GenServer for Lifecycle {
|
||||
}
|
||||
}
|
||||
|
||||
// init -> handle_call -> (drop last ref closes inbox) -> terminate.
|
||||
// init -> handle_call -> shutdown -> terminate.
|
||||
#[test]
|
||||
fn init_and_terminate_run() {
|
||||
let log = Arc::new(Mutex::new(Vec::new()));
|
||||
@@ -96,9 +96,9 @@ fn init_and_terminate_run() {
|
||||
run(move || {
|
||||
let server = start(Lifecycle { log: log2 });
|
||||
server.call(()).unwrap();
|
||||
// Dropping the only ref closes the inbox; the server breaks out of its
|
||||
// recv loop and runs terminate. run() will not return until it has.
|
||||
drop(server);
|
||||
// Refs are addresses: dropping one does not end the server. The
|
||||
// explicit close does, and waits for terminate.
|
||||
server.shutdown();
|
||||
});
|
||||
assert_eq!(*log.lock().unwrap(), vec!["init", "call", "terminate"]);
|
||||
}
|
||||
@@ -702,8 +702,8 @@ fn no_timer_survives_exit() {
|
||||
let _ = server.call(()).unwrap(); // sync: periodic armed
|
||||
smarm::sleep(Duration::from_millis(45)); // a couple of ticks
|
||||
let mon = smarm::monitor(server.pid());
|
||||
drop(server); // inbox closes → loop exits → guard drains timers
|
||||
// Clean Down ⇒ the loop returned without the no-leak assert aborting.
|
||||
server.shutdown(); // loop exits → guard drains timers
|
||||
// Clean Down ⇒ the loop returned without the no-leak assert aborting.
|
||||
assert!(mon.rx.recv().is_ok());
|
||||
let at_exit = f_read.lock().unwrap().len();
|
||||
smarm::sleep(Duration::from_millis(90)); // would be several more ticks
|
||||
|
||||
@@ -0,0 +1,243 @@
|
||||
//! gen_server lifetime is the actor's, not its refs' (OTP: a pid is an
|
||||
//! address, a process lives until it stops, is shut down, or is killed).
|
||||
//!
|
||||
//! - Dropping the last `GenServerRef` does NOT end the server. It ends via
|
||||
//! `StopHandle::stop`, `request_shutdown` / `GenServerRef::shutdown`,
|
||||
//! `request_stop`, or a handler panic.
|
||||
//! - `GenServerBuilder::named(N).run()` runs the loop inline as the *current*
|
||||
//! actor, so a server is a direct `ChildSpec` child: the supervisor's
|
||||
//! shutdown reaches it as `handle_shutdown`, a restart re-binds the name,
|
||||
//! and by-name `call`/`cast` reach whichever incarnation is live.
|
||||
|
||||
use smarm::gen_server::{
|
||||
self, GenServer, GenServerBuilder, GenServerCtx, GenServerName, ShutdownAction, StopHandle,
|
||||
};
|
||||
use smarm::registry::RegisterError;
|
||||
use smarm::supervisor::{ChildSpec, OneForOne, Restart, Shutdown};
|
||||
use smarm::{monitor, request_shutdown, run, sleep, spawn, DownReason};
|
||||
use std::sync::atomic::{AtomicUsize, Ordering};
|
||||
use std::sync::{Arc, Mutex};
|
||||
use std::time::Duration;
|
||||
|
||||
#[derive(Default, Clone)]
|
||||
struct Log(Arc<Mutex<Vec<String>>>);
|
||||
impl Log {
|
||||
fn push(&self, s: impl Into<String>) {
|
||||
self.0.lock().unwrap().push(s.into());
|
||||
}
|
||||
fn get(&self) -> Vec<String> {
|
||||
self.0.lock().unwrap().clone()
|
||||
}
|
||||
}
|
||||
|
||||
struct Counter {
|
||||
log: Log,
|
||||
n: u64,
|
||||
trap: bool,
|
||||
stop: Option<StopHandle<Counter>>,
|
||||
}
|
||||
|
||||
enum Call {
|
||||
Get,
|
||||
}
|
||||
enum Cast {
|
||||
Inc,
|
||||
Stop,
|
||||
}
|
||||
|
||||
impl GenServer for Counter {
|
||||
type Call = Call;
|
||||
type Reply = u64;
|
||||
type Cast = Cast;
|
||||
type Info = ();
|
||||
type Timer = ();
|
||||
|
||||
fn init(&mut self, ctx: &GenServerCtx<Self>) {
|
||||
if self.trap {
|
||||
ctx.trap_exit();
|
||||
}
|
||||
self.stop = Some(ctx.stop_handle());
|
||||
self.log.push("init");
|
||||
}
|
||||
fn handle_call(&mut self, Call::Get: Call) -> u64 {
|
||||
self.n
|
||||
}
|
||||
fn handle_cast(&mut self, c: Cast) {
|
||||
match c {
|
||||
Cast::Inc => self.n += 1,
|
||||
Cast::Stop => self.stop.as_ref().unwrap().stop(),
|
||||
}
|
||||
}
|
||||
fn handle_shutdown(&mut self) -> ShutdownAction {
|
||||
self.log.push("handle_shutdown");
|
||||
ShutdownAction::Exit
|
||||
}
|
||||
fn terminate(&mut self) {
|
||||
self.log.push("terminate");
|
||||
}
|
||||
}
|
||||
|
||||
fn counter(log: &Log, trap: bool) -> Counter {
|
||||
Counter {
|
||||
log: log.clone(),
|
||||
n: 0,
|
||||
trap,
|
||||
stop: None,
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Refs are addresses: dropping the last one does not end the server.
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
#[test]
|
||||
fn dropping_last_ref_does_not_end_server() {
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let srv = gen_server::start(counter(&l, true));
|
||||
let pid = srv.pid();
|
||||
srv.cast(Cast::Inc).unwrap();
|
||||
assert_eq!(srv.call(Call::Get).unwrap(), 1);
|
||||
let mon = monitor(pid);
|
||||
drop(srv);
|
||||
sleep(Duration::from_millis(30));
|
||||
assert!(
|
||||
mon.rx.try_recv().unwrap().is_none(),
|
||||
"server must outlive its last ref"
|
||||
);
|
||||
assert_eq!(l.get(), vec!["init"], "terminate must not have run");
|
||||
// Explicit teardown still works, and is what ends it.
|
||||
request_shutdown(pid);
|
||||
let down = mon.rx.recv().unwrap();
|
||||
assert_eq!(down.reason, DownReason::Exit);
|
||||
});
|
||||
assert_eq!(log.get(), vec!["init", "handle_shutdown", "terminate"]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn ref_shutdown_is_the_explicit_close() {
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let srv = gen_server::start(counter(&l, true));
|
||||
srv.call(Call::Get).unwrap(); // sync: init (and trap_exit) has run
|
||||
srv.shutdown(); // graceful, waits
|
||||
assert_eq!(l.get(), vec!["init", "handle_shutdown", "terminate"]);
|
||||
});
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn forgotten_server_is_shut_down_at_root_exit() {
|
||||
// A ref-less server is not a hung run: root exit shuts it down.
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let srv = gen_server::start(counter(&l, false));
|
||||
drop(srv);
|
||||
});
|
||||
assert_eq!(log.get(), vec!["init", "terminate"]);
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Inline run: a gen_server as a direct ChildSpec child.
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
const COUNTER: GenServerName<Counter> = GenServerName::new("lifetime-counter");
|
||||
|
||||
#[test]
|
||||
fn named_run_is_a_direct_supervised_child_and_gets_shutdown() {
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let l2 = l.clone();
|
||||
let sup = spawn(move || {
|
||||
let l3 = l2.clone();
|
||||
OneForOne::new()
|
||||
.child(
|
||||
ChildSpec::new(Restart::Permanent, move || {
|
||||
GenServerBuilder::new(counter(&l3, true))
|
||||
.named(COUNTER)
|
||||
.run()
|
||||
.expect("name free");
|
||||
})
|
||||
.shutdown(Shutdown::Infinity),
|
||||
)
|
||||
.run();
|
||||
});
|
||||
sleep(Duration::from_millis(10));
|
||||
gen_server::cast(COUNTER, Cast::Inc).unwrap();
|
||||
assert_eq!(gen_server::call(COUNTER, Call::Get).unwrap(), 1);
|
||||
request_shutdown(sup.pid());
|
||||
sup.join()
|
||||
.expect("ordered shutdown, supervisor returns normally");
|
||||
assert_eq!(l.get(), vec!["init", "handle_shutdown", "terminate"]);
|
||||
assert!(gen_server::whereis_server(COUNTER).is_none());
|
||||
});
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn named_run_child_restarts_and_rebinds_name() {
|
||||
let log = Log::default();
|
||||
let inits = Arc::new(AtomicUsize::new(0));
|
||||
let l = log.clone();
|
||||
let i = inits.clone();
|
||||
run(move || {
|
||||
let l2 = l.clone();
|
||||
let i2 = i.clone();
|
||||
let sup = spawn(move || {
|
||||
let l3 = l2.clone();
|
||||
let i3 = i2.clone();
|
||||
OneForOne::new()
|
||||
.child(ChildSpec::new(Restart::Permanent, move || {
|
||||
i3.fetch_add(1, Ordering::SeqCst);
|
||||
GenServerBuilder::new(counter(&l3, false))
|
||||
.named(COUNTER)
|
||||
.run()
|
||||
.expect("name free on (re)start");
|
||||
}))
|
||||
.run();
|
||||
});
|
||||
sleep(Duration::from_millis(10));
|
||||
gen_server::cast(COUNTER, Cast::Inc).unwrap();
|
||||
assert_eq!(gen_server::call(COUNTER, Call::Get).unwrap(), 1);
|
||||
// Normal self-exit → Permanent restarts it, fresh state, same name.
|
||||
gen_server::cast(COUNTER, Cast::Stop).unwrap();
|
||||
sleep(Duration::from_millis(30));
|
||||
assert_eq!(i.load(Ordering::SeqCst), 2, "restarted once");
|
||||
assert_eq!(gen_server::call(COUNTER, Call::Get).unwrap(), 0);
|
||||
request_shutdown(sup.pid());
|
||||
sup.join().unwrap();
|
||||
});
|
||||
assert_eq!(log.get(), vec!["init", "terminate", "init", "terminate"]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn named_run_name_clash_fails_before_init() {
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let first = GenServerBuilder::new(counter(&l, false))
|
||||
.named(COUNTER)
|
||||
.start()
|
||||
.unwrap();
|
||||
let l2 = l.clone();
|
||||
let res = Arc::new(Mutex::new(None));
|
||||
let r2 = res.clone();
|
||||
let first_pid = first.pid();
|
||||
spawn(move || {
|
||||
let r = GenServerBuilder::new(counter(&l2, false))
|
||||
.named(COUNTER)
|
||||
.run();
|
||||
*r2.lock().unwrap() = Some(r);
|
||||
})
|
||||
.join()
|
||||
.unwrap();
|
||||
assert_eq!(
|
||||
*res.lock().unwrap(),
|
||||
Some(Err(RegisterError::NameTaken { holder: first_pid }))
|
||||
);
|
||||
assert_eq!(l.get(), vec!["init"], "clashing server never ran init");
|
||||
first.shutdown();
|
||||
});
|
||||
}
|
||||
@@ -0,0 +1,222 @@
|
||||
//! gen_server graceful shutdown.
|
||||
//!
|
||||
//! - A server that does not opt in (`ctx.trap_exit()` in `init`) is stopped
|
||||
//! outright by `request_shutdown`, exactly as by `request_stop`.
|
||||
//! - A trapping server receives the request as `handle_shutdown`. The default
|
||||
//! returns `ShutdownAction::Exit`: the loop breaks and `terminate` runs on
|
||||
//! the normal (non-unwind) path, so it may block. `Continue` keeps the loop
|
||||
//! dispatching; the state later ends itself with a `StopHandle` — the only
|
||||
//! way for a gen_server to exit *normally* on its own (`request_stop` on
|
||||
//! self is an abnormal `Stopped`, which `Transient` restarts).
|
||||
//! - Other exit signals (linked peers dying) reach a trapping server via
|
||||
//! `handle_exit`.
|
||||
|
||||
use smarm::gen_server::{
|
||||
start, GenServer, GenServerBuilder, GenServerCtx, GenServerRef, ShutdownAction, StopHandle,
|
||||
};
|
||||
use smarm::supervisor::{ChildSpec, OneForOne, Restart};
|
||||
use smarm::{link, monitor, request_shutdown, run, self_pid, sleep, spawn, DownReason, ExitSignal};
|
||||
use std::sync::atomic::{AtomicBool, AtomicUsize, Ordering};
|
||||
use std::sync::{Arc, Mutex};
|
||||
use std::time::Duration;
|
||||
|
||||
#[derive(Default, Clone)]
|
||||
struct Log {
|
||||
events: Arc<Mutex<Vec<&'static str>>>,
|
||||
}
|
||||
impl Log {
|
||||
fn push(&self, e: &'static str) {
|
||||
self.events.lock().unwrap().push(e);
|
||||
}
|
||||
fn get(&self) -> Vec<&'static str> {
|
||||
self.events.lock().unwrap().clone()
|
||||
}
|
||||
}
|
||||
|
||||
/// A server with configurable shutdown behaviour.
|
||||
struct Srv {
|
||||
log: Log,
|
||||
trap: bool,
|
||||
action: ShutdownAction,
|
||||
stop: Option<StopHandle<Srv>>,
|
||||
exits: Arc<Mutex<Vec<ExitSignal>>>,
|
||||
}
|
||||
|
||||
impl Srv {
|
||||
fn new(log: &Log, trap: bool, action: ShutdownAction) -> Self {
|
||||
Srv {
|
||||
log: log.clone(),
|
||||
trap,
|
||||
action,
|
||||
stop: None,
|
||||
exits: Default::default(),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
enum Cast {
|
||||
Note(&'static str),
|
||||
StopNow,
|
||||
}
|
||||
|
||||
impl GenServer for Srv {
|
||||
type Call = ();
|
||||
type Reply = ();
|
||||
type Cast = Cast;
|
||||
type Info = ();
|
||||
type Timer = ();
|
||||
|
||||
fn init(&mut self, ctx: &GenServerCtx<Self>) {
|
||||
if self.trap {
|
||||
ctx.trap_exit();
|
||||
}
|
||||
self.stop = Some(ctx.stop_handle());
|
||||
}
|
||||
fn handle_call(&mut self, _: ()) {}
|
||||
fn handle_cast(&mut self, c: Cast) {
|
||||
match c {
|
||||
Cast::Note(s) => self.log.push(s),
|
||||
Cast::StopNow => self.stop.as_ref().unwrap().stop(),
|
||||
}
|
||||
}
|
||||
fn handle_shutdown(&mut self) -> ShutdownAction {
|
||||
self.log.push("handle_shutdown");
|
||||
self.action
|
||||
}
|
||||
fn handle_exit(&mut self, sig: ExitSignal) {
|
||||
self.log.push("handle_exit");
|
||||
self.exits.lock().unwrap().push(sig);
|
||||
}
|
||||
fn terminate(&mut self) {
|
||||
// Allowed to block on the graceful path.
|
||||
if self.trap {
|
||||
sleep(Duration::from_millis(10));
|
||||
}
|
||||
self.log.push("terminate");
|
||||
}
|
||||
}
|
||||
|
||||
fn spawn_settled<G: GenServer>(state: G) -> GenServerRef<G> {
|
||||
let r = start(state);
|
||||
sleep(Duration::from_millis(20)); // let init (trap_exit) run
|
||||
r
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn non_trapping_server_is_stopped_outright() {
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let r = spawn_settled(Srv::new(&l, false, ShutdownAction::Exit));
|
||||
let mon = monitor(r.pid());
|
||||
request_shutdown(r.pid());
|
||||
let d = mon.rx.recv().unwrap();
|
||||
assert_eq!(d.reason, DownReason::Stopped);
|
||||
});
|
||||
assert_eq!(log.get(), vec!["terminate"]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn trapping_server_exits_normally_via_handle_shutdown() {
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let r = spawn_settled(Srv::new(&l, true, ShutdownAction::Exit));
|
||||
let mon = monitor(r.pid());
|
||||
request_shutdown(r.pid());
|
||||
let d = mon.rx.recv().unwrap();
|
||||
assert_eq!(d.reason, DownReason::Exit);
|
||||
});
|
||||
assert_eq!(log.get(), vec!["handle_shutdown", "terminate"]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn continue_keeps_dispatching_until_stop_handle() {
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let r = spawn_settled(Srv::new(&l, true, ShutdownAction::Continue));
|
||||
let mon = monitor(r.pid());
|
||||
request_shutdown(r.pid());
|
||||
sleep(Duration::from_millis(20));
|
||||
r.cast(Cast::Note("after-shutdown-request")).unwrap();
|
||||
r.cast(Cast::StopNow).unwrap();
|
||||
let d = mon.rx.recv().unwrap();
|
||||
assert_eq!(d.reason, DownReason::Exit);
|
||||
});
|
||||
assert_eq!(
|
||||
log.get(),
|
||||
vec!["handle_shutdown", "after-shutdown-request", "terminate"]
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn stop_handle_is_a_normal_exit_that_transient_does_not_restart() {
|
||||
let starts = Arc::new(AtomicUsize::new(0));
|
||||
let s = starts.clone();
|
||||
run(move || {
|
||||
let s2 = s.clone();
|
||||
let sup = spawn(move || {
|
||||
let s3 = s2.clone();
|
||||
OneForOne::new()
|
||||
.child(ChildSpec::new(Restart::Transient, move || {
|
||||
s3.fetch_add(1, Ordering::SeqCst);
|
||||
let log = Log::default();
|
||||
let r = GenServerBuilder::new(Srv::new(&log, false, ShutdownAction::Exit))
|
||||
.under(self_pid())
|
||||
.start();
|
||||
r.cast(Cast::StopNow).unwrap();
|
||||
// Block until the server is gone; a bare spawn parent
|
||||
// returning would not itself end the server.
|
||||
let mon = monitor(r.pid());
|
||||
let _ = mon.rx.recv();
|
||||
}))
|
||||
.run();
|
||||
});
|
||||
sup.join().unwrap(); // returns only if the child was not restarted forever
|
||||
});
|
||||
assert_eq!(starts.load(Ordering::SeqCst), 1);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn linked_peer_death_reaches_handle_exit() {
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
let alive = Arc::new(AtomicBool::new(false));
|
||||
let a = alive.clone();
|
||||
run(move || {
|
||||
let r = spawn_settled(Srv::new(&l, true, ShutdownAction::Exit));
|
||||
let pid = r.pid();
|
||||
let peer = spawn(move || {
|
||||
link(pid);
|
||||
panic!("peer dies");
|
||||
});
|
||||
let _ = peer.join();
|
||||
sleep(Duration::from_millis(20));
|
||||
r.cast(Cast::Note("still-serving")).unwrap();
|
||||
sleep(Duration::from_millis(20));
|
||||
a.store(true, Ordering::SeqCst);
|
||||
r.shutdown(); // graceful; waits for terminate
|
||||
});
|
||||
assert!(alive.load(Ordering::SeqCst));
|
||||
assert_eq!(
|
||||
log.get(),
|
||||
vec![
|
||||
"handle_exit",
|
||||
"still-serving",
|
||||
"handle_shutdown",
|
||||
"terminate"
|
||||
]
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn gen_server_ref_shutdown_is_graceful_for_a_trapping_server() {
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let r = spawn_settled(Srv::new(&l, true, ShutdownAction::Exit));
|
||||
r.shutdown();
|
||||
});
|
||||
assert_eq!(log.get(), vec!["handle_shutdown", "terminate"]);
|
||||
}
|
||||
@@ -0,0 +1,128 @@
|
||||
//! gen_statem lifetime parity with gen_server: a machine lives until it
|
||||
//! stops, is shut down, or is killed — its refs are addresses. And
|
||||
//! `gen_statem::run_named` runs a machine inline as the current actor, so it
|
||||
//! is a direct `ChildSpec` child addressed by name.
|
||||
|
||||
use smarm::gen_statem::{self, GenStatemName, Reply};
|
||||
use smarm::supervisor::{ChildSpec, OneForOne, Restart, Shutdown};
|
||||
use smarm::{monitor, request_shutdown, run, sleep, spawn, DownReason};
|
||||
use std::sync::{Arc, Mutex};
|
||||
use std::time::Duration;
|
||||
|
||||
#[derive(Default, Clone)]
|
||||
struct Log(Arc<Mutex<Vec<&'static str>>>);
|
||||
impl Log {
|
||||
fn push(&self, e: &'static str) {
|
||||
self.0.lock().unwrap().push(e);
|
||||
}
|
||||
fn get(&self) -> Vec<&'static str> {
|
||||
self.0.lock().unwrap().clone()
|
||||
}
|
||||
}
|
||||
|
||||
#[derive(Clone, Copy, PartialEq, Eq, Debug)]
|
||||
enum S {
|
||||
On,
|
||||
}
|
||||
|
||||
struct D {
|
||||
log: Log,
|
||||
trap: bool,
|
||||
n: u64,
|
||||
}
|
||||
|
||||
enum Cast {
|
||||
Inc,
|
||||
StopNow,
|
||||
}
|
||||
enum Call {
|
||||
Get(Reply<u64>),
|
||||
}
|
||||
|
||||
smarm::gen_statem! {
|
||||
machine: Sm { state: S, data: D };
|
||||
event: Ev { cast: Cast, call: Call, info: () };
|
||||
context(data, prev, cx);
|
||||
|
||||
enter {
|
||||
S::On => { data.log.push("enter"); if data.trap { cx.trap_exit() } },
|
||||
}
|
||||
|
||||
on S::On => {
|
||||
cast Cast::Inc => { data.n += 1; prev },
|
||||
cast Cast::StopNow => stop,
|
||||
call Call::Get(r) => { r.reply(data.n); prev },
|
||||
shutdown => { data.log.push("shutdown"); cx.stop(); prev },
|
||||
state_timeout => unhandled,
|
||||
timeout _ => unhandled,
|
||||
}
|
||||
|
||||
terminate { data.log.push("terminate"); }
|
||||
}
|
||||
|
||||
fn d(log: &Log, trap: bool) -> D {
|
||||
D {
|
||||
log: log.clone(),
|
||||
trap,
|
||||
n: 0,
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn dropping_last_ref_does_not_end_machine() {
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let m = Sm::start(S::On, d(&l, true));
|
||||
let pid = m.pid();
|
||||
m.send(Ev::Cast(Cast::Inc)).unwrap();
|
||||
assert_eq!(m.call(|r| Ev::Call(Call::Get(r))).unwrap(), 1);
|
||||
let mon = monitor(pid);
|
||||
drop(m);
|
||||
sleep(Duration::from_millis(30));
|
||||
assert!(
|
||||
mon.rx.try_recv().unwrap().is_none(),
|
||||
"machine must outlive its refs"
|
||||
);
|
||||
assert_eq!(l.get(), vec!["enter"]);
|
||||
request_shutdown(pid);
|
||||
assert_eq!(mon.rx.recv().unwrap().reason, DownReason::Exit);
|
||||
});
|
||||
assert_eq!(log.get(), vec!["enter", "shutdown", "terminate"]);
|
||||
}
|
||||
|
||||
const SM: GenStatemName<Sm> = GenStatemName::new("lifetime-sm");
|
||||
|
||||
#[test]
|
||||
fn run_named_is_a_direct_supervised_child() {
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let l2 = l.clone();
|
||||
let sup = spawn(move || {
|
||||
let l3 = l2.clone();
|
||||
OneForOne::new()
|
||||
.child(
|
||||
ChildSpec::new(Restart::Permanent, move || {
|
||||
gen_statem::run_named(SM, Sm::new(S::On, d(&l3, true))).expect("name free");
|
||||
})
|
||||
.shutdown(Shutdown::Infinity),
|
||||
)
|
||||
.run();
|
||||
});
|
||||
sleep(Duration::from_millis(10));
|
||||
gen_statem::send(SM, Ev::Cast(Cast::Inc)).unwrap();
|
||||
assert_eq!(gen_statem::call(SM, |r| Ev::Call(Call::Get(r))).unwrap(), 1);
|
||||
// Normal self-exit → Permanent restart → fresh data, same name.
|
||||
gen_statem::send(SM, Ev::Cast(Cast::StopNow)).unwrap();
|
||||
sleep(Duration::from_millis(30));
|
||||
assert_eq!(gen_statem::call(SM, |r| Ev::Call(Call::Get(r))).unwrap(), 0);
|
||||
request_shutdown(sup.pid());
|
||||
sup.join().unwrap();
|
||||
assert!(gen_statem::whereis_machine(SM).is_none());
|
||||
});
|
||||
assert_eq!(
|
||||
log.get(),
|
||||
vec!["enter", "terminate", "enter", "shutdown", "terminate"]
|
||||
);
|
||||
}
|
||||
@@ -0,0 +1,219 @@
|
||||
//! gen_statem graceful shutdown — the gen_server surface, in state-machine
|
||||
//! clothes. Where gen_server routes a shutdown request to a `handle_shutdown`
|
||||
//! method, a gen_statem gets it as an **event** so it can be routed by state:
|
||||
//!
|
||||
//! - A machine that does not opt in (`cx.trap_exit()` in the initial `enter`)
|
||||
//! is stopped outright by `request_shutdown`, exactly as by `request_stop`.
|
||||
//! - A trapping machine sees the request as a `shutdown` row (a unit event
|
||||
//! like `state_timeout`). The macro's default, when a state writes no
|
||||
//! `shutdown` row, is `stop` — the loop breaks and `terminate` runs on the
|
||||
//! normal path. A row may instead transition (e.g. into a Draining state)
|
||||
//! and `stop` later from any row via the `stop` tail keyword.
|
||||
//! - Linked-peer deaths reach a trapping machine as `exit <pat>` rows; an
|
||||
//! unmatched exit is silently dropped, like an unmatched info.
|
||||
//! - `terminate { … }` is an optional macro block, run on every exit path.
|
||||
|
||||
use smarm::gen_statem;
|
||||
use smarm::gen_statem::{GenStatemRef, Reply};
|
||||
use smarm::{link, monitor, request_shutdown, run, sleep, spawn, DownReason, ExitSignal};
|
||||
use std::sync::{Arc, Mutex};
|
||||
use std::time::Duration;
|
||||
|
||||
#[derive(Default, Clone)]
|
||||
struct Log(Arc<Mutex<Vec<&'static str>>>);
|
||||
impl Log {
|
||||
fn push(&self, e: &'static str) {
|
||||
self.0.lock().unwrap().push(e);
|
||||
}
|
||||
fn get(&self) -> Vec<&'static str> {
|
||||
self.0.lock().unwrap().clone()
|
||||
}
|
||||
}
|
||||
|
||||
#[derive(Clone, Copy, PartialEq, Eq, Debug)]
|
||||
enum S {
|
||||
Idle,
|
||||
Draining,
|
||||
}
|
||||
|
||||
struct D {
|
||||
log: Log,
|
||||
trap: bool,
|
||||
exits: Vec<ExitSignal>,
|
||||
}
|
||||
|
||||
enum Cast {
|
||||
Note(&'static str),
|
||||
StopNow,
|
||||
}
|
||||
enum Call {
|
||||
Exits(Reply<usize>),
|
||||
}
|
||||
|
||||
gen_statem! {
|
||||
machine: Sm { state: S, data: D };
|
||||
event: Ev { cast: Cast, call: Call, info: () };
|
||||
context(data, prev, cx);
|
||||
|
||||
enter {
|
||||
S::Idle => if data.trap { cx.trap_exit() },
|
||||
S::Draining => { data.log.push("draining"); cx.state_timeout(Duration::from_millis(30)); },
|
||||
}
|
||||
|
||||
on S::Idle => {
|
||||
// Shutdown in Idle: go drain first, stop later.
|
||||
shutdown => S::Draining,
|
||||
cast Cast::StopNow => stop,
|
||||
state_timeout => unhandled,
|
||||
}
|
||||
|
||||
on S::Draining => {
|
||||
// Drained: end the machine normally.
|
||||
state_timeout => { data.log.push("drained"); cx.stop(); prev },
|
||||
// A second request while draining is ignored.
|
||||
shutdown => unhandled,
|
||||
cast Cast::StopNow => stop,
|
||||
}
|
||||
|
||||
on _ => {
|
||||
cast Cast::Note(s) => { data.log.push(s); prev },
|
||||
call Call::Exits(r) => { r.reply(data.exits.len()); prev },
|
||||
exit sig => { data.log.push("exit"); data.exits.push(sig); prev },
|
||||
timeout _ => unhandled,
|
||||
}
|
||||
|
||||
terminate {
|
||||
data.log.push("terminate");
|
||||
}
|
||||
}
|
||||
|
||||
/// A machine with no `shutdown` rows at all: the macro default applies.
|
||||
#[derive(Clone, Copy, PartialEq, Eq, Debug)]
|
||||
enum P {
|
||||
On,
|
||||
}
|
||||
struct PD {
|
||||
log: Log,
|
||||
}
|
||||
enum PCast {}
|
||||
enum PCall {}
|
||||
|
||||
gen_statem! {
|
||||
machine: Plain { state: P, data: PD };
|
||||
event: PEv { cast: PCast, call: PCall, info: () };
|
||||
context(data, prev, cx);
|
||||
enter { P::On => cx.trap_exit(), }
|
||||
on P::On => {
|
||||
cast _ => unhandled,
|
||||
call _ => unhandled,
|
||||
state_timeout => unhandled,
|
||||
timeout _ => unhandled,
|
||||
}
|
||||
terminate { data.log.push("terminate"); }
|
||||
}
|
||||
|
||||
fn settled(log: &Log, trap: bool) -> GenStatemRef<Sm> {
|
||||
let r = Sm::start(
|
||||
S::Idle,
|
||||
D {
|
||||
log: log.clone(),
|
||||
trap,
|
||||
exits: Vec::new(),
|
||||
},
|
||||
);
|
||||
sleep(Duration::from_millis(20)); // let on_start (trap_exit) run
|
||||
r
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn non_trapping_machine_is_stopped_outright() {
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let r = settled(&l, false);
|
||||
let mon = monitor(r.pid());
|
||||
request_shutdown(r.pid());
|
||||
assert_eq!(mon.rx.recv().unwrap().reason, DownReason::Stopped);
|
||||
});
|
||||
assert_eq!(log.get(), vec!["terminate"]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn shutdown_row_routes_by_state_and_stop_tail_exits_normally() {
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let r = settled(&l, true);
|
||||
let mon = monitor(r.pid());
|
||||
request_shutdown(r.pid());
|
||||
// The second request lands in Draining and is `unhandled` (ignored).
|
||||
sleep(Duration::from_millis(5));
|
||||
request_shutdown(r.pid());
|
||||
assert_eq!(mon.rx.recv().unwrap().reason, DownReason::Exit);
|
||||
});
|
||||
assert_eq!(log.get(), vec!["draining", "drained", "terminate"]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn default_shutdown_is_stop() {
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let r = Plain::start(P::On, PD { log: l });
|
||||
sleep(Duration::from_millis(20));
|
||||
let mon = monitor(r.pid());
|
||||
request_shutdown(r.pid());
|
||||
assert_eq!(mon.rx.recv().unwrap().reason, DownReason::Exit);
|
||||
});
|
||||
assert_eq!(log.get(), vec!["terminate"]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn stop_tail_from_a_cast_is_a_normal_exit() {
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let r = settled(&l, false);
|
||||
let mon = monitor(r.pid());
|
||||
r.send(Ev::Cast(Cast::Note("a"))).unwrap();
|
||||
r.send(Ev::Cast(Cast::StopNow)).unwrap();
|
||||
r.send(Ev::Cast(Cast::Note("after-stop"))).unwrap(); // never dispatched
|
||||
assert_eq!(mon.rx.recv().unwrap().reason, DownReason::Exit);
|
||||
});
|
||||
assert_eq!(log.get(), vec!["a", "terminate"]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn linked_peer_death_reaches_exit_row() {
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let r = settled(&l, true);
|
||||
let pid = r.pid();
|
||||
let peer = spawn(move || {
|
||||
link(pid);
|
||||
panic!("peer dies");
|
||||
});
|
||||
let _ = peer.join();
|
||||
sleep(Duration::from_millis(20));
|
||||
r.send(Ev::Cast(Cast::Note("still-running"))).unwrap();
|
||||
assert_eq!(r.call(|r| Ev::Call(Call::Exits(r))).unwrap(), 1);
|
||||
r.shutdown();
|
||||
});
|
||||
assert_eq!(
|
||||
log.get(),
|
||||
vec!["exit", "still-running", "draining", "drained", "terminate"]
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn ref_shutdown_is_graceful_and_waits() {
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let r = settled(&l, true);
|
||||
r.shutdown();
|
||||
// terminate has run by the time shutdown() returns.
|
||||
assert_eq!(l.get(), vec!["draining", "drained", "terminate"]);
|
||||
});
|
||||
}
|
||||
@@ -0,0 +1,288 @@
|
||||
//! Root exit — the run's initial actor returning means "the program is done".
|
||||
//!
|
||||
//! When the root finalizes, the runtime delivers `request_shutdown` to every
|
||||
//! **forest root**: each live actor whose parent is the run itself (a plain
|
||||
//! `spawn` from the root closure) or is already dead. Nothing below a live
|
||||
//! parent is touched directly — a supervisor gets one request and runs its
|
||||
//! own ordered shutdown per child `Shutdown` policy.
|
||||
//!
|
||||
//! - Non-trapping actors are stopped outright, exactly as by
|
||||
//! `request_shutdown` — a `spawn(|| { sleep(..); work() })` the root did
|
||||
//! not `join` does NOT get to finish. Join it, supervise it, or trap.
|
||||
//! - Trapping actors get `handle_shutdown` / an `ExitSignal{Shutdown}` and
|
||||
//! may keep running (`Continue`, drain, then stop themselves) — timers and
|
||||
//! all; the run ends when they do. There is no second, forcing sweep.
|
||||
//! - A periodic-timer daemon (the classic wedge) never blocks `run()`.
|
||||
|
||||
use smarm::gen_server::{start, GenServer, GenServerCtx, ShutdownAction, StopHandle, TimerHandle};
|
||||
use smarm::supervisor::{ChildSpec, OneForOne, Restart, Shutdown};
|
||||
use smarm::{run, sleep, spawn, trap_exit, DownReason};
|
||||
use std::sync::atomic::{AtomicBool, AtomicUsize, Ordering};
|
||||
use std::sync::{Arc, Mutex};
|
||||
use std::time::{Duration, Instant};
|
||||
|
||||
fn assert_prompt(start: Instant, what: &str) {
|
||||
assert!(
|
||||
start.elapsed() < Duration::from_secs(2),
|
||||
"{what}: run() took {:?}",
|
||||
start.elapsed()
|
||||
);
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Bare (non-gen_server) actors
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/// A non-trapping sleeper the root did not join is stopped, not waited for.
|
||||
#[test]
|
||||
fn unjoined_non_trapping_sleeper_is_stopped() {
|
||||
let finished = Arc::new(AtomicBool::new(false));
|
||||
let f = finished.clone();
|
||||
let t = Instant::now();
|
||||
run(move || {
|
||||
spawn(move || {
|
||||
sleep(Duration::from_secs(5));
|
||||
f.store(true, Ordering::SeqCst);
|
||||
});
|
||||
});
|
||||
assert_prompt(t, "sleeper");
|
||||
assert!(
|
||||
!finished.load(Ordering::SeqCst),
|
||||
"sleeper should have been stopped"
|
||||
);
|
||||
}
|
||||
|
||||
/// A trapping bare actor sees `Shutdown` from the root's exit and may keep
|
||||
/// working — here it sleeps (a timer!) after the signal, then returns. The
|
||||
/// run waits for it: no forcing sweep.
|
||||
#[test]
|
||||
fn trapping_actor_may_finish_after_shutdown_signal() {
|
||||
let finished = Arc::new(AtomicBool::new(false));
|
||||
let f = finished.clone();
|
||||
run(move || {
|
||||
spawn(move || {
|
||||
let inbox = trap_exit();
|
||||
let sig = inbox.recv().expect("shutdown signal");
|
||||
assert_eq!(sig.reason, DownReason::Shutdown);
|
||||
sleep(Duration::from_millis(100));
|
||||
f.store(true, Ordering::SeqCst);
|
||||
});
|
||||
// Trapping is a runtime opt-in: give the actor a chance to run
|
||||
// `trap_exit()`; a not-yet-run actor is non-trapping and is stopped.
|
||||
sleep(Duration::from_millis(20));
|
||||
});
|
||||
assert!(
|
||||
finished.load(Ordering::SeqCst),
|
||||
"trapping actor must be allowed to finish"
|
||||
);
|
||||
}
|
||||
|
||||
/// The classic wedge: a lazily spawned daemon that never returns on its own.
|
||||
#[test]
|
||||
fn parked_forever_daemon_does_not_block_run() {
|
||||
let t = Instant::now();
|
||||
run(|| {
|
||||
let (_tx, rx) = smarm::channel::<()>();
|
||||
spawn(move || {
|
||||
let _ = rx.recv(); // parked forever: sender is held by the root, which returns
|
||||
});
|
||||
// Leak the sender into the daemon's own scope so nothing else drops it.
|
||||
std::mem::forget(_tx);
|
||||
});
|
||||
assert_prompt(t, "daemon");
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// gen_servers
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/// A non-trapping ticker with a periodic timer: the timer wheel is never empty,
|
||||
/// and root exit must still end the run.
|
||||
struct Ticker {
|
||||
ticks: Arc<AtomicUsize>,
|
||||
}
|
||||
impl GenServer for Ticker {
|
||||
type Call = ();
|
||||
type Reply = ();
|
||||
type Cast = ();
|
||||
type Info = ();
|
||||
type Timer = ();
|
||||
fn init(&mut self, ctx: &GenServerCtx<Self>) {
|
||||
ctx.timer().tick_every(Duration::from_millis(5), ());
|
||||
}
|
||||
fn handle_call(&mut self, _: ()) {}
|
||||
fn handle_cast(&mut self, _: ()) {}
|
||||
fn handle_timer(&mut self, _: ()) {
|
||||
self.ticks.fetch_add(1, Ordering::SeqCst);
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn periodic_timer_daemon_does_not_block_run() {
|
||||
let ticks = Arc::new(AtomicUsize::new(0));
|
||||
let tk = ticks.clone();
|
||||
let t = Instant::now();
|
||||
run(move || {
|
||||
let _r = start(Ticker { ticks: tk });
|
||||
sleep(Duration::from_millis(50));
|
||||
});
|
||||
assert_prompt(t, "ticker");
|
||||
assert!(
|
||||
ticks.load(Ordering::SeqCst) >= 3,
|
||||
"ticker should have ticked"
|
||||
);
|
||||
}
|
||||
|
||||
/// A trapping server that answers `Continue`, keeps ticking on its own timer
|
||||
/// (draining), and stops itself later. Root exit must not cut it short.
|
||||
struct Drainer {
|
||||
log: Arc<Mutex<Vec<&'static str>>>,
|
||||
shutdowns: Arc<AtomicUsize>,
|
||||
ticks_after_shutdown: usize,
|
||||
stop: Option<StopHandle<Drainer>>,
|
||||
timer: Option<TimerHandle<Drainer>>,
|
||||
draining: bool,
|
||||
}
|
||||
impl GenServer for Drainer {
|
||||
type Call = ();
|
||||
type Reply = ();
|
||||
type Cast = ();
|
||||
type Info = ();
|
||||
type Timer = ();
|
||||
fn init(&mut self, ctx: &GenServerCtx<Self>) {
|
||||
ctx.trap_exit();
|
||||
self.stop = Some(ctx.stop_handle());
|
||||
self.timer = Some(ctx.timer());
|
||||
}
|
||||
fn handle_call(&mut self, _: ()) {}
|
||||
fn handle_cast(&mut self, _: ()) {}
|
||||
fn handle_shutdown(&mut self) -> ShutdownAction {
|
||||
self.shutdowns.fetch_add(1, Ordering::SeqCst);
|
||||
self.log.lock().unwrap().push("handle_shutdown");
|
||||
self.draining = true;
|
||||
self.timer
|
||||
.as_ref()
|
||||
.unwrap()
|
||||
.tick_every(Duration::from_millis(10), ());
|
||||
ShutdownAction::Continue
|
||||
}
|
||||
fn handle_timer(&mut self, _: ()) {
|
||||
if !self.draining {
|
||||
return;
|
||||
}
|
||||
self.ticks_after_shutdown += 1;
|
||||
if self.ticks_after_shutdown == 3 {
|
||||
self.log.lock().unwrap().push("drained");
|
||||
self.stop.as_ref().unwrap().stop();
|
||||
}
|
||||
}
|
||||
fn terminate(&mut self) {
|
||||
self.log.lock().unwrap().push("terminate");
|
||||
}
|
||||
}
|
||||
|
||||
fn drainer(log: &Arc<Mutex<Vec<&'static str>>>, shutdowns: &Arc<AtomicUsize>) -> Drainer {
|
||||
Drainer {
|
||||
log: log.clone(),
|
||||
shutdowns: shutdowns.clone(),
|
||||
ticks_after_shutdown: 0,
|
||||
stop: None,
|
||||
timer: None,
|
||||
draining: false,
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn trapping_server_drains_with_timers_after_root_exit() {
|
||||
let log = Arc::new(Mutex::new(Vec::new()));
|
||||
let shutdowns = Arc::new(AtomicUsize::new(0));
|
||||
let (l, s) = (log.clone(), shutdowns.clone());
|
||||
run(move || {
|
||||
let r = start(drainer(&l, &s));
|
||||
// A gen_server's lifetime is governed by its refs: dropping the last
|
||||
// one closes the inbox and ends the loop cleanly, which would cut the
|
||||
// drain short for a reason unrelated to root exit. Pin it the way a
|
||||
// registered name would.
|
||||
std::mem::forget(r);
|
||||
sleep(Duration::from_millis(20)); // let init (trap_exit) run
|
||||
});
|
||||
assert_eq!(
|
||||
*log.lock().unwrap(),
|
||||
vec!["handle_shutdown", "drained", "terminate"]
|
||||
);
|
||||
assert_eq!(shutdowns.load(Ordering::SeqCst), 1);
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Supervision trees: only forest roots are addressed
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/// The supervisor gets ONE request and runs its ordered shutdown; a trapping
|
||||
/// child under it sees exactly one `Shutdown` — from the supervisor, not a
|
||||
/// second one from the runtime — and is allowed to finish its drain (a sleep,
|
||||
/// i.e. a timer) under `Shutdown::Infinity`.
|
||||
#[test]
|
||||
fn supervised_children_are_shut_down_only_via_their_supervisor() {
|
||||
let signals = Arc::new(AtomicUsize::new(0));
|
||||
let drained = Arc::new(AtomicBool::new(false));
|
||||
let sup_returned = Arc::new(AtomicBool::new(false));
|
||||
let (sg, dr, sr) = (signals.clone(), drained.clone(), sup_returned.clone());
|
||||
run(move || {
|
||||
spawn(move || {
|
||||
OneForOne::new()
|
||||
.child(
|
||||
ChildSpec::new(Restart::Permanent, move || {
|
||||
let inbox = trap_exit();
|
||||
while let Ok(sig) = inbox.recv() {
|
||||
if sig.reason == DownReason::Shutdown {
|
||||
sg.fetch_add(1, Ordering::SeqCst);
|
||||
sleep(Duration::from_millis(100));
|
||||
// A late second signal would land here.
|
||||
while let Ok(Some(sig)) = inbox.try_recv() {
|
||||
if sig.reason == DownReason::Shutdown {
|
||||
sg.fetch_add(1, Ordering::SeqCst);
|
||||
}
|
||||
}
|
||||
dr.store(true, Ordering::SeqCst);
|
||||
return;
|
||||
}
|
||||
}
|
||||
})
|
||||
.shutdown(Shutdown::Infinity),
|
||||
)
|
||||
.run();
|
||||
sr.store(true, Ordering::SeqCst);
|
||||
});
|
||||
sleep(Duration::from_millis(30)); // let the tree settle
|
||||
});
|
||||
assert!(
|
||||
sup_returned.load(Ordering::SeqCst),
|
||||
"supervisor should return normally"
|
||||
);
|
||||
assert!(
|
||||
drained.load(Ordering::SeqCst),
|
||||
"child should finish its drain"
|
||||
);
|
||||
assert_eq!(signals.load(Ordering::SeqCst), 1);
|
||||
}
|
||||
|
||||
/// A supervised non-trapping child under `Shutdown::Timeout` is stopped by
|
||||
/// the supervisor's policy, and the run ends promptly.
|
||||
#[test]
|
||||
fn supervisor_tree_is_torn_down_promptly_on_root_exit() {
|
||||
let t = Instant::now();
|
||||
run(|| {
|
||||
spawn(|| {
|
||||
OneForOne::new()
|
||||
.child(
|
||||
ChildSpec::new(Restart::Permanent, || loop {
|
||||
sleep(Duration::from_millis(5));
|
||||
})
|
||||
.shutdown(Shutdown::Timeout(Duration::from_millis(50))),
|
||||
)
|
||||
.run();
|
||||
});
|
||||
sleep(Duration::from_millis(30));
|
||||
});
|
||||
assert_prompt(t, "tree");
|
||||
}
|
||||
@@ -0,0 +1,36 @@
|
||||
//! Under `smarm-trace`, every actor the root-exit sweep reaches is recorded
|
||||
//! as a `root_sweep` event — the way to *see* unsupervised leftovers. One test
|
||||
//! per binary: the trace file is process-global.
|
||||
#![cfg(feature = "smarm-trace")]
|
||||
|
||||
use smarm::gen_server::{self, GenServer, GenServerCtx};
|
||||
use smarm::run;
|
||||
|
||||
struct Quiet;
|
||||
impl GenServer for Quiet {
|
||||
type Call = ();
|
||||
type Reply = ();
|
||||
type Cast = ();
|
||||
type Info = ();
|
||||
type Timer = ();
|
||||
fn init(&mut self, _: &GenServerCtx<Self>) {}
|
||||
fn handle_call(&mut self, _: ()) {}
|
||||
fn handle_cast(&mut self, _: ()) {}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn forgotten_server_shows_up_as_root_sweep() {
|
||||
let path = std::env::temp_dir().join(format!("smarm_root_sweep_{}.json", std::process::id()));
|
||||
std::env::set_var("SMARM_TRACE_FILE", &path);
|
||||
run(|| {
|
||||
let srv = gen_server::start(Quiet);
|
||||
srv.call(()).unwrap();
|
||||
drop(srv); // forgotten: nobody supervises it, nobody holds it
|
||||
});
|
||||
let trace = std::fs::read_to_string(&path).expect("trace file written");
|
||||
let _ = std::fs::remove_file(&path);
|
||||
assert!(
|
||||
trace.contains("root_sweep stopped"),
|
||||
"expected a root_sweep line for the non-trapping leftover; got:\n{trace}"
|
||||
);
|
||||
}
|
||||
@@ -0,0 +1,122 @@
|
||||
//! Graceful shutdown — `request_shutdown` (OTP `exit(Pid, shutdown)`).
|
||||
//!
|
||||
//! `request_stop` is `exit(Pid, kill)`: an uncatchable unwind at the target's
|
||||
//! next observation point. `request_shutdown` is the polite form:
|
||||
//! - a target that is NOT trapping exits is stopped exactly as by
|
||||
//! `request_stop` (OTP's rule: don't trap, you die);
|
||||
//! - a target that IS trapping receives an `ExitSignal { reason: Shutdown }`
|
||||
//! on its trap inbox and keeps running — it is expected to wind down and
|
||||
//! exit normally on its own.
|
||||
|
||||
use smarm::{monitor, request_shutdown, run, self_pid, sleep, spawn, trap_exit, DownReason, Pid};
|
||||
use std::sync::atomic::{AtomicBool, Ordering};
|
||||
use std::sync::{mpsc, Arc};
|
||||
use std::thread;
|
||||
use std::time::Duration;
|
||||
|
||||
const WATCHDOG: Duration = Duration::from_secs(10);
|
||||
|
||||
#[test]
|
||||
fn request_shutdown_stops_a_non_trapping_actor() {
|
||||
run(|| {
|
||||
let h = spawn(|| sleep(Duration::from_secs(3600)));
|
||||
let mon = monitor(h.pid());
|
||||
request_shutdown(h.pid());
|
||||
let down = mon.rx.recv().expect("down");
|
||||
assert_eq!(down.reason, DownReason::Stopped);
|
||||
});
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn request_shutdown_is_a_message_to_a_trapping_actor() {
|
||||
let unwound = Arc::new(AtomicBool::new(false));
|
||||
let u = unwound.clone();
|
||||
run(move || {
|
||||
struct Unwound(Arc<AtomicBool>);
|
||||
impl Drop for Unwound {
|
||||
fn drop(&mut self) {
|
||||
if std::thread::panicking() {
|
||||
self.0.store(true, Ordering::SeqCst);
|
||||
}
|
||||
}
|
||||
}
|
||||
let (tx, rx) = smarm::channel::<(Pid, DownReason)>();
|
||||
let (ready_tx, ready_rx) = smarm::channel::<()>();
|
||||
let h = spawn(move || {
|
||||
let _g = Unwound(u);
|
||||
let inbox = trap_exit();
|
||||
let _ = ready_tx.send(());
|
||||
let sig = inbox.recv().expect("exit signal");
|
||||
let _ = tx.send((sig.from, sig.reason));
|
||||
// Keep doing work after the request: shutdown is advisory.
|
||||
sleep(Duration::from_millis(20));
|
||||
});
|
||||
// Trapping is set by the target itself; a request that beats it is a
|
||||
// plain stop (same window as OTP's exit-before-process_flag).
|
||||
ready_rx.recv().expect("ready");
|
||||
let me = self_pid();
|
||||
let mon = monitor(h.pid());
|
||||
request_shutdown(h.pid());
|
||||
let (from, reason) = rx.recv().expect("relayed");
|
||||
assert_eq!(from, me);
|
||||
assert_eq!(reason, DownReason::Shutdown);
|
||||
let down = mon.rx.recv().expect("down");
|
||||
assert_eq!(
|
||||
down.reason,
|
||||
DownReason::Exit,
|
||||
"target exited normally, not stopped"
|
||||
);
|
||||
});
|
||||
assert!(
|
||||
!unwound.load(Ordering::SeqCst),
|
||||
"trapping target must not be unwound"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn request_shutdown_on_dead_pid_is_a_no_op() {
|
||||
run(|| {
|
||||
let h = spawn(|| {});
|
||||
let pid = h.pid();
|
||||
let _ = h.join();
|
||||
request_shutdown(pid); // must not panic
|
||||
});
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn handle_request_shutdown_from_foreign_thread() {
|
||||
let rt = smarm::init(smarm::Config::exact(2));
|
||||
let handle = rt.handle();
|
||||
|
||||
let (pid_tx, pid_rx) = mpsc::channel::<Pid>();
|
||||
let requester = thread::spawn(move || {
|
||||
let pid = pid_rx.recv().expect("pid");
|
||||
thread::sleep(Duration::from_millis(50));
|
||||
handle.request_shutdown(pid);
|
||||
});
|
||||
|
||||
let (done_tx, done_rx) = mpsc::channel();
|
||||
thread::spawn(move || {
|
||||
rt.run(move || {
|
||||
let (tx, rx) = smarm::channel::<DownReason>();
|
||||
let (ready_tx, ready_rx) = smarm::channel::<()>();
|
||||
let h = spawn(move || {
|
||||
let inbox = trap_exit();
|
||||
let _ = ready_tx.send(());
|
||||
let sig = inbox.recv().expect("exit signal");
|
||||
let _ = tx.send(sig.reason);
|
||||
});
|
||||
ready_rx.recv().expect("ready");
|
||||
pid_tx.send(h.pid()).expect("send pid");
|
||||
let reason = rx.recv().expect("relayed");
|
||||
assert_eq!(reason, DownReason::Shutdown);
|
||||
let _ = h.join();
|
||||
});
|
||||
let _ = done_tx.send(());
|
||||
});
|
||||
|
||||
done_rx
|
||||
.recv_timeout(WATCHDOG)
|
||||
.expect("run did not return: foreign-thread request_shutdown never reached the target");
|
||||
requester.join().expect("requester thread");
|
||||
}
|
||||
@@ -0,0 +1,303 @@
|
||||
//! Supervisor shutdown — the OTP child-spec `shutdown` policy.
|
||||
//!
|
||||
//! A supervisor traps exits. A `request_shutdown` reaching it (from its parent
|
||||
//! supervisor, or from the app via `request_shutdown`/`RuntimeHandle`) runs
|
||||
//! the ordered shutdown: children are stopped in reverse start order, each
|
||||
//! per its `Shutdown` policy — `request_shutdown`, wait up to the timeout for
|
||||
//! its termination signal, `request_stop` if it overstays — and then `run()`
|
||||
//! returns normally. Every supervisor-initiated child stop (ordered shutdown,
|
||||
//! OneForAll/RestForOne sibling cycling) goes through the same policy.
|
||||
//!
|
||||
//! A *hard* `request_stop` on a supervisor unwinds it; a drop guard then
|
||||
//! hard-stops its live children so the subtree is never orphaned.
|
||||
|
||||
use smarm::supervisor::{ChildSpec, OneForOne, Restart, Shutdown, Strategy};
|
||||
use smarm::{
|
||||
monitor, request_shutdown, request_stop, run, sleep, spawn, trap_exit, DownReason, JoinHandle,
|
||||
};
|
||||
use std::sync::atomic::{AtomicBool, AtomicUsize, Ordering};
|
||||
use std::sync::{Arc, Mutex};
|
||||
use std::time::{Duration, Instant};
|
||||
|
||||
/// A child that traps exits, records the order it was shut down in, and exits
|
||||
/// normally on the request (after `delay`). Ignores the request if `comply`
|
||||
/// is false — a straggler that must be hard-stopped.
|
||||
fn polite_child(
|
||||
tag: usize,
|
||||
log: &Arc<Mutex<Vec<usize>>>,
|
||||
delay: Duration,
|
||||
comply: bool,
|
||||
) -> impl Fn() + Send + Sync + 'static {
|
||||
let log = log.clone();
|
||||
move || {
|
||||
let inbox = trap_exit();
|
||||
loop {
|
||||
let sig = match inbox.recv() {
|
||||
Ok(s) => s,
|
||||
Err(_) => return,
|
||||
};
|
||||
if sig.reason == DownReason::Shutdown {
|
||||
log.lock().unwrap().push(tag);
|
||||
if comply {
|
||||
sleep(delay);
|
||||
return;
|
||||
}
|
||||
// Not complying: keep running until hard-stopped.
|
||||
loop {
|
||||
sleep(Duration::from_millis(5));
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Spawn `sup`, let its children reach `trap_exit`, return the handle.
|
||||
fn spawn_settled(sup: OneForOne) -> JoinHandle {
|
||||
let h = spawn(move || sup.run());
|
||||
sleep(Duration::from_millis(30));
|
||||
h
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn shutdown_stops_children_in_reverse_order_and_returns_normally() {
|
||||
let log = Arc::new(Mutex::new(Vec::new()));
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let sup = OneForOne::new()
|
||||
.child(ChildSpec::new(
|
||||
Restart::Permanent,
|
||||
polite_child(1, &l, Duration::ZERO, true),
|
||||
))
|
||||
.child(ChildSpec::new(
|
||||
Restart::Permanent,
|
||||
polite_child(2, &l, Duration::ZERO, true),
|
||||
))
|
||||
.child(ChildSpec::new(
|
||||
Restart::Permanent,
|
||||
polite_child(3, &l, Duration::ZERO, true),
|
||||
));
|
||||
let h = spawn_settled(sup);
|
||||
let mon = monitor(h.pid());
|
||||
request_shutdown(h.pid());
|
||||
let down = mon.rx.recv().expect("down");
|
||||
assert_eq!(
|
||||
down.reason,
|
||||
DownReason::Exit,
|
||||
"supervisor exits normally after shutdown"
|
||||
);
|
||||
});
|
||||
assert_eq!(*log.lock().unwrap(), vec![3, 2, 1]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn non_trapping_child_is_simply_stopped() {
|
||||
let dropped = Arc::new(AtomicBool::new(false));
|
||||
let d = dropped.clone();
|
||||
run(move || {
|
||||
struct G(Arc<AtomicBool>);
|
||||
impl Drop for G {
|
||||
fn drop(&mut self) {
|
||||
self.0.store(true, Ordering::SeqCst);
|
||||
}
|
||||
}
|
||||
let sup = OneForOne::new().child(ChildSpec::new(Restart::Permanent, move || {
|
||||
let _g = G(d.clone());
|
||||
loop {
|
||||
sleep(Duration::from_millis(5));
|
||||
}
|
||||
}));
|
||||
let h = spawn_settled(sup);
|
||||
request_shutdown(h.pid());
|
||||
h.join().expect("sup");
|
||||
});
|
||||
assert!(dropped.load(Ordering::SeqCst));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn straggler_is_hard_stopped_after_timeout() {
|
||||
let log = Arc::new(Mutex::new(Vec::new()));
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let sup = OneForOne::new().child(
|
||||
ChildSpec::new(
|
||||
Restart::Permanent,
|
||||
polite_child(1, &l, Duration::ZERO, false),
|
||||
)
|
||||
.shutdown(Shutdown::Timeout(Duration::from_millis(50))),
|
||||
);
|
||||
let h = spawn_settled(sup);
|
||||
let t0 = Instant::now();
|
||||
request_shutdown(h.pid());
|
||||
h.join().expect("sup");
|
||||
let took = t0.elapsed();
|
||||
assert!(
|
||||
took >= Duration::from_millis(50),
|
||||
"returned before the grace period: {took:?}"
|
||||
);
|
||||
assert!(
|
||||
took < Duration::from_secs(2),
|
||||
"did not fall back to a hard stop: {took:?}"
|
||||
);
|
||||
});
|
||||
assert_eq!(
|
||||
*log.lock().unwrap(),
|
||||
vec![1],
|
||||
"the straggler did receive the request"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn infinity_waits_for_a_slow_but_compliant_child() {
|
||||
let log = Arc::new(Mutex::new(Vec::new()));
|
||||
let l = log.clone();
|
||||
let finished = Arc::new(AtomicBool::new(false));
|
||||
let f = finished.clone();
|
||||
run(move || {
|
||||
let f2 = f.clone();
|
||||
let l2 = l.clone();
|
||||
let sup = OneForOne::new().child(
|
||||
ChildSpec::new(Restart::Permanent, move || {
|
||||
let inbox = trap_exit();
|
||||
let _ = inbox.recv();
|
||||
l2.lock().unwrap().push(1);
|
||||
sleep(Duration::from_millis(150));
|
||||
f2.store(true, Ordering::SeqCst); // only reached if not hard-stopped
|
||||
})
|
||||
.shutdown(Shutdown::Infinity),
|
||||
);
|
||||
let h = spawn_settled(sup);
|
||||
request_shutdown(h.pid());
|
||||
h.join().expect("sup");
|
||||
});
|
||||
assert!(
|
||||
finished.load(Ordering::SeqCst),
|
||||
"Infinity must not hard-stop a compliant child"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn brutal_kill_skips_the_request() {
|
||||
let log = Arc::new(Mutex::new(Vec::new()));
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let sup = OneForOne::new().child(
|
||||
ChildSpec::new(
|
||||
Restart::Permanent,
|
||||
polite_child(1, &l, Duration::ZERO, true),
|
||||
)
|
||||
.shutdown(Shutdown::BrutalKill),
|
||||
);
|
||||
let h = spawn_settled(sup);
|
||||
request_shutdown(h.pid());
|
||||
h.join().expect("sup");
|
||||
});
|
||||
assert!(
|
||||
log.lock().unwrap().is_empty(),
|
||||
"a BrutalKill child never sees the request"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn hard_stop_of_supervisor_does_not_orphan_children() {
|
||||
let alive = Arc::new(AtomicUsize::new(0));
|
||||
let a = alive.clone();
|
||||
run(move || {
|
||||
struct Alive(Arc<AtomicUsize>);
|
||||
impl Drop for Alive {
|
||||
fn drop(&mut self) {
|
||||
self.0.fetch_sub(1, Ordering::SeqCst);
|
||||
}
|
||||
}
|
||||
let mk = |a: Arc<AtomicUsize>| {
|
||||
move || {
|
||||
a.fetch_add(1, Ordering::SeqCst);
|
||||
let _g = Alive(a.clone());
|
||||
loop {
|
||||
sleep(Duration::from_millis(5));
|
||||
}
|
||||
}
|
||||
};
|
||||
let sup = OneForOne::new()
|
||||
.child(ChildSpec::new(Restart::Permanent, mk(a.clone())))
|
||||
.child(ChildSpec::new(Restart::Permanent, mk(a.clone())));
|
||||
let h = spawn_settled(sup);
|
||||
assert_eq!(a.load(Ordering::SeqCst), 2);
|
||||
let mon = monitor(h.pid());
|
||||
request_stop(h.pid());
|
||||
let _ = mon.rx.recv();
|
||||
sleep(Duration::from_millis(50));
|
||||
assert_eq!(
|
||||
a.load(Ordering::SeqCst),
|
||||
0,
|
||||
"children orphaned by a hard supervisor stop"
|
||||
);
|
||||
});
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn nested_shutdown_reaches_grandchildren() {
|
||||
let log = Arc::new(Mutex::new(Vec::new()));
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let l_inner = l.clone();
|
||||
let inner = move || {
|
||||
OneForOne::new()
|
||||
.child(ChildSpec::new(
|
||||
Restart::Permanent,
|
||||
polite_child(10, &l_inner, Duration::ZERO, true),
|
||||
))
|
||||
.child(ChildSpec::new(
|
||||
Restart::Permanent,
|
||||
polite_child(11, &l_inner, Duration::ZERO, true),
|
||||
))
|
||||
.run()
|
||||
};
|
||||
let sup = OneForOne::new()
|
||||
.child(ChildSpec::new(
|
||||
Restart::Permanent,
|
||||
polite_child(1, &l, Duration::ZERO, true),
|
||||
))
|
||||
.child(ChildSpec::new(Restart::Permanent, inner).shutdown(Shutdown::Infinity));
|
||||
let h = spawn_settled(sup);
|
||||
request_shutdown(h.pid());
|
||||
h.join().expect("sup");
|
||||
});
|
||||
assert_eq!(*log.lock().unwrap(), vec![11, 10, 1]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn sibling_cycling_uses_graceful_shutdown() {
|
||||
// OneForAll: when child A dies, sibling B (trapping) must receive a
|
||||
// Shutdown request rather than a bare stop.
|
||||
let log = Arc::new(Mutex::new(Vec::new()));
|
||||
let l = log.clone();
|
||||
let a_runs = Arc::new(AtomicUsize::new(0));
|
||||
let ar = a_runs.clone();
|
||||
run(move || {
|
||||
let ar2 = ar.clone();
|
||||
let sup = OneForOne::new()
|
||||
.strategy(Strategy::OneForAll)
|
||||
.intensity(5, Duration::from_secs(60))
|
||||
.child(ChildSpec::new(Restart::Transient, move || {
|
||||
let n = ar2.fetch_add(1, Ordering::SeqCst) + 1;
|
||||
sleep(Duration::from_millis(30));
|
||||
if n == 1 {
|
||||
panic!("first run dies");
|
||||
}
|
||||
// Second run: park until shut down.
|
||||
let inbox = trap_exit();
|
||||
let _ = inbox.recv();
|
||||
}))
|
||||
.child(ChildSpec::new(
|
||||
Restart::Permanent,
|
||||
polite_child(2, &l, Duration::ZERO, true),
|
||||
));
|
||||
let h = spawn(move || sup.run());
|
||||
sleep(Duration::from_millis(150));
|
||||
request_shutdown(h.pid());
|
||||
h.join().expect("sup");
|
||||
});
|
||||
// B was shut down once by the cycle and once by the final shutdown.
|
||||
assert_eq!(*log.lock().unwrap(), vec![2, 2]);
|
||||
assert_eq!(a_runs.load(Ordering::SeqCst), 2);
|
||||
}
|
||||
Reference in New Issue
Block a user