Compare commits
32
Commits
d9addeba5e
...
v0.7.0
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
741c10337b | ||
|
|
e570138da5 | ||
|
|
415effb2e9 | ||
|
|
849a424c8e | ||
|
|
6ceb138f5f | ||
|
|
250f31265b | ||
|
|
9c8f59ca53 | ||
|
|
1002777ef3 | ||
|
|
8f2d513940 | ||
|
|
ca1c98336e | ||
|
|
95306c7f60 | ||
|
|
1262cc30e3 | ||
|
|
461fe4b768 | ||
|
|
b937f1f50f | ||
|
|
301e3463e3 | ||
|
|
410ba33d82 | ||
|
|
5fd8aecf55 | ||
|
|
7d8b9e0310 | ||
|
|
8225716b11 | ||
|
|
3cb64eefc2 | ||
|
|
0fe052bc7e | ||
|
|
a03a7ca01e | ||
|
|
d4839f1d81 | ||
|
|
2854b560d6 | ||
|
|
7b026cfe56 | ||
|
|
006a3283e7 | ||
|
|
8c764e9169 | ||
|
|
41b9d6d056 | ||
|
|
dd845f22fe | ||
|
|
36a0a9832d | ||
|
|
feda6517e5 | ||
|
|
8625ae4c35 |
+17
-1
@@ -2,7 +2,23 @@
|
||||
# smarm pre-commit gate: clippy the library (src/) with warnings as errors.
|
||||
# unwrap_used / expect_used are denied (Cargo.toml [lints.clippy]): library
|
||||
# code must not hide a panic behind unwrap/expect. Tests/examples are not gated.
|
||||
#
|
||||
# Toolchain resolution: prefer an installed cargo-clippy; on machines whose
|
||||
# rust comes without the clippy component (e.g. NixOS home-manager), fall
|
||||
# back to an ephemeral nix-shell toolchain. The fallback uses its own target
|
||||
# dir (target/clippy) because the shell's rustc version may differ from the
|
||||
# default toolchain's — mixed-compiler artifacts in one target dir are an
|
||||
# E0514 hard error. MSRV (Cargo.toml rust-version) keeps the older shell
|
||||
# toolchain a legitimate gate.
|
||||
set -eu
|
||||
[ -f "$HOME/.cargo/env" ] && . "$HOME/.cargo/env"
|
||||
cd "$(git rev-parse --show-toplevel)"
|
||||
cargo clippy --lib -- -D warnings
|
||||
if cargo clippy --version >/dev/null 2>&1; then
|
||||
cargo clippy --lib -- -D warnings
|
||||
elif command -v nix-shell >/dev/null 2>&1; then
|
||||
nix-shell -p clippy -p cargo -p rustc \
|
||||
--run 'CARGO_TARGET_DIR=target/clippy cargo clippy --lib -- -D warnings'
|
||||
else
|
||||
echo "pre-commit: cargo clippy unavailable and no nix-shell fallback" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
+4
-1
@@ -1,6 +1,6 @@
|
||||
[package]
|
||||
name = "smarm"
|
||||
version = "0.4.0"
|
||||
version = "0.7.0"
|
||||
edition = "2021"
|
||||
rust-version = "1.95"
|
||||
|
||||
@@ -39,6 +39,9 @@ rq-mutex = []
|
||||
rq-mpmc = []
|
||||
rq-striped = []
|
||||
|
||||
[build-dependencies]
|
||||
cc = "1"
|
||||
|
||||
[dependencies]
|
||||
libc = "0.2"
|
||||
|
||||
|
||||
@@ -1,35 +1,38 @@
|
||||
# smarm
|
||||
|
||||
> SMARM — Smarm, Marks Actor Runtime Machinery. A proof-of-concept green-thread actor runtime for Rust.
|
||||
> SMARM: Smarm, Marks Actor Runtime Machinery. A proof-of-concept green-thread actor runtime for Rust.
|
||||
|
||||
Implements the core ideas in [`Achitecture.md`](.docs/Architecture.md): green-thread actors on a
|
||||
shared heap, scheduled cooperatively, communicating only by `Send` messages.
|
||||
Erlang's isolation model without Erlang's copying GC, Rust's zero-copy
|
||||
ownership transfers without async's function colouring.
|
||||
SMARM is my attempt to implement the erlang/OTP philosophy in the Rust programming language. This has yielded a fault-tolerant, fast, and scalable runtime. This runtime allows the creation of asynchronous applications in Rust without the function coloring associated with the async/await system. It encourages the writing of simple, synchronous code, and largely elides the need for lifetime annotations.
|
||||
|
||||
The scheduler is multi-threaded — one OS thread per available CPU, all drawing
|
||||
from a shared run queue. The single-threaded `run()` entry point is kept as a
|
||||
convenience wrapper around `runtime::init(Config::exact(1)).run(f)`.
|
||||
|
||||
## What's here
|
||||
## Overview
|
||||
|
||||
SMARM implements green-thread actors on a shared heap, communicating only by `Send` messages. By sharing the heap, SMARM avoids the copying overhead of Erlang, which is safe to do due to Rust's borrow checker.
|
||||
|
||||
On top of the core runtime mechanics, SMARM also provides a library of primitives for making applications closely inspired by erlang/OTP. This includes generic servers (gen_servers), generic state machines (gen_statem), and supervision trees.
|
||||
|
||||
Supervision trees are the core primitive to allow your application to survive an unexpected panic. Supervisors are processes dedicated to monitoring other processes, which can restart these should they fail. This means that when set up properly an application may 'self-heal' when encountering unforeseen circumstances.
|
||||
|
||||
SMARM is not cooperatively scheduled; it uses preemption. This means a heavy task will not starve out other lighter tasks. Everything will make steady progress, which translates to very beneficial behaviour under (over)load: average latency goes up, but tail latency does not blow up.
|
||||
|
||||
To help diagnose these unforeseen circumstances, smarm may be compiled with its `tracing` feature, which emits a full trace using [Perfetto](https://perfetto.dev/).
|
||||
|
||||
Should you want to optimize your application, SMARM is unusually well poised to help. As the runtime functionally controls time, SMARM comes with a built in causal profiler, under the `causal` feature.
|
||||
|
||||
I also built a Phoenix-Framework inspired HTTP 1.1 library on top of SMARM called [URUS](https://git.kalsbeek.dev/Markk116/urus), which implements Pub/Sub, Channels, and basic amenities like Websockets and Server Sent Events.
|
||||
|
||||
## Limitations
|
||||
|
||||
This runtime requires naked assembly to function, and has thus far only been implemented for x86-64 assembly. It expects an operating system that supports virtual address space, and is therefore not (yet) suited for embedded targets. The IO implementation is currently based around the Linux kernel's `epoll` mechanism, meaning it requires a (GNU+)Linux distribution to run.
|
||||
|
||||
The preemption mechanism works by wrapping the memory allocator and checking how many CPU cycles you have used compared to your timeslice budget. This allows preemption to fire in most normal code, but tight zero-allocation loops do not get caught and require manual insertion of `check!()` if you want preemption to function.
|
||||
|
||||
This library is still in its early stages, and while I try my best with loom and tests, stable operation cannot be guaranteed. Therefore it is not (yet) recommended for production use.
|
||||
|
||||
At this moment, stack memory for each green thread is capped. Uncapping this may lead to performance benefits for deeply recursive algorithms that in a traditional async runtime might require pointer-chases through the heap. This is as yet unrealised.
|
||||
|
||||
At this stage, the codebase is largely LLM-generated, which is obvious if you start to read through the internals. While I did the design, and I keep the LLM under tight rein, the codebase is not in a state that I am very happy with. This also goes for the documentation.
|
||||
|
||||
| Module | What it does |
|
||||
|--------------|------------------------------------------------------------------------|
|
||||
| `stack` | `mmap`'d growable stack with guard page; SIGSEGV on overflow |
|
||||
| `context` | `#[naked]` x86-64 context-switch shims, callee-saved regs only |
|
||||
| `preempt` | Allocator-driven preemption; `check!()` macro for no-alloc loops |
|
||||
| `pid` | `(index, generation)` PIDs; stale handles are detectable, not silent |
|
||||
| `actor` | Trampoline + `catch_unwind` boundary at the actor entry point |
|
||||
| `scheduler` | Run queue, slot table, spawn/join, parking, idle path |
|
||||
| `channel` | Unbounded MPSC channel; `recv` parks the actor; `recv_timeout` bounds it; `select`/`select_timeout` park on many receivers at once (ready-index, priority order) |
|
||||
| `mutex` | `Mutex<T>` with mandatory timeout; FIFO waiters; parks the green thread |
|
||||
| `timer` | Min-heap of `(deadline, reason)`; `Sleep` and `WaitTimeout` reasons |
|
||||
| `io` | `block_on_io` for blocking work; `wait_readable`/`wait_writable` + `read`/`write` via epoll |
|
||||
| `supervisor` | `Signal::Exit`/`Panic`/`Stopped` funnelled to a parent; `OneForOne`/`OneForAll`/`RestForOne` strategies + restart-intensity cap |
|
||||
| `monitor` | `monitor(pid)` → `Monitor { id, target, rx }`; one-shot `Down` via `rx`; `demonitor(&m)` tears one registration down; unidirectional death notice |
|
||||
| `link` | bidirectional `link`/`unlink`; abnormal death propagates (cooperative stop, or an `ExitSignal` message under `trap_exit`) |
|
||||
| `gen_server` | `call`/`call_timeout` (sync request-reply) / `cast` (async) over one inbox; `handle_info` over static info arms + `handle_down` via `Watcher`-fed monitors, selected ahead of the inbox; `ServerRef`/`ServerBuilder` + `init`/`terminate` hooks; server-down via channel closure |
|
||||
| `registry` | `register`/`whereis`/`name_of`: name ↔ pid bimap; lazy generation-checked cleanup |
|
||||
|
||||
## Quick taste
|
||||
|
||||
@@ -51,6 +54,29 @@ run(|| {
|
||||
});
|
||||
```
|
||||
|
||||
## Stopping actors
|
||||
|
||||
Two strengths, as in OTP. `request_stop(pid)` is `exit(Pid, kill)`: a cooperative
|
||||
hard stop, unwinding at the actor's next observation point. `request_shutdown(pid)`
|
||||
is `exit(Pid, shutdown)`: an actor that traps exits (`trap_exit()`, or
|
||||
`ctx.trap_exit()` in a gen_server / `cx.trap_exit()` in a gen_statem) receives it
|
||||
as a signal — `handle_shutdown` / a `shutdown` row — and may drain before stopping
|
||||
itself; one that does not trap is stopped outright. Supervisors trap:
|
||||
`request_shutdown(sup)` tears the tree down top-down, each child per its
|
||||
`ChildSpec` `Shutdown` policy (`Timeout(d)`, `Infinity`, `BrutalKill`). The run's
|
||||
root actor returning means "the program is done": every top-level actor gets a
|
||||
`request_shutdown`, and `run()` returns when they are gone. From outside the
|
||||
runtime (a signal thread), `Runtime::handle().request_shutdown(pid)` does the
|
||||
same. `examples/graceful_shutdown.rs` shows all of it.
|
||||
|
||||
A gen_server or gen_statem lives until it stops, is shut down, or is killed; its
|
||||
refs are addresses — dropping them never ends it (a forgotten one is swept at
|
||||
root exit; with `--features smarm-trace` each such sweep is a `root_sweep`
|
||||
trace line). The supervised shape is `GenServerBuilder::named(N).run()` /
|
||||
`gen_statem::run_named(N, m)`: the server runs inline as the `ChildSpec` child
|
||||
itself, so the supervisor's shutdown reaches it directly, a restart re-binds the
|
||||
name, and the program addresses it by name.
|
||||
|
||||
## Layout
|
||||
|
||||
```
|
||||
@@ -67,12 +93,9 @@ benches/
|
||||
|
||||
## Building and running
|
||||
|
||||
Standard Cargo. Requires Rust 1.95 or newer (the `#[naked]` attribute went stable
|
||||
in 1.88; we use a few unrelated post-1.88 features). `master` is x86-64 Linux
|
||||
only. An experimental, **untested** aarch64 context-switch backend lives on the
|
||||
`arm-port` branch (extracted into a `target_arch`-gated `src/arch/`); it has not
|
||||
been validated on hardware yet. macOS remains on the deferred list because of the
|
||||
epoll dependency.
|
||||
Standard Cargo. Requires Rust 1.95 or newer (the `#[naked]` attribute went stable in 1.88; we use a few unrelated post-1.88 features). I have worked hard to keep this library as dependency-free as possible. `master` is x86-64 Linux only. An experimental, **untested** aarch64 context-switch backend lives on the `arm-port` branch (extracted into a `target_arch`-gated `src/arch/`); it has not been validated on hardware yet. macOS remains on the deferred list because of the epoll dependency.
|
||||
|
||||
|
||||
|
||||
```sh
|
||||
cargo test # all tests
|
||||
@@ -80,26 +103,29 @@ cargo test --test mutex # one module
|
||||
cargo bench # primes benchmark vs tokio
|
||||
```
|
||||
|
||||
## What's not here
|
||||
|
||||
See the **Defer** section of `Architecture.md`.
|
||||
`join!` for handle groups, stack growth via remap,
|
||||
hierarchical timer wheel, fd-wait timeouts, `Signal::Timeout`. Each is
|
||||
mechanism we know how to add; none belongs in this iteration.
|
||||
|
||||
## Docs
|
||||
|
||||
| Document | What it covers |
|
||||
|---|---|
|
||||
| [`Architecture.md`](./docs/Architecture.md) | Design intent, runtime model, and deferred work |
|
||||
| [`smarm - Deep Dive.html`](./docs/smarm%20-%20Deep%20Dive.html) | Generated walkthrough of the system; good starting point |
|
||||
| [`smarm - Deep Dive.html`](./docs/smarm%20-%20Deep%20Dive.html) | Generated walkthrough of the system; good starting point if you want to learn about the internals |
|
||||
| [`BENCHMARKS_AND_TUNING.md`](./docs/BENCHMARKS_AND_TUNING.md) | Where smarm wins and loses vs tokio, preemption knob recommendations |
|
||||
| [`benchmarks.md`](./docs/benchmarks.md) | Raw benchmark results, methodology, and tuning experiment log |
|
||||
|
||||
## Coming up
|
||||
|
||||
Clustering: clustering multiple SMARM nodes together is in the pipeline.
|
||||
SMARM-BEAM Interop: Running SMARM as a supervised node under the BEAM via a Rustler NIF works, including message passing and supervision trees that span the runtimes. However, this library is still too unstable to release.
|
||||
SMARM is an interesting platform for implementing a 'dataflow' library, but work on this has not yet started.
|
||||
|
||||
|
||||
## Contributing
|
||||
|
||||
This is a personal proof-of-concept. There's no PR workflow. If you fork it and do something interesting, just send me an email. If it's nice, I'll upstream the changes.
|
||||
This started as a personal proof-of-concept, but it is starting to outgrow that name. If you want to contribute, please get in contact to discuss what you want to work on. Code without prior communication is not welcome.
|
||||
|
||||
|
||||
## A note on open source
|
||||
|
||||
An open source project is a gift, and by giving it, it is no longer mine. I highly enourage you to fork it, to make it your own. This repository, however, is still mine.
|
||||
|
||||
---
|
||||
|
||||
|
||||
|
||||
+48
-4
@@ -77,10 +77,10 @@ Delivered surface:
|
||||
`ServerBuilder::start` untouched; free `call` / `cast` / `whereis_server`;
|
||||
`ServerRef::shutdown` + free `shutdown` as the sys-style synchronous stop.
|
||||
- **Root-exit teardown** (final phase): the run's initial actor is the root;
|
||||
when it exits, the scheduler's idle verdict stops the parked-forever remainder
|
||||
(deferred past the queue drain, so actors with in-flight work finish rather
|
||||
than unwinding on the stop). Closes the "app actor blocks AllDone" stall — see
|
||||
Look into, below.
|
||||
when it exits the run winds down. *(Reworked with the graceful-shutdown work:
|
||||
root exit now delivers `request_shutdown` to every forest root — see
|
||||
"Root exit" below and `tests/root_exit.rs`.)* Closes the "app actor blocks
|
||||
AllDone" stall — see Look into, below.
|
||||
|
||||
Extends — does not retire — the "select exists; a unified per-process mailbox
|
||||
still does not" invariant: 014 adds addressable *delivery*, not a unified inbox;
|
||||
@@ -260,8 +260,52 @@ path the atomic-bool workaround stood in for. Re-check the urus crud repro to
|
||||
confirm the workaround can be retired (the teardown is cooperative — an actor in
|
||||
a tight loop with no observation point still can't be stopped).
|
||||
|
||||
**Update (graceful shutdown):** the RFC 014 sweep was a hard `request_stop` of
|
||||
every live slot, deferred until nothing was runnable — which killed a sleeping
|
||||
actor (timer pending) but drained a queued one, for no principled reason. It is
|
||||
now the OTP semantics: root exit = "the program is done" = `request_shutdown`
|
||||
to every **forest root** (live actor whose parent is the run or is dead), run
|
||||
synchronously on the root's finalize path. Supervisors cascade with their child
|
||||
`Shutdown` policies; trapping actors may `Continue`/drain (timers keep working)
|
||||
and end the run when they stop themselves; non-trapping actors are stopped
|
||||
outright — `join` what you need finished. No forcing sweep follows.
|
||||
|
||||
---
|
||||
|
||||
### Open items from the graceful-shutdown work (not scheduled)
|
||||
- ~~gen_server / gen_statem as a direct supervised child.~~ Done: lifetime is
|
||||
the actor's (refs are addresses, the loop holds an inbox sender);
|
||||
`NamedGenServerBuilder::run` / `gen_statem::run_named` run the loop inline as
|
||||
the `ChildSpec` child; the root-exit sweep traces each leftover as
|
||||
`root_sweep` under `smarm-trace`.
|
||||
- Supervisor `Live` drop-guard sweep is `request_stop` (kill propagates as
|
||||
kill); OTP would deliver a trappable `killed`. Chosen for boundedness.
|
||||
- A root-exit shutdown reaches only actors live *at that instant*; a
|
||||
non-trapping forest root that spawns before it unwinds leaves that spawn
|
||||
to itself (Erlang: an unlinked spawn is nobody's child).
|
||||
- **Supervisor start *order* is not start *readiness*.** `start_child`
|
||||
spawns and moves straight on, so an earlier child is merely *scheduled*,
|
||||
not initialised, when a later sibling starts. A later child that resolves
|
||||
an earlier one by name (`whereis_server`) can therefore miss it — the
|
||||
classic "named registry sibling, then its consumers" tree. Ordered
|
||||
`OneForOne`/`RestForOne` shutdown is unaffected (reverse order is honoured
|
||||
and each stop *is* awaited); this is a start-side gap only.
|
||||
Making `spawn` itself block does NOT fix it — it would only shrink the
|
||||
window to "child has begun executing", while the property callers need is
|
||||
"child has bound its name / opened its socket", which only the child can
|
||||
declare. It would also tax the hot path (one round-trip per accepted
|
||||
connection) and turn every spawn into a context-switch point. OTP has the
|
||||
same async `spawn` and puts the synchronisation one level up:
|
||||
`gen_server:start_link` blocks the caller until `init/1` returns.
|
||||
Fix shape when scheduled: a readiness ack in the supervisor's child-start
|
||||
path (`ChildSpec` variant whose factory receives a ready-signal;
|
||||
`NamedGenServerBuilder::run` acks after its name bind, gen_server default
|
||||
acks after `init`; plain closures ack at spawn as today, i.e. opt-in with
|
||||
no cost to existing children). Until then the workaround is structural:
|
||||
have the registrar spawn its own consumers so the ordering is program
|
||||
order inside one actor, not a cross-actor guarantee (urus v0.3 endpoint
|
||||
does exactly this).
|
||||
|
||||
## Invariants & gotchas (respect these across all cycles)
|
||||
|
||||
- **Shared mutex is non-reentrant.** `Sender::send` can call `unpark` →
|
||||
|
||||
+229
@@ -0,0 +1,229 @@
|
||||
# urus / smarm handoff — updated 2026-08-19 (session 3)
|
||||
|
||||
## TL;DR for the next session
|
||||
**smarm is done for now** (5 unpushed commits on local `master`, see below).
|
||||
**Next = urus v0.3 endpoint refactor.** You should NOT need to read smarm
|
||||
scheduler internals; the contract you build on is fully described here and in
|
||||
`smarm_full/examples/graceful_shutdown.rs` (read that file first — it is the
|
||||
exact shape urus's tree will take) plus `smarm_full/tests/root_exit.rs`.
|
||||
|
||||
### The smarm contract urus builds on (all on local master, verified by tests)
|
||||
- `request_stop(pid)` = kill (cooperative hard stop). `request_shutdown(pid)` =
|
||||
polite: trapping target gets `ExitSignal{reason: Shutdown}`, non-trapping is
|
||||
stopped outright. `RuntimeHandle::{request_stop,request_shutdown}` do the same
|
||||
from any OS thread (signal handler); grab `rt.handle()` before `rt.run`.
|
||||
- Supervisor traps; `request_shutdown(sup)` = ordered reverse-start shutdown,
|
||||
per-child `ChildSpec::shutdown(Shutdown::{Timeout(d)|Infinity|BrutalKill})`
|
||||
(default Timeout(5s)); sup then returns normally. `request_stop(sup)`
|
||||
hard-stops children too (no orphans).
|
||||
- gen_server: `ctx.trap_exit()` in init; `handle_shutdown() -> Exit|Continue`;
|
||||
`handle_exit(sig)`; `ctx.stop_handle().stop()` = normal self-exit;
|
||||
`terminate()` may block only on the graceful path (Exit / stop / inbox close).
|
||||
`GenServerRef::shutdown()` is graceful and waits.
|
||||
- gen_statem: same in event clothes — `cx.trap_exit()` in initial enter,
|
||||
`shutdown` rows (default `stop`), `exit sig` rows, `cx.stop()` / `stop` tail,
|
||||
optional `terminate { }` block. `GenStatemRef::shutdown()`.
|
||||
- **Root exit = program done**: when the root actor returns, the runtime
|
||||
`request_shutdown`s every *forest root* (live actor whose parent is the run
|
||||
or dead). Supervisors cascade; trapping actors may drain (timers keep
|
||||
working) and end the run when they stop; non-trapping are stopped; **no
|
||||
forcing sweep** (`join` what must finish). The old "wait until nothing
|
||||
runnable then kill all" deferral is gone.
|
||||
- **Gotcha for urus:** a gen_server's lifetime is governed by its refs — drop
|
||||
the last `GenServerRef` and the inbox closes → clean exit *even mid-drain*.
|
||||
The endpoint must be pinned (named, or its ref held by the supervisor
|
||||
wrapper) or it will terminate the moment the root drops its ref.
|
||||
- **Known gap (ROADMAP open item):** a gen_server can't be a direct `ChildSpec`
|
||||
child; use the trapping wrapper pattern in `examples/graceful_shutdown.rs::
|
||||
drainer_child` (starts `under(self_pid())`, forwards shutdown, waits). Doing
|
||||
an inline `GenServerBuilder::run()` first may be worth a short smarm detour —
|
||||
decide with Markk.
|
||||
|
||||
### smarm commits this session (local master, NOT pushed, NOT tagged)
|
||||
`250f312` root-exit = graceful shutdown of forest roots (tests/root_exit.rs)
|
||||
`6ceb138` gen_statem shutdown parity (tests/gen_statem_shutdown.rs)
|
||||
`849a424` docs + examples/graceful_shutdown.rs + README "Stopping actors"
|
||||
On top of `1002777` (cross-thread wake) and `9c8f59c` (graceful shutdown).
|
||||
Cargo.toml still `0.6.1`. Release cut (push, tag — v0.7 is justified by the
|
||||
API surface — version bump) is Markk's. Full suite, doc tests, examples,
|
||||
`cargo fmt`, `cargo clippy --lib` all clean. (`clippy --tests` has pre-existing
|
||||
unwrap lints in tests/fd_select.rs, untouched.)
|
||||
|
||||
### Decisions taken this session (Markk)
|
||||
- Root exit means "program done" (Go/tokio/OTP), not "wait for pending work";
|
||||
the previously agreed "sleep(50ms) must finish" test was dropped as encoding
|
||||
the wrong contract (a timer-wheel gate would re-wedge periodic-timer daemons).
|
||||
- No behaviour-preserving deferral, no forcing second sweep.
|
||||
- Examples/docs done in the same session; urus next session.
|
||||
|
||||
---
|
||||
# Previous handoff (still accurate where not superseded above)
|
||||
|
||||
|
||||
## Next-session goal
|
||||
Phase 1 is **done and committed**; Phase 2 is next:
|
||||
1. **smarm v0.6.2** — cross-thread wake root fix. **DONE**, committed on `master`
|
||||
as `1002777`. Not yet tagged, not yet version-bumped (Cargo.toml still reads
|
||||
`0.6.1`), and **not yet pushed to origin** — it exists only in the delivered
|
||||
snapshot zip and the local sandbox clone. Cutting the release (push + tag
|
||||
`v0.6.2` + bump `0.6.1`→`0.6.2`) is Markk's step.
|
||||
2. **urus v0.3** — endpoint refactor. Working against `smarm = { path = "../smarm_full" }`
|
||||
with `git update-index --skip-worktree Cargo.toml` (Markk approved); release commit
|
||||
swaps back to the tag once Markk cuts it. Phase 2 plan below is STALE where it says
|
||||
drain-in-terminate; the endpoint is a trapping GenServer: `handle_shutdown` →
|
||||
`Continue`, enter Draining, `StopHandle::stop()` when the conn set empties.
|
||||
|
||||
Decisions below are locked unless marked *(confirm)*.
|
||||
|
||||
## Reconstruction (the sandbox resets between sessions)
|
||||
A fresh sandbox has an empty home and **no Rust toolchain**. To restore:
|
||||
- Install rustup/cargo. smarm reformats under **rustc 1.97.1**; urus `rust-version`
|
||||
is 1.95. Use 1.97.1.
|
||||
- urus: `git clone https://git.kalsbeek.dev/Markk116/urus` — `origin` is registered
|
||||
and public-read. master `8bdec97` = the v0.2.x line. (Zips in outputs are stale;
|
||||
prefer the remote now.)
|
||||
- smarm: `git clone https://git.kalsbeek.dev/Markk116/smarm`. Latest tag **v0.6.1**
|
||||
(`ca1c983`). The cross-thread wake fix is committed as `1002777` on top of the
|
||||
post-v0.6.1 README commit `8f2d513` (= origin/master). **It is NOT on origin
|
||||
yet** — a fresh clone won't have it until Markk pushes. Restore it from the
|
||||
snapshot zip if working before the push. **v0.6.2 is not yet tagged.**
|
||||
- urus pins smarm by git **tag** in `Cargo.toml` (currently `v0.6.0`). A trivial
|
||||
first commit bumps it to `v0.6.1` (also picks up `try_spawn` + monitor
|
||||
terminal-outcome fixes).
|
||||
|
||||
## Why (context — the finding that drives the plan)
|
||||
urus's shutdown machinery (the `AtomicBool` listener flag + the `SHUTDOWN_POLL`
|
||||
loop in `serve.rs`) is scaffolding around two smarm properties. Their statuses
|
||||
differ, which is the whole point:
|
||||
|
||||
- **Issue A — lossy stop vs a QUEUED actor: ALREADY FIXED in smarm.** Commit
|
||||
`7bab4d2` added an entry-side `check_cancelled()` in `park_current`. A
|
||||
`request_stop` against a listener parked in `wait_readable_timeout` now unwinds
|
||||
cleanly (it parks via `try_select_timeout → park_current`). urus's flag + its
|
||||
stale "smarm's lossy stop-while-QUEUED window" comment can be deleted.
|
||||
- **Issue B — foreign-thread wake is a no-op: FIXED in `1002777` (was present
|
||||
through v0.6.1).** The gap: `unpark`/`unpark_at`/`request_stop` all route through
|
||||
`try_with_runtime`, which reads a thread-local that is `None` on any non-scheduler
|
||||
thread, so no cross-thread wake worked — a signal handler / OS thread could not
|
||||
wake *or* stop a parked actor, which is why `serve.rs` polls the shutdown signal
|
||||
instead of parking on it. Now closed (see Phase 1 below): urus's `SHUTDOWN_POLL`
|
||||
loop can be deleted and its `Handle::shutdown` can park on a handle-driven stop.
|
||||
|
||||
## Phase 1 — smarm cross-thread wake (root fix) — DONE (`1002777`)
|
||||
Shipped as one commit generalizing RFC 018 (a producer reaches the runtime through
|
||||
a `Weak` it holds) from the IO backend to channel senders and a new handle:
|
||||
- **`Runtime::handle() -> RuntimeHandle`** (`Send + Sync`), holding a
|
||||
`Weak<RuntimeInner>`. Grab it before `rt.run` and hand it to the signal thread.
|
||||
- **`RuntimeHandle::request_stop<A>(Pid<A>)`** — upgrades the Weak and calls
|
||||
`request_stop_inner` on the inner; no-op if the runtime is gone. This is the
|
||||
signal-handler-drives-shutdown path; it cascades the ordered stop down the tree
|
||||
exactly like an in-runtime `request_stop`.
|
||||
- **Send-wake:** the receiver captures `scheduler::runtime_weak()` into its
|
||||
`parked_receiver` tuple **at park time** (not at channel creation — the resolved
|
||||
sub-decision; a parked receiver is a live actor so the Weak is provably upgradable,
|
||||
and it scopes the capture to when a wake is possible). `send()` and last-sender
|
||||
`drop` wake via `scheduler::unpark_at_via(pid, epoch, &weak)`: thread-local path
|
||||
when on a scheduler thread (preempt-gated, slot-eligible), captured Weak otherwise.
|
||||
In-runtime timer wakes (recv/select) were left on `scheduler::unpark_at`.
|
||||
|
||||
**API scope decision (signed off):** `RuntimeHandle` exposes **`request_stop` only**.
|
||||
No public `unpark`/`unpark_at` on the handle — send-wake needs no user-facing handle,
|
||||
and "unpark off-runtime" is covered because `request_stop` drives `unpark` on the
|
||||
upgraded inner. No `is_alive()`. Both are one-line additions if a consumer appears.
|
||||
|
||||
**No RFC written** — pattern was already established (RFC 018), agreed not needed.
|
||||
|
||||
Tests: `tests/cross_thread_wake.rs` (foreign-thread send wakes a parked receiver;
|
||||
foreign-thread `request_stop` wakes+stops a parked actor; a lingering handle never
|
||||
blocks all-done and degrades to a no-op once the runtime drops). Full suite green;
|
||||
`cargo fmt` + `cargo clippy --lib` clean.
|
||||
|
||||
**Remaining release step (Markk):** push `master`, tag `v0.6.2`, bump Cargo.toml
|
||||
`0.6.1`→`0.6.2`. Left paired with the tag as the release cut, not done in `1002777`.
|
||||
|
||||
## Phase 1b — smarm graceful shutdown (OTP lift) — DONE (`9c8f59c`, on top of `1002777`)
|
||||
Decided this session (Markk): B — fix at the smarm level rather than a two-stop
|
||||
split in urus. No RFC (Markk: "just implement it"). Shipped, tested, committed on
|
||||
the local `master`, **not pushed, not tagged**. It should ship as the same
|
||||
release as 1002777 (v0.6.2, or v0.7 given the API surface — Markk's call).
|
||||
- `request_shutdown(pid)` / `RuntimeHandle::request_shutdown` = `exit(Pid, shutdown)`;
|
||||
`request_stop` = `exit(Pid, kill)`. Trapping target gets `ExitSignal{reason:
|
||||
DownReason::Shutdown}`; non-trapping is stopped outright.
|
||||
- `ChildSpec::shutdown(Shutdown::{BrutalKill, Timeout(d), Infinity})`, default 5s.
|
||||
Supervisor traps exits; `request_shutdown(sup)` = ordered top-down shutdown,
|
||||
returns normally. **Also fixed**: `request_stop(sup)` used to ORPHAN children
|
||||
(probe-verified; the handoff's "cascade" claim was wrong) — `Live` drop guard now
|
||||
hard-stops them.
|
||||
- gen_server: `ctx.trap_exit()`, `handle_shutdown() -> ShutdownAction::{Exit,
|
||||
Continue}`, `handle_exit(ExitSignal)`, `ctx.stop_handle().stop()` = normal
|
||||
self-exit (`{stop, normal}`; previously impossible — only abnormal `Stopped`).
|
||||
`GenServerRef::shutdown()` is graceful now.
|
||||
- Root finding that forced this: gen_server `terminate()` runs from a Drop guard,
|
||||
mid-unwind on the stop path; any park in it = double panic = abort. So
|
||||
"drain-in-terminate()" (the old Phase 2 plan) was never viable.
|
||||
|
||||
### ~~Next-session smarm work~~ DONE this session (see TL;DR)
|
||||
1. **Root-exit sweep**: make `Pop::RootDrain` also require an empty timer wheel
|
||||
(and it already requires nothing runnable; io_out is only checked for AllDone —
|
||||
check whether it should gate RootDrain too). TDD: an actor in `sleep(50ms)` when
|
||||
the root returns must finish, not be swept. Then a `Reservoir`-style test that a
|
||||
*truly* parked-forever daemon still gets swept.
|
||||
2. **gen_statem parity**: `ctx.trap_exit()`, `handle_shutdown -> ShutdownAction`,
|
||||
`handle_exit`, stop handle. Mirror gen_server; mechanical.
|
||||
3. **Examples review**: `examples/*.rs` predate all of this. Rework where they show
|
||||
shutdown/teardown to use `request_shutdown`, `Shutdown` policies, and
|
||||
`StopHandle`; `named_genserver.rs` first (uses `shutdown`). Also
|
||||
`docs/smarm - Deep Dive.html` says terminate() must be non-blocking — now only
|
||||
true on the unwind paths; and README could use a "Stopping actors" paragraph
|
||||
(request_stop = kill, request_shutdown = shutdown, Shutdown policy).
|
||||
|
||||
### Open smarm items found on the way (noted, not scheduled)
|
||||
- (root-exit sweep and gen_statem parity moved up to the scheduled list.)
|
||||
- Sweep in the supervisor `Live` drop guard is `request_stop` (kill propagates as
|
||||
kill); OTP would deliver a trappable `killed`. Chosen for boundedness.
|
||||
|
||||
## Phase 2 — urus v0.3: endpoint refactor (after v0.6.2 is tagged)
|
||||
Target = the spec's original shape (`urus-spec.md` §2.1/§6: `listener_sup` under the
|
||||
**user's** root supervisor). Deviation to unwind: `serve` owning `rt.run`.
|
||||
- App owns the runtime: `smarm::init(cfg).run(|| root_sup.run())`, root e.g.
|
||||
`RestForOne[ app actors…, urus::endpoint(config, pipeline) ]`. This is what kills
|
||||
the `Arc<OnceLock>` idiom for the right reason (app state born in-runtime as a
|
||||
supervised, ordered child).
|
||||
- `urus::endpoint` = one GenServer child owning the registry + an **internal**
|
||||
listener sub-supervisor + drain-in-`terminate()`. Listeners stay internal, not
|
||||
app-visible peers.
|
||||
- Shutdown = `request_stop` the root supervisor (or via the runtime handle from a
|
||||
signal thread) → cascades down → `endpoint.terminate()` runs the drain
|
||||
(`drain_timeout`, force-stop sweep).
|
||||
- **DELETE:** the `AtomicBool` listener flag (A fixed) and the `SHUTDOWN_POLL` loop
|
||||
+ its apologetic comment (B fixed → park, don't poll).
|
||||
- Nuance: `request_stop` → `Signal::Stopped` is *abnormal* → `Transient` restarts.
|
||||
Stop-without-restart = stop the **supervisor**, not the children.
|
||||
- *(confirm)* Keep `serve`/`serve_with`/`serve_with_shutdown` as thin wrappers that
|
||||
build the one-child tree internally, so the simple case stays one line.
|
||||
- *(confirm)* Keep `Handle`/`ShutdownSignal`? Now that cross-thread wake works,
|
||||
`Handle::shutdown` can map to handle-driven `request_stop` on the endpoint.
|
||||
- Breaking → cut **urus v0.3**; bump the smarm pin to `v0.6.2` here.
|
||||
|
||||
## Working norms
|
||||
- Every bash call: `export PATH=$HOME/.cargo/bin:$PATH` (once the toolchain's in).
|
||||
- **TDD**: failing test first, then implement; keep suites green.
|
||||
- **Hammer ritual** for ANY connection-lifecycle change (the urus #2 shutdown work
|
||||
qualifies): 35× subset (`shutdown timeout reaped slowloris streaming chunked sse
|
||||
stalled ws_ channels session`) + 3× full + 1× trace. `scripts/hammer.sh` does NOT
|
||||
pass feature flags — loop manually with `--features phoenix`. Subset filter must
|
||||
NOT use `--test integration` (session tests live in lib).
|
||||
- Example smoke tests: hold the server's stdin open (`mkfifo` + `sleep > fifo`) or
|
||||
the Enter-to-shutdown thread fires on EOF instantly.
|
||||
- Background procs are reaped BETWEEN bash calls; `pkill -f` matches your own shell.
|
||||
- Artefact store (specs): `curl -H "Authorization: Bearer sk-llmingest-2e45d80c63db24c6781e761eb2a9a58e83d9f48ef77a42185bad311d07c80e68" https://artefacts.kalsbeek.dev/artifacts/<name>`
|
||||
— `urus-spec.md`, `urus-bench-spec.md`, `rfc_008-implementation-notes.md`, …
|
||||
- smarm feature flags: `smarm-trace`, `smarm-causal` (urus re-exports both).
|
||||
|
||||
## Local cross-repo testing — KEEP OUT OF COMMITS
|
||||
To test urus #2 against un-tagged smarm 0.6.2, point urus's `Cargo.toml` smarm dep
|
||||
at a local path (`smarm = { path = "../smarm" }`) instead of the git tag.
|
||||
- Must NOT land in commits. Guard: `git update-index --skip-worktree Cargo.toml`
|
||||
after editing (undo with `--no-skip-worktree`), or stash before committing.
|
||||
- The committed `Cargo.toml` stays pinned to the git tag; restore the tag (bumped to
|
||||
`v0.6.2`) for the release commit.
|
||||
+67
-26
@@ -26,7 +26,9 @@ use std::time::Instant;
|
||||
const ITERS: u32 = 15;
|
||||
|
||||
fn available_threads() -> usize {
|
||||
std::thread::available_parallelism().map(|n| n.get()).unwrap_or(1)
|
||||
std::thread::available_parallelism()
|
||||
.map(|n| n.get())
|
||||
.unwrap_or(1)
|
||||
}
|
||||
|
||||
fn env_sets() -> u32 {
|
||||
@@ -108,17 +110,15 @@ fn bench_chained_smarm(threads: usize) -> (u64, u128) {
|
||||
fn bench_chained_tokio_current() -> (u64, u128) {
|
||||
let counter = Arc::new(AtomicU64::new(0));
|
||||
let c2 = counter.clone();
|
||||
let rt = tokio::runtime::Builder::new_current_thread().build().unwrap();
|
||||
let rt = tokio::runtime::Builder::new_current_thread()
|
||||
.build()
|
||||
.unwrap();
|
||||
let start = Instant::now();
|
||||
let local = tokio::task::LocalSet::new();
|
||||
local.block_on(&rt, async move {
|
||||
// Use a oneshot done channel like tokio's own chained_spawn bench.
|
||||
let (done_tx, done_rx) = tokio::sync::oneshot::channel();
|
||||
fn iter(
|
||||
c: Arc<AtomicU64>,
|
||||
done: tokio::sync::oneshot::Sender<()>,
|
||||
n: u64,
|
||||
) {
|
||||
fn iter(c: Arc<AtomicU64>, done: tokio::sync::oneshot::Sender<()>, n: u64) {
|
||||
if n == 0 {
|
||||
let _ = done.send(());
|
||||
} else {
|
||||
@@ -186,7 +186,9 @@ fn bench_yield_smarm(threads: usize) -> (u64, u128) {
|
||||
}
|
||||
|
||||
fn bench_yield_tokio_current() -> (u64, u128) {
|
||||
let rt = tokio::runtime::Builder::new_current_thread().build().unwrap();
|
||||
let rt = tokio::runtime::Builder::new_current_thread()
|
||||
.build()
|
||||
.unwrap();
|
||||
let start = Instant::now();
|
||||
let local = tokio::task::LocalSet::new();
|
||||
local.block_on(&rt, async move {
|
||||
@@ -235,11 +237,22 @@ const PRIME_N: u64 = 400_000;
|
||||
const PRIME_WORKERS: u64 = 64;
|
||||
|
||||
fn is_prime(n: u64) -> bool {
|
||||
if n < 2 { return false; }
|
||||
if n < 4 { return true; }
|
||||
if n % 2 == 0 { return false; }
|
||||
if n < 2 {
|
||||
return false;
|
||||
}
|
||||
if n < 4 {
|
||||
return true;
|
||||
}
|
||||
if n % 2 == 0 {
|
||||
return false;
|
||||
}
|
||||
let mut i = 3u64;
|
||||
while i * i <= n { if n % i == 0 { return false; } i += 2; }
|
||||
while i * i <= n {
|
||||
if n % i == 0 {
|
||||
return false;
|
||||
}
|
||||
i += 2;
|
||||
}
|
||||
true
|
||||
}
|
||||
|
||||
@@ -250,7 +263,11 @@ fn count_primes(lo: u64, hi: u64) -> u64 {
|
||||
fn primes_slice(w: u64) -> (u64, u64) {
|
||||
let per = PRIME_N / PRIME_WORKERS;
|
||||
let lo = w * per;
|
||||
let hi = if w + 1 == PRIME_WORKERS { PRIME_N } else { lo + per };
|
||||
let hi = if w + 1 == PRIME_WORKERS {
|
||||
PRIME_N
|
||||
} else {
|
||||
lo + per
|
||||
};
|
||||
(lo, hi)
|
||||
}
|
||||
|
||||
@@ -267,7 +284,9 @@ fn bench_primes_smarm(threads: usize) -> (u64, u128) {
|
||||
tc.fetch_add(count_primes(lo, hi), Ordering::Relaxed);
|
||||
}));
|
||||
}
|
||||
for h in handles { h.join().unwrap(); }
|
||||
for h in handles {
|
||||
h.join().unwrap();
|
||||
}
|
||||
});
|
||||
(total.load(Ordering::Relaxed), start.elapsed().as_micros())
|
||||
}
|
||||
@@ -275,7 +294,9 @@ fn bench_primes_smarm(threads: usize) -> (u64, u128) {
|
||||
fn bench_primes_tokio_current() -> (u64, u128) {
|
||||
let total = Arc::new(AtomicU64::new(0));
|
||||
let t2 = total.clone();
|
||||
let rt = tokio::runtime::Builder::new_current_thread().build().unwrap();
|
||||
let rt = tokio::runtime::Builder::new_current_thread()
|
||||
.build()
|
||||
.unwrap();
|
||||
let start = Instant::now();
|
||||
let local = tokio::task::LocalSet::new();
|
||||
local.block_on(&rt, async move {
|
||||
@@ -287,7 +308,9 @@ fn bench_primes_tokio_current() -> (u64, u128) {
|
||||
tc.fetch_add(count_primes(lo, hi), Ordering::Relaxed);
|
||||
}));
|
||||
}
|
||||
for h in handles { let _ = h.await; }
|
||||
for h in handles {
|
||||
let _ = h.await;
|
||||
}
|
||||
});
|
||||
(total.load(Ordering::Relaxed), start.elapsed().as_micros())
|
||||
}
|
||||
@@ -309,7 +332,9 @@ fn bench_primes_tokio_multi() -> (u64, u128) {
|
||||
tc.fetch_add(count_primes(lo, hi), Ordering::Relaxed);
|
||||
}));
|
||||
}
|
||||
for h in handles { let _ = h.await; }
|
||||
for h in handles {
|
||||
let _ = h.await;
|
||||
}
|
||||
});
|
||||
(total.load(Ordering::Relaxed), start.elapsed().as_micros())
|
||||
}
|
||||
@@ -344,7 +369,9 @@ fn bench_pp_smarm(threads: usize) -> (u64, u128) {
|
||||
}
|
||||
|
||||
fn bench_pp_tokio_current() -> (u64, u128) {
|
||||
let rt = tokio::runtime::Builder::new_current_thread().build().unwrap();
|
||||
let rt = tokio::runtime::Builder::new_current_thread()
|
||||
.build()
|
||||
.unwrap();
|
||||
let start = Instant::now();
|
||||
let local = tokio::task::LocalSet::new();
|
||||
local.block_on(&rt, async move {
|
||||
@@ -395,7 +422,6 @@ fn bench_pp_tokio_multi() -> (u64, u128) {
|
||||
// main
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Knob helper — reads SMARM_ALLOC_INTERVAL / SMARM_TIMESLICE_CYCLES env vars
|
||||
// so the sweep script can override the preemption knobs without recompiling.
|
||||
@@ -404,10 +430,14 @@ fn bench_pp_tokio_multi() -> (u64, u128) {
|
||||
fn bench_cfg(threads: usize) -> smarm::runtime::Config {
|
||||
let mut cfg = smarm::runtime::Config::exact(threads);
|
||||
if let Ok(v) = std::env::var("SMARM_ALLOC_INTERVAL") {
|
||||
if let Ok(n) = v.parse::<u32>() { cfg = cfg.alloc_interval(n); }
|
||||
if let Ok(n) = v.parse::<u32>() {
|
||||
cfg = cfg.alloc_interval(n);
|
||||
}
|
||||
}
|
||||
if let Ok(v) = std::env::var("SMARM_TIMESLICE_CYCLES") {
|
||||
if let Ok(n) = v.parse::<u64>() { cfg = cfg.timeslice_cycles(n); }
|
||||
if let Ok(n) = v.parse::<u64>() {
|
||||
cfg = cfg.timeslice_cycles(n);
|
||||
}
|
||||
}
|
||||
cfg
|
||||
}
|
||||
@@ -417,7 +447,10 @@ fn main() {
|
||||
println!("smarm general benchmarks");
|
||||
println!("available parallelism: {n} threads");
|
||||
let sets = env_sets();
|
||||
println!("ITERS={ITERS}×{sets} sets = {} samples (+1 warmup, discarded)", ITERS * sets);
|
||||
println!(
|
||||
"ITERS={ITERS}×{sets} sets = {} samples (+1 warmup, discarded)",
|
||||
ITERS * sets
|
||||
);
|
||||
println!(
|
||||
"CHAIN_DEPTH={CHAIN_DEPTH}, YIELD_TASKS={YIELD_TASKS}×{YIELD_ROUNDS}, \
|
||||
PRIME_N={PRIME_N}/{PRIME_WORKERS} workers, PP_ROUNDS={PP_ROUNDS}"
|
||||
@@ -426,21 +459,29 @@ fn main() {
|
||||
// ---- 1. chained_spawn ----
|
||||
print_header(&format!("chained_spawn: depth {CHAIN_DEPTH}"));
|
||||
run_n("smarm 1-thread", ITERS, || bench_chained_smarm(1));
|
||||
run_n(&format!("smarm {n}-thread"), ITERS, || bench_chained_smarm(n));
|
||||
run_n(&format!("smarm {n}-thread"), ITERS, || {
|
||||
bench_chained_smarm(n)
|
||||
});
|
||||
run_n("tokio current_thread", ITERS, bench_chained_tokio_current);
|
||||
run_n("tokio multi-thread", ITERS, bench_chained_tokio_multi);
|
||||
|
||||
// ---- 2. yield_many ----
|
||||
print_header(&format!("yield_many: {YIELD_TASKS} tasks × {YIELD_ROUNDS} yields"));
|
||||
print_header(&format!(
|
||||
"yield_many: {YIELD_TASKS} tasks × {YIELD_ROUNDS} yields"
|
||||
));
|
||||
run_n("smarm 1-thread", ITERS, || bench_yield_smarm(1));
|
||||
run_n(&format!("smarm {n}-thread"), ITERS, || bench_yield_smarm(n));
|
||||
run_n("tokio current_thread", ITERS, bench_yield_tokio_current);
|
||||
run_n("tokio multi-thread", ITERS, bench_yield_tokio_multi);
|
||||
|
||||
// ---- 3. fan_out_compute ----
|
||||
print_header(&format!("fan_out_compute: primes in [2, {PRIME_N}) across {PRIME_WORKERS}"));
|
||||
print_header(&format!(
|
||||
"fan_out_compute: primes in [2, {PRIME_N}) across {PRIME_WORKERS}"
|
||||
));
|
||||
run_n("smarm 1-thread", ITERS, || bench_primes_smarm(1));
|
||||
run_n(&format!("smarm {n}-thread"), ITERS, || bench_primes_smarm(n));
|
||||
run_n(&format!("smarm {n}-thread"), ITERS, || {
|
||||
bench_primes_smarm(n)
|
||||
});
|
||||
run_n("tokio current_thread", ITERS, bench_primes_tokio_current);
|
||||
run_n("tokio multi-thread", ITERS, bench_primes_tokio_multi);
|
||||
|
||||
|
||||
+90
-41
@@ -64,11 +64,22 @@ const PRIME_N: u64 = 400_000;
|
||||
const WORKERS: u64 = 64;
|
||||
|
||||
fn is_prime(n: u64) -> bool {
|
||||
if n < 2 { return false; }
|
||||
if n < 4 { return true; }
|
||||
if n % 2 == 0 { return false; }
|
||||
if n < 2 {
|
||||
return false;
|
||||
}
|
||||
if n < 4 {
|
||||
return true;
|
||||
}
|
||||
if n % 2 == 0 {
|
||||
return false;
|
||||
}
|
||||
let mut i = 3u64;
|
||||
while i * i <= n { if n % i == 0 { return false; } i += 2; }
|
||||
while i * i <= n {
|
||||
if n % i == 0 {
|
||||
return false;
|
||||
}
|
||||
i += 2;
|
||||
}
|
||||
true
|
||||
}
|
||||
|
||||
@@ -96,7 +107,9 @@ fn bench_primes_smarm(threads: usize) -> (u64, u128) {
|
||||
tc.fetch_add(count_primes(lo, hi), Ordering::Relaxed);
|
||||
}));
|
||||
}
|
||||
for h in handles { h.join().unwrap(); }
|
||||
for h in handles {
|
||||
h.join().unwrap();
|
||||
}
|
||||
});
|
||||
(total.load(Ordering::Relaxed), start.elapsed().as_micros())
|
||||
}
|
||||
@@ -104,7 +117,9 @@ fn bench_primes_smarm(threads: usize) -> (u64, u128) {
|
||||
fn bench_primes_tokio_current() -> (u64, u128) {
|
||||
let total = Arc::new(AtomicU64::new(0));
|
||||
let t2 = total.clone();
|
||||
let rt = tokio::runtime::Builder::new_current_thread().build().unwrap();
|
||||
let rt = tokio::runtime::Builder::new_current_thread()
|
||||
.build()
|
||||
.unwrap();
|
||||
let start = Instant::now();
|
||||
let local = tokio::task::LocalSet::new();
|
||||
local.block_on(&rt, async move {
|
||||
@@ -116,7 +131,9 @@ fn bench_primes_tokio_current() -> (u64, u128) {
|
||||
tc.fetch_add(count_primes(lo, hi), Ordering::Relaxed);
|
||||
}));
|
||||
}
|
||||
for h in handles { let _ = h.await; }
|
||||
for h in handles {
|
||||
let _ = h.await;
|
||||
}
|
||||
});
|
||||
(total.load(Ordering::Relaxed), start.elapsed().as_micros())
|
||||
}
|
||||
@@ -138,17 +155,21 @@ fn bench_primes_tokio_multi() -> (u64, u128) {
|
||||
tc.fetch_add(count_primes(lo, hi), Ordering::Relaxed);
|
||||
}));
|
||||
}
|
||||
for h in handles { let _ = h.await; }
|
||||
for h in handles {
|
||||
let _ = h.await;
|
||||
}
|
||||
});
|
||||
(total.load(Ordering::Relaxed), start.elapsed().as_micros())
|
||||
}
|
||||
|
||||
fn bench_primes_baseline() -> (u64, u128) {
|
||||
let start = Instant::now();
|
||||
let total: u64 = (0..WORKERS).map(|w| {
|
||||
let (lo, hi) = primes_slice(w);
|
||||
count_primes(lo, hi)
|
||||
}).sum();
|
||||
let total: u64 = (0..WORKERS)
|
||||
.map(|w| {
|
||||
let (lo, hi) = primes_slice(w);
|
||||
count_primes(lo, hi)
|
||||
})
|
||||
.sum();
|
||||
(total, start.elapsed().as_micros())
|
||||
}
|
||||
|
||||
@@ -167,15 +188,17 @@ fn bench_pingpong_smarm(threads: usize) -> (u64, u128) {
|
||||
tx_a.send(0).unwrap();
|
||||
loop {
|
||||
let v = rx_b.recv().unwrap();
|
||||
if v >= PING_ROUNDS { break; }
|
||||
if v >= PING_ROUNDS {
|
||||
break;
|
||||
}
|
||||
tx_a.send(v + 1).unwrap();
|
||||
}
|
||||
});
|
||||
let hb = smarm::spawn(move || {
|
||||
loop {
|
||||
let v = rx_a.recv().unwrap();
|
||||
tx_b.send(v + 1).unwrap();
|
||||
if v + 1 >= PING_ROUNDS { break; }
|
||||
let hb = smarm::spawn(move || loop {
|
||||
let v = rx_a.recv().unwrap();
|
||||
tx_b.send(v + 1).unwrap();
|
||||
if v + 1 >= PING_ROUNDS {
|
||||
break;
|
||||
}
|
||||
});
|
||||
ha.join().unwrap();
|
||||
@@ -198,7 +221,9 @@ fn bench_pingpong_tokio_current() -> (u64, u128) {
|
||||
tx_a.send(0).unwrap();
|
||||
loop {
|
||||
let v = rx_b.recv().await.unwrap();
|
||||
if v >= PING_ROUNDS { break; }
|
||||
if v >= PING_ROUNDS {
|
||||
break;
|
||||
}
|
||||
tx_a.send(v + 1).unwrap();
|
||||
}
|
||||
});
|
||||
@@ -206,7 +231,9 @@ fn bench_pingpong_tokio_current() -> (u64, u128) {
|
||||
loop {
|
||||
let v = rx_a.recv().await.unwrap();
|
||||
tx_b.send(v + 1).unwrap();
|
||||
if v + 1 >= PING_ROUNDS { break; }
|
||||
if v + 1 >= PING_ROUNDS {
|
||||
break;
|
||||
}
|
||||
}
|
||||
});
|
||||
let _ = ha.await;
|
||||
@@ -229,7 +256,9 @@ fn bench_pingpong_tokio_multi() -> (u64, u128) {
|
||||
tx_a.send(0).unwrap();
|
||||
loop {
|
||||
let v = rx_b.recv().await.unwrap();
|
||||
if v >= PING_ROUNDS { break; }
|
||||
if v >= PING_ROUNDS {
|
||||
break;
|
||||
}
|
||||
tx_a.send(v + 1).unwrap();
|
||||
}
|
||||
});
|
||||
@@ -237,7 +266,9 @@ fn bench_pingpong_tokio_multi() -> (u64, u128) {
|
||||
loop {
|
||||
let v = rx_a.recv().await.unwrap();
|
||||
tx_b.send(v + 1).unwrap();
|
||||
if v + 1 >= PING_ROUNDS { break; }
|
||||
if v + 1 >= PING_ROUNDS {
|
||||
break;
|
||||
}
|
||||
}
|
||||
});
|
||||
let _ = ha.await;
|
||||
@@ -264,7 +295,9 @@ fn bench_spawn_smarm(threads: usize) -> (u64, u128) {
|
||||
cc.fetch_add(1, Ordering::Relaxed);
|
||||
}));
|
||||
}
|
||||
for h in handles { h.join().unwrap(); }
|
||||
for h in handles {
|
||||
h.join().unwrap();
|
||||
}
|
||||
});
|
||||
(counter.load(Ordering::Relaxed), start.elapsed().as_micros())
|
||||
}
|
||||
@@ -272,7 +305,9 @@ fn bench_spawn_smarm(threads: usize) -> (u64, u128) {
|
||||
fn bench_spawn_tokio_current() -> (u64, u128) {
|
||||
let counter = Arc::new(AtomicU64::new(0));
|
||||
let c = counter.clone();
|
||||
let rt = tokio::runtime::Builder::new_current_thread().build().unwrap();
|
||||
let rt = tokio::runtime::Builder::new_current_thread()
|
||||
.build()
|
||||
.unwrap();
|
||||
let start = Instant::now();
|
||||
let local = tokio::task::LocalSet::new();
|
||||
local.block_on(&rt, async move {
|
||||
@@ -283,7 +318,9 @@ fn bench_spawn_tokio_current() -> (u64, u128) {
|
||||
cc.fetch_add(1, Ordering::Relaxed);
|
||||
}));
|
||||
}
|
||||
for h in handles { let _ = h.await; }
|
||||
for h in handles {
|
||||
let _ = h.await;
|
||||
}
|
||||
});
|
||||
(counter.load(Ordering::Relaxed), start.elapsed().as_micros())
|
||||
}
|
||||
@@ -304,7 +341,9 @@ fn bench_spawn_tokio_multi() -> (u64, u128) {
|
||||
cc.fetch_add(1, Ordering::Relaxed);
|
||||
}));
|
||||
}
|
||||
for h in handles { let _ = h.await; }
|
||||
for h in handles {
|
||||
let _ = h.await;
|
||||
}
|
||||
});
|
||||
(counter.load(Ordering::Relaxed), start.elapsed().as_micros())
|
||||
}
|
||||
@@ -320,24 +359,34 @@ fn main() {
|
||||
println!("PRIME_N={PRIME_N}, WORKERS={WORKERS}, PING_ROUNDS={PING_ROUNDS}, SPAWN_COUNT={SPAWN_COUNT}");
|
||||
|
||||
// ---- Primes ----
|
||||
print_header(&format!("Fan-out/fan-in: count primes in [2, {PRIME_N}) across {WORKERS} workers"));
|
||||
run_n("baseline (serial)", ITERS, bench_primes_baseline);
|
||||
run_n("smarm single-thread", ITERS, || bench_primes_smarm(1));
|
||||
run_n(&format!("smarm {n}-thread"), ITERS, || bench_primes_smarm(n));
|
||||
run_n("tokio current_thread", ITERS, bench_primes_tokio_current);
|
||||
run_n("tokio multi-thread", ITERS, bench_primes_tokio_multi);
|
||||
print_header(&format!(
|
||||
"Fan-out/fan-in: count primes in [2, {PRIME_N}) across {WORKERS} workers"
|
||||
));
|
||||
run_n("baseline (serial)", ITERS, bench_primes_baseline);
|
||||
run_n("smarm single-thread", ITERS, || bench_primes_smarm(1));
|
||||
run_n(&format!("smarm {n}-thread"), ITERS, || {
|
||||
bench_primes_smarm(n)
|
||||
});
|
||||
run_n("tokio current_thread", ITERS, bench_primes_tokio_current);
|
||||
run_n("tokio multi-thread", ITERS, bench_primes_tokio_multi);
|
||||
|
||||
// ---- Ping-pong ----
|
||||
print_header(&format!("Ping-pong: {PING_ROUNDS} round-trips between two actors"));
|
||||
run_n("smarm single-thread", ITERS, || bench_pingpong_smarm(1));
|
||||
run_n(&format!("smarm {n}-thread"), ITERS, || bench_pingpong_smarm(n));
|
||||
run_n("tokio current_thread", ITERS, bench_pingpong_tokio_current);
|
||||
run_n("tokio multi-thread", ITERS, bench_pingpong_tokio_multi);
|
||||
print_header(&format!(
|
||||
"Ping-pong: {PING_ROUNDS} round-trips between two actors"
|
||||
));
|
||||
run_n("smarm single-thread", ITERS, || bench_pingpong_smarm(1));
|
||||
run_n(&format!("smarm {n}-thread"), ITERS, || {
|
||||
bench_pingpong_smarm(n)
|
||||
});
|
||||
run_n("tokio current_thread", ITERS, bench_pingpong_tokio_current);
|
||||
run_n("tokio multi-thread", ITERS, bench_pingpong_tokio_multi);
|
||||
|
||||
// ---- Spawn throughput ----
|
||||
print_header(&format!("Spawn throughput: {SPAWN_COUNT} actors spawned and joined"));
|
||||
run_n("smarm single-thread", ITERS, || bench_spawn_smarm(1));
|
||||
print_header(&format!(
|
||||
"Spawn throughput: {SPAWN_COUNT} actors spawned and joined"
|
||||
));
|
||||
run_n("smarm single-thread", ITERS, || bench_spawn_smarm(1));
|
||||
run_n(&format!("smarm {n}-thread"), ITERS, || bench_spawn_smarm(n));
|
||||
run_n("tokio current_thread", ITERS, bench_spawn_tokio_current);
|
||||
run_n("tokio multi-thread", ITERS, bench_spawn_tokio_multi);
|
||||
run_n("tokio current_thread", ITERS, bench_spawn_tokio_current);
|
||||
run_n("tokio multi-thread", ITERS, bench_spawn_tokio_multi);
|
||||
}
|
||||
|
||||
+24
-7
@@ -16,12 +16,20 @@ const WORKERS: u64 = 16;
|
||||
const ITERATIONS: u32 = 5;
|
||||
|
||||
fn is_prime(n: u64) -> bool {
|
||||
if n < 2 { return false; }
|
||||
if n < 4 { return true; }
|
||||
if n % 2 == 0 { return false; }
|
||||
if n < 2 {
|
||||
return false;
|
||||
}
|
||||
if n < 4 {
|
||||
return true;
|
||||
}
|
||||
if n % 2 == 0 {
|
||||
return false;
|
||||
}
|
||||
let mut i = 3u64;
|
||||
while i * i <= n {
|
||||
if n % i == 0 { return false; }
|
||||
if n % i == 0 {
|
||||
return false;
|
||||
}
|
||||
i += 2;
|
||||
}
|
||||
true
|
||||
@@ -30,7 +38,9 @@ fn is_prime(n: u64) -> bool {
|
||||
fn count_primes_in(lo: u64, hi: u64) -> u64 {
|
||||
let mut count = 0u64;
|
||||
for n in lo..hi {
|
||||
if is_prime(n) { count += 1; }
|
||||
if is_prime(n) {
|
||||
count += 1;
|
||||
}
|
||||
}
|
||||
count
|
||||
}
|
||||
@@ -38,7 +48,11 @@ fn count_primes_in(lo: u64, hi: u64) -> u64 {
|
||||
fn slice(worker: u64) -> (u64, u64) {
|
||||
let per = N / WORKERS;
|
||||
let lo = worker * per;
|
||||
let hi = if worker + 1 == WORKERS { N } else { (worker + 1) * per };
|
||||
let hi = if worker + 1 == WORKERS {
|
||||
N
|
||||
} else {
|
||||
(worker + 1) * per
|
||||
};
|
||||
(lo, hi)
|
||||
}
|
||||
|
||||
@@ -125,7 +139,10 @@ fn main() {
|
||||
"Counting primes in [2, {}) across {} workers, {} iterations each\n",
|
||||
N, WORKERS, ITERATIONS
|
||||
);
|
||||
println!("{:>12} | {:>15} | {:>16} | {:>15} | {:>15}", "runtime", "primes found", "median", "min", "max");
|
||||
println!(
|
||||
"{:>12} | {:>15} | {:>16} | {:>15} | {:>15}",
|
||||
"runtime", "primes found", "median", "min", "max"
|
||||
);
|
||||
println!("{}", "-".repeat(80));
|
||||
|
||||
run_n("baseline", ITERATIONS, bench_baseline);
|
||||
|
||||
+44
-7
@@ -27,12 +27,19 @@ use std::sync::Arc;
|
||||
use std::time::Instant;
|
||||
|
||||
fn env_usize(key: &str, default: usize) -> usize {
|
||||
std::env::var(key).ok().and_then(|v| v.parse().ok()).unwrap_or(default)
|
||||
std::env::var(key)
|
||||
.ok()
|
||||
.and_then(|v| v.parse().ok())
|
||||
.unwrap_or(default)
|
||||
}
|
||||
|
||||
fn env_threads() -> Vec<usize> {
|
||||
std::env::var("SMARM_BENCH_THREADS")
|
||||
.map(|v| v.split_whitespace().filter_map(|t| t.parse().ok()).collect())
|
||||
.map(|v| {
|
||||
v.split_whitespace()
|
||||
.filter_map(|t| t.parse().ok())
|
||||
.collect()
|
||||
})
|
||||
.unwrap_or_else(|_| vec![1, 2, 4])
|
||||
}
|
||||
|
||||
@@ -53,7 +60,11 @@ fn drive<Q: Send + Sync + 'static>(
|
||||
for p in 0..producers {
|
||||
let q = q.clone();
|
||||
// Give the last producer the remainder.
|
||||
let n = if p == producers - 1 { items - per * (producers - 1) } else { per };
|
||||
let n = if p == producers - 1 {
|
||||
items - per * (producers - 1)
|
||||
} else {
|
||||
per
|
||||
};
|
||||
hs.push(std::thread::spawn(move || {
|
||||
let pid = Pid::new(p as u32, 0);
|
||||
for _ in 0..n {
|
||||
@@ -132,7 +143,12 @@ fn main() {
|
||||
for &t in &threads_sweep {
|
||||
for (p, c) in ratios_for(t) {
|
||||
for s in ["mutex", "mpmc", "striped"] {
|
||||
cases.push(Case { structure: s, threads: t, producers: p, consumers: c });
|
||||
cases.push(Case {
|
||||
structure: s,
|
||||
threads: t,
|
||||
producers: p,
|
||||
consumers: c,
|
||||
});
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -147,7 +163,14 @@ fn main() {
|
||||
if case.threads < 2 {
|
||||
drive_single(&*q, MutexQueue::push, MutexQueue::pop, items)
|
||||
} else {
|
||||
drive(q, MutexQueue::push, MutexQueue::pop, case.producers, case.consumers, items)
|
||||
drive(
|
||||
q,
|
||||
MutexQueue::push,
|
||||
MutexQueue::pop,
|
||||
case.producers,
|
||||
case.consumers,
|
||||
items,
|
||||
)
|
||||
}
|
||||
}
|
||||
"mpmc" => {
|
||||
@@ -155,7 +178,14 @@ fn main() {
|
||||
if case.threads < 2 {
|
||||
drive_single(&*q, MpmcRing::push, MpmcRing::pop, items)
|
||||
} else {
|
||||
drive(q, MpmcRing::push, MpmcRing::pop, case.producers, case.consumers, items)
|
||||
drive(
|
||||
q,
|
||||
MpmcRing::push,
|
||||
MpmcRing::pop,
|
||||
case.producers,
|
||||
case.consumers,
|
||||
items,
|
||||
)
|
||||
}
|
||||
}
|
||||
"striped" => {
|
||||
@@ -163,7 +193,14 @@ fn main() {
|
||||
if case.threads < 2 {
|
||||
drive_single(&*q, StripedRing::push, StripedRing::pop, items)
|
||||
} else {
|
||||
drive(q, StripedRing::push, StripedRing::pop, case.producers, case.consumers, items)
|
||||
drive(
|
||||
q,
|
||||
StripedRing::push,
|
||||
StripedRing::pop,
|
||||
case.producers,
|
||||
case.consumers,
|
||||
items,
|
||||
)
|
||||
}
|
||||
}
|
||||
_ => unreachable!(),
|
||||
|
||||
+21
-4
@@ -54,12 +54,19 @@ fn variant() -> &'static str {
|
||||
}
|
||||
|
||||
fn env_usize(key: &str, default: usize) -> usize {
|
||||
std::env::var(key).ok().and_then(|v| v.parse().ok()).unwrap_or(default)
|
||||
std::env::var(key)
|
||||
.ok()
|
||||
.and_then(|v| v.parse().ok())
|
||||
.unwrap_or(default)
|
||||
}
|
||||
|
||||
fn env_threads() -> Vec<usize> {
|
||||
std::env::var("SMARM_BENCH_THREADS")
|
||||
.map(|v| v.split_whitespace().filter_map(|t| t.parse().ok()).collect())
|
||||
.map(|v| {
|
||||
v.split_whitespace()
|
||||
.filter_map(|t| t.parse().ok())
|
||||
.collect()
|
||||
})
|
||||
.unwrap_or_else(|_| vec![1, 2, 4])
|
||||
}
|
||||
|
||||
@@ -238,12 +245,22 @@ fn main() {
|
||||
);
|
||||
println!(
|
||||
"RQCSV,runtime,{},{},{},{},{},{},{}",
|
||||
variant(), slot_str, name, t, work, mid.us, per_s
|
||||
variant(),
|
||||
slot_str,
|
||||
name,
|
||||
t,
|
||||
work,
|
||||
mid.us,
|
||||
per_s
|
||||
);
|
||||
if slot {
|
||||
println!(
|
||||
"RQSLOT,{},{},{},{},{}",
|
||||
variant(), name, t, mid.hits, mid.displacements
|
||||
variant(),
|
||||
name,
|
||||
t,
|
||||
mid.hits,
|
||||
mid.displacements
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
+55
-19
@@ -37,7 +37,9 @@ use std::time::Instant;
|
||||
const ITERS: u32 = 15;
|
||||
|
||||
fn available_threads() -> usize {
|
||||
std::thread::available_parallelism().map(|n| n.get()).unwrap_or(1)
|
||||
std::thread::available_parallelism()
|
||||
.map(|n| n.get())
|
||||
.unwrap_or(1)
|
||||
}
|
||||
|
||||
fn env_sets() -> u32 {
|
||||
@@ -116,7 +118,9 @@ fn bench_recurse_smarm(threads: usize) -> (u64, u128) {
|
||||
fn bench_recurse_tokio_current() -> (u64, u128) {
|
||||
let counter = Arc::new(AtomicU64::new(0));
|
||||
let c2 = counter.clone();
|
||||
let rt = tokio::runtime::Builder::new_current_thread().build().unwrap();
|
||||
let rt = tokio::runtime::Builder::new_current_thread()
|
||||
.build()
|
||||
.unwrap();
|
||||
let start = Instant::now();
|
||||
let local = tokio::task::LocalSet::new();
|
||||
local.block_on(&rt, async move {
|
||||
@@ -199,7 +203,9 @@ fn bench_hot_smarm() -> (u64, u128) {
|
||||
}
|
||||
|
||||
fn bench_hot_tokio_current() -> (u64, u128) {
|
||||
let rt = tokio::runtime::Builder::new_current_thread().build().unwrap();
|
||||
let rt = tokio::runtime::Builder::new_current_thread()
|
||||
.build()
|
||||
.unwrap();
|
||||
let start = Instant::now();
|
||||
let local = tokio::task::LocalSet::new();
|
||||
local.block_on(&rt, async move {
|
||||
@@ -249,7 +255,9 @@ fn bench_unc_smarm() -> (u64, u128) {
|
||||
}
|
||||
|
||||
fn bench_unc_tokio_current() -> (u64, u128) {
|
||||
let rt = tokio::runtime::Builder::new_current_thread().build().unwrap();
|
||||
let rt = tokio::runtime::Builder::new_current_thread()
|
||||
.build()
|
||||
.unwrap();
|
||||
let start = Instant::now();
|
||||
let local = tokio::task::LocalSet::new();
|
||||
local.block_on(&rt, async move {
|
||||
@@ -297,8 +305,12 @@ fn bench_panic_smarm(threads: usize) -> (u64, u128) {
|
||||
}
|
||||
for h in handles {
|
||||
match h.join() {
|
||||
Ok(()) => { ok2.fetch_add(1, Ordering::Relaxed); }
|
||||
Err(_) => { err2.fetch_add(1, Ordering::Relaxed); }
|
||||
Ok(()) => {
|
||||
ok2.fetch_add(1, Ordering::Relaxed);
|
||||
}
|
||||
Err(_) => {
|
||||
err2.fetch_add(1, Ordering::Relaxed);
|
||||
}
|
||||
}
|
||||
}
|
||||
});
|
||||
@@ -312,7 +324,9 @@ fn bench_panic_tokio_current() -> (u64, u128) {
|
||||
let err = Arc::new(AtomicU64::new(0));
|
||||
let ok2 = ok.clone();
|
||||
let err2 = err.clone();
|
||||
let rt = tokio::runtime::Builder::new_current_thread().build().unwrap();
|
||||
let rt = tokio::runtime::Builder::new_current_thread()
|
||||
.build()
|
||||
.unwrap();
|
||||
let prev_hook = std::panic::take_hook();
|
||||
std::panic::set_hook(Box::new(|_| {}));
|
||||
let start = Instant::now();
|
||||
@@ -328,8 +342,12 @@ fn bench_panic_tokio_current() -> (u64, u128) {
|
||||
}
|
||||
for h in handles {
|
||||
match h.await {
|
||||
Ok(()) => { ok2.fetch_add(1, Ordering::Relaxed); }
|
||||
Err(_) => { err2.fetch_add(1, Ordering::Relaxed); }
|
||||
Ok(()) => {
|
||||
ok2.fetch_add(1, Ordering::Relaxed);
|
||||
}
|
||||
Err(_) => {
|
||||
err2.fetch_add(1, Ordering::Relaxed);
|
||||
}
|
||||
}
|
||||
}
|
||||
});
|
||||
@@ -361,8 +379,12 @@ fn bench_panic_tokio_multi() -> (u64, u128) {
|
||||
}
|
||||
for h in handles {
|
||||
match h.await {
|
||||
Ok(()) => { ok2.fetch_add(1, Ordering::Relaxed); }
|
||||
Err(_) => { err2.fetch_add(1, Ordering::Relaxed); }
|
||||
Ok(()) => {
|
||||
ok2.fetch_add(1, Ordering::Relaxed);
|
||||
}
|
||||
Err(_) => {
|
||||
err2.fetch_add(1, Ordering::Relaxed);
|
||||
}
|
||||
}
|
||||
}
|
||||
});
|
||||
@@ -375,7 +397,6 @@ fn bench_panic_tokio_multi() -> (u64, u128) {
|
||||
// main
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Knob helper — reads SMARM_ALLOC_INTERVAL / SMARM_TIMESLICE_CYCLES env vars
|
||||
// so the sweep script can override the preemption knobs without recompiling.
|
||||
@@ -384,10 +405,14 @@ fn bench_panic_tokio_multi() -> (u64, u128) {
|
||||
fn bench_cfg(threads: usize) -> smarm::runtime::Config {
|
||||
let mut cfg = smarm::runtime::Config::exact(threads);
|
||||
if let Ok(v) = std::env::var("SMARM_ALLOC_INTERVAL") {
|
||||
if let Ok(n) = v.parse::<u32>() { cfg = cfg.alloc_interval(n); }
|
||||
if let Ok(n) = v.parse::<u32>() {
|
||||
cfg = cfg.alloc_interval(n);
|
||||
}
|
||||
}
|
||||
if let Ok(v) = std::env::var("SMARM_TIMESLICE_CYCLES") {
|
||||
if let Ok(n) = v.parse::<u64>() { cfg = cfg.timeslice_cycles(n); }
|
||||
if let Ok(n) = v.parse::<u64>() {
|
||||
cfg = cfg.timeslice_cycles(n);
|
||||
}
|
||||
}
|
||||
cfg
|
||||
}
|
||||
@@ -397,7 +422,10 @@ fn main() {
|
||||
println!("smarm smarm-favored benchmarks");
|
||||
println!("available parallelism: {n} threads");
|
||||
let sets = env_sets();
|
||||
println!("ITERS={ITERS}×{sets} sets = {} samples (+1 warmup, discarded)", ITERS * sets);
|
||||
println!(
|
||||
"ITERS={ITERS}×{sets} sets = {} samples (+1 warmup, discarded)",
|
||||
ITERS * sets
|
||||
);
|
||||
println!(
|
||||
"RECURSE_DEPTH={RECURSE_DEPTH}, HOT_YIELDS={HOT_YIELDS}×2, \
|
||||
UNCONT_MSGS={UNCONT_MSGS}, PANIC_TASKS={PANIC_TASKS}"
|
||||
@@ -406,22 +434,30 @@ fn main() {
|
||||
// ---- 9. deep_recursion ----
|
||||
print_header(&format!("deep_recursion: depth {RECURSE_DEPTH}"));
|
||||
run_n("smarm 1-thread", ITERS, || bench_recurse_smarm(1));
|
||||
run_n(&format!("smarm {n}-thread"), ITERS, || bench_recurse_smarm(n));
|
||||
run_n(&format!("smarm {n}-thread"), ITERS, || {
|
||||
bench_recurse_smarm(n)
|
||||
});
|
||||
run_n("tokio current_thread", ITERS, bench_recurse_tokio_current);
|
||||
run_n("tokio multi-thread", ITERS, bench_recurse_tokio_multi);
|
||||
|
||||
// ---- 10. yield_in_hot_loop ----
|
||||
print_header(&format!("yield_in_hot_loop: 2 actors × {HOT_YIELDS} yields (single thread)"));
|
||||
print_header(&format!(
|
||||
"yield_in_hot_loop: 2 actors × {HOT_YIELDS} yields (single thread)"
|
||||
));
|
||||
run_n("smarm 1-thread", ITERS, bench_hot_smarm);
|
||||
run_n("tokio current_thread", ITERS, bench_hot_tokio_current);
|
||||
|
||||
// ---- 11. uncontended_channel ----
|
||||
print_header(&format!("uncontended_channel: 1→1, {UNCONT_MSGS} msgs (single thread)"));
|
||||
print_header(&format!(
|
||||
"uncontended_channel: 1→1, {UNCONT_MSGS} msgs (single thread)"
|
||||
));
|
||||
run_n("smarm 1-thread", ITERS, bench_unc_smarm);
|
||||
run_n("tokio current_thread", ITERS, bench_unc_tokio_current);
|
||||
|
||||
// ---- 12. catch_unwind_panics ----
|
||||
print_header(&format!("catch_unwind_panics: {PANIC_TASKS} tasks, 50% panic"));
|
||||
print_header(&format!(
|
||||
"catch_unwind_panics: {PANIC_TASKS} tasks, 50% panic"
|
||||
));
|
||||
run_n("smarm 1-thread", ITERS, || bench_panic_smarm(1));
|
||||
run_n(&format!("smarm {n}-thread"), ITERS, || bench_panic_smarm(n));
|
||||
run_n("tokio current_thread", ITERS, bench_panic_tokio_current);
|
||||
|
||||
+30
-5
@@ -73,7 +73,10 @@ fn variant() -> &'static str {
|
||||
}
|
||||
|
||||
fn env_usize(key: &str, default: usize) -> usize {
|
||||
std::env::var(key).ok().and_then(|v| v.parse().ok()).unwrap_or(default)
|
||||
std::env::var(key)
|
||||
.ok()
|
||||
.and_then(|v| v.parse().ok())
|
||||
.unwrap_or(default)
|
||||
}
|
||||
|
||||
// --------------------------------------------------------------------------
|
||||
@@ -226,7 +229,11 @@ fn main() {
|
||||
let mean_cyc = pooled_cyc.iter().map(|&v| v as f64).sum::<f64>() / n.max(1) as f64;
|
||||
// Derived effective frequency: cycles per ns = GHz. Cross-checks the two
|
||||
// lenses against the box's known base clock.
|
||||
let derived_ghz = if mean_ns > 0.0 { mean_cyc / mean_ns } else { 0.0 };
|
||||
let derived_ghz = if mean_ns > 0.0 {
|
||||
mean_cyc / mean_ns
|
||||
} else {
|
||||
0.0
|
||||
};
|
||||
|
||||
let p50 = pct(&pooled_ns, 50.0);
|
||||
let p90 = pct(&pooled_ns, 90.0);
|
||||
@@ -241,8 +248,14 @@ fn main() {
|
||||
" rounds={} warmup={} runs={} (instrumentation floor: {} ns / {} cyc, subtracted)",
|
||||
rounds, warmup, runs, floor_ns, floor_cyc
|
||||
);
|
||||
println!(" {:<10} {:<10} {:<10} {:<10} {:<10}", "p50 ns", "p90 ns", "p99 ns", "min ns", "max ns");
|
||||
println!(" {:<10} {:<10} {:<10} {:<10} {:<10}", p50, p90, p99, lo, hi);
|
||||
println!(
|
||||
" {:<10} {:<10} {:<10} {:<10} {:<10}",
|
||||
"p50 ns", "p90 ns", "p99 ns", "min ns", "max ns"
|
||||
);
|
||||
println!(
|
||||
" {:<10} {:<10} {:<10} {:<10} {:<10}",
|
||||
p50, p90, p99, lo, hi
|
||||
);
|
||||
println!(
|
||||
" mean {:.1} ns | mean {:.0} cyc | derived {:.3} GHz",
|
||||
mean_ns, mean_cyc, derived_ghz
|
||||
@@ -251,6 +264,18 @@ fn main() {
|
||||
// Greppable line — same spirit as SPINCSV.
|
||||
println!(
|
||||
"SWITCHCSV,{},{},{},{},{},{},{},{},{},{},{:.1},{:.0},{:.3}",
|
||||
variant(), mode, rounds, runs, n, p50, p90, p99, lo, hi, mean_ns, mean_cyc, derived_ghz
|
||||
variant(),
|
||||
mode,
|
||||
rounds,
|
||||
runs,
|
||||
n,
|
||||
p50,
|
||||
p90,
|
||||
p99,
|
||||
lo,
|
||||
hi,
|
||||
mean_ns,
|
||||
mean_cyc,
|
||||
derived_ghz
|
||||
);
|
||||
}
|
||||
|
||||
+107
-35
@@ -36,7 +36,9 @@ use std::time::{Duration, Instant};
|
||||
const ITERS: u32 = 15;
|
||||
|
||||
fn available_threads() -> usize {
|
||||
std::thread::available_parallelism().map(|n| n.get()).unwrap_or(1)
|
||||
std::thread::available_parallelism()
|
||||
.map(|n| n.get())
|
||||
.unwrap_or(1)
|
||||
}
|
||||
|
||||
fn env_sets() -> u32 {
|
||||
@@ -84,8 +86,8 @@ fn run_n<F: FnMut() -> (u64, u128)>(name: &str, n: u32, mut f: F) {
|
||||
// 5. spawn_storm_busy — workers loaded, then storm of zero-work spawns
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
const STORM_BACKGROUND: u64 = 8; // number of background "busy" actors
|
||||
const STORM_SPAWN: u64 = 10_000; // zero-work spawns to time
|
||||
const STORM_BACKGROUND: u64 = 8; // number of background "busy" actors
|
||||
const STORM_SPAWN: u64 = 10_000; // zero-work spawns to time
|
||||
|
||||
fn bench_storm_smarm(threads: usize) -> (u64, u128) {
|
||||
let counter = Arc::new(AtomicU64::new(0));
|
||||
@@ -114,11 +116,15 @@ fn bench_storm_smarm(threads: usize) -> (u64, u128) {
|
||||
cc.fetch_add(1, Ordering::Relaxed);
|
||||
}));
|
||||
}
|
||||
for h in handles { h.join().unwrap(); }
|
||||
for h in handles {
|
||||
h.join().unwrap();
|
||||
}
|
||||
|
||||
// Tear down background.
|
||||
s2.store(true, Ordering::Relaxed);
|
||||
for h in bg_handles { h.join().unwrap(); }
|
||||
for h in bg_handles {
|
||||
h.join().unwrap();
|
||||
}
|
||||
});
|
||||
(counter.load(Ordering::Relaxed), start.elapsed().as_micros())
|
||||
}
|
||||
@@ -129,7 +135,9 @@ fn bench_storm_tokio_current() -> (u64, u128) {
|
||||
let c2 = counter.clone();
|
||||
let s2 = stop.clone();
|
||||
|
||||
let rt = tokio::runtime::Builder::new_current_thread().build().unwrap();
|
||||
let rt = tokio::runtime::Builder::new_current_thread()
|
||||
.build()
|
||||
.unwrap();
|
||||
let start = Instant::now();
|
||||
let local = tokio::task::LocalSet::new();
|
||||
local.block_on(&rt, async move {
|
||||
@@ -149,9 +157,13 @@ fn bench_storm_tokio_current() -> (u64, u128) {
|
||||
cc.fetch_add(1, Ordering::Relaxed);
|
||||
}));
|
||||
}
|
||||
for h in handles { let _ = h.await; }
|
||||
for h in handles {
|
||||
let _ = h.await;
|
||||
}
|
||||
s2.store(true, Ordering::Relaxed);
|
||||
for h in bg_handles { let _ = h.await; }
|
||||
for h in bg_handles {
|
||||
let _ = h.await;
|
||||
}
|
||||
});
|
||||
(counter.load(Ordering::Relaxed), start.elapsed().as_micros())
|
||||
}
|
||||
@@ -184,9 +196,13 @@ fn bench_storm_tokio_multi() -> (u64, u128) {
|
||||
cc.fetch_add(1, Ordering::Relaxed);
|
||||
}));
|
||||
}
|
||||
for h in handles { let _ = h.await; }
|
||||
for h in handles {
|
||||
let _ = h.await;
|
||||
}
|
||||
s2.store(true, Ordering::Relaxed);
|
||||
for h in bg_handles { let _ = h.await; }
|
||||
for h in bg_handles {
|
||||
let _ = h.await;
|
||||
}
|
||||
});
|
||||
(counter.load(Ordering::Relaxed), start.elapsed().as_micros())
|
||||
}
|
||||
@@ -219,14 +235,21 @@ fn bench_mpsc_smarm(threads: usize) -> (u64, u128) {
|
||||
}
|
||||
let _ = count; // discard; run() closure must return ()
|
||||
});
|
||||
for h in prod_handles { h.join().unwrap(); }
|
||||
for h in prod_handles {
|
||||
h.join().unwrap();
|
||||
}
|
||||
let _ = consumer.join().unwrap();
|
||||
});
|
||||
(MPSC_PRODUCERS * MPSC_PER_PRODUCER, start.elapsed().as_micros())
|
||||
(
|
||||
MPSC_PRODUCERS * MPSC_PER_PRODUCER,
|
||||
start.elapsed().as_micros(),
|
||||
)
|
||||
}
|
||||
|
||||
fn bench_mpsc_tokio_current() -> (u64, u128) {
|
||||
let rt = tokio::runtime::Builder::new_current_thread().build().unwrap();
|
||||
let rt = tokio::runtime::Builder::new_current_thread()
|
||||
.build()
|
||||
.unwrap();
|
||||
let start = Instant::now();
|
||||
let local = tokio::task::LocalSet::new();
|
||||
local.block_on(&rt, async move {
|
||||
@@ -248,10 +271,15 @@ fn bench_mpsc_tokio_current() -> (u64, u128) {
|
||||
}
|
||||
count
|
||||
});
|
||||
for h in prod_handles { let _ = h.await; }
|
||||
for h in prod_handles {
|
||||
let _ = h.await;
|
||||
}
|
||||
let _ = consumer.await;
|
||||
});
|
||||
(MPSC_PRODUCERS * MPSC_PER_PRODUCER, start.elapsed().as_micros())
|
||||
(
|
||||
MPSC_PRODUCERS * MPSC_PER_PRODUCER,
|
||||
start.elapsed().as_micros(),
|
||||
)
|
||||
}
|
||||
|
||||
fn bench_mpsc_tokio_multi() -> (u64, u128) {
|
||||
@@ -279,10 +307,15 @@ fn bench_mpsc_tokio_multi() -> (u64, u128) {
|
||||
}
|
||||
count
|
||||
});
|
||||
for h in prod_handles { let _ = h.await; }
|
||||
for h in prod_handles {
|
||||
let _ = h.await;
|
||||
}
|
||||
let _ = consumer.await;
|
||||
});
|
||||
(MPSC_PRODUCERS * MPSC_PER_PRODUCER, start.elapsed().as_micros())
|
||||
(
|
||||
MPSC_PRODUCERS * MPSC_PER_PRODUCER,
|
||||
start.elapsed().as_micros(),
|
||||
)
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
@@ -308,7 +341,9 @@ fn bench_timers_smarm(threads: usize) -> (u64, u128) {
|
||||
smarm::sleep(Duration::from_millis(ms));
|
||||
}));
|
||||
}
|
||||
for h in handles { h.join().unwrap(); }
|
||||
for h in handles {
|
||||
h.join().unwrap();
|
||||
}
|
||||
});
|
||||
(TIMER_ACTORS, start.elapsed().as_micros())
|
||||
}
|
||||
@@ -328,7 +363,9 @@ fn bench_timers_tokio_current() -> (u64, u128) {
|
||||
tokio::time::sleep(Duration::from_millis(ms)).await;
|
||||
}));
|
||||
}
|
||||
for h in handles { let _ = h.await; }
|
||||
for h in handles {
|
||||
let _ = h.await;
|
||||
}
|
||||
});
|
||||
(TIMER_ACTORS, start.elapsed().as_micros())
|
||||
}
|
||||
@@ -348,7 +385,9 @@ fn bench_timers_tokio_multi() -> (u64, u128) {
|
||||
tokio::time::sleep(Duration::from_millis(ms)).await;
|
||||
}));
|
||||
}
|
||||
for h in handles { let _ = h.await; }
|
||||
for h in handles {
|
||||
let _ = h.await;
|
||||
}
|
||||
});
|
||||
(TIMER_ACTORS, start.elapsed().as_micros())
|
||||
}
|
||||
@@ -361,11 +400,22 @@ const SCALING_N: u64 = 400_000;
|
||||
const SCALING_WORKERS: u64 = 64;
|
||||
|
||||
fn is_prime(n: u64) -> bool {
|
||||
if n < 2 { return false; }
|
||||
if n < 4 { return true; }
|
||||
if n % 2 == 0 { return false; }
|
||||
if n < 2 {
|
||||
return false;
|
||||
}
|
||||
if n < 4 {
|
||||
return true;
|
||||
}
|
||||
if n % 2 == 0 {
|
||||
return false;
|
||||
}
|
||||
let mut i = 3u64;
|
||||
while i * i <= n { if n % i == 0 { return false; } i += 2; }
|
||||
while i * i <= n {
|
||||
if n % i == 0 {
|
||||
return false;
|
||||
}
|
||||
i += 2;
|
||||
}
|
||||
true
|
||||
}
|
||||
|
||||
@@ -376,7 +426,11 @@ fn count_primes(lo: u64, hi: u64) -> u64 {
|
||||
fn scaling_slice(w: u64) -> (u64, u64) {
|
||||
let per = SCALING_N / SCALING_WORKERS;
|
||||
let lo = w * per;
|
||||
let hi = if w + 1 == SCALING_WORKERS { SCALING_N } else { lo + per };
|
||||
let hi = if w + 1 == SCALING_WORKERS {
|
||||
SCALING_N
|
||||
} else {
|
||||
lo + per
|
||||
};
|
||||
(lo, hi)
|
||||
}
|
||||
|
||||
@@ -393,7 +447,9 @@ fn bench_scaling_smarm(threads: usize) -> (u64, u128) {
|
||||
tc.fetch_add(count_primes(lo, hi), Ordering::Relaxed);
|
||||
}));
|
||||
}
|
||||
for h in handles { h.join().unwrap(); }
|
||||
for h in handles {
|
||||
h.join().unwrap();
|
||||
}
|
||||
});
|
||||
(total.load(Ordering::Relaxed), start.elapsed().as_micros())
|
||||
}
|
||||
@@ -415,7 +471,9 @@ fn bench_scaling_tokio_multi(threads: usize) -> (u64, u128) {
|
||||
tc.fetch_add(count_primes(lo, hi), Ordering::Relaxed);
|
||||
}));
|
||||
}
|
||||
for h in handles { let _ = h.await; }
|
||||
for h in handles {
|
||||
let _ = h.await;
|
||||
}
|
||||
});
|
||||
(total.load(Ordering::Relaxed), start.elapsed().as_micros())
|
||||
}
|
||||
@@ -424,7 +482,6 @@ fn bench_scaling_tokio_multi(threads: usize) -> (u64, u128) {
|
||||
// main
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Knob helper — reads SMARM_ALLOC_INTERVAL / SMARM_TIMESLICE_CYCLES env vars
|
||||
// so the sweep script can override the preemption knobs without recompiling.
|
||||
@@ -433,10 +490,14 @@ fn bench_scaling_tokio_multi(threads: usize) -> (u64, u128) {
|
||||
fn bench_cfg(threads: usize) -> smarm::runtime::Config {
|
||||
let mut cfg = smarm::runtime::Config::exact(threads);
|
||||
if let Ok(v) = std::env::var("SMARM_ALLOC_INTERVAL") {
|
||||
if let Ok(n) = v.parse::<u32>() { cfg = cfg.alloc_interval(n); }
|
||||
if let Ok(n) = v.parse::<u32>() {
|
||||
cfg = cfg.alloc_interval(n);
|
||||
}
|
||||
}
|
||||
if let Ok(v) = std::env::var("SMARM_TIMESLICE_CYCLES") {
|
||||
if let Ok(n) = v.parse::<u64>() { cfg = cfg.timeslice_cycles(n); }
|
||||
if let Ok(n) = v.parse::<u64>() {
|
||||
cfg = cfg.timeslice_cycles(n);
|
||||
}
|
||||
}
|
||||
cfg
|
||||
}
|
||||
@@ -446,7 +507,10 @@ fn main() {
|
||||
println!("smarm tokio-favored benchmarks");
|
||||
println!("available parallelism: {n} threads");
|
||||
let sets = env_sets();
|
||||
println!("ITERS={ITERS}×{sets} sets = {} samples (+1 warmup, discarded)", ITERS * sets);
|
||||
println!(
|
||||
"ITERS={ITERS}×{sets} sets = {} samples (+1 warmup, discarded)",
|
||||
ITERS * sets
|
||||
);
|
||||
println!(
|
||||
"STORM_BACKGROUND={STORM_BACKGROUND}, STORM_SPAWN={STORM_SPAWN}, \
|
||||
MPSC={MPSC_PRODUCERS}×{MPSC_PER_PRODUCER}, \
|
||||
@@ -477,7 +541,9 @@ fn main() {
|
||||
"many_timers: {TIMER_ACTORS} actors sleeping {TIMER_MIN_MS}–{TIMER_MAX_MS} ms"
|
||||
));
|
||||
run_n("smarm 1-thread", ITERS, || bench_timers_smarm(1));
|
||||
run_n(&format!("smarm {n}-thread"), ITERS, || bench_timers_smarm(n));
|
||||
run_n(&format!("smarm {n}-thread"), ITERS, || {
|
||||
bench_timers_smarm(n)
|
||||
});
|
||||
run_n("tokio current_thread", ITERS, bench_timers_tokio_current);
|
||||
run_n("tokio multi-thread", ITERS, bench_timers_tokio_multi);
|
||||
|
||||
@@ -487,13 +553,19 @@ fn main() {
|
||||
));
|
||||
let sweep: Vec<usize> = {
|
||||
let mut v = vec![1usize, 2, 4];
|
||||
if n > 4 && !v.contains(&n) { v.push(n); }
|
||||
if n > 4 && !v.contains(&n) {
|
||||
v.push(n);
|
||||
}
|
||||
v.into_iter().filter(|t| *t <= n).collect()
|
||||
};
|
||||
for t in &sweep {
|
||||
run_n(&format!("smarm {t}-thread"), ITERS, || bench_scaling_smarm(*t));
|
||||
run_n(&format!("smarm {t}-thread"), ITERS, || {
|
||||
bench_scaling_smarm(*t)
|
||||
});
|
||||
}
|
||||
for t in &sweep {
|
||||
run_n(&format!("tokio multi {t}-thread"), ITERS, || bench_scaling_tokio_multi(*t));
|
||||
run_n(&format!("tokio multi {t}-thread"), ITERS, || {
|
||||
bench_scaling_tokio_multi(*t)
|
||||
});
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,11 @@
|
||||
fn main() {
|
||||
// RFC 019 §7 test canary (agreed Q3): compiled without stack-clash
|
||||
// protection so its 96 KiB local is a genuine one-displacement guard
|
||||
// jumper; distro-hardened compilers would otherwise probe it page-wise
|
||||
// and defeat the test's purpose.
|
||||
cc::Build::new()
|
||||
.file("canary/canary.c")
|
||||
.flag_if_supported("-fno-stack-clash-protection")
|
||||
.compile("smarm_canary");
|
||||
println!("cargo:rerun-if-changed=canary/canary.c");
|
||||
}
|
||||
@@ -0,0 +1,14 @@
|
||||
/* RFC 019 §7 FFI canary: an honest unprobed C frame with a 96 KiB local,
|
||||
* touched from its LOW end first — the exact "one sub rsp steps over a small
|
||||
* guard" pattern the RFC's motivating incident hit (a cargo-vendored gz
|
||||
* build; cc-invoked builds do not enable -fstack-clash-protection, and this
|
||||
* file pins that off explicitly so the canary stays a canary even on
|
||||
* hardened-default toolchains). */
|
||||
void smarm_canary_burn(void) {
|
||||
volatile char buf[96 * 1024];
|
||||
buf[0] = 1; /* deepest address first */
|
||||
for (unsigned i = 0; i < sizeof buf; i += 4096) {
|
||||
buf[i] = (char)i;
|
||||
}
|
||||
buf[sizeof buf - 1] = 1;
|
||||
}
|
||||
@@ -75,7 +75,7 @@ genuine advantage over tokio's task abort model.
|
||||
|
||||
### Spawn-heavy workloads (19–70×)
|
||||
|
||||
Every smarm actor `mmap`s a 64 KiB stack with a guard page. This is
|
||||
Every smarm actor `mmap`s a 64 KiB stack reserve with a 64 KiB PROT_NONE guard below (both per-actor configurable since RFC 019; the reserve is demand-paged). This is
|
||||
a syscall. Tokio tasks are heap-allocated state machines — no stack,
|
||||
no syscall, ~100 bytes each. For workloads that spawn thousands of
|
||||
short-lived actors per second, this is a structural disadvantage.
|
||||
|
||||
@@ -1620,8 +1620,9 @@
|
||||
wait: <code>select</code> priority is <strong>Down arms › Watcher arm › info channels (declaration order) ›
|
||||
inbox</strong>, rebuilt each turn. A hot inbox can't starve a death notice or a system message;
|
||||
conversely a hot info channel <em>can</em> starve the inbox — deliberately. A closed info arm is
|
||||
silently dropped from the set; a closed <em>inbox</em> (every <code>ServerRef</code> gone) is graceful
|
||||
shutdown.</p>
|
||||
silently dropped from the set. The inbox never closes — the loop holds one sender for its whole
|
||||
life, so a <code>GenServerRef</code> is an address, not an owner: the server ends only by
|
||||
<code>StopHandle::stop</code>, a shutdown, a hard stop, or a panic.</p>
|
||||
|
||||
<h3>Death needs no monitor</h3>
|
||||
<p>Server death detection falls out of channel closure. Already dead → the inbox is closed and
|
||||
@@ -1700,7 +1701,7 @@
|
||||
</div>
|
||||
<div class="module-card">
|
||||
<div class="module-name" style="color:var(--red)">Panics in <code>terminate()</code></div>
|
||||
<p>gen_server's <code>terminate()</code> runs from a drop guard, possibly mid-unwind. A panic inside it during an unwind is a double panic → process abort, no supervision tree to save you. Keep it cheap, non-blocking, non-panicking.</p>
|
||||
<p>gen_server's <code>terminate()</code> runs from a drop guard, possibly mid-unwind. A panic inside it during an unwind is a double panic → process abort, no supervision tree to save you. On the panic and hard-stop paths keep it cheap, non-blocking, non-panicking. Only the graceful path (<code>handle_shutdown → Exit</code>, <code>StopHandle::stop</code>) runs it outside an unwind, where it may do real work.</p>
|
||||
</div>
|
||||
<div class="module-card">
|
||||
<div class="module-name" style="color:var(--yellow)">Cold locks are leaf locks</div>
|
||||
|
||||
@@ -67,7 +67,9 @@ fn main() {
|
||||
println!("calibration: {per_us} work iters/µs");
|
||||
let work_us = move |us: u64| work_iters(us * per_us);
|
||||
|
||||
let cores = std::thread::available_parallelism().map(|n| n.get()).unwrap_or(1);
|
||||
let cores = std::thread::available_parallelism()
|
||||
.map(|n| n.get())
|
||||
.unwrap_or(1);
|
||||
println!("cores: {cores}");
|
||||
if cores < 4 {
|
||||
println!("probe: SKIPPED (needs the stages in parallel)");
|
||||
|
||||
@@ -35,8 +35,8 @@
|
||||
|
||||
#![deny(dead_code, unreachable_patterns)]
|
||||
|
||||
use smarm::gen_statem::{spawn, Cx, GenStatemRef, Machine, Reply, Resolution, Step};
|
||||
use smarm::run;
|
||||
use smarm::gen_statem::{spawn, Cx, Machine, Reply, Resolution, Step, GenStatemRef};
|
||||
|
||||
// === user types ============================================================
|
||||
|
||||
@@ -123,7 +123,11 @@ impl DoorSm {
|
||||
fn start(init: Door) -> GenStatemRef<DoorSm> {
|
||||
spawn(DoorSm {
|
||||
state: init,
|
||||
data: Data { enters: 0, pushes: 0, knocks: 0 },
|
||||
data: Data {
|
||||
enters: 0,
|
||||
pushes: 0,
|
||||
knocks: 0,
|
||||
},
|
||||
})
|
||||
}
|
||||
|
||||
@@ -193,16 +197,12 @@ impl Machine for DoorSm {
|
||||
(Door::Closed, Ev::Cast(Cast::Push | Cast::Unlock(_))) => Resolution::Unhandled,
|
||||
|
||||
// --- Locked (branching row: handler picks within UnlockOutcome) -
|
||||
(Door::Locked, Ev::Cast(Cast::Unlock(key))) => {
|
||||
Resolution::To(on_unlock(key).into())
|
||||
}
|
||||
(Door::Locked, Ev::Cast(Cast::Unlock(key))) => Resolution::To(on_unlock(key).into()),
|
||||
// Routed out in phase 1; listed only to keep this match total.
|
||||
(Door::Locked, Ev::Cast(Cast::Knock)) => {
|
||||
unreachable!("postponed event is replayed, not dispatched here")
|
||||
}
|
||||
(Door::Locked, Ev::Cast(Cast::Push | Cast::Pull | Cast::Lock)) => {
|
||||
Resolution::Unhandled
|
||||
}
|
||||
(Door::Locked, Ev::Cast(Cast::Push | Cast::Pull | Cast::Lock)) => Resolution::Unhandled,
|
||||
|
||||
// --- state-independent queries (reply, then stay) ---------------
|
||||
(_, Ev::Call(Call::GetState(r))) => {
|
||||
|
||||
@@ -18,8 +18,8 @@
|
||||
// dispatch's own unreachable_patterns internally.
|
||||
|
||||
use smarm::gen_statem;
|
||||
use smarm::run;
|
||||
use smarm::gen_statem::Reply;
|
||||
use smarm::run;
|
||||
|
||||
// === user types (identical to gen_statem_expanded.rs) =========================
|
||||
|
||||
@@ -135,7 +135,14 @@ gen_statem! {
|
||||
|
||||
fn main() {
|
||||
run(|| {
|
||||
let door = DoorSm::start(Door::Closed, Data { enters: 0, pushes: 0, knocks: 0 });
|
||||
let door = DoorSm::start(
|
||||
Door::Closed,
|
||||
Data {
|
||||
enters: 0,
|
||||
pushes: 0,
|
||||
knocks: 0,
|
||||
},
|
||||
);
|
||||
|
||||
door.send(Ev::Cast(Cast::Lock)).unwrap(); // Closed -> Locked
|
||||
door.send(Ev::Cast(Cast::Knock)).unwrap(); // Locked: postponed (not yet counted)
|
||||
|
||||
@@ -0,0 +1,139 @@
|
||||
//! Graceful shutdown, end to end: a supervised app tree, a server that
|
||||
//! drains before it exits, and the two ways the whole thing winds down.
|
||||
//!
|
||||
//! Stopping an actor comes in two strengths, as in OTP:
|
||||
//! - `request_stop(pid)` = `exit(Pid, kill)`: cooperative hard stop,
|
||||
//! unwinds at the next observation point.
|
||||
//! - `request_shutdown(pid)` = `exit(Pid, shutdown)`: a trapping target gets
|
||||
//! an `ExitSignal { reason: Shutdown }` and winds
|
||||
//! down on its own terms; a non-trapping one is
|
||||
//! stopped outright.
|
||||
//!
|
||||
//! A supervisor traps exits. `request_shutdown(sup)` runs its ordered
|
||||
//! shutdown — children in reverse start order, each per its `ChildSpec`
|
||||
//! `Shutdown` policy (`Timeout(d)` default 5s, `Infinity`, `BrutalKill`) —
|
||||
//! and the supervisor then returns normally.
|
||||
//!
|
||||
//! Two triggers are shown:
|
||||
//! 1. **Root exit.** The run's root actor returning means "the program is
|
||||
//! done": the runtime delivers `request_shutdown` to every top-level actor
|
||||
//! (here: the supervisor). Trapping actors may keep running to drain and
|
||||
//! end the run when they stop themselves; non-trapping ones are stopped.
|
||||
//! 2. **An outside thread** (e.g. a signal handler) driving it via
|
||||
//! `RuntimeHandle::request_shutdown` on the supervisor — the root then
|
||||
//! just waits for the tree to come down.
|
||||
|
||||
use smarm::gen_server::{
|
||||
GenServer, GenServerBuilder, GenServerCtx, GenServerName, ShutdownAction, StopHandle,
|
||||
TimerHandle,
|
||||
};
|
||||
use smarm::supervisor::{ChildSpec, OneForOne, Restart, Shutdown};
|
||||
use smarm::{sleep, spawn};
|
||||
use std::thread;
|
||||
use std::time::Duration;
|
||||
|
||||
/// A server with in-flight work: on shutdown it stops accepting, finishes what
|
||||
/// it has (simulated with a ticking timer), then ends itself.
|
||||
struct Drainer {
|
||||
pending: u32,
|
||||
stop: Option<StopHandle<Drainer>>,
|
||||
timer: Option<TimerHandle<Drainer>>,
|
||||
}
|
||||
|
||||
impl GenServer for Drainer {
|
||||
type Call = ();
|
||||
type Reply = ();
|
||||
type Cast = ();
|
||||
type Info = ();
|
||||
type Timer = ();
|
||||
|
||||
fn init(&mut self, ctx: &GenServerCtx<Self>) {
|
||||
ctx.trap_exit(); // opt in: shutdown arrives as handle_shutdown
|
||||
self.stop = Some(ctx.stop_handle());
|
||||
self.timer = Some(ctx.timer());
|
||||
}
|
||||
fn handle_call(&mut self, _: ()) {}
|
||||
fn handle_cast(&mut self, _: ()) {}
|
||||
fn handle_shutdown(&mut self) -> ShutdownAction {
|
||||
println!(
|
||||
"drainer: shutdown requested, {} items pending",
|
||||
self.pending
|
||||
);
|
||||
self.timer
|
||||
.as_ref()
|
||||
.unwrap()
|
||||
.tick_every(Duration::from_millis(20), ());
|
||||
ShutdownAction::Continue // keep serving until drained
|
||||
}
|
||||
fn handle_timer(&mut self, _: ()) {
|
||||
self.pending -= 1;
|
||||
if self.pending == 0 {
|
||||
println!("drainer: drained, stopping");
|
||||
self.stop.as_ref().unwrap().stop(); // normal exit
|
||||
}
|
||||
}
|
||||
fn terminate(&mut self) {
|
||||
// Graceful path: this runs on the normal path and may block.
|
||||
println!("drainer: terminate");
|
||||
}
|
||||
}
|
||||
|
||||
/// The server's name: how the rest of the app reaches it (and the only handle
|
||||
/// that survives a restart).
|
||||
const DRAINER: GenServerName<Drainer> = GenServerName::new("drainer");
|
||||
|
||||
fn app_tree() -> OneForOne {
|
||||
OneForOne::new()
|
||||
.child(
|
||||
ChildSpec::new(Restart::Permanent, || {
|
||||
// A plain worker that does not trap: stopped outright on shutdown.
|
||||
loop {
|
||||
sleep(Duration::from_millis(10));
|
||||
}
|
||||
})
|
||||
.shutdown(Shutdown::Timeout(Duration::from_millis(100))),
|
||||
)
|
||||
// A gen_server is a direct child: `named(N).run()` runs the loop as
|
||||
// the child actor itself, so the supervisor's shutdown arrives as
|
||||
// `handle_shutdown` and a restart re-binds the name.
|
||||
.child(
|
||||
ChildSpec::new(Restart::Permanent, || {
|
||||
GenServerBuilder::new(Drainer {
|
||||
pending: 3,
|
||||
stop: None,
|
||||
timer: None,
|
||||
})
|
||||
.named(DRAINER)
|
||||
.run()
|
||||
.expect("drainer name is free");
|
||||
})
|
||||
.shutdown(Shutdown::Infinity),
|
||||
)
|
||||
}
|
||||
|
||||
fn main() {
|
||||
println!("--- 1. root exit drives the shutdown ---");
|
||||
smarm::run(|| {
|
||||
spawn(|| app_tree().run());
|
||||
sleep(Duration::from_millis(50)); // the app "runs" for a while
|
||||
// Returning here asks the supervisor to shut down; the run ends when
|
||||
// the tree — drainer included — is gone.
|
||||
});
|
||||
|
||||
println!("--- 2. an outside thread drives the shutdown ---");
|
||||
let rt = smarm::init(smarm::Config::default());
|
||||
let handle = rt.handle(); // Send + Sync; grab it before run
|
||||
rt.run(move || {
|
||||
let sup = spawn(|| app_tree().run());
|
||||
let sup_pid = sup.pid();
|
||||
// Stand-in for a SIGTERM handler thread.
|
||||
thread::spawn(move || {
|
||||
thread::sleep(Duration::from_millis(50));
|
||||
println!("signal thread: requesting shutdown");
|
||||
handle.request_shutdown(sup_pid);
|
||||
});
|
||||
sup.join()
|
||||
.expect("supervisor returns normally after ordered shutdown");
|
||||
println!("supervisor down; root returns");
|
||||
});
|
||||
}
|
||||
@@ -6,7 +6,9 @@
|
||||
//! every use — so the address keeps working across a supervised restart, with
|
||||
//! no stale [`GenServerRef`] to refresh.
|
||||
|
||||
use smarm::{call, cast, run, whereis_server, GenServer, GenServerBuilder, GenServerName, GenServerRef};
|
||||
use smarm::{
|
||||
call, cast, run, whereis_server, GenServer, GenServerBuilder, GenServerName, GenServerRef,
|
||||
};
|
||||
|
||||
/// A counter server: synchronous `Get`, asynchronous `Inc` / `Add`.
|
||||
struct Counter {
|
||||
@@ -64,6 +66,12 @@ fn main() {
|
||||
let svc: Option<GenServerRef<Counter>> = whereis_server(COUNTER);
|
||||
if let Some(svc) = svc {
|
||||
let _ = svc.call(Query::Get);
|
||||
// A named server is pinned alive by the registry, so dropping refs
|
||||
// does not end it. Stop it explicitly: `shutdown()` asks politely
|
||||
// (a trapping server drains first; this one is stopped outright)
|
||||
// and waits until it is gone. Left running, the root's return
|
||||
// would shut it down the same way — see examples/graceful_shutdown.rs.
|
||||
svc.shutdown();
|
||||
}
|
||||
});
|
||||
}
|
||||
|
||||
+13
-3
@@ -15,7 +15,9 @@
|
||||
//! `call`, nothing more.
|
||||
|
||||
use smarm::observer::{self, ObserverReply, ObserverRequest};
|
||||
use smarm::{channel, register, run, spawn, ActorState, Name, RuntimeSnapshot, RuntimeTree, TreeNode};
|
||||
use smarm::{
|
||||
channel, register, run, spawn, ActorState, Name, RuntimeSnapshot, RuntimeTree, TreeNode,
|
||||
};
|
||||
|
||||
const ECHO: Name<u64> = Name::new("echo");
|
||||
|
||||
@@ -31,7 +33,11 @@ fn state_glyph(s: ActorState) -> &'static str {
|
||||
|
||||
/// A `ps`-style table over the flat snapshot.
|
||||
fn print_snapshot(snap: &RuntimeSnapshot) {
|
||||
println!("snapshot (format v{}, {} actors)", snap.format_version, snap.actors.len());
|
||||
println!(
|
||||
"snapshot (format v{}, {} actors)",
|
||||
snap.format_version,
|
||||
snap.actors.len()
|
||||
);
|
||||
println!(
|
||||
" {:<10} {:<9} {:<10} {:>4} {:>4} {:>4} {:>4} {:>5} {}",
|
||||
"pid", "state", "parent", "mon", "lnk", "joi", "mbox", "msgs", "names"
|
||||
@@ -52,7 +58,11 @@ fn print_snapshot(snap: &RuntimeSnapshot) {
|
||||
a.joiners,
|
||||
a.mailbox_depth,
|
||||
a.messages_received,
|
||||
if a.names.is_empty() { "-".to_string() } else { a.names.join(",") },
|
||||
if a.names.is_empty() {
|
||||
"-".to_string()
|
||||
} else {
|
||||
a.names.join(",")
|
||||
},
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
+1
-1
@@ -99,7 +99,7 @@ pub extern "C-unwind" fn trampoline() {
|
||||
};
|
||||
|
||||
let outcome = match panic::catch_unwind(panic::AssertUnwindSafe(b)) {
|
||||
Ok(()) => Outcome::Exit,
|
||||
Ok(()) => Outcome::Exit,
|
||||
Err(payload) => {
|
||||
if payload.is::<StopSentinel>() {
|
||||
Outcome::Stopped
|
||||
|
||||
+20
-10
@@ -342,8 +342,7 @@ mod inner {
|
||||
// Count the loss in would-be delta terms so the audit's columns
|
||||
// compare directly against `injected_cycles`.
|
||||
DISCARD_OVERMAX_N.fetch_add(1, Ordering::Relaxed);
|
||||
DISCARD_OVERMAX_CYCLES
|
||||
.fetch_add(interval.saturating_mul(pct) / 100, Ordering::Relaxed);
|
||||
DISCARD_OVERMAX_CYCLES.fetch_add(interval.saturating_mul(pct) / 100, Ordering::Relaxed);
|
||||
return;
|
||||
}
|
||||
let delta = interval.saturating_mul(pct) / 100;
|
||||
@@ -419,8 +418,7 @@ mod inner {
|
||||
let gap = preempt::rdtsc()
|
||||
.saturating_sub(desched_tsc)
|
||||
.min(MAX_SAMPLE_CYCLES);
|
||||
OFFCPU_IN_SITE_CYCLES
|
||||
.fetch_add(gap.saturating_mul(pct) / 100, Ordering::Relaxed);
|
||||
OFFCPU_IN_SITE_CYCLES.fetch_add(gap.saturating_mul(pct) / 100, Ordering::Relaxed);
|
||||
OFFCPU_IN_SITE_N.fetch_add(1, Ordering::Relaxed);
|
||||
}
|
||||
}
|
||||
@@ -533,7 +531,9 @@ mod inner {
|
||||
park_forgiven_cycles: self
|
||||
.park_forgiven_cycles
|
||||
.saturating_sub(before.park_forgiven_cycles),
|
||||
drop_park_cycles: self.drop_park_cycles.saturating_sub(before.drop_park_cycles),
|
||||
drop_park_cycles: self
|
||||
.drop_park_cycles
|
||||
.saturating_sub(before.drop_park_cycles),
|
||||
drop_park_n: self.drop_park_n.saturating_sub(before.drop_park_n),
|
||||
drop_yield_cycles: self
|
||||
.drop_yield_cycles
|
||||
@@ -542,12 +542,18 @@ mod inner {
|
||||
discard_overmax_cycles: self
|
||||
.discard_overmax_cycles
|
||||
.saturating_sub(before.discard_overmax_cycles),
|
||||
discard_overmax_n: self.discard_overmax_n.saturating_sub(before.discard_overmax_n),
|
||||
discard_unarmed_n: self.discard_unarmed_n.saturating_sub(before.discard_unarmed_n),
|
||||
discard_overmax_n: self
|
||||
.discard_overmax_n
|
||||
.saturating_sub(before.discard_overmax_n),
|
||||
discard_unarmed_n: self
|
||||
.discard_unarmed_n
|
||||
.saturating_sub(before.discard_unarmed_n),
|
||||
offcpu_in_site_cycles: self
|
||||
.offcpu_in_site_cycles
|
||||
.saturating_sub(before.offcpu_in_site_cycles),
|
||||
offcpu_in_site_n: self.offcpu_in_site_n.saturating_sub(before.offcpu_in_site_n),
|
||||
offcpu_in_site_n: self
|
||||
.offcpu_in_site_n
|
||||
.saturating_sub(before.offcpu_in_site_n),
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -795,7 +801,9 @@ mod inner {
|
||||
let cell = results
|
||||
.iter()
|
||||
.find(|r| r.site == site && r.speedup_pct == speedup_pct)?;
|
||||
let base = results.iter().find(|r| r.site == site && r.speedup_pct == 0)?;
|
||||
let base = results
|
||||
.iter()
|
||||
.find(|r| r.site == site && r.speedup_pct == 0)?;
|
||||
let rate = normalized_rate(cell, point)?;
|
||||
let b = normalized_rate(base, point)?;
|
||||
if b <= 0.0 {
|
||||
@@ -946,7 +954,9 @@ macro_rules! progress {
|
||||
macro_rules! causal_site {
|
||||
($name:literal) => {{
|
||||
static __SMARM_SITE: ::std::sync::OnceLock<u32> = ::std::sync::OnceLock::new();
|
||||
$crate::causal::SiteGuard::enter(*__SMARM_SITE.get_or_init(|| $crate::causal::site_id($name)))
|
||||
$crate::causal::SiteGuard::enter(
|
||||
*__SMARM_SITE.get_or_init(|| $crate::causal::site_id($name)),
|
||||
)
|
||||
}};
|
||||
}
|
||||
|
||||
|
||||
+358
-251
@@ -1,38 +1,103 @@
|
||||
//! Unbounded MPSC channels.
|
||||
//! Unbounded multi-producer, single-consumer channels: how actors talk to
|
||||
//! each other.
|
||||
//!
|
||||
//! Inner state is `Arc<RawMutex<Inner<T>>>` so channels can be sent across OS
|
||||
//! threads (required for the multi-scheduler runtime where a sender and
|
||||
//! receiver may run on different scheduler threads simultaneously).
|
||||
//! A channel is a queue with a typed [`Sender`] on one end and a typed
|
||||
//! [`Receiver`] on the other. Any number of actors can hold a clone of the
|
||||
//! `Sender` and push messages onto the same queue; exactly one [`Receiver`]
|
||||
//! reads them back out, in the order they arrived. This is the basic wiring
|
||||
//! smarm's other actor primitives (`gen_server`, `pg`, the registry) are all
|
||||
//! built out of, and it is directly usable on its own for a worker that just
|
||||
//! needs an inbox.
|
||||
//!
|
||||
//! ## Why `RawMutex` (Channel class), not `std::sync::Mutex`
|
||||
//! ## A first channel
|
||||
//!
|
||||
//! An actor holding a guard with preemption *enabled* can be timesliced
|
||||
//! inside the critical section and resume on a different OS thread — the
|
||||
//! pthread mutex would then be released from a thread that didn't lock it,
|
||||
//! which is UB (Linux futexes happen to tolerate it, but it's not
|
||||
//! guaranteed). `RawMutex` disables preemption for the guard's span and is
|
||||
//! cross-thread-release sound by construction, closing the hole. It also
|
||||
//! cannot poison. Channel locks form their own [`LockClass::Channel`]
|
||||
//! (raw_mutex.rs): they may be taken under a cold (Leaf) lock — finalize and
|
||||
//! `monitor()` clone senders that live in slots — but nothing may be locked
|
||||
//! under them, which the debug build enforces. `recv_match` runs its user
|
||||
//! predicate under this lock: keep it cheap, pure, and channel-free.
|
||||
//! ```
|
||||
//! use smarm::{channel, run, spawn};
|
||||
//!
|
||||
//! Semantics:
|
||||
//! - Senders are clonable; the last sender drop closes the channel.
|
||||
//! - `Receiver::recv` on an empty open channel parks the receiver.
|
||||
//! - `Receiver::recv` on an empty closed channel returns `Err(RecvError)`.
|
||||
//! - `Sender::send` on an open channel always succeeds.
|
||||
//! - `Sender::send` on a closed channel (receiver dropped) returns
|
||||
//! `Err(SendError(value))`.
|
||||
//! - When a send pushes to a previously empty queue and a receiver is
|
||||
//! parked, the receiver is unparked.
|
||||
//! run(|| {
|
||||
//! let (tx, rx) = channel::<u64>();
|
||||
//!
|
||||
//! let worker = spawn(move || {
|
||||
//! // Blocks until a message arrives.
|
||||
//! let n = rx.recv().unwrap();
|
||||
//! assert_eq!(n, 42);
|
||||
//!
|
||||
//! // Once every Sender is dropped, recv() reports the channel closed
|
||||
//! // instead of blocking forever.
|
||||
//! assert!(rx.recv().is_err());
|
||||
//! });
|
||||
//!
|
||||
//! tx.send(42).unwrap();
|
||||
//! drop(tx); // last sender gone: the channel is now closed
|
||||
//! worker.join().unwrap();
|
||||
//! });
|
||||
//! ```
|
||||
//!
|
||||
//! ## Sending
|
||||
//!
|
||||
//! [`Sender`] is cheaply clonable: hand a clone to every actor that needs to
|
||||
//! push messages into this queue. The channel stays open as long as at least
|
||||
//! one clone exists; [`Sender::send`] never blocks and always succeeds while
|
||||
//! the channel is open, since the queue is unbounded. Once the [`Receiver`]
|
||||
//! has been dropped, `send` returns the message back to you in
|
||||
//! [`SendError`] instead of delivering it.
|
||||
//!
|
||||
//! ## Receiving
|
||||
//!
|
||||
//! There is exactly one [`Receiver`] per channel (it is not clonable).
|
||||
//! [`Receiver::recv`] returns the next message in arrival order, parking the
|
||||
//! calling actor if the queue is currently empty. Once every `Sender` has
|
||||
//! been dropped and the queue has been drained, `recv` stops parking and
|
||||
//! returns [`RecvError`] instead, so a receiver never blocks forever waiting
|
||||
//! on senders that are never coming back.
|
||||
//!
|
||||
//! Beyond plain `recv`, three variants cover the common needs:
|
||||
//!
|
||||
//! - [`Receiver::try_recv`]: never parks: reports an empty-but-open channel
|
||||
//! as `Ok(None)` instead of waiting.
|
||||
//! - [`Receiver::recv_timeout`]: parks, but gives up and returns
|
||||
//! [`RecvTimeoutError::Timeout`] if no message arrives before a deadline.
|
||||
//! - [`Receiver::recv_match`] / [`Receiver::try_recv_match`]: selective
|
||||
//! receive. Instead of taking whatever is at the front of the queue, pick
|
||||
//! out the first message matching a predicate, leaving the rest queued in
|
||||
//! order. Handy for an actor that wants to prioritise one kind of message
|
||||
//! over others already waiting.
|
||||
//!
|
||||
//! ## Waiting on several channels: `select`
|
||||
//!
|
||||
//! [`select`] parks an actor across several receivers at once and reports
|
||||
//! the index of the first one that is ready (has a message queued, or has
|
||||
//! been closed). [`select_timeout`] adds a deadline, the way `recv_timeout`
|
||||
//! does for a single channel. See their docs for the full contract,
|
||||
//! including the priority-order and no-fairness guarantee.
|
||||
//!
|
||||
//! ## Implementation notes
|
||||
//!
|
||||
//! The queue and its bookkeeping live behind `Arc<RawMutex<Inner<T>>>`
|
||||
//! rather than a `std::sync::Mutex`, so that a channel can be freely shared
|
||||
//! and sent across the OS threads backing the multi-scheduler runtime.
|
||||
//! `RawMutex` matters here for a subtler reason too: an ordinary pthread
|
||||
//! mutex can be released from a different OS thread than the one that took
|
||||
//! it (smarm's preemption can migrate a timesliced actor between scheduler
|
||||
//! threads mid-critical-section), and doing that to a `std::sync::Mutex` is
|
||||
//! undefined behavior. `RawMutex` disables preemption for the guard's short
|
||||
//! lifetime instead, so the release always happens on the thread that
|
||||
//! acquired it, and it has no poisoning to worry about besides. Channel
|
||||
//! locks are cheap and are never held across another lock acquisition or a
|
||||
//! blocking call; the predicate passed to `recv_match` runs under this lock,
|
||||
//! which is why it needs to stay cheap, pure, and must not call back into
|
||||
//! the same channel.
|
||||
|
||||
use crate::pid::Pid;
|
||||
use crate::raw_mutex::RawMutex;
|
||||
use crate::runtime::RuntimeInner;
|
||||
use std::collections::VecDeque;
|
||||
use std::sync::Arc;
|
||||
use std::sync::{Arc, Weak};
|
||||
|
||||
/// Create a new channel and return its `(Sender, Receiver)` halves.
|
||||
///
|
||||
/// The channel is unbounded (no capacity limit) and open until every
|
||||
/// `Sender` has been dropped.
|
||||
pub fn channel<T>() -> (Sender<T>, Receiver<T>) {
|
||||
let inner = Arc::new(RawMutex::new_channel(Inner {
|
||||
queue: VecDeque::new(),
|
||||
@@ -40,34 +105,53 @@ pub fn channel<T>() -> (Sender<T>, Receiver<T>) {
|
||||
senders: 1,
|
||||
receiver_alive: true,
|
||||
}));
|
||||
(Sender { inner: inner.clone() }, Receiver { inner })
|
||||
(
|
||||
Sender {
|
||||
inner: inner.clone(),
|
||||
},
|
||||
Receiver { inner },
|
||||
)
|
||||
}
|
||||
|
||||
struct Inner<T> {
|
||||
queue: VecDeque<T>,
|
||||
/// The parked receiver's `(pid, park-epoch)`. The epoch is the slot
|
||||
/// word's runtime-wide wait identity (see slot_state.rs): wakers call
|
||||
/// `unpark_at(pid, epoch)`, so an entry left over from an already-woken
|
||||
/// wait — a `select` loser arm, a satisfied `recv_timeout`'s timer — is
|
||||
/// inert: the wake fails the word's epoch CAS and no-ops. This replaces
|
||||
/// the old per-channel `cur_wait`/`next_wait_seq`/`timed_out` trio: wait
|
||||
/// identity now exists exactly once, in the slot word.
|
||||
parked_receiver: Option<(Pid, u32)>,
|
||||
/// The parked receiver's `(pid, park-epoch, runtime)`, if one is currently
|
||||
/// waiting. The epoch identifies exactly which wait this is, so a waker
|
||||
/// left over from a wait that already ended (a losing `select` arm, a
|
||||
/// `recv_timeout` whose timer fired after it was already satisfied) is
|
||||
/// inert and does nothing when it fires. The `Weak<RuntimeInner>` is the
|
||||
/// receiver's runtime, captured while it parked (so provably alive then);
|
||||
/// it lets a sender on a foreign OS thread wake the receiver without the
|
||||
/// `RUNTIME` thread-local, which is unset off a scheduler thread.
|
||||
parked_receiver: Option<(Pid, u32, Weak<RuntimeInner>)>,
|
||||
senders: usize,
|
||||
receiver_alive: bool,
|
||||
}
|
||||
|
||||
/// The sending half of a channel, created by [`channel`]. Clonable: every
|
||||
/// clone pushes onto the same queue, and the channel stays open as long as
|
||||
/// any clone is alive. Dropping the last `Sender` closes the channel, which
|
||||
/// wakes a parked [`Receiver`] so it can observe the closure.
|
||||
pub struct Sender<T> {
|
||||
inner: Arc<RawMutex<Inner<T>>>,
|
||||
}
|
||||
|
||||
/// The receiving half of a channel, created by [`channel`]. Not clonable:
|
||||
/// a channel has exactly one receiver. Reads messages in the order they
|
||||
/// were sent, via [`recv`](Receiver::recv) and its variants.
|
||||
pub struct Receiver<T> {
|
||||
inner: Arc<RawMutex<Inner<T>>>,
|
||||
}
|
||||
|
||||
/// Returned by [`Sender::send`] when the channel's [`Receiver`] has already
|
||||
/// been dropped. Carries the message back so it is never silently lost;
|
||||
/// recover it with `.0` or by matching.
|
||||
#[derive(Debug, PartialEq, Eq)]
|
||||
pub struct SendError<T>(pub T);
|
||||
|
||||
/// Returned by [`Receiver::recv`] (and the other receive methods, in their
|
||||
/// own error types) when the channel is closed: every `Sender` has been
|
||||
/// dropped and no message is left queued.
|
||||
#[derive(Debug, PartialEq, Eq, Clone, Copy)]
|
||||
pub struct RecvError;
|
||||
|
||||
@@ -84,8 +168,8 @@ impl std::error::Error for RecvError {}
|
||||
pub enum RecvTimeoutError {
|
||||
/// The deadline passed with no message available.
|
||||
Timeout,
|
||||
/// All senders dropped with no message available — the bounded analogue
|
||||
/// of [`RecvError`].
|
||||
/// Every sender was dropped with no message available. The
|
||||
/// timeout-aware counterpart of plain [`RecvError`].
|
||||
Disconnected,
|
||||
}
|
||||
|
||||
@@ -103,7 +187,9 @@ impl std::error::Error for RecvTimeoutError {}
|
||||
impl<T> Clone for Sender<T> {
|
||||
fn clone(&self) -> Self {
|
||||
self.inner.lock().senders += 1;
|
||||
Sender { inner: self.inner.clone() }
|
||||
Sender {
|
||||
inner: self.inner.clone(),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -115,8 +201,8 @@ impl<T> Drop for Sender<T> {
|
||||
// Wake the parked receiver on the last sender drop regardless of
|
||||
// whether the queue is empty. A plain `recv` only ever parks on an
|
||||
// empty queue (so this is unchanged for it), but a selective
|
||||
// `recv_match` may be parked on a *non-empty* queue holding only
|
||||
// non-matching messages — it must wake to observe closure and
|
||||
// `recv_match` may be parked on a non-empty queue holding only
|
||||
// non-matching messages. It must wake to observe closure and
|
||||
// return Err rather than sleep forever.
|
||||
if g.senders == 0 {
|
||||
g.parked_receiver.take()
|
||||
@@ -124,8 +210,8 @@ impl<T> Drop for Sender<T> {
|
||||
None
|
||||
}
|
||||
};
|
||||
if let Some((pid, epoch)) = unpark {
|
||||
crate::scheduler::unpark_at(pid, epoch);
|
||||
if let Some((pid, epoch, rt)) = unpark {
|
||||
crate::scheduler::unpark_at_via(pid, epoch, &rt);
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -133,14 +219,15 @@ impl<T> Drop for Sender<T> {
|
||||
impl<T> Drop for Receiver<T> {
|
||||
fn drop(&mut self) {
|
||||
// The only consumer is gone: queued messages can never be delivered.
|
||||
// Drop them now instead of stranding them until the last Sender goes
|
||||
// away (a registry entry under lazy prune can keep a Sender — and thus
|
||||
// the Arc<Inner> — alive long after the server exits). Dropping a queued
|
||||
// Envelope::Call drops its reply_tx, waking any caller parked in `call`
|
||||
// with ServerDown, so the documented guarantee holds on *every* teardown
|
||||
// path, not only the all-senders-drop one. Drain under the lock, then
|
||||
// run item destructors after releasing it (a reply_tx drop reaches into
|
||||
// a *different* channel's lock + the scheduler, so it must not nest).
|
||||
// Drop them now instead of leaving them queued until the last Sender
|
||||
// happens to go away, which can be long after this receiver's owner
|
||||
// has exited if some other part of the runtime is still holding a
|
||||
// clone of the Sender. Draining runs each queued message's own drop
|
||||
// glue, which matters for a gen_server call: dropping a queued call
|
||||
// envelope drops its reply channel too, which wakes the caller with
|
||||
// an error instead of leaving it parked forever. Drain under the
|
||||
// lock, then run the drops after releasing it, since a message's
|
||||
// drop glue may itself touch a different channel or the scheduler.
|
||||
let drained = {
|
||||
let mut g = self.inner.lock();
|
||||
g.receiver_alive = false;
|
||||
@@ -151,14 +238,17 @@ impl<T> Drop for Receiver<T> {
|
||||
}
|
||||
|
||||
impl<T> Sender<T> {
|
||||
/// Number of messages currently queued behind this channel. Introspection
|
||||
/// only (RFC 016 mailbox depth); takes the channel lock, so callers reach
|
||||
/// it under the registry Leaf (Leaf → Channel) via the erased probe in
|
||||
/// `registry.rs`, never on a hot path.
|
||||
/// Number of messages currently queued and not yet received. For
|
||||
/// introspection and monitoring; takes the channel's internal lock, so
|
||||
/// avoid calling it from a hot path.
|
||||
pub(crate) fn queued_len(&self) -> usize {
|
||||
self.inner.lock().queue.len()
|
||||
}
|
||||
|
||||
/// Push `value` onto the channel. Succeeds unconditionally as long as
|
||||
/// the [`Receiver`] is still alive: the queue has no capacity limit, so
|
||||
/// this never blocks and never fails except when the channel is closed,
|
||||
/// in which case `value` comes back in [`SendError`].
|
||||
pub fn send(&self, value: T) -> Result<(), SendError<T>> {
|
||||
let unpark = {
|
||||
let mut g = self.inner.lock();
|
||||
@@ -168,17 +258,29 @@ impl<T> Sender<T> {
|
||||
g.queue.push_back(value);
|
||||
g.parked_receiver.take()
|
||||
};
|
||||
if let Some((pid, epoch)) = unpark {
|
||||
crate::te!(crate::trace::Event::Send { sender: crate::actor::current_pid().unwrap_or(crate::pid::Pid::new(u32::MAX, u32::MAX)), receiver: Some(pid) });
|
||||
crate::scheduler::unpark_at(pid, epoch);
|
||||
if let Some((pid, epoch, rt)) = unpark {
|
||||
crate::te!(crate::trace::Event::Send {
|
||||
sender: crate::actor::current_pid()
|
||||
.unwrap_or(crate::pid::Pid::new(u32::MAX, u32::MAX)),
|
||||
receiver: Some(pid)
|
||||
});
|
||||
crate::scheduler::unpark_at_via(pid, epoch, &rt);
|
||||
} else {
|
||||
crate::te!(crate::trace::Event::Send { sender: crate::actor::current_pid().unwrap_or(crate::pid::Pid::new(u32::MAX, u32::MAX)), receiver: None });
|
||||
crate::te!(crate::trace::Event::Send {
|
||||
sender: crate::actor::current_pid()
|
||||
.unwrap_or(crate::pid::Pid::new(u32::MAX, u32::MAX)),
|
||||
receiver: None
|
||||
});
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
}
|
||||
|
||||
impl<T> Receiver<T> {
|
||||
/// Block until a message is available and return it. Messages come back
|
||||
/// in the order they were sent. If the queue is empty and every
|
||||
/// [`Sender`] has already been dropped, returns [`RecvError`] instead of
|
||||
/// blocking forever.
|
||||
pub fn recv(&self) -> Result<T, RecvError> {
|
||||
loop {
|
||||
{
|
||||
@@ -195,48 +297,43 @@ impl<T> Receiver<T> {
|
||||
None => panic!("smarm: recv() called outside an actor"),
|
||||
};
|
||||
debug_assert!(
|
||||
g.parked_receiver.is_none_or(|(p, _)| p == me),
|
||||
g.parked_receiver.as_ref().is_none_or(|(p, _, _)| *p == me),
|
||||
"channel has more than one receiver"
|
||||
);
|
||||
// begin_wait is lock-free — legal under the Channel lock;
|
||||
// begin_wait is lock-free, so it's legal under the Channel lock;
|
||||
// registering in the same critical section makes the epoch
|
||||
// atomic with the senders' view of the registration.
|
||||
g.parked_receiver = Some((me, crate::scheduler::begin_wait()));
|
||||
g.parked_receiver = Some((
|
||||
me,
|
||||
crate::scheduler::begin_wait(),
|
||||
crate::scheduler::runtime_weak(),
|
||||
));
|
||||
crate::te!(crate::trace::Event::RecvPark(me));
|
||||
}
|
||||
// Release the lock before parking — the unparker will need it.
|
||||
// Release the lock before parking: the unparker will need it.
|
||||
crate::scheduler::park_current();
|
||||
// Woken up — record it before looping to check the queue.
|
||||
crate::te!(crate::trace::Event::RecvWake(match crate::actor::current_pid() {
|
||||
Some(p) => p,
|
||||
None => panic!("smarm: RecvWake outside an actor (core corrupt)"),
|
||||
}));
|
||||
// Woken up. Record it before looping to check the queue.
|
||||
crate::te!(crate::trace::Event::RecvWake(
|
||||
match crate::actor::current_pid() {
|
||||
Some(p) => p,
|
||||
None => panic!("smarm: RecvWake outside an actor (core corrupt)"),
|
||||
}
|
||||
));
|
||||
}
|
||||
}
|
||||
|
||||
/// Bounded receive: like [`recv`](Self::recv), but gives up once
|
||||
/// `timeout` has elapsed, returning [`RecvTimeoutError::Timeout`].
|
||||
/// Like [`recv`](Self::recv), but gives up and returns
|
||||
/// [`RecvTimeoutError::Timeout`] if no message has arrived by the time
|
||||
/// `timeout` elapses.
|
||||
///
|
||||
/// Built on the same timer machinery as `Mutex::lock_timeout`: the wait
|
||||
/// registers a `WaitTimeout` entry stamped with the wait's park-epoch;
|
||||
/// on expiry the channel (as the
|
||||
/// [`TimerTarget`](crate::timer::TimerTarget)) checks whether *this*
|
||||
/// wait is still parked and, only then, cancels it. A wake that races
|
||||
/// the deadline resolves message-first: if a message is available when
|
||||
/// the receiver runs, it is delivered even if the timer had already
|
||||
/// fired. A satisfied or abandoned wait leaves its timer entry to expire
|
||||
/// as a no-op (registration gone; epoch consumed), per the
|
||||
/// no-cancellation convention in `timer.rs`.
|
||||
/// If a message arrives at essentially the same moment the deadline
|
||||
/// passes, the message wins: you get `Ok` rather than `Timeout`. If
|
||||
/// every sender is dropped before a message arrives or the deadline
|
||||
/// passes, you get [`RecvTimeoutError::Disconnected`].
|
||||
///
|
||||
/// The wake is classified from state alone — wakes are precise (the only
|
||||
/// stamped wakers of this wait are a send, the last-sender drop, and the
|
||||
/// timer; a stop wake unwinds out of `park_current` and never reaches
|
||||
/// the classification), so: message queued → `Ok`; `senders == 0` →
|
||||
/// `Disconnected`; neither → it was the timer → `Timeout`.
|
||||
///
|
||||
/// `Duration::ZERO` is a valid timeout: it parks until the immediately-
|
||||
/// due timer is drained, then reports `Timeout` unless a message was
|
||||
/// already queued.
|
||||
/// `Duration::ZERO` is a valid timeout: it still gives any
|
||||
/// already-queued message a chance to be returned, and only then
|
||||
/// reports `Timeout`.
|
||||
pub fn recv_timeout(&self, timeout: std::time::Duration) -> Result<T, RecvTimeoutError>
|
||||
where
|
||||
T: Send + 'static,
|
||||
@@ -258,27 +355,29 @@ impl<T> Receiver<T> {
|
||||
return Err(RecvTimeoutError::Disconnected);
|
||||
}
|
||||
debug_assert!(
|
||||
g.parked_receiver.is_none_or(|(p, _)| p == me),
|
||||
g.parked_receiver.as_ref().is_none_or(|(p, _, _)| *p == me),
|
||||
"channel has more than one receiver"
|
||||
);
|
||||
epoch = crate::scheduler::begin_wait();
|
||||
g.parked_receiver = Some((me, epoch));
|
||||
g.parked_receiver = Some((me, epoch, crate::scheduler::runtime_weak()));
|
||||
crate::te!(crate::trace::Event::RecvPark(me));
|
||||
}
|
||||
|
||||
// Arm the timer after releasing the channel lock (insert takes the
|
||||
// timers lock; never nest under a Channel lock). A send or even the
|
||||
// timer itself may unpark us before we park — the RunningNotified
|
||||
// timer itself may unpark us before we park; the runtime's wake
|
||||
// protocol makes the park below return immediately in that case.
|
||||
let deadline = crate::timer::deadline_from_now(timeout);
|
||||
let target: std::sync::Arc<dyn crate::timer::TimerTarget> = self.inner.clone();
|
||||
crate::scheduler::insert_wait_timer(deadline, me, target, epoch);
|
||||
|
||||
crate::scheduler::park_current();
|
||||
crate::te!(crate::trace::Event::RecvWake(match crate::actor::current_pid() {
|
||||
Some(p) => p,
|
||||
None => panic!("smarm: RecvWake outside an actor (core corrupt)"),
|
||||
}));
|
||||
crate::te!(crate::trace::Event::RecvWake(
|
||||
match crate::actor::current_pid() {
|
||||
Some(p) => p,
|
||||
None => panic!("smarm: RecvWake outside an actor (core corrupt)"),
|
||||
}
|
||||
));
|
||||
let mut g = self.inner.lock();
|
||||
if let Some(v) = g.queue.pop_front() {
|
||||
crate::preempt::note_message_received();
|
||||
@@ -290,16 +389,23 @@ impl<T> Receiver<T> {
|
||||
Err(RecvTimeoutError::Timeout)
|
||||
}
|
||||
|
||||
/// Selective receive: remove and return the first queued message for which
|
||||
/// `pred` holds, leaving the rest in arrival order. If no queued message
|
||||
/// matches, parks and re-scans on every send (a selective receiver may park
|
||||
/// on a *non-empty* queue). Returns `Err(RecvError)` only once the channel
|
||||
/// is closed and no queued message matches.
|
||||
/// Selective receive: find and return the first queued message for
|
||||
/// which `pred` returns `true`, leaving every other message in the
|
||||
/// queue untouched and in order. Useful when an actor's inbox mixes
|
||||
/// message kinds and it wants to handle one kind out of turn, without
|
||||
/// discarding the rest.
|
||||
///
|
||||
/// `pred` is run while the channel lock is held: keep it cheap and pure,
|
||||
/// and do not call back into this channel from inside it. It is modelled as
|
||||
/// `Fn` (not `FnMut`) deliberately — it is re-run from scratch on every
|
||||
/// scan, so a stateful predicate would observe surprising re-counting.
|
||||
/// If nothing queued matches, this blocks and re-checks every time a new
|
||||
/// message arrives, the same way [`recv`](Self::recv) blocks on an empty
|
||||
/// queue: a selective receiver can be waiting even while the queue holds
|
||||
/// messages, just none that match yet. Returns [`RecvError`] only once
|
||||
/// the channel is closed and still nothing matches.
|
||||
///
|
||||
/// `pred` runs while the channel is locked, so keep it cheap, side
|
||||
/// effect free, and make sure it never calls back into this same
|
||||
/// channel. It takes `&T` and is called fresh on every scan (not `FnMut`
|
||||
/// with running state), so it should judge each message purely on its
|
||||
/// own content.
|
||||
pub fn recv_match<F>(&self, pred: F) -> Result<T, RecvError>
|
||||
where
|
||||
F: Fn(&T) -> bool,
|
||||
@@ -325,25 +431,33 @@ impl<T> Receiver<T> {
|
||||
None => panic!("smarm: recv_match() called outside an actor"),
|
||||
};
|
||||
debug_assert!(
|
||||
g.parked_receiver.is_none_or(|(p, _)| p == me),
|
||||
g.parked_receiver.as_ref().is_none_or(|(p, _, _)| *p == me),
|
||||
"channel has more than one receiver"
|
||||
);
|
||||
g.parked_receiver = Some((me, crate::scheduler::begin_wait()));
|
||||
g.parked_receiver = Some((
|
||||
me,
|
||||
crate::scheduler::begin_wait(),
|
||||
crate::scheduler::runtime_weak(),
|
||||
));
|
||||
crate::te!(crate::trace::Event::RecvPark(me));
|
||||
}
|
||||
// Release the lock before parking — the unparker will need it.
|
||||
// Release the lock before parking: the unparker will need it.
|
||||
crate::scheduler::park_current();
|
||||
crate::te!(crate::trace::Event::RecvWake(match crate::actor::current_pid() {
|
||||
Some(p) => p,
|
||||
None => panic!("smarm: RecvWake outside an actor (core corrupt)"),
|
||||
}));
|
||||
crate::te!(crate::trace::Event::RecvWake(
|
||||
match crate::actor::current_pid() {
|
||||
Some(p) => p,
|
||||
None => panic!("smarm: RecvWake outside an actor (core corrupt)"),
|
||||
}
|
||||
));
|
||||
}
|
||||
}
|
||||
|
||||
/// Non-blocking selective receive. `Ok(Some(v))` if a queued message
|
||||
/// matched `pred` (removed, rest left in order), `Ok(None)` if the channel
|
||||
/// is open but nothing matched, `Err(RecvError)` if closed and nothing
|
||||
/// matched. Same predicate contract as [`recv_match`](Self::recv_match).
|
||||
/// The non-blocking counterpart of [`recv_match`](Self::recv_match):
|
||||
/// returns immediately either way. `Ok(Some(v))` if a queued message
|
||||
/// matched `pred` (removed; the rest stay queued in order), `Ok(None)`
|
||||
/// if the channel is open but nothing currently matches, `Err(RecvError)`
|
||||
/// if the channel is closed and nothing matches. Same predicate contract
|
||||
/// as `recv_match`.
|
||||
pub fn try_recv_match<F>(&self, pred: F) -> Result<Option<T>, RecvError>
|
||||
where
|
||||
F: Fn(&T) -> bool,
|
||||
@@ -363,8 +477,10 @@ impl<T> Receiver<T> {
|
||||
Ok(None)
|
||||
}
|
||||
|
||||
/// Non-blocking. `Ok(Some(v))` if a message was available, `Ok(None)` if
|
||||
/// the channel is empty but open, `Err(RecvError)` if closed and drained.
|
||||
/// The non-blocking counterpart of [`recv`](Self::recv): returns
|
||||
/// immediately either way. `Ok(Some(v))` if a message was queued,
|
||||
/// `Ok(None)` if the channel is open but currently empty, `Err(RecvError)`
|
||||
/// if the channel is closed and the queue is drained.
|
||||
pub fn try_recv(&self) -> Result<Option<T>, RecvError> {
|
||||
let mut g = self.inner.lock();
|
||||
if let Some(v) = g.queue.pop_front() {
|
||||
@@ -379,28 +495,29 @@ impl<T> Receiver<T> {
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// TimerTarget — the expiry half of recv_timeout
|
||||
// TimerTarget: the expiry half of recv_timeout
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
impl<T: Send + 'static> crate::timer::TimerTarget for RawMutex<Inner<T>> {
|
||||
fn on_timeout(&self, pid: Pid, epoch: u32) {
|
||||
// Cancel the wait only if THIS wait (epoch match) is still
|
||||
// registered. If a sender already took `parked_receiver`, the
|
||||
// receiver is waking with a message — message wins, the timer
|
||||
// receiver is waking with a message: message wins, the timer
|
||||
// no-ops. If a later wait by the same receiver is registered, the
|
||||
// epoch mismatches — stale entry, no-op. (The unpark_at would fail
|
||||
// its word CAS in either case anyway; checking under the lock keeps
|
||||
// the registration bookkeeping exact.)
|
||||
// epoch mismatches: stale entry, no-op. (unpark_at would fail its
|
||||
// internal check in either case anyway; checking under the lock
|
||||
// keeps the registration bookkeeping exact.)
|
||||
let unpark = {
|
||||
let mut g = self.lock();
|
||||
if g.parked_receiver == Some((pid, epoch)) {
|
||||
g.parked_receiver = None;
|
||||
true
|
||||
} else {
|
||||
false
|
||||
match g.parked_receiver {
|
||||
Some((p, e, _)) if p == pid && e == epoch => {
|
||||
g.parked_receiver = None;
|
||||
true
|
||||
}
|
||||
_ => false,
|
||||
}
|
||||
};
|
||||
// Unpark outside the channel lock — it may take the run-queue lock;
|
||||
// Unpark outside the channel lock: it may take the run-queue lock;
|
||||
// legal under a Channel lock, but pointless to nest.
|
||||
if unpark {
|
||||
crate::scheduler::unpark_at(pid, epoch);
|
||||
@@ -409,7 +526,7 @@ impl<T: Send + 'static> crate::timer::TimerTarget for RawMutex<Inner<T>> {
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// select — ready-index wait over multiple receivers
|
||||
// select: ready-index wait over multiple receivers
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
pub(crate) mod sealed {
|
||||
@@ -417,33 +534,34 @@ pub(crate) mod sealed {
|
||||
}
|
||||
impl<T> sealed::Sealed for Receiver<T> {}
|
||||
|
||||
/// An arm of a [`select`]. Implemented by [`Receiver`]; sealed, because the
|
||||
/// registration contract below is part of the runtime's wake protocol.
|
||||
/// An arm of a [`select`]: something you can wait on alongside other arms
|
||||
/// and be told when it becomes ready. Implemented by [`Receiver`]; sealed
|
||||
/// (cannot be implemented outside this crate), since the registration
|
||||
/// contract below is part of the runtime's internal wake protocol.
|
||||
///
|
||||
/// Contract (all under the arm's own lock): `sel_register` checks-or-
|
||||
/// registers atomically — if the arm is ready it does NOT register and
|
||||
/// registers atomically. If the arm is ready it does not register and
|
||||
/// returns `Ok(false)`; otherwise it publishes `(pid, epoch)` where its
|
||||
/// wakers will find it and returns `Ok(true)`. "Ready" means a receive
|
||||
/// would not park: a message is queued, or the arm is closed. `Err` means
|
||||
/// would not block: a message is queued, or the arm is closed. `Err` means
|
||||
/// the arm could not register at all (only fd arms can fail; channel
|
||||
/// registration is infallible) — the wait must be retired and earlier
|
||||
/// registration always succeeds), and the wait must be retired and earlier
|
||||
/// eager-cleanup arms unregistered.
|
||||
pub trait Selectable: sealed::Sealed {
|
||||
#[doc(hidden)]
|
||||
fn sel_register(&self, pid: Pid, epoch: u32) -> std::io::Result<bool>;
|
||||
#[doc(hidden)]
|
||||
fn sel_ready(&self) -> bool;
|
||||
/// Remove this arm's `(pid, epoch)` registration if — and only if — it
|
||||
/// is still in place. Default no-op: a losing channel arm's stale
|
||||
/// registration is inert (its wakers die at the epoch CAS; the next
|
||||
/// wait overwrites the slot). Fd arms override this: their staleness
|
||||
/// poisons the fd (waiters entry + kernel-side ONESHOT registration)
|
||||
/// and needs an eager cleanup pass.
|
||||
/// Remove this arm's `(pid, epoch)` registration if, and only if, it is
|
||||
/// still in place. Default no-op: a losing channel arm's stale
|
||||
/// registration is harmless and self-cleans. Fd arms override this:
|
||||
/// their staleness would otherwise leave the fd unusable for future
|
||||
/// selects, so they need an eager cleanup pass.
|
||||
#[doc(hidden)]
|
||||
fn sel_unregister(&self, _pid: Pid, _epoch: u32) {}
|
||||
/// Whether this arm requires the eager cleanup pass at all. Gates the
|
||||
/// post-wake `sel_unregister` sweep so channel-only selects keep
|
||||
/// today's zero-cancellation hot path.
|
||||
/// post-wake `sel_unregister` sweep so channel-only selects keep their
|
||||
/// cheap, cleanup-free path.
|
||||
#[doc(hidden)]
|
||||
fn sel_eager_cleanup(&self) -> bool {
|
||||
false
|
||||
@@ -457,10 +575,10 @@ impl<T> Selectable for Receiver<T> {
|
||||
return Ok(false);
|
||||
}
|
||||
debug_assert!(
|
||||
g.parked_receiver.is_none_or(|(p, _)| p == pid),
|
||||
g.parked_receiver.as_ref().is_none_or(|(p, _, _)| *p == pid),
|
||||
"channel has more than one receiver"
|
||||
);
|
||||
g.parked_receiver = Some((pid, epoch));
|
||||
g.parked_receiver = Some((pid, epoch, crate::scheduler::runtime_weak()));
|
||||
Ok(true)
|
||||
}
|
||||
|
||||
@@ -470,38 +588,35 @@ impl<T> Selectable for Receiver<T> {
|
||||
}
|
||||
}
|
||||
|
||||
/// Park on every arm at once; return the index of the first ready one.
|
||||
/// Wait on several channels at once and return the index of the first one
|
||||
/// that is ready, instead of blocking on just one with [`Receiver::recv`].
|
||||
///
|
||||
/// "Ready" means a receive on that arm would not park: a message is queued,
|
||||
/// or the arm is **closed** (so the caller's `try_recv` observes the
|
||||
/// disconnect — a dead arm is an event, not a hang). The caller consumes the
|
||||
/// arm itself, typically via [`Receiver::try_recv`]; single-receiver
|
||||
/// channels guarantee nothing can steal the message in between.
|
||||
/// "Ready" means a receive on that arm would not block: a message is
|
||||
/// queued, or the arm is closed (so the caller's own `try_recv` observes
|
||||
/// the disconnect: a dead arm is something to react to, not something to
|
||||
/// hang on). `select` only tells you which arm is ready; read the actual
|
||||
/// message yourself, typically with [`Receiver::try_recv`] on that arm.
|
||||
///
|
||||
/// A closed arm stays ready *forever*: once its disconnect has been
|
||||
/// observed, drop it from the arm set — under priority order it would
|
||||
/// otherwise win every subsequent call and starve every higher-indexed arm.
|
||||
/// A closed arm stays ready forever. Once you have observed its disconnect,
|
||||
/// drop it from the arm set you pass in next time: otherwise, under the
|
||||
/// priority order below, it would win every subsequent call and starve
|
||||
/// every arm listed after it.
|
||||
///
|
||||
/// Arms are scanned **in order**: index 0 is the highest priority, both on
|
||||
/// the immediate-ready path and after a wake. This is a documented
|
||||
/// guarantee (compose like BEAM receive clauses: put control channels
|
||||
/// first), not an accident — and therefore there is NO fairness promise; a
|
||||
/// saturated arm 0 starves arm 1 by design.
|
||||
/// Arms are checked **in order**: index 0 is the highest priority, both
|
||||
/// when checking immediately and after being woken. This is a deliberate,
|
||||
/// documented guarantee, not an accident of implementation: put a control
|
||||
/// or shutdown channel first so it is always noticed promptly. The
|
||||
/// flip side is that there is **no fairness guarantee**: a busy arm 0 can
|
||||
/// starve arm 1 indefinitely by design.
|
||||
///
|
||||
/// One actor may select on a channel and later `recv` on it (or select on
|
||||
/// overlapping sets) freely. What stays illegal is what was always illegal:
|
||||
/// two *different* actors receiving on one channel.
|
||||
/// One actor can `select` on a channel and later plain `recv` on it (or
|
||||
/// `select` again on an overlapping set of arms) with no restriction. What
|
||||
/// stays illegal is what was always illegal for a channel: two *different*
|
||||
/// actors receiving on the same one.
|
||||
///
|
||||
/// Built on the consuming-wake protocol (see slot_state.rs): all arms are
|
||||
/// registered under one wait epoch; the winning wake consumes it, so losing
|
||||
/// arms' registrations are inert and need no cancellation pass — they
|
||||
/// self-clean at their wakers' failed CAS, or get overwritten by this
|
||||
/// receiver's next wait on that channel.
|
||||
///
|
||||
/// Panics if `arms` is empty, when called outside an actor, or if an fd
|
||||
/// arm fails to register (EBADF, EMFILE, a second waiter on one fd —
|
||||
/// see [`try_select`] for the fallible form; channel-only selects cannot
|
||||
/// fail).
|
||||
/// Panics if `arms` is empty, if called outside an actor, or if an fd arm
|
||||
/// fails to register (see [`try_select`] for the fallible form; a
|
||||
/// channel-only `select` can never fail).
|
||||
pub fn select(arms: &[&dyn Selectable]) -> usize {
|
||||
match try_select(arms) {
|
||||
Ok(i) => i,
|
||||
@@ -509,9 +624,10 @@ pub fn select(arms: &[&dyn Selectable]) -> usize {
|
||||
}
|
||||
}
|
||||
|
||||
/// [`select`], fallible: `Err` when an arm fails to register (only fd
|
||||
/// arms can — EBADF, EMFILE on the epoll set, or a second waiter on an
|
||||
/// fd that already has one). On `Err` the wait is fully retired and no
|
||||
/// The fallible form of [`select`]: `Err` when an arm fails to register.
|
||||
/// Only fd arms can fail this way (for example, the file descriptor is
|
||||
/// invalid, or something else is already waiting on it); a channel-only
|
||||
/// select can never fail. On `Err` the wait is fully retired and no
|
||||
/// registration is left behind: every arm registered before the failing
|
||||
/// one has been unregistered.
|
||||
pub fn try_select(arms: &[&dyn Selectable]) -> std::io::Result<usize> {
|
||||
@@ -527,15 +643,19 @@ pub fn try_select(arms: &[&dyn Selectable]) -> std::io::Result<usize> {
|
||||
}
|
||||
|
||||
// Stale fd registrations are not harmless (a losing fd arm's
|
||||
// waiters entry poisons the fd with AlreadyExists and its
|
||||
// kernel-side ONESHOT registration can fire arbitrarily late), so
|
||||
// selects containing fd arms run an eager cleanup pass after the
|
||||
// park — including when a terminal stop unwinds out of it, via
|
||||
// the guard. Channel-only selects skip all of it: `eager` is
|
||||
// false, the guard is disarmed, and the loser-arm self-cleaning
|
||||
// story is unchanged.
|
||||
// leftover registration can make the fd unusable for the next
|
||||
// select until a kernel event happens to clear it), so selects
|
||||
// containing fd arms run an eager cleanup pass after the park,
|
||||
// including when a terminal stop unwinds out of it, via the guard.
|
||||
// Channel-only selects skip all of it: `eager` is false, the guard
|
||||
// is disarmed, and the loser-arm self-cleaning story is unchanged.
|
||||
let eager = arms.iter().any(|a| a.sel_eager_cleanup());
|
||||
let mut guard = UnregisterGuard { arms, me, epoch, armed: eager };
|
||||
let mut guard = UnregisterGuard {
|
||||
arms,
|
||||
me,
|
||||
epoch,
|
||||
armed: eager,
|
||||
};
|
||||
|
||||
crate::scheduler::park_current();
|
||||
|
||||
@@ -546,22 +666,22 @@ pub fn try_select(arms: &[&dyn Selectable]) -> std::io::Result<usize> {
|
||||
drop(guard);
|
||||
|
||||
// Woken precisely: an arm's send (message) or last-sender drop
|
||||
// (closure) consumed our epoch, and both leave their arm ready —
|
||||
// return the first one, in priority order (which may be a
|
||||
// (closure) is what woke us, and both leave their arm ready.
|
||||
// Return the first ready one, in priority order (which may be a
|
||||
// different, higher-priority arm than the one that woke us; its
|
||||
// message stays queued and re-reports ready on the next call).
|
||||
// Fd arms classify by a fresh zero-timeout poll, so they too are
|
||||
// a pure function of state — independent of the registration the
|
||||
// cleanup pass just removed.
|
||||
// a pure function of current state, independent of the
|
||||
// registration the cleanup pass just removed.
|
||||
for (i, arm) in arms.iter().enumerate() {
|
||||
if arm.sel_ready() {
|
||||
return Ok(i);
|
||||
}
|
||||
}
|
||||
// Unreachable by protocol (a stop wake unwinds out of
|
||||
// park_current). Defensive: re-open the wait and re-register —
|
||||
// stale own-registrations are overwritten (channels) or were
|
||||
// removed by the cleanup pass above (fds).
|
||||
// Unreachable in practice (a stop wake unwinds out of
|
||||
// park_current before we get here). Defensive: re-open the wait
|
||||
// and re-register; stale own-registrations are overwritten
|
||||
// (channels) or were removed by the cleanup pass above (fds).
|
||||
}
|
||||
}
|
||||
|
||||
@@ -576,11 +696,11 @@ fn unregister_arms(arms: &[&dyn Selectable], me: Pid, epoch: u32) {
|
||||
}
|
||||
}
|
||||
|
||||
/// Stop-unwind twin of the explicit cleanup pass: a terminal stop unwinds
|
||||
/// out of `park_current`, and a registered fd arm must not outlive its
|
||||
/// actor (the generalization of `wait_fd`'s `Dereg`). Disarmed on the
|
||||
/// normal path after the explicit pass runs; never armed when no fd arm
|
||||
/// registered, keeping the channel-only path guard-free in effect.
|
||||
// Stop-unwind twin of the explicit cleanup pass: a terminal stop unwinds
|
||||
// out of `park_current`, and a registered fd arm must not outlive its
|
||||
// actor. Disarmed on the normal path after the explicit pass runs; never
|
||||
// armed when no fd arm is registered, keeping the channel-only path
|
||||
// guard-free in effect.
|
||||
struct UnregisterGuard<'a> {
|
||||
arms: &'a [&'a dyn Selectable],
|
||||
me: Pid,
|
||||
@@ -596,25 +716,17 @@ impl Drop for UnregisterGuard<'_> {
|
||||
}
|
||||
}
|
||||
|
||||
/// The registration pass shared by [`select`] and [`select_timeout`]:
|
||||
/// check-or-register each arm, in priority order, each atomically under its
|
||||
/// own lock. Cross-arm atomicity is unnecessary: an arm becoming ready
|
||||
/// right after its registration wakes the caller through the protocol (the
|
||||
/// prep-to-park window is closed by RunningNotified).
|
||||
///
|
||||
/// `Ok(Some(i))` = arm `i` was ready, the pass stopped, and the wait has
|
||||
/// been RETIRED (no park may follow): earlier arms hold live-epoch
|
||||
/// registrations, so earlier *fd* arms are unregistered eagerly, then the
|
||||
/// epoch is bumped, a landed notification eaten, and a pending stop
|
||||
/// re-observed — without which a stale arm wake could fault a later
|
||||
/// one-shot park. `Err` = an arm failed to register; identical unwind
|
||||
/// (earlier fd arms unregistered, wait retired). `Ok(None)` = every arm
|
||||
/// registered; the caller parks.
|
||||
fn register_arms(
|
||||
me: Pid,
|
||||
epoch: u32,
|
||||
arms: &[&dyn Selectable],
|
||||
) -> std::io::Result<Option<usize>> {
|
||||
// The registration pass shared by `select` and `select_timeout`: check-or-
|
||||
// register each arm, in priority order, each atomically under its own lock.
|
||||
// Cross-arm atomicity is unnecessary: an arm becoming ready right after its
|
||||
// registration still wakes the caller through the normal wake path.
|
||||
//
|
||||
// `Ok(Some(i))` = arm `i` was already ready, the pass stopped, and the wait
|
||||
// has been fully retired (no park may follow): earlier fd arms are
|
||||
// unregistered eagerly so none are left dangling. `Err` = an arm failed to
|
||||
// register; same unwind (earlier fd arms unregistered, wait retired).
|
||||
// `Ok(None)` = every arm registered successfully; the caller parks.
|
||||
fn register_arms(me: Pid, epoch: u32, arms: &[&dyn Selectable]) -> std::io::Result<Option<usize>> {
|
||||
for (i, arm) in arms.iter().enumerate() {
|
||||
let registered = match arm.sel_register(me, epoch) {
|
||||
Ok(r) => r,
|
||||
@@ -633,10 +745,10 @@ fn register_arms(
|
||||
Ok(None)
|
||||
}
|
||||
|
||||
/// The [`select_timeout`] timer target: stateless, because precise wakes
|
||||
/// make classification a pure function of channel state. The entry is
|
||||
/// stamped with the select's epoch; if an arm already won, this unpark dies
|
||||
/// at the word's epoch CAS (the no-cancellation convention in `timer.rs`).
|
||||
// The `select_timeout` timer target: stateless, because a wake's cause can
|
||||
// always be read back off plain channel state (an arm ready, or not). If
|
||||
// an arm already won before the deadline, this timer's fire is simply
|
||||
// ignored, the way any other stale wakeup is.
|
||||
struct SelectTimeout;
|
||||
impl crate::timer::TimerTarget for SelectTimeout {
|
||||
fn on_timeout(&self, pid: Pid, epoch: u32) {
|
||||
@@ -644,31 +756,22 @@ impl crate::timer::TimerTarget for SelectTimeout {
|
||||
}
|
||||
}
|
||||
|
||||
/// [`select`] with a deadline: returns `Some(index)` like `select`, or
|
||||
/// `None` once `timeout` elapses with no arm ready.
|
||||
/// Like [`select`], but gives up and returns `None` if no arm becomes
|
||||
/// ready before `timeout` elapses.
|
||||
///
|
||||
/// All of `select`'s semantics carry over (priority order, closed arms
|
||||
/// permanently ready, no fairness promise). The timeout is one more stamped
|
||||
/// waker on the same wait epoch — nothing is registered in any arm for it,
|
||||
/// so there is nothing to cancel or leak: an arm winning leaves the timer
|
||||
/// entry to expire as a stale-epoch no-op; the timer winning leaves the
|
||||
/// arms' registrations to self-clean exactly as a `select` loser's would.
|
||||
/// All of `select`'s semantics carry over: arms are still checked in
|
||||
/// priority order, a closed arm is still permanently ready, and there is
|
||||
/// still no fairness guarantee across arms. A message that arrives at
|
||||
/// essentially the same moment the deadline passes still wins, the same
|
||||
/// way [`Receiver::recv_timeout`] resolves that race.
|
||||
///
|
||||
/// The wake is classified from state alone (wakes are precise): some arm
|
||||
/// ready → `Some` of the first, in priority order; none ready → the timer
|
||||
/// was the only remaining stamped waker → `None`. A message that races the
|
||||
/// deadline resolves message-first, as `recv_timeout` does.
|
||||
/// `Duration::ZERO` is a valid timeout: it still gives an already-ready arm
|
||||
/// a chance to be reported before falling through to `None`.
|
||||
///
|
||||
/// `Duration::ZERO` is a valid timeout: it parks until the immediately-due
|
||||
/// timer is drained, then reports `None` unless an arm was already ready.
|
||||
///
|
||||
/// Panics if `arms` is empty, when called outside an actor, or if an fd
|
||||
/// arm fails to register (see [`try_select_timeout`] for the fallible
|
||||
/// form; channel-only selects cannot fail).
|
||||
pub fn select_timeout(
|
||||
arms: &[&dyn Selectable],
|
||||
timeout: std::time::Duration,
|
||||
) -> Option<usize> {
|
||||
/// Panics if `arms` is empty, if called outside an actor, or if an fd arm
|
||||
/// fails to register (see [`try_select_timeout`] for the fallible form; a
|
||||
/// channel-only select can never fail).
|
||||
pub fn select_timeout(arms: &[&dyn Selectable], timeout: std::time::Duration) -> Option<usize> {
|
||||
match try_select_timeout(arms, timeout) {
|
||||
Ok(r) => r,
|
||||
Err(e) => panic!(
|
||||
@@ -677,9 +780,9 @@ pub fn select_timeout(
|
||||
}
|
||||
}
|
||||
|
||||
/// [`select_timeout`], fallible: `Err` when an arm fails to register
|
||||
/// (only fd arms can). On `Err` the wait is fully retired and no
|
||||
/// registration — arm-side or kernel-side — is left behind.
|
||||
/// The fallible form of [`select_timeout`]: `Err` when an arm fails to
|
||||
/// register (only fd arms can). On `Err` the wait is fully retired and no
|
||||
/// registration is left behind on any arm.
|
||||
pub fn try_select_timeout(
|
||||
arms: &[&dyn Selectable],
|
||||
timeout: std::time::Duration,
|
||||
@@ -694,19 +797,23 @@ pub fn try_select_timeout(
|
||||
return Ok(Some(i)); // ready now: the timer was never armed
|
||||
}
|
||||
|
||||
// Arm the timer after the registration pass, outside every Channel
|
||||
// lock (insert takes the timers lock).
|
||||
// Arm the timer after the registration pass, outside every channel
|
||||
// lock (inserting a timer takes the timers lock).
|
||||
let deadline = crate::timer::deadline_from_now(timeout);
|
||||
let target: std::sync::Arc<dyn crate::timer::TimerTarget> = std::sync::Arc::new(SelectTimeout);
|
||||
crate::scheduler::insert_wait_timer(deadline, me, target, epoch);
|
||||
|
||||
// Same eager-cleanup story as `try_select`: the timer arm needs none
|
||||
// (stateless, stale entries die at the epoch CAS), channel arms need
|
||||
// none, fd arms do — and a timer win in particular leaves every fd
|
||||
// arm's registration behind, which without this pass would poison
|
||||
// those fds until a kernel event happened to fire.
|
||||
// Same eager-cleanup story as `try_select`: a timer win in particular
|
||||
// leaves every fd arm's registration behind, which without this pass
|
||||
// would leave those fds unusable until a kernel event happened to
|
||||
// clear them.
|
||||
let eager = arms.iter().any(|a| a.sel_eager_cleanup());
|
||||
let mut guard = UnregisterGuard { arms, me, epoch, armed: eager };
|
||||
let mut guard = UnregisterGuard {
|
||||
arms,
|
||||
me,
|
||||
epoch,
|
||||
armed: eager,
|
||||
};
|
||||
|
||||
crate::scheduler::park_current();
|
||||
|
||||
|
||||
+26
-11
@@ -16,10 +16,18 @@ thread_local! {
|
||||
static ACTOR_SP: Cell<usize> = const { Cell::new(0) };
|
||||
}
|
||||
|
||||
fn get_scheduler_sp() -> usize { SCHEDULER_SP.with(|c| c.get()) }
|
||||
fn set_scheduler_sp(v: usize) { SCHEDULER_SP.with(|c| c.set(v)) }
|
||||
pub fn get_actor_sp() -> usize { ACTOR_SP.with(|c| c.get()) }
|
||||
pub fn set_actor_sp(v: usize) { ACTOR_SP.with(|c| c.set(v)) }
|
||||
fn get_scheduler_sp() -> usize {
|
||||
SCHEDULER_SP.with(|c| c.get())
|
||||
}
|
||||
fn set_scheduler_sp(v: usize) {
|
||||
SCHEDULER_SP.with(|c| c.set(v))
|
||||
}
|
||||
pub fn get_actor_sp() -> usize {
|
||||
ACTOR_SP.with(|c| c.get())
|
||||
}
|
||||
pub fn set_actor_sp(v: usize) {
|
||||
ACTOR_SP.with(|c| c.set(v))
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Initial stack layout
|
||||
@@ -49,13 +57,20 @@ pub fn set_actor_sp(v: usize) { ACTOR_SP.with(|c| c.set(v)) }
|
||||
pub fn init_actor_stack(top: *mut u8, entry: extern "C-unwind" fn()) -> usize {
|
||||
unsafe {
|
||||
let mut sp = (top as usize & !15) - 8;
|
||||
sp -= 8; (sp as *mut usize).write(entry as usize); // ret target
|
||||
sp -= 8; (sp as *mut usize).write(0); // rbx
|
||||
sp -= 8; (sp as *mut usize).write(0); // rbp
|
||||
sp -= 8; (sp as *mut usize).write(0); // r12
|
||||
sp -= 8; (sp as *mut usize).write(0); // r13
|
||||
sp -= 8; (sp as *mut usize).write(0); // r14
|
||||
sp -= 8; (sp as *mut usize).write(0); // r15
|
||||
sp -= 8;
|
||||
(sp as *mut usize).write(entry as usize); // ret target
|
||||
sp -= 8;
|
||||
(sp as *mut usize).write(0); // rbx
|
||||
sp -= 8;
|
||||
(sp as *mut usize).write(0); // rbp
|
||||
sp -= 8;
|
||||
(sp as *mut usize).write(0); // r12
|
||||
sp -= 8;
|
||||
(sp as *mut usize).write(0); // r13
|
||||
sp -= 8;
|
||||
(sp as *mut usize).write(0); // r14
|
||||
sp -= 8;
|
||||
(sp as *mut usize).write(0); // r15
|
||||
sp
|
||||
}
|
||||
}
|
||||
|
||||
+316
-66
@@ -127,16 +127,51 @@
|
||||
//! - [`GenServer::init`] runs once before the first message. Use it to start
|
||||
//! timers or set up monitors; see the [`GenServerCtx`] it receives.
|
||||
//! - [`GenServer::terminate`] runs when the server is about to exit. It fires
|
||||
//! on every exit path (all `GenServerRef`s dropped, a handler panic, or an
|
||||
//! explicit [`GenServerRef::shutdown`]), not only on clean shutdown. Keep it
|
||||
//! short and non-blocking: if `terminate` panics while the server is already
|
||||
//! unwinding from a handler panic, the process aborts.
|
||||
//! on every exit path (a self-stop, a graceful shutdown, a cooperative hard
|
||||
//! stop, or a handler panic), not only on clean shutdown.
|
||||
//! Keep it non-panicking: on the panic and hard-stop paths it runs
|
||||
//! mid-unwind, where a second panic aborts the process and where it must
|
||||
//! not block (any park re-observes the stop). Only on the graceful path
|
||||
//! (see below) may it do real work.
|
||||
//!
|
||||
//! ## When the server stops
|
||||
//!
|
||||
//! The server runs as long as at least one [`GenServerRef`] exists. When the last
|
||||
//! one is dropped, the inbox closes and the loop exits gracefully. To stop a
|
||||
//! server explicitly and wait for it to finish, call [`GenServerRef::shutdown`].
|
||||
//! A server lives until it stops, is shut down, or is killed — as an OTP
|
||||
//! process does. A [`GenServerRef`] is an *address*: cloning and dropping it
|
||||
//! never changes the server's lifetime, and a ref that nobody holds is not a
|
||||
//! leak — a forgotten server idles until the run ends, when the root-exit
|
||||
//! shutdown (see [`Runtime::run`](crate::Runtime::run)) takes it down with
|
||||
//! every other unsupervised actor. Anything meant to live long should be
|
||||
//! supervised (see *Supervised servers* below); [`start`] / [`start_under`]
|
||||
//! are for scripts, tests and short-lived helpers, and the explicit close is
|
||||
//! [`GenServerRef::shutdown`].
|
||||
//!
|
||||
//! A server can end itself: clone a [`StopHandle`] from
|
||||
//! [`GenServerCtx::stop_handle`] in `init` and call [`StopHandle::stop`] from
|
||||
//! any handler — the loop breaks after the current message and exits
|
||||
//! *normally* (OTP's `{stop, normal}`). This is distinct from
|
||||
//! `request_stop(self_pid())`, which is an abnormal `Stopped` and gets a
|
||||
//! `Transient` child restarted.
|
||||
//!
|
||||
//! ## Graceful shutdown
|
||||
//!
|
||||
//! From outside, [`GenServerRef::shutdown`] (or a plain
|
||||
//! [`request_shutdown`](crate::request_shutdown), which is what a supervisor
|
||||
//! sends) asks the server to stop. What happens next is the server's choice:
|
||||
//!
|
||||
//! - By default a server does not trap exits, and the request stops it
|
||||
//! outright at its next observation point — `terminate` runs mid-unwind.
|
||||
//! - A server that calls [`GenServerCtx::trap_exit`] in `init` receives the
|
||||
//! request as [`GenServer::handle_shutdown`]. Return
|
||||
//! [`ShutdownAction::Exit`] (the default) to have the loop break and
|
||||
//! `terminate` run on the normal path, where it may block; return
|
||||
//! [`ShutdownAction::Continue`] to keep serving — e.g. to drain in-flight
|
||||
//! work — and end the server later with a [`StopHandle`]. The supervisor's
|
||||
//! [`Shutdown`](crate::supervisor::Shutdown) policy bounds how long that
|
||||
//! may take before it falls back to a hard stop.
|
||||
//!
|
||||
//! A trapping server also receives the deaths of its linked peers as
|
||||
//! [`GenServer::handle_exit`] messages instead of dying with them.
|
||||
//!
|
||||
//! If the server panics inside a handler, the panic unwinds the server thread.
|
||||
//! Any caller currently waiting in `call` sees `Err(ServerDown)`: the reply
|
||||
@@ -167,9 +202,23 @@
|
||||
//! Names and registration: a server can be given a static name so other
|
||||
//! actors can reach it without holding a `GenServerRef`. Use
|
||||
//! [`GenServerBuilder::named`] to register on start, and the free functions
|
||||
//! [`call`], [`cast`], and [`whereis_server`] to address it by name. Registered
|
||||
//! servers are a natural fit for supervision; see `supervisor` for how to
|
||||
//! build a tree that restarts servers on failure.
|
||||
//! [`call`], [`cast`], and [`whereis_server`] to address it by name.
|
||||
//!
|
||||
//! ## Supervised servers
|
||||
//!
|
||||
//! The supervised shape is [`NamedGenServerBuilder::run`]: it runs the loop
|
||||
//! **inline, as the current actor**, so the closure of a
|
||||
//! [`ChildSpec`](crate::supervisor::ChildSpec) *is* the server — the
|
||||
//! supervisor's shutdown arrives as [`GenServer::handle_shutdown`], a restart
|
||||
//! runs the factory again and re-binds the name, and the rest of the program
|
||||
//! addresses it by name (a held ref would go stale on restart anyway).
|
||||
//!
|
||||
//! ```ignore
|
||||
//! const COUNTER: GenServerName<Counter> = GenServerName::new("counter");
|
||||
//! OneForOne::new().child(ChildSpec::new(Restart::Permanent, || {
|
||||
//! GenServerBuilder::new(Counter::default()).named(COUNTER).run().unwrap();
|
||||
//! }));
|
||||
//! ```
|
||||
//!
|
||||
//! ## Limitations
|
||||
//!
|
||||
@@ -178,11 +227,15 @@
|
||||
//! from any handler via [`Watcher::watch`]) because monitors are inherently
|
||||
//! created at runtime. The idle window is set once, in `init`.
|
||||
|
||||
use crate::channel::{channel, select, select_timeout, Receiver, RecvTimeoutError, Selectable, Sender};
|
||||
use crate::channel::{
|
||||
channel, select, select_timeout, Receiver, RecvTimeoutError, Selectable, Sender,
|
||||
};
|
||||
use crate::link::ExitSignal;
|
||||
use crate::monitor::DownReason;
|
||||
use crate::monitor::{demonitor, monitor, Down, Monitor};
|
||||
use crate::pid::Pid;
|
||||
use crate::registry::{register_with, resolve_named_sender, RegisterError};
|
||||
use crate::scheduler::{cancel_timer, request_stop, send_after_to, spawn, spawn_under};
|
||||
use crate::scheduler::{cancel_timer, request_shutdown, send_after_to};
|
||||
use crate::timer::TimerId;
|
||||
use std::cell::Cell;
|
||||
use std::collections::HashMap;
|
||||
@@ -252,10 +305,34 @@ pub trait GenServer: Send + 'static {
|
||||
/// Default: no-op.
|
||||
fn handle_idle(&mut self) {}
|
||||
|
||||
/// A graceful shutdown request (a [`request_shutdown`](crate::request_shutdown)
|
||||
/// reaching this server), delivered only if `init` called
|
||||
/// [`GenServerCtx::trap_exit`]. Return [`ShutdownAction::Exit`] to stop
|
||||
/// now (the default), or [`ShutdownAction::Continue`] to keep serving and
|
||||
/// end the server later with a [`StopHandle`].
|
||||
fn handle_shutdown(&mut self) -> ShutdownAction {
|
||||
ShutdownAction::Exit
|
||||
}
|
||||
|
||||
/// A linked peer's abnormal death (an [`ExitSignal`] that is not a
|
||||
/// shutdown request), delivered only if `init` called
|
||||
/// [`GenServerCtx::trap_exit`]. Default: drop it.
|
||||
fn handle_exit(&mut self, _sig: ExitSignal) {}
|
||||
|
||||
/// Runs as the server actor exits, on any exit path (see module docs).
|
||||
fn terminate(&mut self) {}
|
||||
}
|
||||
|
||||
/// What a server does with a shutdown request; see [`GenServer::handle_shutdown`].
|
||||
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
|
||||
pub enum ShutdownAction {
|
||||
/// Break the loop now. `terminate` runs on the normal path and may block.
|
||||
Exit,
|
||||
/// Keep dispatching. The server is expected to end itself with a
|
||||
/// [`StopHandle`] once it is done winding down.
|
||||
Continue,
|
||||
}
|
||||
|
||||
/// What travels the server's single inbox channel: a synchronous call (with a
|
||||
/// reply sender) or an asynchronous cast. Private — callers use [`GenServerRef`].
|
||||
enum Envelope<G: GenServer> {
|
||||
@@ -263,9 +340,9 @@ enum Envelope<G: GenServer> {
|
||||
Cast(G::Cast),
|
||||
}
|
||||
|
||||
/// A clonable handle to a running server. Cloning yields another sender to the
|
||||
/// same inbox; the server lives until the last `GenServerRef` is dropped, at which
|
||||
/// point its inbox closes and the loop exits normally.
|
||||
/// A clonable handle to a running server: an *address*, not an owner. Cloning
|
||||
/// yields another sender to the same inbox; dropping refs never ends the
|
||||
/// server (see the module docs, *When the server stops*).
|
||||
pub struct GenServerRef<G: GenServer> {
|
||||
tx: Sender<Envelope<G>>,
|
||||
pid: Pid,
|
||||
@@ -273,7 +350,10 @@ pub struct GenServerRef<G: GenServer> {
|
||||
|
||||
impl<G: GenServer> Clone for GenServerRef<G> {
|
||||
fn clone(&self) -> Self {
|
||||
GenServerRef { tx: self.tx.clone(), pid: self.pid }
|
||||
GenServerRef {
|
||||
tx: self.tx.clone(),
|
||||
pid: self.pid,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -357,18 +437,22 @@ impl<G: GenServer> GenServerRef<G> {
|
||||
|
||||
/// Stop the server and block until it has fully exited.
|
||||
///
|
||||
/// Sends a cooperative stop signal to the server actor and waits for it to
|
||||
/// exit, so [`GenServer::terminate`] has run by the time this returns.
|
||||
/// Returns immediately if the server is already gone.
|
||||
/// Asks the server to shut down (a [`request_shutdown`](crate::request_shutdown))
|
||||
/// and waits for it to exit, so [`GenServer::terminate`] has run by the
|
||||
/// time this returns. A trapping server gets to wind down via
|
||||
/// [`GenServer::handle_shutdown`]; any other is stopped outright. Returns
|
||||
/// immediately if the server is already gone. Waits as long as the server
|
||||
/// takes — the caller, not the server, decides whether that is acceptable;
|
||||
/// a supervisor uses its child's [`Shutdown`](crate::supervisor::Shutdown)
|
||||
/// policy to bound it.
|
||||
///
|
||||
/// This is the right teardown for a server kept alive by a registered
|
||||
/// [`GenServerName`], where dropping every external `GenServerRef` is not enough
|
||||
/// to close the inbox. Like all cooperative cancellation, it is best-effort:
|
||||
/// This is the explicit close: dropping refs never ends a server. Like
|
||||
/// all cooperative cancellation, it is best-effort:
|
||||
/// a server wedged in a tight loop with no observation point cannot be
|
||||
/// stopped this way. Panics if called outside `Runtime::run()`.
|
||||
pub fn shutdown(&self) {
|
||||
let mon = monitor(self.pid);
|
||||
request_stop(self.pid);
|
||||
request_shutdown(self.pid);
|
||||
// The Down lands when the server finalizes; an already-dead target makes
|
||||
// `monitor` deliver NoProc immediately, so this never blocks forever.
|
||||
let _ = mon.rx.recv();
|
||||
@@ -391,6 +475,8 @@ enum Sys<G: GenServer> {
|
||||
/// the payload factory, dispatches it to [`GenServer::handle_timer`], and
|
||||
/// re-arms the next tick before returning.
|
||||
Tick(crate::timer::TimerId),
|
||||
/// The state asked to end the server (via [`StopHandle::stop`]).
|
||||
Stop,
|
||||
}
|
||||
|
||||
/// The server loop's runtime hook, passed to [`GenServer::init`]. Hands out the
|
||||
@@ -406,13 +492,35 @@ pub struct GenServerCtx<G: GenServer> {
|
||||
/// because `init` holds only `&ctx`; not `Send`, but `GenServerCtx` is only ever
|
||||
/// borrowed on the actor's own stack during `init`, never sent.
|
||||
idle: Cell<Option<Duration>>,
|
||||
/// Whether the loop should trap exits (set via [`trap_exit`](Self::trap_exit)
|
||||
/// during `init`, read by the loop after).
|
||||
trap: Cell<bool>,
|
||||
}
|
||||
|
||||
impl<G: GenServer> GenServerCtx<G> {
|
||||
/// Trap exits for the server's lifetime: shutdown requests then arrive as
|
||||
/// [`GenServer::handle_shutdown`] and linked-peer deaths as
|
||||
/// [`GenServer::handle_exit`], instead of stopping the server outright.
|
||||
/// Call this once during `init`.
|
||||
pub fn trap_exit(&self) {
|
||||
self.trap.set(true);
|
||||
}
|
||||
|
||||
/// A clonable handle that lets the state end the server from any handler
|
||||
/// (a normal exit; see the module docs). Store it on the state during
|
||||
/// `init`.
|
||||
pub fn stop_handle(&self) -> StopHandle<G> {
|
||||
StopHandle {
|
||||
sys_tx: self.sys_tx.clone(),
|
||||
}
|
||||
}
|
||||
|
||||
/// A clonable handle to the loop's monitor intake. Store it in the state
|
||||
/// during `init` to watch monitors from later handlers.
|
||||
pub fn watcher(&self) -> Watcher<G> {
|
||||
Watcher { tx: self.sys_tx.clone() }
|
||||
Watcher {
|
||||
tx: self.sys_tx.clone(),
|
||||
}
|
||||
}
|
||||
|
||||
/// Shorthand for `ctx.watcher().watch(m)` when watching during `init`.
|
||||
@@ -426,7 +534,10 @@ impl<G: GenServer> GenServerCtx<G> {
|
||||
/// [`tick_every`](TimerHandle::tick_every) /
|
||||
/// [`cancel`](TimerHandle::cancel) from any later handler.
|
||||
pub fn timer(&self) -> TimerHandle<G> {
|
||||
TimerHandle { sys_tx: self.sys_tx.clone(), reg: self.reg.clone() }
|
||||
TimerHandle {
|
||||
sys_tx: self.sys_tx.clone(),
|
||||
reg: self.reg.clone(),
|
||||
}
|
||||
}
|
||||
|
||||
/// Set a quiet-period window: if the loop goes `after` without dispatching
|
||||
@@ -442,6 +553,30 @@ impl<G: GenServer> GenServerCtx<G> {
|
||||
}
|
||||
}
|
||||
|
||||
/// Lets a server's state end the server, cloned from
|
||||
/// [`GenServerCtx::stop_handle`] during `init`. [`stop`](Self::stop) makes the
|
||||
/// loop break after the current message and exit normally; `terminate` runs on
|
||||
/// the normal path.
|
||||
pub struct StopHandle<G: GenServer> {
|
||||
sys_tx: Sender<Sys<G>>,
|
||||
}
|
||||
|
||||
impl<G: GenServer> Clone for StopHandle<G> {
|
||||
fn clone(&self) -> Self {
|
||||
StopHandle {
|
||||
sys_tx: self.sys_tx.clone(),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
impl<G: GenServer> StopHandle<G> {
|
||||
/// End the server after the current message. Idempotent; a no-op once the
|
||||
/// server is gone.
|
||||
pub fn stop(&self) {
|
||||
let _ = self.sys_tx.send(Sys::Stop);
|
||||
}
|
||||
}
|
||||
|
||||
/// Per-server timer bookkeeping, shared between the loop and every
|
||||
/// [`TimerHandle`] clone. A gen_server actor is single-threaded — handlers
|
||||
/// and the loop never run concurrently — so this `Mutex` is always
|
||||
@@ -518,7 +653,10 @@ pub struct TimerHandle<G: GenServer> {
|
||||
// Manual Clone for the same reason as `Watcher`: no `G: Clone` needed.
|
||||
impl<G: GenServer> Clone for TimerHandle<G> {
|
||||
fn clone(&self) -> Self {
|
||||
TimerHandle { sys_tx: self.sys_tx.clone(), reg: self.reg.clone() }
|
||||
TimerHandle {
|
||||
sys_tx: self.sys_tx.clone(),
|
||||
reg: self.reg.clone(),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -569,8 +707,18 @@ impl<G: GenServer> TimerHandle<G> {
|
||||
// First instance fires after `every`; the payload is produced loop-side
|
||||
// from `make` on fire, so the tick carries only the stable id.
|
||||
let sub = send_after_to(every, self.sys_tx.clone(), Sys::Tick(local));
|
||||
reg.periodics.insert(local, Periodic { every, live: sub, make });
|
||||
debug_assert!(reg.rearm_tx.is_some(), "rearm_tx must be Some while periodics is non-empty");
|
||||
reg.periodics.insert(
|
||||
local,
|
||||
Periodic {
|
||||
every,
|
||||
live: sub,
|
||||
make,
|
||||
},
|
||||
);
|
||||
debug_assert!(
|
||||
reg.rearm_tx.is_some(),
|
||||
"rearm_tx must be Some while periodics is non-empty"
|
||||
);
|
||||
local
|
||||
}
|
||||
|
||||
@@ -617,7 +765,9 @@ pub struct Watcher<G: GenServer> {
|
||||
// regardless of the server type (it clones only the inner sender).
|
||||
impl<G: GenServer> Clone for Watcher<G> {
|
||||
fn clone(&self) -> Self {
|
||||
Watcher { tx: self.tx.clone() }
|
||||
Watcher {
|
||||
tx: self.tx.clone(),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -643,11 +793,17 @@ pub struct GenServerBuilder<G: GenServer> {
|
||||
state: G,
|
||||
infos: Vec<Receiver<G::Info>>,
|
||||
supervisor: Option<Pid>,
|
||||
stack_opts: crate::scheduler::SpawnOpts,
|
||||
}
|
||||
|
||||
impl<G: GenServer> GenServerBuilder<G> {
|
||||
pub fn new(state: G) -> Self {
|
||||
GenServerBuilder { state, infos: Vec::new(), supervisor: None }
|
||||
GenServerBuilder {
|
||||
state,
|
||||
infos: Vec::new(),
|
||||
supervisor: None,
|
||||
stack_opts: crate::scheduler::SpawnOpts::default(),
|
||||
}
|
||||
}
|
||||
|
||||
/// Add an out-of-band channel; messages arriving on it are dispatched to
|
||||
@@ -665,9 +821,18 @@ impl<G: GenServer> GenServerBuilder<G> {
|
||||
self
|
||||
}
|
||||
|
||||
/// Spawn the server actor and hand back its [`GenServerRef`]. The server's
|
||||
/// lifetime is governed by its refs, not by joining, so the backing join
|
||||
/// handle is dropped.
|
||||
/// Stack shape for the server actor (RFC 019) — see
|
||||
/// [`SpawnOpts`](crate::SpawnOpts). Useful for servers that recurse
|
||||
/// deeply or call into FFI with large C frames.
|
||||
pub fn stack_opts(mut self, opts: crate::scheduler::SpawnOpts) -> Self {
|
||||
self.stack_opts = opts;
|
||||
self
|
||||
}
|
||||
|
||||
/// Spawn the server actor and hand back its [`GenServerRef`] (an address;
|
||||
/// the server's lifetime is its own, see the module docs). The backing
|
||||
/// join handle is dropped. For a supervised server use
|
||||
/// [`named`](Self::named) + [`NamedGenServerBuilder::run`] instead.
|
||||
pub fn start(self) -> GenServerRef<G> {
|
||||
self.spawn_server()
|
||||
}
|
||||
@@ -677,7 +842,10 @@ impl<G: GenServer> GenServerBuilder<G> {
|
||||
/// live server). Consumes the builder, carrying its `with_info` / `under`
|
||||
/// configuration through.
|
||||
pub fn named(self, name: GenServerName<G>) -> NamedGenServerBuilder<G> {
|
||||
NamedGenServerBuilder { builder: self, name: name.as_str() }
|
||||
NamedGenServerBuilder {
|
||||
builder: self,
|
||||
name: name.as_str(),
|
||||
}
|
||||
}
|
||||
|
||||
/// Private shared body behind [`start`](Self::start) and
|
||||
@@ -686,12 +854,25 @@ impl<G: GenServer> GenServerBuilder<G> {
|
||||
/// under the name before returning.
|
||||
fn spawn_server(self) -> GenServerRef<G> {
|
||||
let (tx, rx) = channel::<Envelope<G>>();
|
||||
let GenServerBuilder { state, infos, supervisor } = self;
|
||||
let GenServerBuilder {
|
||||
state,
|
||||
infos,
|
||||
supervisor,
|
||||
stack_opts,
|
||||
} = self;
|
||||
let keep = tx.clone();
|
||||
let handle = match supervisor {
|
||||
Some(sup) => spawn_under(sup, move || server_loop::<G>(rx, state, infos)),
|
||||
None => spawn(move || server_loop::<G>(rx, state, infos)),
|
||||
Some(sup) => crate::scheduler::spawn_under_with(sup, stack_opts, move || {
|
||||
server_loop::<G>(keep, rx, state, infos)
|
||||
}),
|
||||
None => crate::scheduler::spawn_with(stack_opts, move || {
|
||||
server_loop::<G>(keep, rx, state, infos)
|
||||
}),
|
||||
};
|
||||
GenServerRef { tx, pid: handle.pid() }
|
||||
GenServerRef {
|
||||
tx,
|
||||
pid: handle.pid(),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -719,7 +900,10 @@ impl<G> GenServerName<G> {
|
||||
/// associated constants at call sites.
|
||||
#[inline]
|
||||
pub const fn new(name: &'static str) -> Self {
|
||||
Self { name, _marker: PhantomData }
|
||||
Self {
|
||||
name,
|
||||
_marker: PhantomData,
|
||||
}
|
||||
}
|
||||
|
||||
/// The underlying registry key.
|
||||
@@ -758,6 +942,12 @@ impl<G: GenServer> NamedGenServerBuilder<G> {
|
||||
self
|
||||
}
|
||||
|
||||
/// Stack shape for the server actor (see [`GenServerBuilder::stack_opts`]).
|
||||
pub fn stack_opts(mut self, opts: crate::scheduler::SpawnOpts) -> Self {
|
||||
self.builder = self.builder.stack_opts(opts);
|
||||
self
|
||||
}
|
||||
|
||||
/// Spawn the server and bind its name in one step. Fallible: returns
|
||||
/// [`RegisterError::NameTaken`] if the name is already held by a different
|
||||
/// live server.
|
||||
@@ -765,19 +955,38 @@ impl<G: GenServer> NamedGenServerBuilder<G> {
|
||||
/// The inbox sender is published under the name **from the parent side,
|
||||
/// before this returns**, so a by-name `call` / `cast` resolves the instant
|
||||
/// `start()` returns — no race with the server body. On a name clash the
|
||||
/// just-spawned server is wound down (its only ref is dropped, closing the
|
||||
/// inbox), so a failed bind leaks no actor.
|
||||
/// just-spawned server is stopped, so a failed bind leaks no actor.
|
||||
pub fn start(self) -> Result<GenServerRef<G>, RegisterError> {
|
||||
let NamedGenServerBuilder { builder, name } = self;
|
||||
let server = builder.spawn_server();
|
||||
match register_with::<Envelope<G>>(server.pid, name, server.tx.clone()) {
|
||||
Ok(()) => Ok(server),
|
||||
Err(e) => {
|
||||
drop(server); // inbox closes → loop exits gracefully
|
||||
crate::scheduler::request_stop(server.pid); // never ran init
|
||||
Err(e)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Run the server **inline, as the current actor**, bound to its name.
|
||||
/// This is the supervised shape: the closure of a
|
||||
/// [`ChildSpec`](crate::supervisor::ChildSpec) *is* the server, so the
|
||||
/// supervisor's shutdown reaches it as [`GenServer::handle_shutdown`], a
|
||||
/// restart runs the factory again and re-binds the name, and clients
|
||||
/// address it by name ([`call`], [`cast`], [`whereis_server`]). Returns
|
||||
/// when the server exits; [`RegisterError::NameTaken`] (before `init`) if
|
||||
/// the name is held by a different live actor.
|
||||
///
|
||||
/// `under` / `stack_opts` are spawn options and do not apply here — the
|
||||
/// actor already exists.
|
||||
pub fn run(self) -> Result<(), RegisterError> {
|
||||
let NamedGenServerBuilder { builder, name } = self;
|
||||
let GenServerBuilder { state, infos, .. } = builder;
|
||||
let (tx, rx) = channel::<Envelope<G>>();
|
||||
register_with::<Envelope<G>>(crate::scheduler::self_pid(), name, tx.clone())?;
|
||||
server_loop::<G>(tx, rx, state, infos);
|
||||
Ok(())
|
||||
}
|
||||
}
|
||||
|
||||
/// Resolve a [`GenServerName`] to a [`GenServerRef`] when you want a handle to hold or
|
||||
@@ -836,10 +1045,17 @@ pub fn start_under<G: GenServer>(supervisor: Pid, state: G) -> GenServerRef<G> {
|
||||
}
|
||||
|
||||
fn server_loop<G: GenServer>(
|
||||
keep: Sender<Envelope<G>>,
|
||||
rx: Receiver<Envelope<G>>,
|
||||
state: G,
|
||||
mut infos: Vec<Receiver<G::Info>>,
|
||||
) {
|
||||
// The loop holds one inbox sender for its whole life: the inbox never
|
||||
// closes, so refs are addresses and the server's lifetime is the actor's
|
||||
// (stop handle, shutdown, stop, panic). The `Disconnected` arms below are
|
||||
// defensive only.
|
||||
let _keep = keep;
|
||||
|
||||
// Drop guard — owns the server state and the timer registry.
|
||||
//
|
||||
// Why a guard rather than code after the loop:
|
||||
@@ -907,9 +1123,18 @@ fn server_loop<G: GenServer>(
|
||||
// Bind the ctx so the idle window set during init can be read back, then
|
||||
// drop it — that drops the loop's own Sys sender, so a state that cloned no
|
||||
// Watcher/TimerHandle lets the arm auto-close (the unused-ctx behaviour).
|
||||
let ctx = GenServerCtx { sys_tx, reg: reg.clone(), idle: Cell::new(None) };
|
||||
let ctx = GenServerCtx {
|
||||
sys_tx,
|
||||
reg: reg.clone(),
|
||||
idle: Cell::new(None),
|
||||
trap: Cell::new(false),
|
||||
};
|
||||
guard.0.init(&ctx);
|
||||
let idle = ctx.idle.get();
|
||||
// Trapping is opted into during init and fixed for the loop's life. The
|
||||
// inbox is armed only when set: an untrapped server keeps the fast path,
|
||||
// and a shutdown request simply stops it as `request_stop` would.
|
||||
let exits: Option<Receiver<ExitSignal>> = ctx.trap.get().then(crate::link::trap_exit);
|
||||
drop(ctx);
|
||||
|
||||
let mut monitors: Vec<Monitor> = Vec::new();
|
||||
@@ -925,7 +1150,7 @@ fn server_loop<G: GenServer>(
|
||||
};
|
||||
|
||||
loop {
|
||||
if monitors.is_empty() && !sys_open && infos.is_empty() {
|
||||
if exits.is_none() && monitors.is_empty() && !sys_open && infos.is_empty() {
|
||||
// Fast path: no extra arms, no select overhead — park directly on
|
||||
// the inbox. Mirrors the inbox arm of the select path below; any
|
||||
// change there must be applied here too.
|
||||
@@ -942,7 +1167,7 @@ fn server_loop<G: GenServer>(
|
||||
guard.0.handle_idle();
|
||||
reset_idle(&mut idle_deadline);
|
||||
}
|
||||
// All ServerRefs dropped → inbox closed → shutdown.
|
||||
// Defensive: the loop holds a sender, so unreachable.
|
||||
Err(RecvTimeoutError::Disconnected) => break,
|
||||
}
|
||||
}
|
||||
@@ -953,17 +1178,21 @@ fn server_loop<G: GenServer>(
|
||||
}
|
||||
} else {
|
||||
// Slow path: one or more extra arms live — build the arm slice and
|
||||
// select. Arm order encodes priority: downs → system → infos →
|
||||
// inbox. The slice is rebuilt each iteration because the monitor
|
||||
// and info sets shrink/grow. Mirrors the fast-path inbox park
|
||||
// above; keep them in sync.
|
||||
let nd = monitors.len(); // monitor band: [0, nd)
|
||||
let nw = sys_open as usize; // system arm: [nd, nd+nw)
|
||||
// info band: [nd+nw, nd+nw+ni)
|
||||
// inbox arm: [nd+nw+ni]
|
||||
// select. Arm order encodes priority: exits → downs → system →
|
||||
// infos → inbox (a shutdown request is noticed under any load).
|
||||
// The slice is rebuilt each iteration because the monitor and info
|
||||
// sets shrink/grow. Mirrors the fast-path inbox park above; keep
|
||||
// them in sync.
|
||||
let ne = exits.is_some() as usize; // exit arm: [0, ne)
|
||||
let nd = ne + monitors.len(); // monitor band: [ne, nd)
|
||||
let nw = sys_open as usize; // system arm: [nd, nd+nw)
|
||||
// info band: [nd+nw, nd+nw+ni)
|
||||
// inbox arm: [nd+nw+ni]
|
||||
let sel = {
|
||||
let mut arms: Vec<&dyn Selectable> =
|
||||
Vec::with_capacity(nd + nw + infos.len() + 1);
|
||||
let mut arms: Vec<&dyn Selectable> = Vec::with_capacity(nd + nw + infos.len() + 1);
|
||||
if let Some(e) = &exits {
|
||||
arms.push(e);
|
||||
}
|
||||
for m in &monitors {
|
||||
arms.push(&m.rx);
|
||||
}
|
||||
@@ -975,9 +1204,7 @@ fn server_loop<G: GenServer>(
|
||||
}
|
||||
arms.push(&rx);
|
||||
match idle_deadline {
|
||||
Some(dl) => {
|
||||
select_timeout(&arms, dl.saturating_duration_since(Instant::now()))
|
||||
}
|
||||
Some(dl) => select_timeout(&arms, dl.saturating_duration_since(Instant::now())),
|
||||
None => Some(select(&arms)),
|
||||
}
|
||||
};
|
||||
@@ -991,10 +1218,25 @@ fn server_loop<G: GenServer>(
|
||||
continue;
|
||||
}
|
||||
};
|
||||
if i < nd {
|
||||
if i < ne {
|
||||
// Exit arm: a shutdown request or a linked peer's death.
|
||||
// The inbox lives for the loop's life, so it never closes.
|
||||
let sig = exits.as_ref().and_then(|e| e.try_recv().ok().flatten());
|
||||
if let Some(sig) = sig {
|
||||
if sig.reason == DownReason::Shutdown {
|
||||
match guard.0.handle_shutdown() {
|
||||
ShutdownAction::Exit => break,
|
||||
ShutdownAction::Continue => {}
|
||||
}
|
||||
} else {
|
||||
guard.0.handle_exit(sig);
|
||||
}
|
||||
reset_idle(&mut idle_deadline);
|
||||
}
|
||||
} else if i < nd {
|
||||
// Monitor band: a Down retires its arm either way (one-shot)
|
||||
// or closes without delivering (defensive; shouldn't happen).
|
||||
let m = monitors.remove(i);
|
||||
let m = monitors.remove(i - ne);
|
||||
if let Ok(Some(down)) = m.rx.try_recv() {
|
||||
guard.0.handle_down(down);
|
||||
reset_idle(&mut idle_deadline);
|
||||
@@ -1003,13 +1245,19 @@ fn server_loop<G: GenServer>(
|
||||
match sys_rx.try_recv() {
|
||||
// Control intake, not a dispatched message: no idle reset.
|
||||
Ok(Some(Sys::Watch(m))) => monitors.push(m),
|
||||
// The state ended the server: a normal exit.
|
||||
Ok(Some(Sys::Stop)) => break,
|
||||
Ok(Some(Sys::Timer(id, msg))) => {
|
||||
// The one-shot fired: retire its registry entry so the
|
||||
// live set tracks only still-pending timers, then
|
||||
// dispatch.
|
||||
match reg.lock() {
|
||||
Ok(mut g) => { g.oneshots.remove(&id); }
|
||||
Err(e) => panic!("smarm: gen_server reg lock poisoned (core corrupt): {e}"),
|
||||
Ok(mut g) => {
|
||||
g.oneshots.remove(&id);
|
||||
}
|
||||
Err(e) => {
|
||||
panic!("smarm: gen_server reg lock poisoned (core corrupt): {e}")
|
||||
}
|
||||
}
|
||||
guard.0.handle_timer(msg);
|
||||
reset_idle(&mut idle_deadline);
|
||||
@@ -1023,7 +1271,9 @@ fn server_loop<G: GenServer>(
|
||||
let msg = {
|
||||
let mut g = match reg.lock() {
|
||||
Ok(g) => g,
|
||||
Err(e) => panic!("smarm: gen_server reg lock poisoned (core corrupt): {e}"),
|
||||
Err(e) => panic!(
|
||||
"smarm: gen_server reg lock poisoned (core corrupt): {e}"
|
||||
),
|
||||
};
|
||||
let r = &mut *g;
|
||||
if let Some(p) = r.periodics.get_mut(&id) {
|
||||
@@ -1031,9 +1281,9 @@ fn server_loop<G: GenServer>(
|
||||
let msg = (p.make)();
|
||||
let tx = match r.rearm_tx.as_ref() {
|
||||
Some(tx) => tx.clone(),
|
||||
None => panic!(
|
||||
"smarm: live periodic without rearm_tx (logic bug)"
|
||||
),
|
||||
None => {
|
||||
panic!("smarm: live periodic without rearm_tx (logic bug)")
|
||||
}
|
||||
};
|
||||
p.live = send_after_to(every, tx, Sys::Tick(id));
|
||||
Some(msg)
|
||||
|
||||
+503
-72
@@ -68,10 +68,38 @@
|
||||
//! inbox or timer event; a replayed event may postpone again (it re-queues for
|
||||
//! the next transition). See the macro docs for the row surface and [`Step`]
|
||||
//! for how a postpone surfaces to the loop.
|
||||
//!
|
||||
//! ## Stopping, shutdown, and exits
|
||||
//!
|
||||
//! A machine lives until it stops, is shut down, or is killed; a
|
||||
//! [`GenStatemRef`] is an address, and dropping refs never ends it (the
|
||||
//! gen_server rule — see its *When the server stops*). The supervised shape
|
||||
//! is [`run_named`], which runs the machine inline as the current actor so it
|
||||
//! is a direct `ChildSpec` child addressed by [`GenStatemName`]. A machine
|
||||
//! can end itself: any row
|
||||
//! body may call [`cx.stop()`](Cx::stop) (or use the `stop` tail keyword,
|
||||
//! sugar for `{ cx.stop(); prev }`) — the loop breaks after that event and the
|
||||
//! actor exits *normally* (OTP's `{stop, normal}`; a `Transient` child is not
|
||||
//! restarted). [`terminate`](Machine::terminate) — the optional `terminate
|
||||
//! { … }` macro block — runs on every exit path.
|
||||
//!
|
||||
//! From outside, [`GenStatemRef::shutdown`] (or a plain
|
||||
//! [`request_shutdown`](crate::request_shutdown), which is what a supervisor
|
||||
//! sends) asks the machine to stop. By default a machine does not trap exits
|
||||
//! and the request stops it outright. A machine that calls
|
||||
//! [`cx.trap_exit()`](Cx::trap_exit) in its initial `enter` instead receives
|
||||
//! it as the **`shutdown` event**, routed by state like any other — a
|
||||
//! `Connected` state may transition into `Draining` and stop later from a
|
||||
//! timeout row, while a state with no `shutdown` row takes the macro's
|
||||
//! default, `stop`. Linked-peer deaths reach a trapping machine as
|
||||
//! `exit <pat>` events; an unmatched one is dropped like an unmatched info.
|
||||
|
||||
use crate::channel::{channel, select, Receiver, Sender};
|
||||
use crate::channel::{channel, select, Receiver, Selectable, Sender};
|
||||
use crate::link::ExitSignal;
|
||||
use crate::monitor::{monitor, DownReason};
|
||||
use crate::pid::Pid;
|
||||
use crate::scheduler::{cancel_timer, send_after_to, spawn as spawn_actor};
|
||||
use crate::registry::{register_with, resolve_named_sender, RegisterError};
|
||||
use crate::scheduler::{cancel_timer, request_shutdown, send_after_to};
|
||||
use crate::timer::TimerId;
|
||||
use std::collections::{HashMap, VecDeque};
|
||||
use std::marker::PhantomData;
|
||||
@@ -119,6 +147,33 @@ pub trait Machine: Send + 'static {
|
||||
/// state's `enter`, and returns [`Step::Transitioned`]; a stay or unmatched
|
||||
/// event returns [`Step::Stayed`].
|
||||
fn handle(&mut self, ev: Self::Ev, cx: &mut Cx<Self::Ev>) -> Step<Self::Ev>;
|
||||
|
||||
/// Wrap a graceful shutdown request into this machine's event, so a
|
||||
/// trapping machine (see [`Cx::trap_exit`]) can route it **by state**. The
|
||||
/// macro generates it as `Ev::Shutdown` and matches it in `shutdown` rows;
|
||||
/// its default for a state that writes no such row is `stop`. A
|
||||
/// hand-written machine that returns `None` (the default) is simply
|
||||
/// stopped — the loop breaks and `terminate` runs on the normal path.
|
||||
fn shutdown_ev() -> Option<Self::Ev> {
|
||||
None
|
||||
}
|
||||
|
||||
/// Wrap a linked peer's death (an [`ExitSignal`] that is not a shutdown
|
||||
/// request, delivered only when trapping) into this machine's event. The
|
||||
/// macro generates it as `Ev::Exit(sig)` and matches it in `exit <pat>`
|
||||
/// rows; an unmatched exit is silently dropped, like an unmatched info.
|
||||
/// A hand-written machine that returns `None` (the default) drops it.
|
||||
fn exit_ev(_sig: ExitSignal) -> Option<Self::Ev> {
|
||||
None
|
||||
}
|
||||
|
||||
/// Runs as the machine actor exits, on any exit path (a `stop`, a
|
||||
/// graceful shutdown, a handler panic, a hard stop). Like
|
||||
/// `gen_server::terminate`: on the panic and hard-stop paths it runs
|
||||
/// mid-unwind — do not panic or park there; only on the normal path
|
||||
/// (`stop`, shutdown rows) may it do real work. The macro's optional
|
||||
/// `terminate { … }` block generates it.
|
||||
fn terminate(&mut self) {}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
@@ -219,7 +274,11 @@ struct Timers {
|
||||
|
||||
impl Timers {
|
||||
fn new() -> Self {
|
||||
Timers { next_local: 0, state: None, named: HashMap::new() }
|
||||
Timers {
|
||||
next_local: 0,
|
||||
state: None,
|
||||
named: HashMap::new(),
|
||||
}
|
||||
}
|
||||
|
||||
fn mint(&mut self) -> u64 {
|
||||
@@ -243,12 +302,41 @@ impl Timers {
|
||||
pub struct Cx<Ev> {
|
||||
sys_tx: Sender<Sys>,
|
||||
reg: Arc<Mutex<Timers>>,
|
||||
/// Set by [`trap_exit`](Self::trap_exit) during `on_start`; read once by
|
||||
/// the loop right after, fixed for the machine's life.
|
||||
trap: bool,
|
||||
/// Set by [`stop`](Self::stop); the loop breaks after the current event.
|
||||
stop: bool,
|
||||
_ev: PhantomData<fn() -> Ev>,
|
||||
}
|
||||
|
||||
impl<Ev> Cx<Ev> {
|
||||
fn new(sys_tx: Sender<Sys>, reg: Arc<Mutex<Timers>>) -> Self {
|
||||
Cx { sys_tx, reg, _ev: PhantomData }
|
||||
Cx {
|
||||
sys_tx,
|
||||
reg,
|
||||
trap: false,
|
||||
stop: false,
|
||||
_ev: PhantomData,
|
||||
}
|
||||
}
|
||||
|
||||
/// Trap exits for the machine's lifetime: a shutdown request then arrives
|
||||
/// as the `shutdown` event (routed by state) and linked-peer deaths as
|
||||
/// `exit` events, instead of stopping the machine outright. Call it in the
|
||||
/// initial state's `enter` (i.e. during `on_start`); later calls have no
|
||||
/// effect.
|
||||
pub fn trap_exit(&mut self) {
|
||||
self.trap = true;
|
||||
}
|
||||
|
||||
/// End the machine after the current event: the loop breaks and the actor
|
||||
/// exits *normally* (OTP's `{stop, normal}`); `terminate` runs on the
|
||||
/// normal path. Anything still queued or postponed is dropped. The `stop`
|
||||
/// tail keyword in a macro row is sugar for `{ cx.stop(); prev }`.
|
||||
/// Mirrors gen_server's [`StopHandle`](crate::gen_server::StopHandle).
|
||||
pub fn stop(&mut self) {
|
||||
self.stop = true;
|
||||
}
|
||||
|
||||
/// Arm the **state timeout**: fire a `state_timeout` event after `after` in
|
||||
@@ -378,8 +466,7 @@ pub enum SendError {
|
||||
}
|
||||
|
||||
/// A clonable handle to a running machine. Cloning yields another sender to the
|
||||
/// same inbox; the machine lives until the last `GenStatemRef` is dropped, at which
|
||||
/// point its inbox closes and the loop exits.
|
||||
/// same inbox. An address, not an owner: dropping refs never ends the machine.
|
||||
pub struct GenStatemRef<M: Machine> {
|
||||
tx: Sender<M::Ev>,
|
||||
pid: Pid,
|
||||
@@ -387,7 +474,10 @@ pub struct GenStatemRef<M: Machine> {
|
||||
|
||||
impl<M: Machine> Clone for GenStatemRef<M> {
|
||||
fn clone(&self) -> Self {
|
||||
GenStatemRef { tx: self.tx.clone(), pid: self.pid }
|
||||
GenStatemRef {
|
||||
tx: self.tx.clone(),
|
||||
pid: self.pid,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -422,6 +512,19 @@ impl<M: Machine> GenStatemRef<M> {
|
||||
self.send(ev).map_err(|_| CallError::Down)?;
|
||||
rx.recv().map_err(|_| CallError::Down)
|
||||
}
|
||||
|
||||
/// Ask the machine to shut down and block until it has fully exited, so
|
||||
/// [`Machine::terminate`] has run by the time this returns. A trapping
|
||||
/// machine winds down through its `shutdown` rows; any other is stopped
|
||||
/// outright. Returns immediately if the machine is already gone. Waits as
|
||||
/// long as the machine takes — a supervisor bounds that with its child's
|
||||
/// [`Shutdown`](crate::supervisor::Shutdown) policy. Mirrors
|
||||
/// [`GenServerRef::shutdown`](crate::gen_server::GenServerRef::shutdown).
|
||||
pub fn shutdown(&self) {
|
||||
let mon = monitor(self.pid);
|
||||
request_shutdown(self.pid);
|
||||
let _ = mon.rx.recv();
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
@@ -430,42 +533,204 @@ impl<M: Machine> GenStatemRef<M> {
|
||||
|
||||
/// Spawn `machine` as an actor and hand back its [`GenStatemRef`]. Shape mirrors
|
||||
/// `gen_server::start`: make the inbox, spawn the loop, return the ref; the
|
||||
/// backing join handle is dropped (lifetime is governed by refs, not joining).
|
||||
/// backing join handle is dropped (the machine's lifetime is its own). For a
|
||||
/// supervised machine use [`run_named`].
|
||||
///
|
||||
/// Panics if called outside `Runtime::run()`.
|
||||
pub fn spawn<M: Machine>(machine: M) -> GenStatemRef<M> {
|
||||
spawn_with(crate::scheduler::SpawnOpts::default(), machine)
|
||||
}
|
||||
|
||||
/// [`spawn`] with per-actor stack shape overrides (RFC 019) for the machine's
|
||||
/// actor — see [`SpawnOpts`](crate::SpawnOpts). gen_statem has no builder
|
||||
/// (its one-shot `spawn(machine)` shape predates RFC 019), so the opts ride
|
||||
/// a `_with` variant like the scheduler's own spawns.
|
||||
///
|
||||
/// Panics if called outside `Runtime::run()`.
|
||||
pub fn spawn_with<M: Machine>(opts: crate::scheduler::SpawnOpts, machine: M) -> GenStatemRef<M> {
|
||||
let (tx, rx) = channel::<M::Ev>();
|
||||
let handle = spawn_actor(move || statem_loop(rx, machine));
|
||||
GenStatemRef { tx, pid: handle.pid() }
|
||||
let keep = tx.clone();
|
||||
let handle = crate::scheduler::spawn_with(opts, move || statem_loop(keep, rx, machine));
|
||||
GenStatemRef {
|
||||
tx,
|
||||
pid: handle.pid(),
|
||||
}
|
||||
}
|
||||
|
||||
/// A typed, static name for a gen_statem, used to address a machine through
|
||||
/// the registry without holding a [`GenStatemRef`]. Mirrors
|
||||
/// [`GenServerName`](crate::gen_server::GenServerName): declare it as a
|
||||
/// constant and bind it with [`run_named`].
|
||||
pub struct GenStatemName<M> {
|
||||
name: &'static str,
|
||||
_marker: PhantomData<fn() -> M>,
|
||||
}
|
||||
|
||||
impl<M> GenStatemName<M> {
|
||||
/// Bind a static string as a machine name.
|
||||
#[inline]
|
||||
pub const fn new(name: &'static str) -> Self {
|
||||
Self {
|
||||
name,
|
||||
_marker: PhantomData,
|
||||
}
|
||||
}
|
||||
|
||||
/// The underlying registry key.
|
||||
#[inline]
|
||||
pub const fn as_str(self) -> &'static str {
|
||||
self.name
|
||||
}
|
||||
}
|
||||
|
||||
impl<M> Copy for GenStatemName<M> {}
|
||||
impl<M> Clone for GenStatemName<M> {
|
||||
fn clone(&self) -> Self {
|
||||
*self
|
||||
}
|
||||
}
|
||||
|
||||
/// Run `machine` **inline, as the current actor**, bound to `name`. The
|
||||
/// supervised shape: the closure of a
|
||||
/// [`ChildSpec`](crate::supervisor::ChildSpec) *is* the machine, so the
|
||||
/// supervisor's shutdown reaches it as a `shutdown` row, a restart runs the
|
||||
/// factory again and re-binds the name, and clients address it by name
|
||||
/// ([`send`], [`call`], [`whereis_machine`]). Returns when the machine exits;
|
||||
/// [`RegisterError::NameTaken`] (before `on_start`) if the name is held by a
|
||||
/// different live actor. Mirrors
|
||||
/// [`NamedGenServerBuilder::run`](crate::gen_server::NamedGenServerBuilder::run).
|
||||
pub fn run_named<M: Machine>(name: GenStatemName<M>, machine: M) -> Result<(), RegisterError> {
|
||||
let (tx, rx) = channel::<M::Ev>();
|
||||
register_with::<M::Ev>(crate::scheduler::self_pid(), name.as_str(), tx.clone())?;
|
||||
statem_loop(tx, rx, machine);
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Resolve a [`GenStatemName`] to a [`GenStatemRef`]; `None` if no live
|
||||
/// machine holds the name. Panics if called outside `Runtime::run()`.
|
||||
pub fn whereis_machine<M: Machine>(name: GenStatemName<M>) -> Option<GenStatemRef<M>> {
|
||||
resolve_named_sender::<M::Ev>(name.as_str()).map(|(pid, tx)| GenStatemRef { tx, pid })
|
||||
}
|
||||
|
||||
/// Push an event to the machine registered under `name`, resolving per send.
|
||||
/// [`SendError::Down`] if no live machine holds the name.
|
||||
pub fn send<M: Machine>(name: GenStatemName<M>, ev: M::Ev) -> Result<(), SendError> {
|
||||
match whereis_machine(name) {
|
||||
Some(m) => m.send(ev),
|
||||
None => Err(SendError::Down),
|
||||
}
|
||||
}
|
||||
|
||||
/// Synchronous request-reply to the machine registered under `name`,
|
||||
/// resolving per call (a machine restarted under the same name is reached
|
||||
/// transparently). [`CallError::Down`] if no live machine holds the name.
|
||||
pub fn call<M, T, F>(name: GenStatemName<M>, make: F) -> Result<T, CallError>
|
||||
where
|
||||
M: Machine,
|
||||
T: Send + 'static,
|
||||
F: FnOnce(Reply<T>) -> M::Ev,
|
||||
{
|
||||
match whereis_machine(name) {
|
||||
Some(m) => m.call(make),
|
||||
None => Err(CallError::Down),
|
||||
}
|
||||
}
|
||||
|
||||
/// Shut down the machine registered under `name` and wait for it (see
|
||||
/// [`GenStatemRef::shutdown`]). A no-op if no live machine holds the name.
|
||||
pub fn shutdown<M: Machine>(name: GenStatemName<M>) {
|
||||
if let Some(m) = whereis_machine(name) {
|
||||
m.shutdown();
|
||||
}
|
||||
}
|
||||
|
||||
/// The machine actor body: `on_start`, then one `handle` per event until the
|
||||
/// inbox closes (all refs dropped → graceful shutdown).
|
||||
/// row resolves to `stop`, a shutdown row stops it, or the actor is stopped
|
||||
/// from outside.
|
||||
///
|
||||
/// Two intake sources are selected each iteration with the **timer arm above
|
||||
/// the inbox**, so a timeout fire is never starved by inbox traffic: `sys_rx`
|
||||
/// carries timer fires armed through `cx`, `rx` is the user inbox. A fire is
|
||||
/// turned into the matching internal event (`state_timeout` / `timeout(name)`)
|
||||
/// and run through the same `handle` dispatch as an inbox event — the
|
||||
/// gen_statem model, where timeouts surface as ordinary events.
|
||||
/// Intake arms are selected each iteration in priority order — **exits**
|
||||
/// (only when trapping) above **timers** above the **inbox** — so a shutdown
|
||||
/// request or a timeout fire is never starved by inbox traffic. `sys_rx`
|
||||
/// carries timer fires armed through `cx`; a fire is turned into the matching
|
||||
/// internal event (`state_timeout` / `timeout(name)`) and run through the same
|
||||
/// `handle` dispatch as an inbox event — the gen_statem model, where timeouts
|
||||
/// (and, when trapping, shutdown and exits) surface as ordinary events.
|
||||
///
|
||||
/// The loop owns the **postpone queue**: a `handle` that defers its event hands
|
||||
/// it back ([`Step::Postponed`]) for the queue; a `handle` that transitions
|
||||
/// ([`Step::Transitioned`]) triggers a [`replay`] of the queue in the new
|
||||
/// state, ahead of the next intake.
|
||||
fn statem_loop<M: Machine>(rx: Receiver<M::Ev>, mut machine: M) {
|
||||
fn statem_loop<M: Machine>(keep: Sender<M::Ev>, rx: Receiver<M::Ev>, machine: M) {
|
||||
// One inbox sender lives with the loop: the inbox never closes, refs are
|
||||
// addresses, the machine's lifetime is the actor's (stop row, shutdown,
|
||||
// stop, panic). The `Disconnected` inbox arm below is defensive only.
|
||||
let _keep = keep;
|
||||
// Drop guard — owns the machine and the timer registry, so `terminate`
|
||||
// fires on every exit path (clean close, `stop`, a handler panic, a hard
|
||||
// stop) and the timer drain is sequenced before it. Same shape as
|
||||
// gen_server's guard; see the rationale there.
|
||||
struct Terminate<M: Machine>(M, Arc<Mutex<Timers>>);
|
||||
impl<M: Machine> Drop for Terminate<M> {
|
||||
fn drop(&mut self) {
|
||||
{
|
||||
let mut reg = match self.1.lock() {
|
||||
Ok(g) => g,
|
||||
Err(e) => panic!("smarm: gen_statem reg lock poisoned (core corrupt): {e}"),
|
||||
};
|
||||
if let Some((_, sub)) = reg.state.take() {
|
||||
cancel_timer(sub);
|
||||
}
|
||||
for (_, (_, sub)) in reg.named.drain() {
|
||||
cancel_timer(sub);
|
||||
}
|
||||
}
|
||||
self.0.terminate();
|
||||
}
|
||||
}
|
||||
|
||||
let (sys_tx, sys_rx) = channel::<Sys>();
|
||||
let reg = Arc::new(Mutex::new(Timers::new()));
|
||||
let mut guard = Terminate(machine, reg.clone());
|
||||
// The loop owns `cx` (and through it a `sys_tx` clone) for its whole life,
|
||||
// so the sys arm never closes from under us — no auto-close dance needed.
|
||||
let mut cx = Cx::new(sys_tx, reg.clone());
|
||||
// Events deferred by `postpone` rows, replayed FIFO on the next transition.
|
||||
let mut postpone: VecDeque<M::Ev> = VecDeque::new();
|
||||
machine.on_start(&mut cx);
|
||||
guard.0.on_start(&mut cx);
|
||||
// Trapping is opted into during on_start and fixed for the loop's life.
|
||||
// The inbox is armed only when set: an untrapped machine keeps the
|
||||
// two-arm select, and a shutdown request simply stops it as
|
||||
// `request_stop` would.
|
||||
let exits: Option<Receiver<ExitSignal>> = cx.trap.then(crate::link::trap_exit);
|
||||
loop {
|
||||
// Timer arm first: a ready fire is taken in preference to the inbox.
|
||||
let i = select(&[&sys_rx, &rx]);
|
||||
if i == 0 {
|
||||
// Arm order encodes priority: exits → timers → inbox.
|
||||
let ne = exits.is_some() as usize;
|
||||
let i = {
|
||||
let mut arms: Vec<&dyn Selectable> = Vec::with_capacity(3);
|
||||
if let Some(e) = &exits {
|
||||
arms.push(e);
|
||||
}
|
||||
arms.push(&sys_rx);
|
||||
arms.push(&rx);
|
||||
select(&arms)
|
||||
};
|
||||
if i < ne {
|
||||
// Exit arm: a shutdown request or a linked peer's death. The trap
|
||||
// inbox lives for the loop's life, so it never closes.
|
||||
let sig = exits.as_ref().and_then(|e| e.try_recv().ok().flatten());
|
||||
match sig {
|
||||
Some(sig) if sig.reason == DownReason::Shutdown => match M::shutdown_ev() {
|
||||
Some(ev) => dispatch(&mut guard.0, &mut cx, &mut postpone, ev),
|
||||
None => cx.stop(),
|
||||
},
|
||||
Some(sig) => {
|
||||
if let Some(ev) = M::exit_ev(sig) {
|
||||
dispatch(&mut guard.0, &mut cx, &mut postpone, ev);
|
||||
}
|
||||
}
|
||||
None => {}
|
||||
}
|
||||
} else if i == ne {
|
||||
match sys_rx.try_recv() {
|
||||
Ok(Some(fire)) => {
|
||||
// Confirm the fire is still the live one before dispatching:
|
||||
@@ -475,7 +740,9 @@ fn statem_loop<M: Machine>(rx: Receiver<M::Ev>, mut machine: M) {
|
||||
Sys::StateTimeout(local) => {
|
||||
let mut t = match reg.lock() {
|
||||
Ok(g) => g,
|
||||
Err(e) => panic!("smarm: gen_statem reg lock poisoned (core corrupt): {e}"),
|
||||
Err(e) => panic!(
|
||||
"smarm: gen_statem reg lock poisoned (core corrupt): {e}"
|
||||
),
|
||||
};
|
||||
match t.state {
|
||||
Some((live, _)) if live == local => {
|
||||
@@ -488,7 +755,9 @@ fn statem_loop<M: Machine>(rx: Receiver<M::Ev>, mut machine: M) {
|
||||
Sys::Timeout(name, local) => {
|
||||
let mut t = match reg.lock() {
|
||||
Ok(g) => g,
|
||||
Err(e) => panic!("smarm: gen_statem reg lock poisoned (core corrupt): {e}"),
|
||||
Err(e) => panic!(
|
||||
"smarm: gen_statem reg lock poisoned (core corrupt): {e}"
|
||||
),
|
||||
};
|
||||
match t.named.get(name) {
|
||||
Some(&(live, _)) if live == local => {
|
||||
@@ -500,7 +769,7 @@ fn statem_loop<M: Machine>(rx: Receiver<M::Ev>, mut machine: M) {
|
||||
}
|
||||
};
|
||||
if let Some(ev) = ev {
|
||||
dispatch(&mut machine, &mut cx, &mut postpone, ev);
|
||||
dispatch(&mut guard.0, &mut cx, &mut postpone, ev);
|
||||
}
|
||||
}
|
||||
// Single-receiver: nothing can drain the arm between select's
|
||||
@@ -512,12 +781,16 @@ fn statem_loop<M: Machine>(rx: Receiver<M::Ev>, mut machine: M) {
|
||||
}
|
||||
} else {
|
||||
match rx.try_recv() {
|
||||
Ok(Some(ev)) => dispatch(&mut machine, &mut cx, &mut postpone, ev),
|
||||
Ok(Some(ev)) => dispatch(&mut guard.0, &mut cx, &mut postpone, ev),
|
||||
Ok(None) => debug_assert!(false, "ready inbox was empty"),
|
||||
// All GenStatemRefs dropped → inbox closed → shutdown.
|
||||
// Defensive: the loop holds a sender, so unreachable.
|
||||
Err(_) => break,
|
||||
}
|
||||
}
|
||||
// A handler (or a replay) asked to stop: a normal exit.
|
||||
if cx.stop {
|
||||
break;
|
||||
}
|
||||
// Observation point so a machine fed a hot inbox stays preemptible and
|
||||
// cancellable.
|
||||
crate::check!();
|
||||
@@ -526,7 +799,8 @@ fn statem_loop<M: Machine>(rx: Receiver<M::Ev>, mut machine: M) {
|
||||
|
||||
/// Run one event through `handle` and act on its [`Step`]: stash a deferred
|
||||
/// event on the postpone queue, or — on a real transition — [`replay`] the
|
||||
/// queue in the new state. A stay/unmatched event needs nothing further.
|
||||
/// queue in the new state. A stay/unmatched event needs nothing further. A
|
||||
/// [`Cx::stop`] raised by the handler skips the replay; the loop breaks next.
|
||||
fn dispatch<M: Machine>(
|
||||
machine: &mut M,
|
||||
cx: &mut Cx<M::Ev>,
|
||||
@@ -536,14 +810,19 @@ fn dispatch<M: Machine>(
|
||||
match machine.handle(ev, cx) {
|
||||
Step::Postponed(ev) => postpone.push_back(ev),
|
||||
Step::Stayed => {}
|
||||
Step::Transitioned => replay(machine, cx, postpone),
|
||||
Step::Transitioned => {
|
||||
if !cx.stop {
|
||||
replay(machine, cx, postpone)
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Replay deferred events after a real transition: each goes back through
|
||||
/// `handle` in FIFO order, in the now-current state. An event that postpones
|
||||
/// again re-queues (to wait for the *next* transition); one that transitions
|
||||
/// re-arms the replay, so a later state can in turn drain what is still pending.
|
||||
/// re-arms the replay, so a later state can in turn drain what is still pending;
|
||||
/// one that raises [`Cx::stop`] ends the replay (and the machine) at once.
|
||||
/// Subsequent events in a batch already see the post-transition state, since
|
||||
/// `handle` reads the live state cell — the outer loop only re-runs to give
|
||||
/// re-queued events another pass once a transition has occurred within a batch.
|
||||
@@ -562,6 +841,9 @@ fn replay<M: Machine>(machine: &mut M, cx: &mut Cx<M::Ev>, postpone: &mut VecDeq
|
||||
Step::Stayed => {}
|
||||
Step::Transitioned => transitioned = true,
|
||||
}
|
||||
if cx.stop {
|
||||
return;
|
||||
}
|
||||
}
|
||||
if !transitioned {
|
||||
return;
|
||||
@@ -663,8 +945,9 @@ fn replay<M: Machine>(machine: &mut M, cx: &mut Cx<M::Ev>, postpone: &mut VecDeq
|
||||
/// // The transition table. Group rows by current state with `on <pat>`.
|
||||
/// // A row is: <kind> <event-pattern> [if <guard>] => <tail> ,
|
||||
/// // where <kind> is one of `cast`, `call`, `info`, `state_timeout`
|
||||
/// // (no pattern — it is a unit event), or `timeout <name-pattern>`,
|
||||
/// // and the tail is one of:
|
||||
/// // (no pattern — it is a unit event), `timeout <name-pattern>`,
|
||||
/// // `shutdown` (unit; trapping machines only) or `exit <sig-pattern>`
|
||||
/// // (trapping only), and the tail is one of:
|
||||
/// // * a state tag `Door::Closed` (transition, or "stay"
|
||||
/// // if it equals current)
|
||||
/// // * a block ending in one `{ data.enters += 1; Door::Closed }`
|
||||
@@ -673,6 +956,9 @@ fn replay<M: Machine>(machine: &mut M, cx: &mut Cx<M::Ev>, postpone: &mut VecDeq
|
||||
/// // * the keyword `postpone` (defer until next
|
||||
/// // transition; cast/call/
|
||||
/// // info only)
|
||||
/// // * the keyword `stop` (end the machine
|
||||
/// // normally; sugar for
|
||||
/// // `{ cx.stop(); prev }`)
|
||||
/// on Door::Open => {
|
||||
/// cast DoorCast::Push => Door::Closed,
|
||||
/// // An armed state-timeout surfaces as an ordinary event:
|
||||
@@ -696,6 +982,9 @@ fn replay<M: Machine>(machine: &mut M, cx: &mut Cx<M::Ev>, postpone: &mut VecDeq
|
||||
/// on _ => {
|
||||
/// call DoorCall::GetState(r) => { r.reply(prev); prev },
|
||||
/// }
|
||||
///
|
||||
/// // Optional: runs as the machine exits, on every exit path.
|
||||
/// terminate { data.enters = 0; }
|
||||
/// }
|
||||
/// ```
|
||||
///
|
||||
@@ -720,6 +1009,12 @@ fn replay<M: Machine>(machine: &mut M, cx: &mut Cx<M::Ev>, postpone: &mut VecDeq
|
||||
/// ignores info writes no `info` rows at all. State-timeouts and named
|
||||
/// timeouts have **no** such default: a state that can see one must handle it
|
||||
/// (or `unhandled` it) or the match is non-exhaustive.
|
||||
/// * **`shutdown` defaults to `stop`, `exit` to a silent drop.** Both reach a
|
||||
/// machine only if its initial `enter` called `cx.trap_exit()`; a
|
||||
/// non-trapping machine is stopped outright by a shutdown request. Write
|
||||
/// `shutdown => …` rows only in the states that want to wind down first
|
||||
/// (transition into a draining state, `stop` later); write `exit sig => …`
|
||||
/// rows to react to linked-peer deaths. Neither is postponable.
|
||||
/// * **Stay** = return the current tag. The `prev` you named in `context` is
|
||||
/// bound to the pre-handler state for exactly this — handy in any-state
|
||||
/// (`on _`) rows where there is no single literal tag to write.
|
||||
@@ -745,12 +1040,12 @@ fn replay<M: Machine>(machine: &mut M, cx: &mut Cx<M::Ev>, postpone: &mut VecDeq
|
||||
/// # What it emits
|
||||
///
|
||||
/// The unified `enum $Ev` (the `Cast`/`Call`/`Info` wrappers plus the internal
|
||||
/// `StateTimeout` / `Timeout` events), `struct $Sm { state, data }`,
|
||||
/// `$Sm::start(init, data) -> GenStatemRef<$Sm>`, and the `Machine` impl:
|
||||
/// `on_start` runs the initial `enter`; `handle` is the dispatch match plus the
|
||||
/// stay/transition/unhandled apply-tail (the cell's sole writer, which also
|
||||
/// auto-resets the state-timeout on every real transition); and the `enter`
|
||||
/// dispatch.
|
||||
/// `StateTimeout` / `Timeout` / `Shutdown` / `Exit` events), `struct $Sm {
|
||||
/// state, data }`, `$Sm::start(init, data) -> GenStatemRef<$Sm>`, and the
|
||||
/// `Machine` impl: `on_start` runs the initial `enter`; `handle` is the
|
||||
/// dispatch match plus the stay/transition/unhandled apply-tail (the cell's
|
||||
/// sole writer, which also auto-resets the state-timeout on every real
|
||||
/// transition); the `enter` dispatch; and `terminate` when the block is given.
|
||||
///
|
||||
/// # Limitation
|
||||
///
|
||||
@@ -766,9 +1061,10 @@ macro_rules! gen_statem {
|
||||
context ( $data:ident , $cur:ident , $cx:ident ) ;
|
||||
enter { $( $est:pat => $ebody:expr ),+ $(,)? }
|
||||
$( on $st:pat => { $($rows:tt)* } )+
|
||||
$( terminate { $($tbody:tt)* } )?
|
||||
) => {
|
||||
/// Unified inbox payload: the user's `cast`/`call`/`info` enums folded
|
||||
/// together with the runtime's internal timeout events.
|
||||
/// together with the runtime's internal events.
|
||||
enum $Ev {
|
||||
Cast($Cast),
|
||||
Call($Call),
|
||||
@@ -779,6 +1075,12 @@ macro_rules! gen_statem {
|
||||
StateTimeout,
|
||||
/// A named timeout fired (matched in `timeout <pat>` rows).
|
||||
Timeout(&'static str),
|
||||
/// A graceful shutdown request reached this (trapping) machine
|
||||
/// (matched in `shutdown` rows). A state with no such row `stop`s.
|
||||
Shutdown,
|
||||
/// A linked peer died (trapping only; matched in `exit <pat>`
|
||||
/// rows). An unmatched exit is silently dropped.
|
||||
Exit($crate::ExitSignal),
|
||||
}
|
||||
|
||||
struct $sm {
|
||||
@@ -787,8 +1089,14 @@ macro_rules! gen_statem {
|
||||
}
|
||||
|
||||
impl $sm {
|
||||
/// The machine value, for [`gen_statem::run_named`]
|
||||
/// (`$crate::gen_statem::run_named`) or `spawn`.
|
||||
fn new(init: $State, data: $Data) -> $sm {
|
||||
$sm { state: init, data }
|
||||
}
|
||||
|
||||
fn start(init: $State, data: $Data) -> $crate::gen_statem::GenStatemRef<$sm> {
|
||||
$crate::gen_statem::spawn($sm { state: init, data })
|
||||
$crate::gen_statem::spawn($sm::new(init, data))
|
||||
}
|
||||
|
||||
#[allow(unused_variables)]
|
||||
@@ -812,11 +1120,27 @@ macro_rules! gen_statem {
|
||||
$Ev::Timeout(name)
|
||||
}
|
||||
|
||||
fn shutdown_ev() -> Option<$Ev> {
|
||||
Some($Ev::Shutdown)
|
||||
}
|
||||
|
||||
fn exit_ev(sig: $crate::ExitSignal) -> Option<$Ev> {
|
||||
Some($Ev::Exit(sig))
|
||||
}
|
||||
|
||||
fn on_start(&mut self, $cx: &mut $crate::gen_statem::Cx<$Ev>) {
|
||||
let s = self.state;
|
||||
self.enter(s, $cx);
|
||||
}
|
||||
|
||||
$(
|
||||
#[allow(unused_variables)]
|
||||
fn terminate(&mut self) {
|
||||
let $data = &mut self.data;
|
||||
$($tbody)*
|
||||
}
|
||||
)?
|
||||
|
||||
#[allow(unused_variables)]
|
||||
#[deny(unreachable_patterns)] // conflicting rows must fail even though
|
||||
// this match is external-macro-expanded
|
||||
@@ -832,7 +1156,7 @@ macro_rules! gen_statem {
|
||||
// untouched), then the consuming `match (state, event)` whose
|
||||
// value is this `Resolution`.
|
||||
let next: $crate::gen_statem::Resolution<$State> =
|
||||
$crate::gen_statem!(@arms ($Ev) ($cur, ev) [ ] [ ]
|
||||
$crate::gen_statem!(@arms ($Ev) ($cur, ev, $cx) [ ] [ ]
|
||||
$( on $st => { $($rows)* } )+);
|
||||
match next {
|
||||
$crate::gen_statem::Resolution::To(s) if s == $cur => {
|
||||
@@ -868,7 +1192,7 @@ macro_rules! gen_statem {
|
||||
// is the `Info` silent-drop — last and broadest, so per-state `info` rows
|
||||
// stay reachable; cast/call/timeouts get no fallback, so a forgotten pair is
|
||||
// still E0004).
|
||||
(@arms ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ]) => {
|
||||
(@arms ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ]) => {
|
||||
{
|
||||
// Phase 1 — postpone routing (borrow-only). A guard on a postpone
|
||||
// row runs here, against by-ref bindings, so it must not depend on
|
||||
@@ -885,140 +1209,247 @@ macro_rules! gen_statem {
|
||||
// returned for it.
|
||||
match ($ss, $se) {
|
||||
$($arms)*
|
||||
// Macro-injected defaults, last and broadest so per-state rows
|
||||
// stay reachable; a user's own catch-all row may shadow them
|
||||
// entirely, hence the allow.
|
||||
#[allow(unreachable_patterns)]
|
||||
(_, $Ev::Info(_)) => $crate::gen_statem::Resolution::Unhandled,
|
||||
#[allow(unreachable_patterns)]
|
||||
(_, $Ev::Exit(_)) => $crate::gen_statem::Resolution::Unhandled,
|
||||
#[allow(unreachable_patterns)]
|
||||
(_, $Ev::Shutdown) => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) },
|
||||
}
|
||||
}
|
||||
};
|
||||
// Open an on-block: remember its state pat, drain its rows, then continue.
|
||||
(@arms ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ]
|
||||
(@arms ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ]
|
||||
on $st:pat => { $($rows:tt)* } $($more:tt)*
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se) [ $($arms)* ] [ $($post)* ] ($st)
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx) [ $($arms)* ] [ $($post)* ] ($st)
|
||||
{ $($rows)* } { $($more)* })
|
||||
};
|
||||
|
||||
// ===== @rows: drain one on-block's rows, threading both accs =============
|
||||
// cast, explicit refusal
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ cast $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se)
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Cast($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// cast, stop (end the machine normally after this event)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ cast $ev:pat $(if $g:expr)? => stop , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Cast($ev)) $(if $g)? => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) }, ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// cast, postpone (defer the event; the replay in a later state handles it)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ cast $ev:pat $(if $g:expr)? => postpone , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se)
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Cast($ev)) $(if $g)? => unreachable!("postponed event is replayed, not dispatched here"), ]
|
||||
[ $($post)* ($st, $Ev::Cast($ev)) $(if $g)? => true, ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// cast, transition / stay / branch
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ cast $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se)
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Cast($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// call, explicit refusal
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ call $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se)
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Call($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// call, stop (end the machine normally after this event)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ call $ev:pat $(if $g:expr)? => stop , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Call($ev)) $(if $g)? => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) }, ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// call, postpone (the Reply rides inside the event onto the queue)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ call $ev:pat $(if $g:expr)? => postpone , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se)
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Call($ev)) $(if $g)? => unreachable!("postponed event is replayed, not dispatched here"), ]
|
||||
[ $($post)* ($st, $Ev::Call($ev)) $(if $g)? => true, ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// call, transition / stay / branch
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ call $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se)
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Call($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// info, explicit refusal
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ info $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se)
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Info($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// info, stop (end the machine normally after this event)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ info $ev:pat $(if $g:expr)? => stop , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Info($ev)) $(if $g)? => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) }, ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// info, postpone
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ info $ev:pat $(if $g:expr)? => postpone , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se)
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Info($ev)) $(if $g)? => unreachable!("postponed event is replayed, not dispatched here"), ]
|
||||
[ $($post)* ($st, $Ev::Info($ev)) $(if $g)? => true, ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// info, transition / stay / branch
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ info $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se)
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Info($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// state_timeout, explicit refusal (unit event — no pattern; not postponable)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ state_timeout $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se)
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::StateTimeout) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// state_timeout, stop (end the machine normally after this event)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ state_timeout $(if $g:expr)? => stop , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::StateTimeout) $(if $g)? => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) }, ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// state_timeout, transition / stay / branch
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ state_timeout $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se)
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::StateTimeout) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// timeout, explicit refusal (pattern matches the name; not postponable)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ timeout $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se)
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Timeout($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// timeout, stop (end the machine normally after this event)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ timeout $ev:pat $(if $g:expr)? => stop , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Timeout($ev)) $(if $g)? => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) }, ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// timeout, transition / stay / branch
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ timeout $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se)
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Timeout($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// shutdown, explicit refusal (unit event — no pattern; not postponable; a state with no shutdown row defaults to `stop`)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ shutdown $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Shutdown) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// shutdown, stop (end the machine normally after this event)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ shutdown $(if $g:expr)? => stop , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Shutdown) $(if $g)? => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) }, ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// shutdown, transition / stay / branch
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ shutdown $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Shutdown) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// exit, explicit refusal (pattern matches the ExitSignal; not postponable)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ exit $ev:pat $(if $g:expr)? => unhandled , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Exit($ev)) $(if $g)? => $crate::gen_statem::Resolution::Unhandled, ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// exit, stop (end the machine normally after this event)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ exit $ev:pat $(if $g:expr)? => stop , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Exit($ev)) $(if $g)? => { $cx.stop(); $crate::gen_statem::Resolution::To($ss) }, ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// exit, transition / stay / branch
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ exit $ev:pat $(if $g:expr)? => $tail:expr , $($rows:tt)* } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@rows ($Ev) ($ss, $se, $cx)
|
||||
[ $($arms)* ($st, $Ev::Exit($ev)) $(if $g)? => $crate::gen_statem::Resolution::To($tail.into()), ]
|
||||
[ $($post)* ]
|
||||
($st) { $($rows)* } { $($more)* })
|
||||
};
|
||||
// this block is drained: hand the remaining on-blocks back to @arms
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
(@rows ($Ev:ident) ($ss:expr, $se:expr, $cx:ident) [ $($arms:tt)* ] [ $($post:tt)* ] ($st:pat)
|
||||
{ } { $($more:tt)* }
|
||||
) => {
|
||||
$crate::gen_statem!(@arms ($Ev) ($ss, $se) [ $($arms)* ] [ $($post)* ] $($more)*)
|
||||
$crate::gen_statem!(@arms ($Ev) ($ss, $se, $cx) [ $($arms)* ] [ $($post)* ] $($more)*)
|
||||
};
|
||||
}
|
||||
|
||||
+251
-87
@@ -1,26 +1,85 @@
|
||||
//! RFC 016 — runtime introspection (Chunk 1: the read primitive).
|
||||
//! Inspect what is running right now: which actors exist, what state each one
|
||||
//! is in, and how they are related.
|
||||
//!
|
||||
//! A synchronous, internal read of the slab that returns *owned* data. This is
|
||||
//! the mechanism the whole RFC hangs off: tests, the future observer
|
||||
//! gen_server (Chunk 4), and a later control plane (RFC 003) are all consumers
|
||||
//! of [`snapshot`] / [`actor_info`], never of the runtime internals directly.
|
||||
//! This is the tool for questions like "is my server still alive", "how many
|
||||
//! actors are currently parked waiting on something", or "what does the spawn
|
||||
//! tree look like". It is meant for debugging, test assertions, a health check
|
||||
//! endpoint, or a monitoring dashboard: anywhere you want to look at the
|
||||
//! runtime from the outside without stopping it or coupling your code to its
|
||||
//! internals.
|
||||
//!
|
||||
//! ## Consistency (DECISION D2 — per-slot tearing, `ps` semantics)
|
||||
//! Three entry points, in order of scope:
|
||||
//!
|
||||
//! [`snapshot`] is point-in-time and mildly racy *across* actors: each slot's
|
||||
//! scheduling state is a lock-free word load, so an actor reported `Running`
|
||||
//! may already be `Parked`, and an actor can die mid-scan. This is the cheap,
|
||||
//! useful model (a coherent stop-the-world cut is expensive and rarely wanted).
|
||||
//! [`actor_info`] is coherent for the single actor it names.
|
||||
//! - [`snapshot`] returns every actor that currently exists, as a plain
|
||||
//! owned `Vec`, so you can filter, count, or search it however you like.
|
||||
//! - [`actor_info`] returns a coherent view of exactly one actor, by pid.
|
||||
//! Cheaper than filtering a whole snapshot down to one entry, and more
|
||||
//! precise (see "Consistency" below).
|
||||
//! - [`tree`] returns the same actors as [`snapshot`], folded into a
|
||||
//! parent/child forest that mirrors who spawned whom.
|
||||
//!
|
||||
//! ## Locking
|
||||
//! ```
|
||||
//! use smarm::{actor_info, channel, run, snapshot, spawn, ActorState};
|
||||
//!
|
||||
//! The lock order is **Leaf → Channel, at most one of each** (`raw_mutex.rs`);
|
||||
//! cold locks, the registry, and the free list are all Leaves, so we may never
|
||||
//! hold two at once. The read is therefore phased: first a single registry-leaf
|
||||
//! pass for names and mailbox depth (the per-channel length read is a Channel
|
||||
//! lock taken under that Leaf — legal), released before the slab scan takes any
|
||||
//! per-slot cold Leaf.
|
||||
//! run(|| {
|
||||
//! let (ready_tx, ready_rx) = channel::<()>();
|
||||
//! let (gate_tx, gate_rx) = channel::<()>();
|
||||
//!
|
||||
//! let worker = spawn(move || {
|
||||
//! ready_tx.send(()).unwrap();
|
||||
//! gate_rx.recv().unwrap(); // blocks here until released
|
||||
//! });
|
||||
//! ready_rx.recv().unwrap();
|
||||
//!
|
||||
//! // `snapshot` sees every actor, including this one and the worker.
|
||||
//! let snap = snapshot();
|
||||
//! assert!(snap.actors.len() >= 2);
|
||||
//!
|
||||
//! // `actor_info` gives a coherent view of just the worker. It is
|
||||
//! // blocked on the gate channel, so it must be Parked.
|
||||
//! let pid = worker.pid();
|
||||
//! let info = actor_info(pid).expect("worker is still alive");
|
||||
//! assert_eq!(info.state, ActorState::Parked);
|
||||
//!
|
||||
//! gate_tx.send(()).unwrap();
|
||||
//! worker.join().unwrap();
|
||||
//!
|
||||
//! // Once joined, the pid no longer names a live actor.
|
||||
//! assert!(actor_info(pid).is_none());
|
||||
//! });
|
||||
//! ```
|
||||
//!
|
||||
//! ## Consistency
|
||||
//!
|
||||
//! [`snapshot`] is not a single atomic pause-the-world freeze: it walks every
|
||||
//! actor's state one after another, so it is a series of independent,
|
||||
//! cheap, lock-free reads rather than one coherent moment in time. Between
|
||||
//! reading actor A and actor B, either one can change state, and an actor can
|
||||
//! even finish and disappear mid-scan. In practice this is exactly what you
|
||||
//! want: a coherent stop-the-world snapshot would mean pausing every actor in
|
||||
//! the runtime just to look at it, which is expensive and rarely necessary
|
||||
//! for a dashboard, a test assertion, or a debugging session.
|
||||
//!
|
||||
//! [`actor_info`], in contrast, is coherent for the one actor it names: all of
|
||||
//! its fields describe the same instant for that actor, because a single
|
||||
//! actor's data cannot tear the way a scan across many actors can.
|
||||
//!
|
||||
//! ## Implementation notes
|
||||
//!
|
||||
//! These details matter if you are working on smarm itself; they are not part
|
||||
//! of the public contract.
|
||||
//!
|
||||
//! The read never stops the scheduler and never holds a lock across the whole
|
||||
//! scan. Each actor's scheduling state is a single lock-free word load
|
||||
//! (hence the possible tearing described above). Reading the rest of an
|
||||
//! actor's cold data (its supervisor, monitors, links, and so on) takes a
|
||||
//! brief per-actor lock, just long enough to copy those fields out; nothing
|
||||
//! is held across actors. Locking follows the crate-wide rule that at most
|
||||
//! one "leaf" lock (a per-actor lock, the registry lock, or the free list
|
||||
//! lock) is held at a time, with no leaf lock held while acquiring another.
|
||||
//! The read is phased accordingly: first one pass over the registry to
|
||||
//! collect every actor's registered names and mailbox depth, released before
|
||||
//! the per-actor scan begins.
|
||||
|
||||
use crate::pid::Pid;
|
||||
use crate::registry::MailboxInfo;
|
||||
@@ -31,15 +90,28 @@ use crate::slot_state::{
|
||||
};
|
||||
use std::collections::HashMap;
|
||||
|
||||
/// Snapshot wire-format version (DECISION D1). [`RuntimeSnapshot`] is treated as
|
||||
/// a stable type from day one: it becomes the observer protocol (Chunk 4) and
|
||||
/// crosses a version boundary the moment a remote observer attaches (RFC 011),
|
||||
/// so the version travels with the data from the start.
|
||||
/// The format version carried by every [`RuntimeSnapshot`] and
|
||||
/// [`RuntimeTree`], as [`RuntimeSnapshot::format_version`] /
|
||||
/// [`RuntimeTree::format_version`]. If you serialize a snapshot (for example
|
||||
/// to send it somewhere else, or to compare snapshots taken with different
|
||||
/// versions of smarm) check this field: a change in its value means the shape
|
||||
/// of [`ActorInfo`] or its neighbors has changed and old and new snapshots
|
||||
/// should not be assumed compatible. If you only ever read a snapshot
|
||||
/// in-process in the same version of smarm that produced it, you can ignore
|
||||
/// this field.
|
||||
pub const SNAPSHOT_FORMAT_VERSION: u16 = 1;
|
||||
|
||||
/// Fine-grained scheduling state, mapped from the packed slot word with no new
|
||||
/// storage. `RunningNotified` collapses into `Notified` — a wake landed while
|
||||
/// the actor was on-CPU and it will re-queue when it yields.
|
||||
/// What an actor is doing right now, from the scheduler's point of view.
|
||||
///
|
||||
/// - `Queued`: runnable, waiting for a scheduler thread to pick it up.
|
||||
/// - `Running`: currently executing on a scheduler thread.
|
||||
/// - `Notified`: was running and got woken up (for example, a message
|
||||
/// arrived) before it had a chance to yield or park; it will be re-queued
|
||||
/// as soon as it does.
|
||||
/// - `Parked`: blocked, waiting on something such as a channel receive, a
|
||||
/// mutex, a timer, or an IO event.
|
||||
/// - `Done`: has finished (returned or panicked) but its slot has not been
|
||||
/// reclaimed for reuse yet, so it is still visible to introspection.
|
||||
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
|
||||
pub enum ActorState {
|
||||
Queued,
|
||||
@@ -49,8 +121,8 @@ pub enum ActorState {
|
||||
Done,
|
||||
}
|
||||
|
||||
/// Classify a packed state word. `None` for a Vacant slot (skipped by the scan)
|
||||
/// — the only state that is not an actor.
|
||||
/// Classify a packed state word. `None` for a Vacant slot (skipped by the
|
||||
/// scan): a vacant slot holds no actor at all, live or done.
|
||||
fn classify(w: u64) -> Option<ActorState> {
|
||||
Some(match word_state(w) {
|
||||
ST_QUEUED => ActorState::Queued,
|
||||
@@ -62,61 +134,104 @@ fn classify(w: u64) -> Option<ActorState> {
|
||||
})
|
||||
}
|
||||
|
||||
/// Owned, point-in-time view of one actor — no borrows of runtime internals, so
|
||||
/// it is safe to hand to any consumer.
|
||||
/// An owned, self-contained view of one actor at (approximately) one moment.
|
||||
/// It borrows nothing from the runtime, so you can keep it, send it
|
||||
/// elsewhere, or print it long after the actor it describes has changed
|
||||
/// state or even exited.
|
||||
#[derive(Debug, Clone)]
|
||||
pub struct ActorInfo {
|
||||
pub pid: Pid,
|
||||
/// Registered names, inverted from the registry (usually 0 or 1).
|
||||
/// Names this actor is currently registered under (see the
|
||||
/// [`registry`](crate::registry) module). Usually empty or one name;
|
||||
/// an actor can have more if it registered several.
|
||||
pub names: Vec<&'static str>,
|
||||
pub state: ActorState,
|
||||
/// Spawn-time parent edge (DECISION D9): `spawn_under` sets it to the
|
||||
/// supervisor, plain `spawn` to the spawning actor — so it is parentage,
|
||||
/// not necessarily a supervision relationship. `ROOT_PID` for the run's
|
||||
/// root actor and for `Done` tombstones (whose `Actor` is already gone).
|
||||
/// The actor that spawned this one: whoever called `spawn` or
|
||||
/// `spawn_under` to create it. This is a parentage record, not
|
||||
/// necessarily a supervision relationship: `spawn_under` records the
|
||||
/// supervisor you asked for, while plain `spawn` records the spawning
|
||||
/// actor itself, whether or not it supervises anything. It is the
|
||||
/// runtime's root pid for the run's own root actor, and for a `Done`
|
||||
/// actor whose bookkeeping has already been cleared.
|
||||
pub supervisor: Pid,
|
||||
pub trap_exit: bool,
|
||||
pub monitors: u32,
|
||||
pub links: u32,
|
||||
pub joiners: u32,
|
||||
/// Queued messages summed over the actor's *published* channels (register /
|
||||
/// install / spawn_addr / gen_server). 0 for an actor that holds only a
|
||||
/// private `channel()` receiver — those are invisible to the registry.
|
||||
/// Messages currently queued and not yet delivered, summed across every
|
||||
/// channel this actor has published (via `register`, `install`,
|
||||
/// `spawn_addr`, or starting a gen_server). This is 0 for an actor that
|
||||
/// only holds a private, unpublished `channel()` receiver, since nothing
|
||||
/// outside the actor can see that channel exists.
|
||||
pub mailbox_depth: u32,
|
||||
/// Timeslice overruns tallied for this incarnation (RFC 016 Chunk 2): how
|
||||
/// many times the actor was preempted for exceeding its slice. Resets on
|
||||
/// restart (per-incarnation, D7).
|
||||
/// How many times this actor has been preempted for running past its
|
||||
/// scheduling timeslice. Counts only since the actor's current start (a
|
||||
/// supervisor restart begins a fresh count).
|
||||
pub overruns: u64,
|
||||
/// Messages this actor has received (dequeued) this incarnation (RFC 016
|
||||
/// Chunk 2) — answers "is this actor a hotspot / draining slower than its
|
||||
/// mailbox fills." Counts received, not sent (D4). Per-incarnation (D7).
|
||||
/// How many messages this actor has received (taken off its inbox), since
|
||||
/// its current start. Useful for spotting an actor whose mailbox is
|
||||
/// filling up faster than it can drain it: compare this against
|
||||
/// `mailbox_depth` over time.
|
||||
pub messages_received: u64,
|
||||
/// Approximate on-CPU cycles this incarnation has consumed (RFC 016 Chunk 2)
|
||||
/// — a reductions-like work metric for relative comparison. Always 0 unless
|
||||
/// the `budget-accounting` feature is enabled (it costs an RDTSC per resume,
|
||||
/// D6). Per-incarnation (D7).
|
||||
/// Approximate CPU cycles this actor has spent running, since its current
|
||||
/// start. A relative measure for comparing actors against each other, not
|
||||
/// an absolute or wall-clock figure. Always 0 unless the crate's
|
||||
/// `budget-accounting` feature is enabled, since measuring it costs a
|
||||
/// timestamp read on every resume.
|
||||
pub budget_cycles: u64,
|
||||
/// RFC 019 §8 — this actor's stack, as the runtime sees it. All fields
|
||||
/// are lock-free atomic reads, coherent for this incarnation via the
|
||||
/// same generation check as the counters above. Exact RSS is
|
||||
/// deliberately absent: `mincore` is debug tooling, never a runtime
|
||||
/// path.
|
||||
pub stack: StackInfo,
|
||||
}
|
||||
|
||||
/// A whole-runtime snapshot. See the module docs for the D2 tearing model.
|
||||
/// RFC 019 §8 — per-actor stack introspection. Sizes are page-rounded, as
|
||||
/// [`Stack::new`](crate::stack::Stack::new) rounds them.
|
||||
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
|
||||
pub struct StackInfo {
|
||||
/// Usable stack size ([`SpawnOpts::stack_reserve`]
|
||||
/// (crate::SpawnOpts::stack_reserve) or the Config/default).
|
||||
pub reserve: usize,
|
||||
/// PROT_NONE guard below the usable region.
|
||||
pub guard: usize,
|
||||
/// Sampled high-water depth in bytes: `top − lowest saved sp`. Sampled,
|
||||
/// not exact — the context save at yields/parks/preemptions is the
|
||||
/// sampler (RFC 019 §2), so a spike the actor never yielded inside is
|
||||
/// invisible. 0 depth means "never descheduled at any depth", not
|
||||
/// "never ran".
|
||||
pub depth_high_water: usize,
|
||||
/// Parks on this incarnation since its last shrink (or since install if
|
||||
/// it has never shrunk) — the §3 cooldown counter, live.
|
||||
pub parks_since_shrink: u32,
|
||||
/// §3 shrinks performed on this incarnation.
|
||||
pub shrinks: u32,
|
||||
}
|
||||
|
||||
/// A snapshot of every actor in the runtime at (approximately) one moment.
|
||||
/// See the module docs' "Consistency" section for what "approximately" means
|
||||
/// here.
|
||||
#[derive(Debug, Clone)]
|
||||
pub struct RuntimeSnapshot {
|
||||
pub format_version: u16,
|
||||
pub actors: Vec<ActorInfo>,
|
||||
}
|
||||
|
||||
/// Snapshot every live (and `Done`-but-not-yet-reclaimed) actor on the slab.
|
||||
/// O(n) over the slot table, running with preemption disabled (like every
|
||||
/// runtime primitive) but holding no lock across the scan. Panics outside
|
||||
/// `Runtime::run()`; callable from actor code and the run thread.
|
||||
/// Every actor that currently exists: running, queued, parked, or finished
|
||||
/// but not yet cleaned up. Cheap and lock-free per actor; see the module
|
||||
/// docs for what "approximately one moment" means for the result as a whole.
|
||||
/// Panics if called outside [`run`](crate::run).
|
||||
pub fn snapshot() -> RuntimeSnapshot {
|
||||
with_runtime(|inner| {
|
||||
// Phase A: one registry-leaf pass for names + mailbox depth, released
|
||||
// before any cold leaf (no two Leaves at once).
|
||||
// First pass: one registry lock to collect names + mailbox depth for
|
||||
// every actor, released before touching any per-actor lock below.
|
||||
let mail = inner.registry.lock().introspect_map();
|
||||
|
||||
// Phase B: lock-free slab scan; per-slot cold leaf only to copy cold
|
||||
// fields. Tearing across slots is intentional (D2).
|
||||
// Second pass: walk the actor table. Each actor's scheduling state is
|
||||
// a lock-free word load; only copying its other fields takes a brief
|
||||
// per-actor lock. Tearing across actors is expected here (see the
|
||||
// module docs' "Consistency" section).
|
||||
let mut actors = Vec::new();
|
||||
for (idx, slot) in inner.slots.iter().enumerate() {
|
||||
let idx = idx as u32;
|
||||
@@ -124,12 +239,34 @@ pub fn snapshot() -> RuntimeSnapshot {
|
||||
actors.push(info);
|
||||
}
|
||||
}
|
||||
RuntimeSnapshot { format_version: SNAPSHOT_FORMAT_VERSION, actors }
|
||||
RuntimeSnapshot {
|
||||
format_version: SNAPSHOT_FORMAT_VERSION,
|
||||
actors,
|
||||
}
|
||||
})
|
||||
}
|
||||
|
||||
/// A coherent view of exactly one actor, or `None` if `pid` does not name a
|
||||
/// currently-live entry: it is stale (that actor has already exited and its
|
||||
/// slot was reused by another), out of range, or was never a real pid at
|
||||
/// all. Unlike [`snapshot`], every field of the result describes the same
|
||||
/// instant, since there is only one actor to read.
|
||||
/// The stack shape `(reserve, guard)` of a live actor, page-rounded — the
|
||||
/// RFC 019 introspection surface's first field (depth sampling and shrink
|
||||
/// counters land with the shrink machinery). `None` if `pid` no longer names
|
||||
/// a live actor. Takes the actor's cold lock briefly; debugging/assertion
|
||||
/// use, not a hot-path call.
|
||||
pub fn stack_shape(pid: Pid) -> Option<(usize, usize)> {
|
||||
with_runtime(|inner| {
|
||||
let slot = inner.slot_at(pid)?;
|
||||
let cold = slot.cold.lock();
|
||||
if slot.generation() != pid.generation() {
|
||||
return None;
|
||||
}
|
||||
cold.actor.as_ref().map(|a| a.stack.shape())
|
||||
})
|
||||
}
|
||||
|
||||
/// Coherent view of a single actor, or `None` if the pid is stale, out of
|
||||
/// range, or names a Vacant slot.
|
||||
pub fn actor_info(pid: Pid) -> Option<ActorInfo> {
|
||||
with_runtime(|inner| {
|
||||
let slot = inner.slot_at(pid)?;
|
||||
@@ -141,11 +278,12 @@ pub fn actor_info(pid: Pid) -> Option<ActorInfo> {
|
||||
})
|
||||
}
|
||||
|
||||
/// Build one `ActorInfo` for slot `idx`, or `None` if Vacant or
|
||||
/// racing-reclaimed. State is classified from a lock-free word load (the torn
|
||||
/// read); the cold lock then pins the generation (reclaim bumps it under that
|
||||
/// same lock) so the cold fields are coherent for this incarnation. `mail` is
|
||||
/// this slot's registry entry, if any.
|
||||
/// Build one `ActorInfo` for slot `idx`, or `None` if the slot is empty or
|
||||
/// was reclaimed while this read was in progress. The scheduling state comes
|
||||
/// from a lock-free word load (the source of the tearing described in the
|
||||
/// module docs); the per-actor lock then confirms the actor has not since
|
||||
/// exited and been replaced, so the rest of the fields are coherent for this
|
||||
/// exact actor. `mail` is this slot's registry entry, if any.
|
||||
fn read_slot(slot: &Slot, idx: u32, mail: Option<&MailboxInfo>) -> Option<ActorInfo> {
|
||||
let w = slot.state_word();
|
||||
let state = classify(w)?;
|
||||
@@ -153,10 +291,11 @@ fn read_slot(slot: &Slot, idx: u32, mail: Option<&MailboxInfo>) -> Option<ActorI
|
||||
let pid = Pid::new(idx, gen);
|
||||
|
||||
let cold = slot.cold.lock();
|
||||
// If the generation moved between the lock-free load and acquiring the cold
|
||||
// lock, the slot was reclaimed (and maybe reused) — drop it rather than mix
|
||||
// one incarnation's state with another's cold data. (ps semantics: a racing
|
||||
// actor may simply be missed mid-scan.)
|
||||
// If the generation moved between the lock-free load and acquiring the
|
||||
// per-actor lock, this actor exited (and the slot may already hold a new
|
||||
// one). Drop it rather than mix one actor's state with another's data; a
|
||||
// racing actor may simply be missed by this scan, which is expected (see
|
||||
// the module docs' "Consistency" section).
|
||||
if word_gen(slot.state_word()) != gen {
|
||||
return None;
|
||||
}
|
||||
@@ -172,7 +311,15 @@ fn read_slot(slot: &Slot, idx: u32, mail: Option<&MailboxInfo>) -> Option<ActorI
|
||||
let joiners = cold.waiters.len() as u32;
|
||||
drop(cold);
|
||||
|
||||
// Counters are hot-region atomics, read lock-free (RFC 016 Chunk 2).
|
||||
// Counters are plain atomics, read lock-free.
|
||||
let (reserve, guard, top, hwm, parks_since_shrink, shrinks) = slot.stack_introspect();
|
||||
let stack = StackInfo {
|
||||
reserve,
|
||||
guard,
|
||||
depth_high_water: top.saturating_sub(hwm),
|
||||
parks_since_shrink,
|
||||
shrinks,
|
||||
};
|
||||
let overruns = slot.overruns();
|
||||
let messages_received = slot.messages_received();
|
||||
let budget_cycles = slot.budget_cycles();
|
||||
@@ -198,45 +345,55 @@ fn read_slot(slot: &Slot, idx: u32, mail: Option<&MailboxInfo>) -> Option<ActorI
|
||||
overruns,
|
||||
messages_received,
|
||||
budget_cycles,
|
||||
stack,
|
||||
})
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Chunk 3 — tree view (pure derivation over a Chunk-1 snapshot)
|
||||
// Tree view: a pure derivation over a snapshot
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/// One node in the parentage forest. `children` are the actors whose recorded
|
||||
/// parent edge points at this node's pid.
|
||||
/// One node in the parentage forest returned by [`tree`]. `children` are the
|
||||
/// actors whose recorded parent (see [`ActorInfo::supervisor`]) points at
|
||||
/// this node's actor.
|
||||
#[derive(Debug, Clone)]
|
||||
pub struct TreeNode {
|
||||
pub info: ActorInfo,
|
||||
/// The actor's recorded parent was absent from the snapshot (already
|
||||
/// Done/Vacant, or itself a tombstone), so it was re-rooted under the forest
|
||||
/// sentinel rather than dropped — the tree stays total (DECISION D8).
|
||||
/// True if this actor's recorded parent was not found in the snapshot
|
||||
/// (it had already exited, or was itself missing), so this node was
|
||||
/// placed at the top of the forest instead of being dropped. This keeps
|
||||
/// every actor in the snapshot visible somewhere in the tree, even one
|
||||
/// whose parent is gone.
|
||||
pub orphaned: bool,
|
||||
pub children: Vec<TreeNode>,
|
||||
}
|
||||
|
||||
/// The parentage forest. Roots are actors parented at `ROOT_PID` (genuine
|
||||
/// roots) plus re-rooted orphans. The edge is *spawned-by / parent*, not
|
||||
/// necessarily supervision (DECISION D9) — see [`ActorInfo::supervisor`].
|
||||
/// The parentage forest: every actor from a snapshot, arranged by who spawned
|
||||
/// whom. Roots are actors with no parent in the snapshot (including the
|
||||
/// run's own root actor) plus any orphaned actors (see [`TreeNode::orphaned`]).
|
||||
/// This mirrors spawn parentage, not necessarily a supervision tree; see
|
||||
/// [`ActorInfo::supervisor`].
|
||||
#[derive(Debug, Clone)]
|
||||
pub struct RuntimeTree {
|
||||
pub format_version: u16,
|
||||
pub roots: Vec<TreeNode>,
|
||||
}
|
||||
|
||||
/// Take a live [`snapshot`] and fold it into the parentage forest.
|
||||
/// Take a fresh [`snapshot`] and fold it into the parentage forest.
|
||||
pub fn tree() -> RuntimeTree {
|
||||
tree_from(snapshot())
|
||||
}
|
||||
|
||||
/// Fold an existing snapshot into a forest by grouping each actor under its
|
||||
/// parent pid — a single O(n) pass, no new reads. Exposed separately so a
|
||||
/// consumer that already holds a snapshot (or a synthetic one, in tests) can
|
||||
/// derive the tree without a second scan.
|
||||
/// Fold an existing snapshot into a parentage forest by grouping each actor
|
||||
/// under its parent, without taking a new snapshot. Useful if you already
|
||||
/// have one (for example, one built in a test, or one you took earlier and
|
||||
/// want to inspect again) and want the tree view of it without re-reading
|
||||
/// the runtime.
|
||||
pub fn tree_from(snap: RuntimeSnapshot) -> RuntimeTree {
|
||||
let RuntimeSnapshot { format_version, actors } = snap;
|
||||
let RuntimeSnapshot {
|
||||
format_version,
|
||||
actors,
|
||||
} = snap;
|
||||
|
||||
let mut index_of: HashMap<Pid, usize> = HashMap::with_capacity(actors.len());
|
||||
for (i, a) in actors.iter().enumerate() {
|
||||
@@ -254,7 +411,7 @@ pub fn tree_from(snap: RuntimeSnapshot) -> RuntimeTree {
|
||||
children_of.entry(parent).or_default().push(i);
|
||||
} else {
|
||||
// Parent is the forest sentinel (genuine root) or absent from the
|
||||
// snapshot (orphan, D8) — either way a root of the forest.
|
||||
// snapshot (orphan): either way, a root of the forest.
|
||||
orphaned[i] = parent != ROOT_PID;
|
||||
roots.push(i);
|
||||
}
|
||||
@@ -267,7 +424,10 @@ pub fn tree_from(snap: RuntimeSnapshot) -> RuntimeTree {
|
||||
.into_iter()
|
||||
.filter_map(|i| build_node(i, &children_of, &orphaned, &mut slots))
|
||||
.collect();
|
||||
RuntimeTree { format_version, roots: root_nodes }
|
||||
RuntimeTree {
|
||||
format_version,
|
||||
roots: root_nodes,
|
||||
}
|
||||
}
|
||||
|
||||
fn build_node(
|
||||
@@ -285,5 +445,9 @@ fn build_node(
|
||||
.collect()
|
||||
})
|
||||
.unwrap_or_default();
|
||||
Some(TreeNode { info, orphaned: orphaned[i], children })
|
||||
Some(TreeNode {
|
||||
info,
|
||||
orphaned: orphaned[i],
|
||||
children,
|
||||
})
|
||||
}
|
||||
|
||||
@@ -13,44 +13,68 @@
|
||||
//! leaves the actor, no copying through an intermediary thread. Built on
|
||||
//! these are the conveniences `read(fd, &mut buf)` and `write(fd, &buf)`.
|
||||
//!
|
||||
//! Architecture
|
||||
//! ============
|
||||
//! Per `run()`, two OS threads:
|
||||
//! - **epoll thread**: owns the epollfd. Loops in `epoll_wait`. On a
|
||||
//! ready fd, pushes `Completion::FdReady { pid, fd, events }` to the
|
||||
//! shared completion queue and writes the scheduler-wake pipe. On the
|
||||
//! shutdown pipe (also registered in epollfd), exits.
|
||||
//! - **pool thread**: blocks on the request mpsc. Runs the closure
|
||||
//! inside `catch_unwind`, pushes `Completion::Blocking { pid, result }`,
|
||||
//! writes the scheduler-wake pipe.
|
||||
//! Architecture (RFC 018: driver-enqueues)
|
||||
//! =======================================
|
||||
//! Per `run()`, two OS threads, each a *producer* behind the runtime's
|
||||
//! two-call contract — make the actor runnable (`unpark_at`, whose enqueue
|
||||
//! tail wakes a parked scheduler), nothing else:
|
||||
//!
|
||||
//! Both threads share a single `completions: Arc<Mutex<VecDeque<Completion>>>`
|
||||
//! and the same scheduler-wake pipe.
|
||||
//! - **epoll thread**: owns `epoll_wait` on the epollfd. On a ready fd it
|
||||
//! removes the parked waiter from the shared `waiters` map and DELs the
|
||||
//! fd (both under the waiters lock — see below), then unparks the
|
||||
//! actor directly. On the shutdown pipe (also registered in the
|
||||
//! epollfd), exits.
|
||||
//! - **pool thread**: blocks on the request mpsc. Runs the closure inside
|
||||
//! `catch_unwind`, stashes the result in the actor's slot
|
||||
//! (`pending_io_result`, under the cold lock, generation-checked),
|
||||
//! decrements the runtime's `io_outstanding`, and unparks the actor.
|
||||
//!
|
||||
//! `epoll_ctl` (register/unregister fd interest) is called by the
|
||||
//! scheduler thread *directly* on the epollfd. That's well-defined per
|
||||
//! `epoll_ctl(2)`: a thread may be calling `epoll_wait` on the epollfd
|
||||
//! while another thread calls `epoll_ctl`. Avoids needing a second mpsc
|
||||
//! and a second wake mechanism.
|
||||
//! There is no shared completion queue and no wake pipe: each producer
|
||||
//! routes its own completion, so the whole byte-vs-completion visibility
|
||||
//! discipline of the drain era — and the stranded-completion hazards it
|
||||
//! defended against — is unrepresentable. Producers reach the runtime
|
||||
//! through a `Weak<RuntimeInner>`: upgraded per completion (the path is
|
||||
//! syscall-bound; the refcount op is noise) and avoiding an Arc cycle
|
||||
//! through `RuntimeInner::io`.
|
||||
//!
|
||||
//! `epoll_ctl` (register fd interest) is called by the scheduler thread
|
||||
//! directly on the epollfd. That's well-defined per `epoll_ctl(2)`: a
|
||||
//! thread may be calling `epoll_wait` on the epollfd while another thread
|
||||
//! calls `epoll_ctl`.
|
||||
//!
|
||||
//! Epoll mode
|
||||
//! ==========
|
||||
//! Level-triggered with EPOLLONESHOT. After a wakeup the kernel
|
||||
//! auto-disarms the fd, so we never get two wakeups for one
|
||||
//! `wait_readable` call. The scheduler explicitly `EPOLL_CTL_DEL`s the fd
|
||||
//! on completion to free the slot for re-registration. Net effect: each
|
||||
//! `wait_readable` call. The epoll thread explicitly `EPOLL_CTL_DEL`s the
|
||||
//! fd on readiness to free the slot for re-registration. Net effect: each
|
||||
//! `wait_readable(fd)` is one ADD, one wakeup, one DEL — symmetric and
|
||||
//! stateless between calls.
|
||||
//!
|
||||
//! ## The waiters lock is the ADD/DEL serialization
|
||||
//!
|
||||
//! Registration (scheduler thread: check-vacant, defensive DEL, ADD,
|
||||
//! insert) and readiness consumption (epoll thread: remove, DEL) each run
|
||||
//! entirely under the `waiters` mutex. This is what makes the
|
||||
//! oneshot-rearm race unrepresentable: a woken actor re-registering the
|
||||
//! same fd cannot interleave with the epoll thread's DEL for the *previous*
|
||||
//! registration — whichever takes the lock second sees a consistent
|
||||
//! kernel-side state. Lock order: `io` (the runtime's outer mutex, held by
|
||||
//! scheduler-side callers) → `waiters` → slot/queue leaves via `unpark_at`.
|
||||
//! The epoll thread takes `waiters` without `io` — it must never take
|
||||
//! `io`, both for lock-order hygiene and because teardown holds `io` while
|
||||
//! joining it.
|
||||
//!
|
||||
//! Fd hygiene
|
||||
//! ==========
|
||||
//! An actor stopped while waiting on an fd unwinds out of `wait_fd`'s park;
|
||||
//! a drop guard there (armed after a successful register, forgotten on a
|
||||
//! normal wake) removes the `waiters` entry iff it is still that wait's
|
||||
//! `(pid, epoch)` and only then `EPOLL_CTL_DEL`s the fd — an entry already
|
||||
//! consumed by a racing `FdReady` means the fd may carry someone else's
|
||||
//! fresh registration, which must be left alone. `epoll_register` keeps a
|
||||
//! defensive bare DEL before ADD as belt-and-braces.
|
||||
//! normal wake) calls [`IoThread::cancel_waiter`], which removes the
|
||||
//! `waiters` entry iff it is still that wait's `(pid, epoch)` and only then
|
||||
//! `EPOLL_CTL_DEL`s the fd — an entry already consumed by the epoll thread
|
||||
//! means the fd may carry someone else's fresh registration, which must be
|
||||
//! left alone. `epoll_register` keeps a defensive bare DEL before ADD as
|
||||
//! belt-and-braces.
|
||||
//!
|
||||
//! Buffers used with `read`/`write` should be on fds opened with
|
||||
//! `O_NONBLOCK`. If they aren't, the syscall may block the scheduler
|
||||
@@ -68,13 +92,14 @@
|
||||
//! they have no equivalent panic-propagation path.
|
||||
|
||||
use crate::pid::Pid;
|
||||
use crate::runtime::RuntimeInner;
|
||||
use std::any::Any;
|
||||
use std::collections::{HashMap, VecDeque};
|
||||
use std::collections::HashMap;
|
||||
use std::io;
|
||||
use std::os::fd::RawFd;
|
||||
use std::panic;
|
||||
use std::sync::mpsc;
|
||||
use std::sync::{Arc, Mutex};
|
||||
use std::sync::atomic::Ordering;
|
||||
use std::sync::{mpsc, Arc, Mutex, Weak};
|
||||
use std::thread::JoinHandle as OsJoinHandle;
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
@@ -86,45 +111,31 @@ use std::thread::JoinHandle as OsJoinHandle;
|
||||
pub type IoResult = Result<Box<dyn Any + Send>, Box<dyn Any + Send>>;
|
||||
|
||||
struct Request {
|
||||
/// The submitter's park-epoch — carried through to the `Blocking`
|
||||
/// completion so the wake is epoch-matched.
|
||||
/// The submitter's park-epoch — the eventual wake is epoch-matched.
|
||||
epoch: u32,
|
||||
pid: Pid,
|
||||
/// The work to perform. Returns the wire-form result directly.
|
||||
work: Box<dyn FnOnce() -> IoResult + Send>,
|
||||
}
|
||||
|
||||
/// Completion message from either IO thread back to the scheduler.
|
||||
pub enum Completion {
|
||||
/// A `block_on_io` closure has finished (Ok = return value, Err = panic
|
||||
/// payload).
|
||||
Blocking { pid: Pid, epoch: u32, result: IoResult },
|
||||
/// An fd registered via `wait_readable`/`wait_writable` is ready. The
|
||||
/// scheduler looks up the parked pid in `waiters`, unparks it, and
|
||||
/// removes the entry. `pid` isn't in this variant because the epoll
|
||||
/// thread doesn't have access to the `waiters` map; the scheduler
|
||||
/// thread owns that.
|
||||
FdReady { fd: RawFd, events: u32 },
|
||||
}
|
||||
/// The parked-waiter map, shared between scheduler-side registration and
|
||||
/// the epoll thread's readiness consumption. See the module docs on why
|
||||
/// this single lock is the ADD/DEL serialization.
|
||||
type Waiters = Arc<Mutex<HashMap<RawFd, (Pid, u32)>>>;
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// IoThread — created per `run()`, owned by `SchedulerState`.
|
||||
// IoThread — created per `run()`, owned by `RuntimeInner::io`.
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
pub struct IoThread {
|
||||
// ----- Channels & queues -----
|
||||
|
||||
/// Submission queue into the blocking-work pool.
|
||||
tx: mpsc::Sender<Request>,
|
||||
/// Shared completion queue, fed by both the pool and the epoll thread.
|
||||
completions: Arc<Mutex<VecDeque<Completion>>>,
|
||||
/// Pipe the scheduler polls in its idle path. Both IO threads write to
|
||||
/// `wake_write` after pushing a completion.
|
||||
wake_read: RawFd,
|
||||
wake_write: RawFd,
|
||||
/// One parked actor per registered fd. Populated by `epoll_register`,
|
||||
/// consumed by the epoll thread on readiness or `cancel_waiter` on an
|
||||
/// unwound wait.
|
||||
waiters: Waiters,
|
||||
|
||||
// ----- Epoll machinery -----
|
||||
|
||||
/// The epollfd, owned by `IoThread`. Callable cross-thread via
|
||||
/// `epoll_ctl` per the man page.
|
||||
epollfd: RawFd,
|
||||
@@ -133,39 +144,24 @@ pub struct IoThread {
|
||||
/// shutdown.
|
||||
shutdown_read: RawFd,
|
||||
shutdown_write: RawFd,
|
||||
/// One parked actor per registered fd. Populated by `wait_readable` /
|
||||
/// `wait_writable` and drained by the scheduler when a `FdReady`
|
||||
/// completion is processed.
|
||||
pub waiters: HashMap<RawFd, (Pid, u32)>,
|
||||
|
||||
// ----- Threads -----
|
||||
|
||||
pool_thread: Option<OsJoinHandle<()>>,
|
||||
epoll_thread: Option<OsJoinHandle<()>>,
|
||||
|
||||
/// Number of `block_on_io` requests in-flight. Used by the scheduler's
|
||||
/// idle path to decide whether to wait on the pipe or exit. Fd waits
|
||||
/// are not counted here; they're counted by `waiters.len()`.
|
||||
pub outstanding: u32,
|
||||
}
|
||||
|
||||
impl IoThread {
|
||||
pub fn start() -> io::Result<Self> {
|
||||
// Scheduler-facing wake pipe.
|
||||
let (wake_read, wake_write) = make_pipe()?;
|
||||
// Pool submission channel + shared completion queue.
|
||||
/// Start the pool and epoll threads. `rt` is the producers' route back
|
||||
/// into the runtime (slot table + unpark protocol); a `Weak` so the
|
||||
/// `RuntimeInner → IoThread → RuntimeInner` cycle never forms.
|
||||
pub(crate) fn start(rt: Weak<RuntimeInner>) -> io::Result<Self> {
|
||||
// Pool submission channel.
|
||||
let (tx, rx) = mpsc::channel::<Request>();
|
||||
let completions: Arc<Mutex<VecDeque<Completion>>> =
|
||||
Arc::new(Mutex::new(VecDeque::new()));
|
||||
let waiters: Waiters = Arc::new(Mutex::new(HashMap::new()));
|
||||
|
||||
// Epoll machinery.
|
||||
let epollfd = unsafe { libc::epoll_create1(libc::EPOLL_CLOEXEC) };
|
||||
if epollfd < 0 {
|
||||
// Best-effort fd cleanup before bailing.
|
||||
unsafe {
|
||||
libc::close(wake_read);
|
||||
libc::close(wake_write);
|
||||
}
|
||||
return Err(io::Error::last_os_error());
|
||||
}
|
||||
|
||||
@@ -174,8 +170,6 @@ impl IoThread {
|
||||
Err(e) => {
|
||||
unsafe {
|
||||
libc::close(epollfd);
|
||||
libc::close(wake_read);
|
||||
libc::close(wake_write);
|
||||
}
|
||||
return Err(e);
|
||||
}
|
||||
@@ -202,42 +196,37 @@ impl IoThread {
|
||||
libc::close(epollfd);
|
||||
libc::close(shutdown_read);
|
||||
libc::close(shutdown_write);
|
||||
libc::close(wake_read);
|
||||
libc::close(wake_write);
|
||||
}
|
||||
return Err(e);
|
||||
}
|
||||
|
||||
// Spawn pool thread.
|
||||
let pool_comps = completions.clone();
|
||||
let pool_rt = rt.clone();
|
||||
let pool_thread = std::thread::Builder::new()
|
||||
.name("smarm-io-pool".into())
|
||||
.spawn(move || pool_loop(rx, pool_comps, wake_write))?;
|
||||
.spawn(move || pool_loop(rx, pool_rt))?;
|
||||
|
||||
// Spawn epoll thread.
|
||||
let epoll_comps = completions.clone();
|
||||
let epoll_waiters = waiters.clone();
|
||||
let epoll_thread = std::thread::Builder::new()
|
||||
.name("smarm-io-epoll".into())
|
||||
.spawn(move || epoll_loop(epollfd, epoll_comps, wake_write))?;
|
||||
.spawn(move || epoll_loop(epollfd, epoll_waiters, rt))?;
|
||||
|
||||
Ok(Self {
|
||||
tx,
|
||||
completions,
|
||||
wake_read,
|
||||
wake_write,
|
||||
waiters,
|
||||
epollfd,
|
||||
shutdown_read,
|
||||
shutdown_write,
|
||||
waiters: HashMap::new(),
|
||||
pool_thread: Some(pool_thread),
|
||||
epoll_thread: Some(epoll_thread),
|
||||
outstanding: 0,
|
||||
})
|
||||
}
|
||||
|
||||
/// Hand a request to the pool. Increments `outstanding`.
|
||||
/// Hand a request to the pool. The caller (scheduler.rs) increments
|
||||
/// `io_outstanding` BEFORE calling — the pool decrements on completion,
|
||||
/// and an increment that trailed the completion would underflow.
|
||||
pub fn submit(&mut self, pid: Pid, epoch: u32, work: Box<dyn FnOnce() -> IoResult + Send>) {
|
||||
self.outstanding += 1;
|
||||
// Send can only fail if the pool has hung up, which only happens
|
||||
// on shutdown. submit during shutdown is a bug.
|
||||
if self.tx.send(Request { pid, epoch, work }).is_err() {
|
||||
@@ -245,39 +234,13 @@ impl IoThread {
|
||||
}
|
||||
}
|
||||
|
||||
/// Drain every available completion. Caller (the scheduler) routes the
|
||||
/// results and updates `outstanding` / `waiters` accordingly.
|
||||
pub fn drain_completions(&mut self) -> Vec<Completion> {
|
||||
let mut q = match self.completions.lock() {
|
||||
Ok(g) => g,
|
||||
Err(e) => panic!("smarm: io completions lock poisoned (core corrupt): {e}"),
|
||||
};
|
||||
let mut out = Vec::with_capacity(q.len());
|
||||
while let Some(c) = q.pop_front() {
|
||||
out.push(c);
|
||||
}
|
||||
out
|
||||
}
|
||||
|
||||
pub fn wake_fd(&self) -> RawFd {
|
||||
self.wake_read
|
||||
}
|
||||
|
||||
/// Write the wake pipe directly: rouse every scheduler thread blocked in
|
||||
/// its idle `poll_wake`. Used by the terminal (AllDone) path — an idle
|
||||
/// sibling may be blocked on a snapshot that nothing will ever refresh
|
||||
/// (an orphaned timer deadline, or `io_outstanding` from a waiter that
|
||||
/// was stop-cancelled and so never produces a completion).
|
||||
pub fn wake(&self) {
|
||||
wake_scheduler(self.wake_write);
|
||||
}
|
||||
|
||||
/// Register interest in `fd` becoming readable/writable; record `pid`
|
||||
/// as the parked waiter. The epoll thread will push a `FdReady`
|
||||
/// completion when the kernel signals.
|
||||
/// as the parked waiter. The epoll thread unparks it on readiness.
|
||||
/// The caller increments `io_fd_waiters` BEFORE calling (mirror of
|
||||
/// `submit`'s contract) and decrements it again if this errors.
|
||||
///
|
||||
/// EPOLLONESHOT: one wakeup per registration. The scheduler must
|
||||
/// `epoll_del` on completion to free the slot for re-registration.
|
||||
/// EPOLLONESHOT: one wakeup per registration; the epoll thread DELs on
|
||||
/// readiness, `cancel_waiter` DELs on an unwound wait.
|
||||
pub fn epoll_register(
|
||||
&mut self,
|
||||
fd: RawFd,
|
||||
@@ -286,20 +249,24 @@ impl IoThread {
|
||||
readable: bool,
|
||||
writable: bool,
|
||||
) -> io::Result<()> {
|
||||
let mut waiters = match self.waiters.lock() {
|
||||
Ok(g) => g,
|
||||
Err(e) => panic!("smarm: io waiters lock poisoned (core corrupt): {e}"),
|
||||
};
|
||||
// Two actors waiting on the same fd would be a misuse: the kernel
|
||||
// delivers exactly one EPOLLONESHOT wakeup, so the second waiter
|
||||
// would hang. Reject up front.
|
||||
if self.waiters.contains_key(&fd) {
|
||||
if waiters.contains_key(&fd) {
|
||||
return Err(io::Error::new(
|
||||
io::ErrorKind::AlreadyExists,
|
||||
"fd already has a parked waiter",
|
||||
));
|
||||
}
|
||||
|
||||
// Belt-and-braces: the unwind guard in `wait_fd` is responsible for
|
||||
// cleaning up a stopped waiter's registration, but a bare DEL is
|
||||
// harmless if the fd isn't registered (ENOENT) and removes any leak
|
||||
// a path we haven't thought of might leave behind.
|
||||
// Belt-and-braces: `cancel_waiter` is responsible for cleaning up a
|
||||
// stopped waiter's registration, but a bare DEL is harmless if the
|
||||
// fd isn't registered (ENOENT) and removes any leak a path we
|
||||
// haven't thought of might leave behind.
|
||||
unsafe {
|
||||
libc::epoll_ctl(self.epollfd, libc::EPOLL_CTL_DEL, fd, std::ptr::null_mut());
|
||||
}
|
||||
@@ -315,25 +282,34 @@ impl IoThread {
|
||||
events,
|
||||
u64: fd as u64,
|
||||
};
|
||||
let r = unsafe {
|
||||
libc::epoll_ctl(self.epollfd, libc::EPOLL_CTL_ADD, fd, &mut ev as *mut _)
|
||||
};
|
||||
let r =
|
||||
unsafe { libc::epoll_ctl(self.epollfd, libc::EPOLL_CTL_ADD, fd, &mut ev as *mut _) };
|
||||
if r < 0 {
|
||||
return Err(io::Error::last_os_error());
|
||||
}
|
||||
self.waiters.insert(fd, (pid, epoch));
|
||||
waiters.insert(fd, (pid, epoch));
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Remove `fd` from the epollfd. Called by the scheduler after a
|
||||
/// `FdReady` completion, so the next `wait_readable(fd)` can ADD again.
|
||||
///
|
||||
/// Does NOT touch `waiters` — that's the scheduler's bookkeeping; this
|
||||
/// is purely the kernel-side cleanup.
|
||||
pub fn epoll_deregister(&mut self, fd: RawFd) {
|
||||
// EPOLL_CTL_DEL of an already-removed fd returns ENOENT; ignore.
|
||||
unsafe {
|
||||
libc::epoll_ctl(self.epollfd, libc::EPOLL_CTL_DEL, fd, std::ptr::null_mut());
|
||||
/// Remove `fd`'s waiter iff it is still `(pid, epoch)`, DELing the fd
|
||||
/// from the epollfd in the same critical section. Returns whether the
|
||||
/// entry was removed (the caller then decrements `io_fd_waiters`).
|
||||
/// `false` means the epoll thread consumed the registration first —
|
||||
/// the fd may already carry someone else's fresh ADD; hands off.
|
||||
pub fn cancel_waiter(&mut self, fd: RawFd, pid: Pid, epoch: u32) -> bool {
|
||||
let mut waiters = match self.waiters.lock() {
|
||||
Ok(g) => g,
|
||||
Err(e) => panic!("smarm: io waiters lock poisoned (core corrupt): {e}"),
|
||||
};
|
||||
if waiters.get(&fd) == Some(&(pid, epoch)) {
|
||||
waiters.remove(&fd);
|
||||
// EPOLL_CTL_DEL of an already-removed fd returns ENOENT; ignore.
|
||||
unsafe {
|
||||
libc::epoll_ctl(self.epollfd, libc::EPOLL_CTL_DEL, fd, std::ptr::null_mut());
|
||||
}
|
||||
true
|
||||
} else {
|
||||
false
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -354,7 +330,10 @@ impl Drop for IoThread {
|
||||
let real_tx = std::mem::replace(&mut self.tx, dead_tx);
|
||||
drop(real_tx);
|
||||
|
||||
// 3. Join both threads.
|
||||
// 3. Join both threads. Safe even while the caller holds the
|
||||
// runtime's `io` mutex: neither thread ever takes it (they reach
|
||||
// the runtime through a Weak they upgrade per completion, and
|
||||
// the epoll thread's only lock is `waiters`).
|
||||
if let Some(h) = self.epoll_thread.take() {
|
||||
let _ = h.join();
|
||||
}
|
||||
@@ -367,8 +346,6 @@ impl Drop for IoThread {
|
||||
libc::close(self.epollfd);
|
||||
libc::close(self.shutdown_read);
|
||||
libc::close(self.shutdown_write);
|
||||
libc::close(self.wake_read);
|
||||
libc::close(self.wake_write);
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -379,36 +356,38 @@ impl Drop for IoThread {
|
||||
const SHUTDOWN_EPOLL_TOKEN: u64 = u64::MAX;
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Pool loop
|
||||
// Pool loop (producer: Blocking completions)
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
fn pool_loop(
|
||||
rx: mpsc::Receiver<Request>,
|
||||
completions: Arc<Mutex<VecDeque<Completion>>>,
|
||||
wake_write: RawFd,
|
||||
) {
|
||||
fn pool_loop(rx: mpsc::Receiver<Request>, rt: Weak<RuntimeInner>) {
|
||||
while let Ok(Request { pid, epoch, work }) = rx.recv() {
|
||||
let result: IoResult = match panic::catch_unwind(panic::AssertUnwindSafe(work)) {
|
||||
Ok(r) => r,
|
||||
Err(payload) => Err(payload),
|
||||
};
|
||||
match completions.lock() {
|
||||
Ok(mut g) => g.push_back(Completion::Blocking { pid, epoch, result }),
|
||||
Err(e) => panic!("smarm: io completions lock poisoned (core corrupt): {e}"),
|
||||
let Some(inner) = rt.upgrade() else { return };
|
||||
// Stash the result under the cold lock (generation-checked: an
|
||||
// actor stopped with the op in flight discards it), decrement the
|
||||
// in-flight count, then wake through the epoch-matched unpark. The
|
||||
// unpark's enqueue tail wakes a parked scheduler; the actor stays
|
||||
// `live` until it resumes and finalizes, so the decrement's
|
||||
// ordering against the termination verdict is not load-bearing.
|
||||
if let Some(slot) = inner.slot_at(pid) {
|
||||
let mut cold = slot.cold.lock();
|
||||
if slot.generation() == pid.generation() {
|
||||
cold.pending_io_result = Some(result);
|
||||
}
|
||||
}
|
||||
wake_scheduler(wake_write);
|
||||
inner.io_outstanding.fetch_sub(1, Ordering::AcqRel);
|
||||
inner.unpark_at(pid, epoch);
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Epoll loop
|
||||
// Epoll loop (producer: FdReady completions)
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
fn epoll_loop(
|
||||
epollfd: RawFd,
|
||||
completions: Arc<Mutex<VecDeque<Completion>>>,
|
||||
wake_write: RawFd,
|
||||
) {
|
||||
fn epoll_loop(epollfd: RawFd, waiters: Waiters, rt: Weak<RuntimeInner>) {
|
||||
// Buffer for epoll_wait. 64 is plenty for our scale; if a real load
|
||||
// appears that needs more, this is a one-line change.
|
||||
const MAX_EVENTS: usize = 64;
|
||||
@@ -416,12 +395,7 @@ fn epoll_loop(
|
||||
|
||||
loop {
|
||||
let n = unsafe {
|
||||
libc::epoll_wait(
|
||||
epollfd,
|
||||
events.as_mut_ptr(),
|
||||
MAX_EVENTS as libc::c_int,
|
||||
-1,
|
||||
)
|
||||
libc::epoll_wait(epollfd, events.as_mut_ptr(), MAX_EVENTS as libc::c_int, -1)
|
||||
};
|
||||
|
||||
if n < 0 {
|
||||
@@ -436,29 +410,36 @@ fn epoll_loop(
|
||||
}
|
||||
|
||||
let mut shutdown_requested = false;
|
||||
let mut pushed_any = false;
|
||||
{
|
||||
let mut q = match completions.lock() {
|
||||
Ok(g) => g,
|
||||
Err(e) => panic!("smarm: io completions lock poisoned (core corrupt): {e}"),
|
||||
for ev in events.iter().take(n as usize) {
|
||||
if ev.u64 == SHUTDOWN_EPOLL_TOKEN {
|
||||
shutdown_requested = true;
|
||||
continue;
|
||||
}
|
||||
let fd = ev.u64 as RawFd;
|
||||
// Consume the registration: remove + DEL under the waiters
|
||||
// lock (the ADD/DEL serialization — see module docs). A
|
||||
// vanished entry means `cancel_waiter` beat us: the wake is
|
||||
// already moot.
|
||||
let entry = {
|
||||
let mut w = match waiters.lock() {
|
||||
Ok(g) => g,
|
||||
Err(e) => {
|
||||
panic!("smarm: io waiters lock poisoned (core corrupt): {e}")
|
||||
}
|
||||
};
|
||||
let entry = w.remove(&fd);
|
||||
if entry.is_some() {
|
||||
unsafe {
|
||||
libc::epoll_ctl(epollfd, libc::EPOLL_CTL_DEL, fd, std::ptr::null_mut());
|
||||
}
|
||||
}
|
||||
entry
|
||||
};
|
||||
for ev in events.iter().take(n as usize) {
|
||||
if ev.u64 == SHUTDOWN_EPOLL_TOKEN {
|
||||
shutdown_requested = true;
|
||||
continue;
|
||||
}
|
||||
let fd = ev.u64 as RawFd;
|
||||
let evs = ev.events;
|
||||
q.push_back(Completion::FdReady {
|
||||
fd,
|
||||
events: evs,
|
||||
});
|
||||
pushed_any = true;
|
||||
if let Some((pid, epoch)) = entry {
|
||||
let Some(inner) = rt.upgrade() else { return };
|
||||
inner.io_fd_waiters.fetch_sub(1, Ordering::AcqRel);
|
||||
inner.unpark_at(pid, epoch);
|
||||
}
|
||||
}
|
||||
|
||||
if pushed_any {
|
||||
wake_scheduler(wake_write);
|
||||
}
|
||||
if shutdown_requested {
|
||||
return;
|
||||
@@ -466,27 +447,8 @@ fn epoll_loop(
|
||||
}
|
||||
}
|
||||
|
||||
/// Write one byte to the scheduler's wake pipe. Retries on EINTR; ignores
|
||||
/// EAGAIN (pipe full means there's already an outstanding wake we haven't
|
||||
/// consumed yet, which is sufficient).
|
||||
fn wake_scheduler(wake_write: RawFd) {
|
||||
let buf: [u8; 1] = [0];
|
||||
unsafe {
|
||||
loop {
|
||||
let n = libc::write(wake_write, buf.as_ptr() as *const _, 1);
|
||||
if n < 0 {
|
||||
let e = *libc::__errno_location();
|
||||
if e == libc::EINTR {
|
||||
continue;
|
||||
}
|
||||
}
|
||||
break;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Pipe helpers (unchanged from v0.2)
|
||||
// Pipe helper
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
fn make_pipe() -> io::Result<(RawFd, RawFd)> {
|
||||
@@ -497,50 +459,3 @@ fn make_pipe() -> io::Result<(RawFd, RawFd)> {
|
||||
}
|
||||
Ok((fds[0], fds[1]))
|
||||
}
|
||||
|
||||
/// Drain pending bytes from the wake pipe. Nonblocking (pipe is O_NONBLOCK).
|
||||
///
|
||||
/// DISCIPLINE: called only by the phase-1 drain-lock winner, immediately
|
||||
/// before `drain_completions`. Bytes are the notification channel for
|
||||
/// completions; consuming one anywhere else can strand the completion it
|
||||
/// announces (see the lost-wakeup note at the call site in `schedule_loop`).
|
||||
pub fn drain_wake_pipe(fd: RawFd) {
|
||||
let mut buf = [0u8; 64];
|
||||
loop {
|
||||
let n = unsafe { libc::read(fd, buf.as_mut_ptr() as *mut _, buf.len()) };
|
||||
if n <= 0 {
|
||||
break;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Block on `fd` for up to `timeout`, returning when either there's data
|
||||
/// to read or the timeout elapses. `None` for `timeout` means wait forever.
|
||||
pub fn poll_wake(fd: RawFd, timeout: Option<std::time::Duration>) {
|
||||
let timeout_ms: libc::c_int = match timeout {
|
||||
None => -1,
|
||||
Some(d) => {
|
||||
let ms = d.as_millis();
|
||||
if ms > i32::MAX as u128 {
|
||||
i32::MAX
|
||||
} else {
|
||||
ms as i32
|
||||
}
|
||||
}
|
||||
};
|
||||
let mut pfd = libc::pollfd {
|
||||
fd,
|
||||
events: libc::POLLIN,
|
||||
revents: 0,
|
||||
};
|
||||
loop {
|
||||
let r = unsafe { libc::poll(&mut pfd as *mut _, 1, timeout_ms) };
|
||||
if r < 0 {
|
||||
let e = unsafe { *libc::__errno_location() };
|
||||
if e == libc::EINTR {
|
||||
continue;
|
||||
}
|
||||
}
|
||||
break;
|
||||
}
|
||||
}
|
||||
|
||||
+41
-33
@@ -11,34 +11,36 @@
|
||||
//!
|
||||
//! See `LOOM.md` for the design intent and the deferred-for-later list.
|
||||
|
||||
pub mod stack;
|
||||
pub mod context;
|
||||
pub mod preempt;
|
||||
pub mod pid;
|
||||
pub mod actor;
|
||||
pub mod causal;
|
||||
pub mod channel;
|
||||
pub mod scheduler;
|
||||
pub mod supervisor;
|
||||
pub mod timer;
|
||||
pub mod io;
|
||||
pub mod mutex;
|
||||
pub mod monitor;
|
||||
pub mod registry;
|
||||
pub mod pg;
|
||||
pub mod link;
|
||||
pub mod context;
|
||||
pub mod gen_server;
|
||||
pub mod gen_statem;
|
||||
pub mod introspect;
|
||||
pub mod io;
|
||||
pub mod link;
|
||||
pub mod monitor;
|
||||
pub mod mutex;
|
||||
#[cfg(feature = "observer")]
|
||||
pub mod observer;
|
||||
pub mod runtime;
|
||||
pub(crate) mod park;
|
||||
pub mod pg;
|
||||
pub mod pid;
|
||||
pub mod preempt;
|
||||
pub(crate) mod raw_mutex;
|
||||
pub(crate) mod slot_state;
|
||||
pub(crate) mod sync_shim;
|
||||
pub mod registry;
|
||||
#[doc(hidden)] // pub only so benches/rq_micro.rs can drive the raw structures
|
||||
pub mod run_queue;
|
||||
pub mod runtime;
|
||||
pub mod scheduler;
|
||||
pub(crate) mod signal;
|
||||
pub(crate) mod slot_state;
|
||||
pub mod stack;
|
||||
pub mod supervisor;
|
||||
pub(crate) mod sync_shim;
|
||||
pub mod timer;
|
||||
pub mod trace;
|
||||
pub mod causal;
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Global allocator
|
||||
@@ -57,35 +59,41 @@ pub use channel::{
|
||||
};
|
||||
pub use gen_server::{
|
||||
call, cast, shutdown, whereis_server, CallError, CallTimeoutError, CastError, GenServer,
|
||||
NamedGenServerBuilder, GenServerBuilder, GenServerCtx, GenServerName, GenServerRef, TimerHandle, Watcher,
|
||||
GenServerBuilder, GenServerCtx, GenServerName, GenServerRef, NamedGenServerBuilder,
|
||||
ShutdownAction, StopHandle, TimerHandle, Watcher,
|
||||
};
|
||||
pub use gen_statem::{
|
||||
CallError as GenStatemCallError, Cx, Machine, Reply, Resolution, SendError as GenStatemSendError,
|
||||
GenStatemRef,
|
||||
CallError as GenStatemCallError, Cx, GenStatemName, GenStatemRef, Machine, Reply, Resolution,
|
||||
SendError as GenStatemSendError,
|
||||
};
|
||||
pub use introspect::{
|
||||
actor_info, snapshot, tree, tree_from, ActorInfo, ActorState, RuntimeSnapshot, RuntimeTree,
|
||||
TreeNode, SNAPSHOT_FORMAT_VERSION,
|
||||
StackInfo, TreeNode, SNAPSHOT_FORMAT_VERSION,
|
||||
};
|
||||
pub use link::{link, trap_exit, unlink, ExitSignal};
|
||||
pub use monitor::{
|
||||
demonitor, mark_watchable, monitor, terminal_reason, Down, DownReason, Monitor, MonitorId,
|
||||
};
|
||||
pub use mutex::{LockTimeout, Mutex, MutexGuard};
|
||||
#[cfg(feature = "observer")]
|
||||
pub use observer::{ObserverReply, ObserverRequest};
|
||||
pub use link::{link, trap_exit, unlink, ExitSignal};
|
||||
pub use monitor::{demonitor, monitor, Down, DownReason, Monitor, MonitorId};
|
||||
pub use mutex::{LockTimeout, Mutex, MutexGuard};
|
||||
pub use pg::{
|
||||
dispatch, join, leave, members, members_as, pick, pick_as, Incarnation, Member, NodeId,
|
||||
};
|
||||
pub use pid::{Addressable, Erased, Name, Pid, RawPid};
|
||||
pub use pg::{dispatch, join, leave, members, members_as, pick, pick_as, Incarnation, Member, NodeId};
|
||||
pub use registry::{
|
||||
install, lookup_as, register, send, send_dyn, send_to, unregister, whereis, RegisterError,
|
||||
SendError,
|
||||
install, lookup_as, register, resolve_name, send, send_dyn, send_to, unregister, whereis,
|
||||
NameResolution, RegisterError, SendError,
|
||||
};
|
||||
pub use runtime::{init, Config, Runtime};
|
||||
pub use runtime::{init, Config, Runtime, RuntimeHandle};
|
||||
pub use scheduler::{
|
||||
block_on_io, cancel_timer, request_stop, run, self_pid, send_after, send_after_named,
|
||||
send_after_named_wall, send_after_wall, sleep, sleep_wall,
|
||||
spawn, spawn_addr, spawn_under, wait_readable, wait_readable_timeout, wait_writable,
|
||||
wait_writable_timeout, yield_now, FdArm, JoinError, JoinHandle,
|
||||
block_on_io, cancel_timer, request_shutdown, request_stop, run, self_pid, send_after,
|
||||
send_after_named, send_after_named_wall, send_after_wall, sleep, sleep_wall, spawn, spawn_addr,
|
||||
spawn_addr_with, spawn_under, spawn_under_with, spawn_with, try_spawn, try_spawn_under_with,
|
||||
wait_readable, wait_readable_timeout, wait_writable, wait_writable_timeout, yield_now, FdArm,
|
||||
JoinError, JoinHandle, SpawnError, SpawnOpts,
|
||||
};
|
||||
pub use supervisor::{ChildSpec, OneForOne, Restart, Signal, Strategy};
|
||||
pub use supervisor::{ChildSpec, OneForOne, Restart, Shutdown, Signal, Strategy};
|
||||
pub use timer::TimerId;
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
+4
-1
@@ -157,7 +157,10 @@ pub fn link<A>(target: Pid<A>) {
|
||||
});
|
||||
match my_trap {
|
||||
Some(tx) => {
|
||||
let _ = tx.send(ExitSignal { from: target, reason: DownReason::NoProc });
|
||||
let _ = tx.send(ExitSignal {
|
||||
from: target,
|
||||
reason: DownReason::NoProc,
|
||||
});
|
||||
}
|
||||
None => request_stop(me),
|
||||
}
|
||||
|
||||
+170
-64
@@ -1,49 +1,85 @@
|
||||
//! Process monitors.
|
||||
//! Find out when another actor dies, without it knowing or caring that you're
|
||||
//! watching.
|
||||
//!
|
||||
//! `monitor(target)` asks the runtime to deliver a single [`Down`] when
|
||||
//! `target` terminates, and hands back a [`Monitor`] — the [`Receiver`] to read
|
||||
//! it from, plus the identity (`id`, `target`) needed to take the registration
|
||||
//! back down with [`demonitor`]. A monitor is:
|
||||
//! Say one actor manages a pool of workers and needs to know when a worker
|
||||
//! exits, so it can replace it. The worker does not need to know it is being
|
||||
//! watched, and nothing about the worker's own behavior should change because
|
||||
//! someone is watching it. That is what [`monitor`] is for: call
|
||||
//! `monitor(target)` to get a [`Monitor`], and read exactly one [`Down`]
|
||||
//! message off `monitor.rx` whenever `target` terminates, however it
|
||||
//! terminates.
|
||||
//!
|
||||
//! - **unidirectional** — the watcher learns of the target's death, but the
|
||||
//! target learns nothing of the watcher, and the watcher is unaffected by
|
||||
//! the death beyond the notification (contrast a *link*, which propagates
|
||||
//! failure);
|
||||
//! - **one-shot** — exactly one `Down` is ever sent for a given monitor.
|
||||
//! The returned channel closes afterwards, so a second `recv()` yields
|
||||
//! `Err(RecvError)`.
|
||||
//! ```
|
||||
//! use smarm::{monitor, run, spawn, DownReason};
|
||||
//!
|
||||
//! This generalizes the older single-`supervisor_channel` mechanism: a
|
||||
//! supervisor is just a hard-wired monitor that the parent installs at spawn
|
||||
//! time. Here any actor may monitor any pid, any number of times.
|
||||
//! run(|| {
|
||||
//! let worker = spawn(|| {
|
||||
//! // does some work, then returns
|
||||
//! });
|
||||
//! let pid = worker.pid();
|
||||
//!
|
||||
//! ## Reasons
|
||||
//! let m = monitor(pid);
|
||||
//! let _ = worker.join();
|
||||
//!
|
||||
//! [`DownReason`] is deliberately payload-free. A panicking actor's payload
|
||||
//! has a single owner and is delivered to whoever `join()`s the actor (as
|
||||
//! `JoinError`); a monitor only learns *that* it panicked, not the value.
|
||||
//! Monitoring a pid that is already gone (reclaimed, or never alive) yields
|
||||
//! [`DownReason::NoProc`] immediately, mirroring Erlang's `noproc`.
|
||||
//! let down = m.rx.recv().expect("monitor channel closed before Down");
|
||||
//! assert_eq!(down.pid, pid);
|
||||
//! assert_eq!(down.reason, DownReason::Exit);
|
||||
//! });
|
||||
//! ```
|
||||
//!
|
||||
//! ## Demonitoring
|
||||
//! A monitor is one-directional and one-shot:
|
||||
//!
|
||||
//! Each `monitor()` registration is tagged with a process-unique [`MonitorId`].
|
||||
//! [`demonitor`] removes the registration named by a [`Monitor`] from its
|
||||
//! target's slot, returning `Some(id)` if a live registration was found or
|
||||
//! `None` if it had already fired (or the target is gone). Dropping the
|
||||
//! [`Monitor`] afterwards discards any `Down` that the target had *already*
|
||||
//! queued — the equivalent of Erlang's `demonitor(Ref, [flush])`.
|
||||
//! - **One-directional**: the watcher learns that the target died, but the
|
||||
//! target is completely unaffected. It never learns it was being watched,
|
||||
//! and its own behavior and lifetime do not change because of the monitor.
|
||||
//! This is the opposite of a [`link`](mod@crate::link), which is bidirectional:
|
||||
//! linking two actors means an abnormal death on either side can bring the
|
||||
//! other down too. Reach for a monitor when you just want to *know*; reach
|
||||
//! for a link when a peer's crash should actually stop you.
|
||||
//! - **One-shot**: you get exactly one [`Down`] per `monitor()` call, then the
|
||||
//! channel closes. Calling `monitor` again on the same target (or a
|
||||
//! different one) gives you an independent registration with its own
|
||||
//! [`Monitor`] and its own one-shot channel; nothing stops you from
|
||||
//! monitoring the same actor many times over; each call is watched and
|
||||
//! fires on its own.
|
||||
//!
|
||||
//! ## Races
|
||||
//! ## Why a monitor never hands you the panic value
|
||||
//!
|
||||
//! Registration (below) and `finalize_actor` (in `runtime`) both run under the
|
||||
//! shared-state mutex, so a target that is still alive when its monitor is
|
||||
//! registered is guaranteed to deliver a real `Down`; there is no window in
|
||||
//! which the death slips between the liveness check and the registration.
|
||||
//! `demonitor` is protected by the generation half of the pid: if the target
|
||||
//! has died and its slot index been recycled, `slot_mut(target)` fails the
|
||||
//! generation check and `demonitor` is a clean no-op — it can never strip a
|
||||
//! *different* actor's monitor that happens to share the slot index.
|
||||
//! If the target panicked, [`Down`] tells you *that* it panicked
|
||||
//! ([`DownReason::Panic`]), but not the panic's payload. The payload has a
|
||||
//! single owner: it is handed to whichever caller `join()`s the actor's
|
||||
//! [`JoinHandle`](crate::JoinHandle), as a `JoinError`. A monitor only needs
|
||||
//! to know that something went wrong, not reproduce the exact value that
|
||||
//! caused it, so it gets the reason and nothing else.
|
||||
//!
|
||||
//! Monitoring a target that is already gone (it finished and was cleaned up,
|
||||
//! or the pid never pointed at a real actor) is not an error: you get a
|
||||
//! [`Down`] with [`DownReason::NoProc`] right away, instead of waiting
|
||||
//! forever for something that already happened.
|
||||
//!
|
||||
//! ## Stopping a monitor early
|
||||
//!
|
||||
//! [`demonitor`] cancels a monitor before it fires. If the registration was
|
||||
//! still live, it removes it and returns `Some` of the monitor's id: no
|
||||
//! `Down` will arrive on that channel from here on. If the target had already
|
||||
//! died and its `Down` already sent, there is nothing left to cancel and
|
||||
//! `demonitor` returns `None`; the `Down` you already have (or that is
|
||||
//! already sitting in the channel) is unaffected.
|
||||
//!
|
||||
//! If you want to cancel *and* make sure a `Down` that already arrived is
|
||||
//! discarded without reading it, just drop the [`Monitor`]: dropping it closes
|
||||
//! its receiver, and any queued `Down` is dropped along with it.
|
||||
//!
|
||||
//! ## Correctness notes for implementers
|
||||
//!
|
||||
//! A target that is still alive at the moment `monitor()` registers is
|
||||
//! guaranteed to eventually produce a real `Down`: registration and the
|
||||
//! target's own termination bookkeeping run under the same lock, so there is
|
||||
//! no window in which the target could die without the just-added
|
||||
//! registration seeing it. `demonitor` is similarly race-free against a target
|
||||
//! that has since died and had its slot reused by a new, unrelated actor: it
|
||||
//! is checked against the exact monitored incarnation, so it can never remove
|
||||
//! a different actor's registration by accident, it simply reports `None`.
|
||||
|
||||
use crate::channel::{channel, Receiver, Sender};
|
||||
use crate::pid::Pid;
|
||||
@@ -51,8 +87,8 @@ use crate::scheduler::with_runtime;
|
||||
|
||||
/// Why a monitored actor went down.
|
||||
///
|
||||
/// `Copy` because it carries no payload — see the module docs for why the
|
||||
/// panic payload is *not* included here.
|
||||
/// Carries no payload: see the module docs for why a monitor never receives
|
||||
/// the panic value itself.
|
||||
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
|
||||
pub enum DownReason {
|
||||
/// The target returned normally.
|
||||
@@ -62,6 +98,11 @@ pub enum DownReason {
|
||||
Panic,
|
||||
/// The target was cooperatively cancelled via `request_stop`.
|
||||
Stopped,
|
||||
/// A graceful shutdown was requested via `request_shutdown`. Only ever
|
||||
/// appears in an [`ExitSignal`](crate::link::ExitSignal) delivered to a
|
||||
/// trapping actor — never in a [`Down`]: a target that honours the request
|
||||
/// exits *normally*, one that does not trap is `Stopped`.
|
||||
Shutdown,
|
||||
/// The target was already gone (finished and reclaimed, or never alive)
|
||||
/// at the moment `monitor()` was called.
|
||||
NoProc,
|
||||
@@ -76,21 +117,22 @@ pub struct Down {
|
||||
pub reason: DownReason,
|
||||
}
|
||||
|
||||
/// A process-unique identifier for one `monitor()` registration.
|
||||
/// A unique identifier for one [`monitor`] registration.
|
||||
///
|
||||
/// Opaque and `Copy`. Allocated from a monotonic counter in shared state, so
|
||||
/// it is never reused for the lifetime of the runtime — distinct `monitor()`
|
||||
/// calls on the same target get distinct ids, which is what lets [`demonitor`]
|
||||
/// tear down exactly one of several monitors on a target.
|
||||
/// Opaque and `Copy`. Never reused for the life of the runtime, so if you
|
||||
/// monitor the same target more than once, each call's id is distinct. This
|
||||
/// is what lets [`demonitor`] tear down exactly one of several monitors on
|
||||
/// the same target without disturbing the others.
|
||||
#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash)]
|
||||
pub struct MonitorId(pub(crate) u64);
|
||||
|
||||
/// A live monitor: the receiving end of the one-shot [`Down`] channel, plus the
|
||||
/// identity needed to [`demonitor`] it.
|
||||
///
|
||||
/// Read the notification from [`Monitor::rx`]. Not `Clone` (the receiver is a
|
||||
/// single consumer). Dropping it closes the receiving end; if a `Down` was
|
||||
/// already queued it is discarded with the channel.
|
||||
/// Read the notification from [`Monitor::rx`]. Not `Clone`, since only one
|
||||
/// side is meant to consume it. Dropping a `Monitor` closes the receiving
|
||||
/// end; if a `Down` had already arrived but was never read, it is discarded
|
||||
/// along with it.
|
||||
pub struct Monitor {
|
||||
/// This registration's process-unique id.
|
||||
pub id: MonitorId,
|
||||
@@ -110,10 +152,11 @@ pub fn monitor<A>(target: Pid<A>) -> Monitor {
|
||||
let target = target.erase();
|
||||
let (tx, rx) = channel::<Down>();
|
||||
|
||||
// Register under the target's cold lock. `tx.clone()` takes the channel's
|
||||
// own lock — a Channel-class RawMutex, explicitly permitted *under* a Leaf
|
||||
// (cold) lock by the lock order (see raw_mutex.rs). We must still not
|
||||
// *send* under the lock, as `Sender::send` can unpark a parked receiver,
|
||||
// Implementation note: registration happens under the target's cold
|
||||
// lock. `tx.clone()` takes the channel's own lock, a Channel-class
|
||||
// RawMutex, which is explicitly permitted under a Leaf (cold) lock by
|
||||
// the lock order documented in raw_mutex.rs. We must still not *send*
|
||||
// under the lock, since `Sender::send` can unpark a parked receiver,
|
||||
// and there's no reason to nest that.
|
||||
let (id, registered) = with_runtime(|inner| {
|
||||
let id = inner.alloc_monitor_id();
|
||||
@@ -133,26 +176,89 @@ pub fn monitor<A>(target: Pid<A>) -> Monitor {
|
||||
});
|
||||
|
||||
if !registered {
|
||||
let _ = tx.send(Down { pid: target, reason: DownReason::NoProc });
|
||||
let _ = tx.send(Down {
|
||||
pid: target,
|
||||
reason: DownReason::NoProc,
|
||||
});
|
||||
}
|
||||
|
||||
Monitor { id, target, rx }
|
||||
}
|
||||
|
||||
/// Cancel the monitor `m`. Returns `Some(id)` if a live registration was found
|
||||
/// on the target's slot and removed, or `None` if there was nothing to remove
|
||||
/// — the target already fired its `Down` (the registration is drained on
|
||||
/// finalize), was never alive (`NoProc`), or has been reclaimed.
|
||||
/// Flag `target`'s tenancy as watchable: its death will stamp the slot's
|
||||
/// terminal record (see [`terminal_reason`]), exactly as registering a name
|
||||
/// does. The bridge calls this wherever a smarm pid is *encoded across the
|
||||
/// boundary* — a contract reply, an introspection listing — because BEAM can
|
||||
/// only watch pids it holds, and can only hold pids that crossed. Keeping the
|
||||
/// bit rare is what keeps the record alive: anonymous never-exported churn
|
||||
/// (holder threads, egress tasks) stays ineligible and cannot evict a
|
||||
/// watchable tenancy's record from a LIFO-recycled slot.
|
||||
///
|
||||
/// This stops any *future* `Down`. To also discard a `Down` the target may have
|
||||
/// *already* queued (the finalize-races-demonitor case), drop `m` afterwards;
|
||||
/// dropping the [`Monitor`] closes its receiver and the queued notice goes with
|
||||
/// it — the analogue of Erlang's `demonitor(Ref, [flush])`.
|
||||
/// Generation-checked and live-screened: marking a pid whose tenancy already
|
||||
/// ended is a no-op — its record either exists (it was flagged before dying)
|
||||
/// or is honestly unknowable. Same `Runtime::run()` context contract as
|
||||
/// [`monitor`].
|
||||
pub fn mark_watchable<A>(target: Pid<A>) {
|
||||
let target = target.erase();
|
||||
with_runtime(|inner| {
|
||||
if let Some(slot) = inner.slot_at(target) {
|
||||
// Cold lock FIRST: finalize publishes Done and checks the
|
||||
// watchable bit under this same lock, so the mark either lands
|
||||
// before finalize reads it (the death stamps) or observes the
|
||||
// tenancy already dead (no-op). No lost-stamp window between an
|
||||
// unlocked liveness read and the flag set.
|
||||
let mut cold = slot.cold.lock();
|
||||
if slot.is_live_for(target) {
|
||||
cold.watchable = true;
|
||||
}
|
||||
}
|
||||
});
|
||||
}
|
||||
|
||||
/// The terminal [`DownReason`] of the tenancy `target` names, if that tenancy
|
||||
/// ever registered a name and is the *most recent named* death of its slot:
|
||||
/// finalize stamps the slot with `(generation, reason)` for once-registered
|
||||
/// tenancies (anonymous green-thread churn does not stamp — nor evict), and
|
||||
/// the record survives reclaim and the next tenant's install, until the next
|
||||
/// *named* tenant of the slot itself dies. `None` means the pid never lived,
|
||||
/// is still alive, never held a name, or its record was overwritten by a
|
||||
/// later named tenancy's death — callers fall back to `NoProc` semantics.
|
||||
///
|
||||
/// This exists for watch-installers that raced their target's death (bridge
|
||||
/// soak signature 4): a `NoProc` observed at install time can be upgraded to
|
||||
/// the real reason while the record still matches, which is exactly what an
|
||||
/// install that had won the race would have delivered. It does NOT change
|
||||
/// [`monitor`]'s own semantics — monitoring a stale pid still queues `NoProc`,
|
||||
/// the same shape Erlang gives — the upgrade is the caller's deliberate act.
|
||||
/// Same context contract as [`monitor`]: must run inside `Runtime::run()`.
|
||||
pub fn terminal_reason<A>(target: Pid<A>) -> Option<DownReason> {
|
||||
let target = target.erase();
|
||||
with_runtime(|inner| {
|
||||
let slot = inner.slot_at(target)?;
|
||||
let cold = slot.cold.lock();
|
||||
match cold.terminal {
|
||||
Some((generation, reason)) if generation == target.generation() => Some(reason),
|
||||
_ => None,
|
||||
}
|
||||
})
|
||||
}
|
||||
|
||||
/// Cancel the monitor `m`. Returns `Some(id)` if a live registration was found
|
||||
/// and removed, so no `Down` will arrive on `m.rx` from here on. Returns
|
||||
/// `None` if there was nothing left to remove: the target had already gone
|
||||
/// down and its `Down` was already sent (or is already sitting in the
|
||||
/// channel, unread).
|
||||
///
|
||||
/// This only stops a *future* `Down`. If you also want to discard a `Down`
|
||||
/// that already arrived (or is about to, in a race with this call), drop `m`
|
||||
/// instead of, or in addition to, calling this: dropping the [`Monitor`]
|
||||
/// closes its receiver and any queued notice is discarded with it.
|
||||
pub fn demonitor(m: &Monitor) -> Option<MonitorId> {
|
||||
// Remove the registration under the target's cold lock, but move the
|
||||
// `Sender` *out* and let it drop only after the lock is released:
|
||||
// dropping the last sender runs `Sender::drop`, which may unpark a parked
|
||||
// receiver — legal under a cold lock, but pointless to nest.
|
||||
// Implementation note: the registration is removed under the target's
|
||||
// cold lock, but the `Sender` is moved *out* and dropped only after the
|
||||
// lock is released. Dropping the last sender runs `Sender::drop`, which
|
||||
// may unpark a parked receiver; legal under a cold lock, but pointless
|
||||
// to nest.
|
||||
let removed: Option<(MonitorId, Sender<Down>)> = with_runtime(|inner| {
|
||||
let slot = inner.slot_at(m.target)?;
|
||||
let mut cold = slot.cold.lock();
|
||||
|
||||
+149
-19
@@ -1,12 +1,89 @@
|
||||
//! Actor-aware mutex with mandatory timeout.
|
||||
//! Shared mutable state across actors, when a channel is overkill.
|
||||
//!
|
||||
//! `Mutex<T>` parks the calling *green* thread on contention rather than
|
||||
//! blocking the OS thread. Every lock attempt is bounded by a timeout.
|
||||
//! smarm actors normally coordinate by sending messages, and for a piece of
|
||||
//! owned state the right tool is usually a `gen_server`: one actor holds the
|
||||
//! data and everyone else talks to it. Sometimes that is more machinery than
|
||||
//! you need, and plain shared, lockable state is simpler: [`Mutex<T>`] is
|
||||
//! that escape hatch. It behaves like `std::sync::Mutex<T>`, guarding a value
|
||||
//! of type `T` behind a guard that gives you `&mut T` while held, but it is
|
||||
//! built for smarm's actors rather than OS threads.
|
||||
//!
|
||||
//! Internals use `Arc<std::sync::Mutex<...>>` so the type is genuinely
|
||||
//! `Send + Sync` and can be shared across scheduler threads.
|
||||
//! The key difference from `std::sync::Mutex` is what happens on contention.
|
||||
//! [`Mutex::lock`] parks the calling actor (a cooperatively scheduled green
|
||||
//! thread) rather than blocking the underlying OS thread, so other actors on
|
||||
//! the same OS thread keep running while it waits. And every lock attempt is
|
||||
//! bounded by a timeout: an actor that hangs on to the lock forever (stuck in
|
||||
//! a bug, or just slow) would otherwise wedge every other actor waiting on
|
||||
//! it, so smarm makes the wait bounded by default instead of leaving it up
|
||||
//! to you to remember.
|
||||
//!
|
||||
//! Fairness: FIFO. Poisoning: none. Reentrance: deadlock (caller bug).
|
||||
//! ## A first lock
|
||||
//!
|
||||
//! ```
|
||||
//! use smarm::{run, spawn, Mutex};
|
||||
//!
|
||||
//! run(|| {
|
||||
//! let counter = Mutex::new(0u32);
|
||||
//!
|
||||
//! // Mutex::clone() is cheap and hands out another handle to the SAME
|
||||
//! // underlying value, much like Arc::clone: every clone shares one lock
|
||||
//! // and one value, so mutations through one are visible through all.
|
||||
//! let a = counter.clone();
|
||||
//! let b = counter.clone();
|
||||
//!
|
||||
//! let h1 = spawn(move || {
|
||||
//! let mut guard = a.lock().unwrap();
|
||||
//! *guard += 1;
|
||||
//! });
|
||||
//! let h2 = spawn(move || {
|
||||
//! let mut guard = b.lock().unwrap();
|
||||
//! *guard += 1;
|
||||
//! });
|
||||
//! h1.join().unwrap();
|
||||
//! h2.join().unwrap();
|
||||
//!
|
||||
//! assert_eq!(*counter.lock().unwrap(), 2);
|
||||
//! });
|
||||
//! ```
|
||||
//!
|
||||
//! ## Choosing a timeout
|
||||
//!
|
||||
//! [`Mutex::lock`] waits up to [`DEFAULT_TIMEOUT`] (30 seconds) before giving
|
||||
//! up with [`LockTimeout`]. To use a different bound for one call, use
|
||||
//! [`Mutex::lock_timeout`] instead; to change the default for every future
|
||||
//! `lock()` call on this mutex (including through its clones), use
|
||||
//! [`Mutex::set_default_timeout`]. If you never want to wait at all, use
|
||||
//! [`Mutex::try_lock`], which returns immediately whether or not the lock was
|
||||
//! free.
|
||||
//!
|
||||
//! ## Fairness and panics
|
||||
//!
|
||||
//! Waiters are granted the lock in the order they started waiting (FIFO), so
|
||||
//! no actor can be starved by later arrivals repeatedly cutting in line.
|
||||
//!
|
||||
//! This mutex never poisons. `std::sync::Mutex` marks itself poisoned if a
|
||||
//! thread panics while holding the lock, because a partly mutated value might
|
||||
//! be left behind for the next lock holder to see. smarm's actors already
|
||||
//! rely on `Drop` running during unwinding to release the lock, so if a
|
||||
//! holder panics, [`MutexGuard::drop`] still runs and the next waiter is
|
||||
//! granted the lock normally. It is the same tradeoff `std::sync::Mutex`
|
||||
//! offers you if you choose to ignore poisoning: you may see a value left
|
||||
//! mid-update by the panicking actor, so a panic inside a critical section is
|
||||
//! still a bug worth fixing, just not one that also wedges every future lock
|
||||
//! attempt.
|
||||
//!
|
||||
//! Locking a mutex you already hold (on the same actor) does not queue
|
||||
//! behind yourself: it deadlocks, the same way relocking a non-reentrant
|
||||
//! `std::sync::Mutex` does. Don't call `lock` while already holding a guard
|
||||
//! from the same `Mutex`.
|
||||
//!
|
||||
//! ## Outside the runtime
|
||||
//!
|
||||
//! `Mutex<T>` also works when called from plain code that is not running as
|
||||
//! a smarm actor (for example, in a test's setup code before calling
|
||||
//! [`run`](crate::run)). There, an actor's cooperative park has no meaning,
|
||||
//! so a lock attempt instead blocks the calling OS thread directly until the
|
||||
//! mutex is free; there is no timeout on this path.
|
||||
|
||||
use crate::pid::Pid;
|
||||
use crate::scheduler;
|
||||
@@ -15,8 +92,14 @@ use std::collections::VecDeque;
|
||||
use std::sync::{Arc, Mutex as StdMutex};
|
||||
use std::time::Duration;
|
||||
|
||||
/// How long [`Mutex::lock`] waits for the lock before giving up, unless
|
||||
/// overridden per-mutex with [`Mutex::set_default_timeout`] or per-call with
|
||||
/// [`Mutex::lock_timeout`].
|
||||
pub const DEFAULT_TIMEOUT: Duration = Duration::from_secs(30);
|
||||
|
||||
/// Returned by [`Mutex::lock`] / [`Mutex::lock_timeout`] when the timeout
|
||||
/// elapses before the lock became available. The lock attempt is abandoned;
|
||||
/// nothing was acquired, and the mutex's value is unaffected.
|
||||
#[derive(Debug, PartialEq, Eq, Clone, Copy)]
|
||||
pub struct LockTimeout;
|
||||
|
||||
@@ -70,12 +153,16 @@ impl TimerTarget for MutexCore {
|
||||
};
|
||||
// Remove from waiters only if still there with matching epoch.
|
||||
// If the lock was already granted (holder == Some(pid)), the
|
||||
// timer fired after the grant — treat as no-op; the actor
|
||||
// timer fired after the grant: treat as no-op; the actor
|
||||
// will see `is_holder == true` and return Ok.
|
||||
if st.holder == Some(pid) {
|
||||
return;
|
||||
}
|
||||
match st.waiters.iter().position(|w| w.pid == pid && w.epoch == epoch) {
|
||||
match st
|
||||
.waiters
|
||||
.iter()
|
||||
.position(|w| w.pid == pid && w.epoch == epoch)
|
||||
{
|
||||
Some(pos) => {
|
||||
st.waiters.remove(pos);
|
||||
true
|
||||
@@ -100,6 +187,8 @@ pub struct Mutex<T> {
|
||||
}
|
||||
|
||||
impl<T> Mutex<T> {
|
||||
/// Wrap `value` in a new mutex, initially unlocked, with the default
|
||||
/// lock timeout ([`DEFAULT_TIMEOUT`]).
|
||||
pub fn new(value: T) -> Self {
|
||||
Self {
|
||||
core: Arc::new(MutexCore::new(DEFAULT_TIMEOUT)),
|
||||
@@ -107,6 +196,11 @@ impl<T> Mutex<T> {
|
||||
}
|
||||
}
|
||||
|
||||
/// Change how long future [`lock`](Self::lock) calls on this mutex wait
|
||||
/// before giving up. Applies to every clone of this `Mutex` (they share
|
||||
/// one underlying lock), and to `lock` calls already in progress that
|
||||
/// have not yet started waiting. Does not affect [`lock_timeout`](Self::lock_timeout)
|
||||
/// calls, which always use the timeout passed in.
|
||||
pub fn set_default_timeout(&self, timeout: Duration) {
|
||||
match self.core.state.lock() {
|
||||
Ok(mut st) => st.default_timeout = timeout,
|
||||
@@ -114,6 +208,12 @@ impl<T> Mutex<T> {
|
||||
}
|
||||
}
|
||||
|
||||
/// Acquire the lock, waiting up to this mutex's default timeout
|
||||
/// ([`DEFAULT_TIMEOUT`], or whatever [`set_default_timeout`](Self::set_default_timeout)
|
||||
/// last set) if it is currently held elsewhere. Returns a [`MutexGuard`]
|
||||
/// that releases the lock when dropped, or [`LockTimeout`] if the
|
||||
/// deadline passes first. To use a one-off timeout instead of the
|
||||
/// mutex's default, call [`lock_timeout`](Self::lock_timeout) directly.
|
||||
pub fn lock(&self) -> Result<MutexGuard<'_, T>, LockTimeout> {
|
||||
let timeout = match self.core.state.lock() {
|
||||
Ok(st) => st.default_timeout,
|
||||
@@ -122,6 +222,10 @@ impl<T> Mutex<T> {
|
||||
self.lock_timeout(timeout)
|
||||
}
|
||||
|
||||
/// Acquire the lock, waiting up to `timeout` (ignoring this mutex's
|
||||
/// default) if it is currently held elsewhere. Returns a [`MutexGuard`]
|
||||
/// that releases the lock when dropped, or [`LockTimeout`] if `timeout`
|
||||
/// elapses first with the lock still unavailable.
|
||||
pub fn lock_timeout(&self, timeout: Duration) -> Result<MutexGuard<'_, T>, LockTimeout> {
|
||||
// Outside the runtime (e.g. in tests, after run() returns) there is no
|
||||
// current actor PID. Fall back to a blocking std::sync::Mutex acquire.
|
||||
@@ -146,7 +250,10 @@ impl<T> Mutex<T> {
|
||||
Some(v) => v,
|
||||
None => panic!("smarm: Mutex value missing on free fast path (core corrupt)"),
|
||||
};
|
||||
return Ok(MutexGuard { mutex: self, value: Some(value) });
|
||||
return Ok(MutexGuard {
|
||||
mutex: self,
|
||||
value: Some(value),
|
||||
});
|
||||
}
|
||||
}
|
||||
|
||||
@@ -157,7 +264,7 @@ impl<T> Mutex<T> {
|
||||
Ok(g) => g,
|
||||
Err(e) => panic!("smarm: mutex state lock poisoned (core corrupt): {e}"),
|
||||
};
|
||||
// begin_wait is lock-free — legal under the state lock; this
|
||||
// begin_wait is lock-free (legal under the state lock); this
|
||||
// makes the epoch atomic with the registration's visibility to
|
||||
// grants and timeouts.
|
||||
let epoch = scheduler::begin_wait();
|
||||
@@ -170,7 +277,7 @@ impl<T> Mutex<T> {
|
||||
scheduler::insert_wait_timer(deadline, me, target, epoch);
|
||||
scheduler::park_current();
|
||||
|
||||
// Resumed — precisely: only our grant or our timer can wake this
|
||||
// Resumed, precisely: only our grant or our timer can wake this
|
||||
// wait (both epoch-stamped; a stop wake unwinds out of
|
||||
// park_current). The one-shot interpretation below is therefore
|
||||
// exhaustive. Are we the holder?
|
||||
@@ -187,12 +294,18 @@ impl<T> Mutex<T> {
|
||||
Some(v) => v,
|
||||
None => panic!("smarm: Mutex value missing after grant (core corrupt)"),
|
||||
};
|
||||
Ok(MutexGuard { mutex: self, value: Some(value) })
|
||||
Ok(MutexGuard {
|
||||
mutex: self,
|
||||
value: Some(value),
|
||||
})
|
||||
} else {
|
||||
Err(LockTimeout)
|
||||
}
|
||||
}
|
||||
|
||||
/// Acquire the lock only if it is immediately available: never parks and
|
||||
/// never waits. Returns `Some` with a [`MutexGuard`] if the lock was
|
||||
/// free, `None` if it is currently held elsewhere.
|
||||
pub fn try_lock(&self) -> Option<MutexGuard<'_, T>> {
|
||||
let me = crate::actor::current_pid()?;
|
||||
let mut st = match self.core.state.lock() {
|
||||
@@ -212,7 +325,10 @@ impl<T> Mutex<T> {
|
||||
Some(v) => v,
|
||||
None => panic!("smarm: Mutex value missing on try_lock free path (core corrupt)"),
|
||||
};
|
||||
Some(MutexGuard { mutex: self, value: Some(value) })
|
||||
Some(MutexGuard {
|
||||
mutex: self,
|
||||
value: Some(value),
|
||||
})
|
||||
}
|
||||
|
||||
/// Blocking fallback used when called outside the smarm runtime.
|
||||
@@ -226,16 +342,28 @@ impl<T> Mutex<T> {
|
||||
Ok(mut g) => g.take(),
|
||||
Err(e) => panic!("smarm: mutex value lock poisoned (core corrupt): {e}"),
|
||||
};
|
||||
if let Some(v) = v { break v; }
|
||||
if let Some(v) = v {
|
||||
break v;
|
||||
}
|
||||
std::thread::yield_now();
|
||||
};
|
||||
Ok(MutexGuard { mutex: self, value: Some(value) })
|
||||
Ok(MutexGuard {
|
||||
mutex: self,
|
||||
value: Some(value),
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
impl<T> Clone for Mutex<T> {
|
||||
/// Cheap: hands back another handle to the same underlying lock and
|
||||
/// value, the way `Arc::clone` does. All clones of a `Mutex` share one
|
||||
/// lock and one protected value; locking through any clone excludes
|
||||
/// every other clone.
|
||||
fn clone(&self) -> Self {
|
||||
Self { core: self.core.clone(), value: self.value.clone() }
|
||||
Self {
|
||||
core: self.core.clone(),
|
||||
value: self.value.clone(),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -247,6 +375,10 @@ unsafe impl<T: Send> Sync for Mutex<T> {}
|
||||
// Guard
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/// Grants access to the value inside a [`Mutex`] while the lock is held.
|
||||
/// Dereferences to `&T` and `&mut T`. Dropping the guard releases the lock
|
||||
/// and, if another actor is waiting, wakes the next one in arrival order.
|
||||
/// Returned by [`Mutex::lock`], [`Mutex::lock_timeout`], and [`Mutex::try_lock`].
|
||||
pub struct MutexGuard<'a, T> {
|
||||
mutex: &'a Mutex<T>,
|
||||
value: Option<T>,
|
||||
@@ -277,9 +409,7 @@ impl<T: std::fmt::Debug> std::fmt::Debug for MutexGuard<'_, T> {
|
||||
Some(v) => v,
|
||||
None => panic!("smarm: MutexGuard value missing (core corrupt)"),
|
||||
};
|
||||
f.debug_tuple("MutexGuard")
|
||||
.field(value)
|
||||
.finish()
|
||||
f.debug_tuple("MutexGuard").field(value).finish()
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
+1017
File diff suppressed because it is too large
Load Diff
@@ -221,7 +221,9 @@ pub(crate) struct ProcessGroups {
|
||||
|
||||
impl ProcessGroups {
|
||||
pub(crate) fn new() -> Self {
|
||||
Self { groups: HashMap::new() }
|
||||
Self {
|
||||
groups: HashMap::new(),
|
||||
}
|
||||
}
|
||||
|
||||
/// Insert `ms` into `group`. Idempotent on the *member*: if the member is
|
||||
@@ -323,20 +325,33 @@ impl ProcessGroups {
|
||||
fn members_where(&self, group: &str, mut is_live: impl FnMut(Pid) -> bool) -> Vec<Pid> {
|
||||
self.groups
|
||||
.get(group)
|
||||
.map(|v| v.iter().map(|e| e.member.pid).filter(|&p| is_live(p)).collect())
|
||||
.map(|v| {
|
||||
v.iter()
|
||||
.map(|e| e.member.pid)
|
||||
.filter(|&p| is_live(p))
|
||||
.collect()
|
||||
})
|
||||
.unwrap_or_default()
|
||||
}
|
||||
|
||||
/// The first live member of `group` in insertion order — stateless
|
||||
/// first-live `pick`, with the same read-path backstop as `members_where`.
|
||||
fn first_member_where(&self, group: &str, mut is_live: impl FnMut(Pid) -> bool) -> Option<Pid> {
|
||||
self.groups.get(group)?.iter().map(|e| e.member.pid).find(|&p| is_live(p))
|
||||
self.groups
|
||||
.get(group)?
|
||||
.iter()
|
||||
.map(|e| e.member.pid)
|
||||
.find(|&p| is_live(p))
|
||||
}
|
||||
}
|
||||
|
||||
/// Build the full member identity for `pid` from runtime identity.
|
||||
fn member_for(inner: &crate::runtime::RuntimeInner, pid: Pid) -> Member {
|
||||
Member { node: inner.node_id, incarnation: inner.incarnation, pid }
|
||||
Member {
|
||||
node: inner.node_id,
|
||||
incarnation: inner.incarnation,
|
||||
pid,
|
||||
}
|
||||
}
|
||||
|
||||
/// Is `pid` a live actor right now? Generation-checked atomic slot-word read,
|
||||
@@ -367,7 +382,10 @@ pub fn join<A>(group: impl Into<String>, pid: Pid<A>) -> bool {
|
||||
let mon = monitor(pid);
|
||||
|
||||
let (rejected, reaped) = with_runtime(|inner| {
|
||||
let ms = Membership { member: member_for(inner, pid), monitor: mon };
|
||||
let ms = Membership {
|
||||
member: member_for(inner, pid),
|
||||
monitor: mon,
|
||||
};
|
||||
let mut pg = inner.process_groups.lock();
|
||||
let reaped = pg.reap_group(&group);
|
||||
let rejected = pg.join(&group, ms);
|
||||
@@ -507,7 +525,11 @@ mod tests {
|
||||
let (tx, rx) = channel::<Down>();
|
||||
let ms = Membership {
|
||||
member: member(index, generation),
|
||||
monitor: Monitor { id: MonitorId(0), target: pid, rx },
|
||||
monitor: Monitor {
|
||||
id: MonitorId(0),
|
||||
target: pid,
|
||||
rx,
|
||||
},
|
||||
};
|
||||
(ms, tx)
|
||||
}
|
||||
@@ -518,7 +540,10 @@ mod tests {
|
||||
let (a, _ta) = synth(1, 0);
|
||||
let (b, _tb) = synth(1, 0);
|
||||
assert!(pg.join("workers", a).is_none(), "first join inserts");
|
||||
assert!(pg.join("workers", b).is_some(), "second identical join is handed back");
|
||||
assert!(
|
||||
pg.join("workers", b).is_some(),
|
||||
"second identical join is handed back"
|
||||
);
|
||||
assert_eq!(pg.members_of("workers"), vec![member(1, 0)]);
|
||||
}
|
||||
|
||||
@@ -542,7 +567,10 @@ mod tests {
|
||||
let (a, _ta) = synth(1, 0);
|
||||
let (b, _tb) = synth(1, 1);
|
||||
assert!(pg.join("g", a).is_none());
|
||||
assert!(pg.join("g", b).is_none(), "different generation is a distinct member");
|
||||
assert!(
|
||||
pg.join("g", b).is_none(),
|
||||
"different generation is a distinct member"
|
||||
);
|
||||
assert_eq!(pg.members_of("g"), vec![member(1, 0), member(1, 1)]);
|
||||
}
|
||||
|
||||
@@ -555,16 +583,27 @@ mod tests {
|
||||
pg.join("g", b);
|
||||
assert!(pg.leave("g", member(1, 0)).is_some());
|
||||
assert_eq!(pg.members_of("g"), vec![member(2, 0)]);
|
||||
assert!(pg.leave("g", member(1, 0)).is_none(), "second leave finds nothing");
|
||||
assert!(
|
||||
pg.leave("g", member(1, 0)).is_none(),
|
||||
"second leave finds nothing"
|
||||
);
|
||||
assert!(pg.leave("g", member(2, 0)).is_some());
|
||||
assert!(pg.members_of("g").is_empty(), "group is now empty");
|
||||
assert!(pg.leave("never", member(9, 0)).is_none(), "leaving an unknown group is a no-op");
|
||||
assert!(
|
||||
pg.leave("never", member(9, 0)).is_none(),
|
||||
"leaving an unknown group is a no-op"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn remove_where_sweeps_every_group() {
|
||||
let mut pg = ProcessGroups::new();
|
||||
for (g, (m, _t)) in [("a", synth(1, 0)), ("a", synth(2, 0)), ("b", synth(1, 0)), ("c", synth(3, 0))] {
|
||||
for (g, (m, _t)) in [
|
||||
("a", synth(1, 0)),
|
||||
("a", synth(2, 0)),
|
||||
("b", synth(1, 0)),
|
||||
("c", synth(3, 0)),
|
||||
] {
|
||||
pg.join(g, m);
|
||||
}
|
||||
// Death of pid index 1 (any generation) evicts it everywhere.
|
||||
@@ -582,8 +621,16 @@ mod tests {
|
||||
let pid = Pid::new(1, 0);
|
||||
let (tx, rx) = channel::<Down>();
|
||||
let dead = Membership {
|
||||
member: Member { node: DEFAULT_NODE_ID, incarnation: Incarnation::new(7), pid },
|
||||
monitor: Monitor { id: MonitorId(0), target: pid, rx },
|
||||
member: Member {
|
||||
node: DEFAULT_NODE_ID,
|
||||
incarnation: Incarnation::new(7),
|
||||
pid,
|
||||
},
|
||||
monitor: Monitor {
|
||||
id: MonitorId(0),
|
||||
target: pid,
|
||||
rx,
|
||||
},
|
||||
};
|
||||
let _keep = tx;
|
||||
let (live, _tl) = synth(2, 0);
|
||||
@@ -614,9 +661,17 @@ mod tests {
|
||||
pg.join("b", b1);
|
||||
// pid 1 dies: its group-a monitor receives a Down. Its group-b monitor
|
||||
// has not — reap must still sweep pid 1 out of b by the pid predicate.
|
||||
ta1.send(Down { pid: Pid::new(1, 0), reason: DownReason::Exit }).unwrap();
|
||||
ta1.send(Down {
|
||||
pid: Pid::new(1, 0),
|
||||
reason: DownReason::Exit,
|
||||
})
|
||||
.unwrap();
|
||||
let evicted = pg.reap_group("a");
|
||||
assert_eq!(evicted.len(), 2, "pid 1's memberships in both a and b are evicted");
|
||||
assert_eq!(
|
||||
evicted.len(),
|
||||
2,
|
||||
"pid 1's memberships in both a and b are evicted"
|
||||
);
|
||||
assert_eq!(pg.members_of("a"), vec![member(2, 0)]);
|
||||
assert!(pg.members_of("b").is_empty(), "swept from b too; pruned");
|
||||
}
|
||||
@@ -646,8 +701,16 @@ mod tests {
|
||||
let dead = Pid::new(1, 0);
|
||||
let oracle = |pid: Pid| pid != dead;
|
||||
|
||||
assert_eq!(pg.members_where("g", oracle), vec![Pid::new(2, 0)], "dead pid filtered from read");
|
||||
assert_eq!(pg.first_member_where("g", oracle), Some(Pid::new(2, 0)), "pick skips the dead first member");
|
||||
assert_eq!(
|
||||
pg.members_where("g", oracle),
|
||||
vec![Pid::new(2, 0)],
|
||||
"dead pid filtered from read"
|
||||
);
|
||||
assert_eq!(
|
||||
pg.first_member_where("g", oracle),
|
||||
Some(Pid::new(2, 0)),
|
||||
"pick skips the dead first member"
|
||||
);
|
||||
|
||||
// Backstop does not evict — that stays the monitor's job; raw storage
|
||||
// still holds both until reap runs.
|
||||
|
||||
+12
-3
@@ -79,7 +79,10 @@ impl Pid<Erased> {
|
||||
/// here; typing happens at typed-actor boundaries via [`Pid::from_raw`].
|
||||
#[inline]
|
||||
pub const fn new(index: u32, generation: u32) -> Self {
|
||||
Self { raw: RawPid::new(index, generation), _marker: PhantomData }
|
||||
Self {
|
||||
raw: RawPid::new(index, generation),
|
||||
_marker: PhantomData,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -90,7 +93,10 @@ impl<A> Pid<A> {
|
||||
/// resolution paths.
|
||||
#[inline]
|
||||
pub(crate) const fn from_raw(raw: RawPid) -> Self {
|
||||
Self { raw, _marker: PhantomData }
|
||||
Self {
|
||||
raw,
|
||||
_marker: PhantomData,
|
||||
}
|
||||
}
|
||||
|
||||
/// The raw identity, dropping the actor type — the key for identity-only
|
||||
@@ -192,7 +198,10 @@ impl<M> Name<M> {
|
||||
/// associated constants at call sites.
|
||||
#[inline]
|
||||
pub const fn new(name: &'static str) -> Self {
|
||||
Self { name, _marker: PhantomData }
|
||||
Self {
|
||||
name,
|
||||
_marker: PhantomData,
|
||||
}
|
||||
}
|
||||
|
||||
/// The underlying registry key.
|
||||
|
||||
+7
-4
@@ -98,10 +98,13 @@ pub(crate) fn clear_current_slot() {
|
||||
CURRENT_SLOT.with(|c| c.set(std::ptr::null()));
|
||||
}
|
||||
|
||||
/// RFC 007 (`smarm-causal`) — raw pointer to the on-CPU actor's slot, null on
|
||||
/// the scheduler's own stack. Same lifetime argument as `note_overrun`: the
|
||||
/// slot is never reclaimed while its actor is on-CPU.
|
||||
#[cfg(feature = "smarm-causal")]
|
||||
/// Raw pointer to the on-CPU actor's slot, null on the scheduler's own
|
||||
/// stack. Same lifetime argument as `note_overrun`: the slot is never
|
||||
/// reclaimed while its actor is on-CPU. Consumers: the `smarm-causal`
|
||||
/// profiler (RFC 007) and — unconditionally — the SIGSEGV classifier
|
||||
/// (RFC 019 §7), which additionally relies on this being a plain load of a
|
||||
/// const-initialized TLS Cell (no lazy init, no allocation, no dtor): safe
|
||||
/// from a signal handler.
|
||||
#[inline]
|
||||
pub(crate) fn current_slot_ptr() -> *const crate::runtime::Slot {
|
||||
CURRENT_SLOT.with(|c| c.get())
|
||||
|
||||
+4
-1
@@ -166,7 +166,10 @@ impl<T> RawMutex<T> {
|
||||
{
|
||||
self.lock_slow();
|
||||
}
|
||||
RawMutexGuard { m: self, prev_preempt }
|
||||
RawMutexGuard {
|
||||
m: self,
|
||||
prev_preempt,
|
||||
}
|
||||
}
|
||||
|
||||
#[cold]
|
||||
|
||||
+348
-181
@@ -1,55 +1,107 @@
|
||||
//! Named mailbox registry — resolve a name (or pid) to a *messageable* actor.
|
||||
//! Give an actor a name so other actors can find it and message it.
|
||||
//!
|
||||
//! ## What changed (RFC 014)
|
||||
//! Without the registry, the only way to reach an actor is to already be
|
||||
//! holding its [`Pid`], usually because you spawned it yourself or someone
|
||||
//! passed it to you. That is fine for a worker you just created, but it does
|
||||
//! not work for a well-known service that arbitrary parts of your program
|
||||
//! need to find independently, like a logger, a config store, or a
|
||||
//! connection pool. The registry solves this: an actor claims a name once,
|
||||
//! and from then on any other actor can look that name up, or send to it
|
||||
//! directly, without ever having been handed a `Pid`.
|
||||
//!
|
||||
//! The old registry was a `name <-> pid` bimap: `whereis` handed back a `Pid`
|
||||
//! you could not send to, because a pid is just `(index, generation)` with no
|
||||
//! delivery endpoint. This rework makes resolution yield something messageable.
|
||||
//! ```
|
||||
//! use smarm::{channel, register, run, send, spawn, unregister, whereis, Name};
|
||||
//!
|
||||
//! Two facts shape the structure:
|
||||
//! const COUNTER: Name<u64> = Name::new("counter");
|
||||
//!
|
||||
//! 1. **A name resolves to a single actor.** Many actors under one label is
|
||||
//! what *process groups* (`pg`) are for; the registry is one-name-one-actor
|
||||
//! (several names *may* point at the same actor).
|
||||
//! 2. **Channels are typed**, so an actor has no single untyped mailbox. An
|
||||
//! actor instead owns a *set* of typed channels — one [`Sender`] per message
|
||||
//! type it accepts. So the registry maps name/pid to a [`Mailbox`]: a small
|
||||
//! structure holding that actor's pid plus all of its typed channels, keyed
|
||||
//! by message [`TypeId`].
|
||||
//! run(|| {
|
||||
//! let (ready_tx, ready_rx) = channel::<()>();
|
||||
//! let (tx, rx) = channel::<u64>();
|
||||
//!
|
||||
//! Resolution is therefore: `name -> pid` (single actor) `-> Mailbox -> the
|
||||
//! channel for message type M`. A `Name<Cmd>` and a `Name<Admin>` on the *same*
|
||||
//! actor select *different* channels purely by their type parameter, so
|
||||
//! capability separation (RFC 014 §4.7) needs no extra machinery.
|
||||
//! let worker = spawn(move || {
|
||||
//! // Claim the name for this actor's inbox. Any actor holding
|
||||
//! // `COUNTER` can now reach this one by name.
|
||||
//! register(COUNTER, tx).unwrap();
|
||||
//! ready_tx.send(()).unwrap();
|
||||
//! assert_eq!(rx.recv().unwrap(), 42);
|
||||
//! });
|
||||
//!
|
||||
//! ## Type erasure is contained
|
||||
//! ready_rx.recv().unwrap(); // wait for the worker to register
|
||||
//!
|
||||
//! Each stored channel is a `Box<dyn Any + Send>` that is concretely a
|
||||
//! `Sender<M>`, filed under `TypeId::of::<M>()`. A resolve for `M` looks up
|
||||
//! that exact `TypeId` and downcasts to `Sender<M>` — keyed by the very type we
|
||||
//! downcast to, so the downcast cannot fail on correct data; a failure is a
|
||||
//! smarm bug, asserted in debug. The phantom `M` on [`Name`] re-imposes the
|
||||
//! type at the call site, so callers never touch the erasure.
|
||||
//! // Look the name up, or just send to it directly.
|
||||
//! assert_eq!(whereis("counter"), Some(worker.pid()));
|
||||
//! send(COUNTER, 42).unwrap();
|
||||
//!
|
||||
//! ## Cleanup is lazy (prune-on-contact)
|
||||
//! worker.join().unwrap();
|
||||
//!
|
||||
//! As before, there is no `finalize` hook and no name field on the slot. Every
|
||||
//! operation that touches a binding checks the target pid's liveness via the
|
||||
//! generation-checked slot word; a binding to a dead actor behaves as absent
|
||||
//! and is pruned on contact (its [`Mailbox`] and every name pointing at it are
|
||||
//! dropped). The cost is a dead binding lingering until something looks at it;
|
||||
//! the payoff is zero coupling to the actor lifecycle.
|
||||
//! // The name dies with the actor: nobody holds it anymore.
|
||||
//! assert_eq!(whereis("counter"), None);
|
||||
//! });
|
||||
//! ```
|
||||
//!
|
||||
//! ## Locking
|
||||
//! ## Names carry a message type
|
||||
//!
|
||||
//! One `RawMutex` (Leaf class) in `RuntimeInner`, exactly like the old
|
||||
//! registry. The fold (name index *and* handles under the one lock) is what
|
||||
//! keeps a name-addressed `send` on a single Leaf — `raw_mutex` panics on a
|
||||
//! second Leaf acquired while one is held. The send path clones the `Sender`
|
||||
//! **under** the Leaf lock (a `Sender::clone` takes a Channel lock, permitted
|
||||
//! under a Leaf), then **releases** the Leaf and only *then* sends — a send can
|
||||
//! unpark a receiver, and wakeup-bearing work runs outside the Leaf. Order is
|
||||
//! **Leaf -> Channel**, as `pg`/`finalize`.
|
||||
//! A [`Name<M>`] is a plain string plus a type parameter `M`: the message
|
||||
//! type that name expects to receive. [`Name::new`] is `const`, so the usual
|
||||
//! pattern is a module-level constant like `COUNTER` above, shared by every
|
||||
//! caller. The type parameter means a name is only ever sent the kind of
|
||||
//! message it was declared for. If two different constants share the same
|
||||
//! string but have different message types, they still address two
|
||||
//! independent channels on the same actor: registering both just gives that
|
||||
//! actor two ways to be reached, one per message type. This is how you give
|
||||
//! one actor a "public" channel and a separate, differently-typed "admin"
|
||||
//! channel under related names, without inventing an enum to merge them.
|
||||
//!
|
||||
//! ## One actor per name, looked up fresh every time
|
||||
//!
|
||||
//! A name always points at exactly one actor at a time (contrast a *process
|
||||
//! group*, from the [`pg`](crate::pg) module, which is one name mapping to
|
||||
//! many actors). Unlike a plain [`Pid`], which names one specific actor
|
||||
//! forever and stops working the moment that actor dies, a name is
|
||||
//! re-resolved on every [`send`]: if the actor holding it dies and a new one
|
||||
//! registers under the same name, the next `send` reaches the new holder
|
||||
//! automatically. Use a name for a long-lived service whose exact identity
|
||||
//! you do not want to track by hand; use a `Pid` when you already have one
|
||||
//! and want to talk to that exact actor.
|
||||
//!
|
||||
//! ## Registration ends when the actor does
|
||||
//!
|
||||
//! There is no separate step to clean up a name when its actor exits: dying
|
||||
//! is enough. The next operation that touches a dead binding (a [`whereis`],
|
||||
//! a [`send`], or another actor's [`register`] of the same name) notices the
|
||||
//! actor is gone and clears the stale entry as a side effect, so the name
|
||||
//! becomes free again. [`unregister`] is only for a live actor voluntarily
|
||||
//! giving up a name it no longer wants; nothing has to call it on the way
|
||||
//! out.
|
||||
//!
|
||||
//! ## Implementation notes
|
||||
//!
|
||||
//! These details matter if you are working on smarm itself; they are not
|
||||
//! part of the public contract.
|
||||
//!
|
||||
//! Internally, each live actor that has published at least one channel owns
|
||||
//! a `Mailbox`: its pid plus a set of typed channels, keyed by the message
|
||||
//! type's `TypeId`. A stored channel is a `Box<dyn Any + Send>` that
|
||||
//! is concretely a `Sender<M>`; resolving for `M` looks up that exact
|
||||
//! `TypeId` and downcasts, so the downcast cannot fail on correct data (a
|
||||
//! failure would be a bug in the registry itself, checked in debug builds).
|
||||
//! Registering a name therefore means: find or create the actor's mailbox,
|
||||
//! insert the channel under its type, and point the name at the actor's pid.
|
||||
//!
|
||||
//! There is no callback when an actor exits. Every operation that touches a
|
||||
//! binding checks the target pid's liveness directly against the scheduler's
|
||||
//! slot table (which also tracks a generation counter, so a dead actor's
|
||||
//! reused slot index is never mistaken for the same actor). A binding to a
|
||||
//! dead actor is treated as absent and dropped right there. This keeps the
|
||||
//! registry decoupled from actor teardown, at the cost of a dead binding
|
||||
//! lingering until something happens to look at it.
|
||||
//!
|
||||
//! The whole registry (both the name index and the per-actor mailboxes) sits
|
||||
//! behind one lock, which is what lets a name-addressed [`send`] resolve and
|
||||
//! clone the target's sender in a single critical section. The sender is
|
||||
//! cloned while that lock is held, then the lock is released before the
|
||||
//! actual send, since delivering a message can wake a parked receiver and
|
||||
//! that wakeup work should not run while the registry is locked.
|
||||
|
||||
use crate::channel::Sender;
|
||||
use crate::pid::{Addressable, Name, Pid};
|
||||
@@ -80,28 +132,33 @@ impl std::fmt::Display for RegisterError {
|
||||
|
||||
impl std::error::Error for RegisterError {}
|
||||
|
||||
/// Why a name-addressed [`send`] did not deliver. Carries the message back so
|
||||
/// the caller never loses it (mirrors [`crate::channel::SendError`]).
|
||||
/// Why a send did not deliver. Every variant carries the undelivered message
|
||||
/// back, mirroring [`crate::channel::SendError`], so a failed send never
|
||||
/// silently drops what you tried to send.
|
||||
///
|
||||
/// `Debug`/`Display` are hand-written so neither demands `M: Debug` — the
|
||||
/// payload is returned, not printed.
|
||||
/// `Debug` and `Display` are hand-written so neither requires `M: Debug`,
|
||||
/// since the payload is handed back to you, not printed.
|
||||
pub enum SendError<M> {
|
||||
/// No live actor is currently registered under this name. Name-addressed
|
||||
/// [`send`] only; the pid-addressed counterpart is [`SendError::Dead`].
|
||||
/// No live actor is currently registered under this name. Returned only
|
||||
/// by name-addressed [`send`]; the pid-addressed counterpart of "nothing
|
||||
/// there" is [`SendError::Dead`].
|
||||
Unresolved(M),
|
||||
/// The pid-addressed actor is no longer the live incarnation this pid names
|
||||
/// — it has died, even if its slot now holds a *different* actor (a direct
|
||||
/// `Pid<A>` send never redirects; contrast name-addressed [`send`]). Pid
|
||||
/// paths ([`send_to`] / [`send_dyn`]) only.
|
||||
/// The actor this pid identifies has died, even if its slot has since
|
||||
/// been taken over by a different, live actor. A direct `Pid<A>` send
|
||||
/// never redirects to that new occupant; contrast name-addressed
|
||||
/// [`send`], which would reach it. Returned by the pid-addressed sends,
|
||||
/// [`send_to`] and [`send_dyn`].
|
||||
Dead(M),
|
||||
/// The actor is live but exposes no channel for this message type.
|
||||
/// The actor is live but has not published a channel for this message
|
||||
/// type.
|
||||
NoChannel(M),
|
||||
/// The actor's channel for this message type is closed (its receiver is gone).
|
||||
/// The actor's channel for this message type is closed (its receiver has
|
||||
/// been dropped).
|
||||
Closed(M),
|
||||
/// No live member to deliver to — a [`dispatch`](crate::dispatch) over an
|
||||
/// empty (or all-dead) process group. Group-addressed dispatch only; the
|
||||
/// name-addressed counterpart is [`SendError::Unresolved`]. The message is
|
||||
/// handed back undelivered.
|
||||
/// No live member was available to deliver to: returned by
|
||||
/// [`dispatch`](crate::dispatch) when the target process group is empty
|
||||
/// or every member in it has died. The name-addressed counterpart of
|
||||
/// this case is [`SendError::Unresolved`].
|
||||
NoMember(M),
|
||||
}
|
||||
|
||||
@@ -138,7 +195,9 @@ impl<M> std::fmt::Display for SendError<M> {
|
||||
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
|
||||
match self {
|
||||
SendError::Unresolved(_) => write!(f, "no live actor registered under that name"),
|
||||
SendError::Dead(_) => write!(f, "the addressed actor is no longer the live incarnation"),
|
||||
SendError::Dead(_) => {
|
||||
write!(f, "the addressed actor is no longer the live incarnation")
|
||||
}
|
||||
SendError::NoChannel(_) => write!(f, "actor has no channel for this message type"),
|
||||
SendError::Closed(_) => write!(f, "the actor's channel for this type is closed"),
|
||||
SendError::NoMember(_) => write!(f, "no live member in the process group"),
|
||||
@@ -150,9 +209,9 @@ impl<M> std::error::Error for SendError<M> {}
|
||||
|
||||
/// A registry-stored channel, type-erased over its message type. The stored
|
||||
/// object must serve two readers: `clone_sender` (downcast back to the concrete
|
||||
/// `Sender<M>`) and the RFC 016 snapshot (queued length without knowing `M`).
|
||||
/// A bare `Box<dyn Any>` gives the first but not the second, so we erase behind
|
||||
/// this small trait instead.
|
||||
/// `Sender<M>`) and the runtime introspection snapshot (queued length without
|
||||
/// knowing `M`). A bare `Box<dyn Any>` gives the first but not the second, so
|
||||
/// we erase behind this small trait instead.
|
||||
trait ErasedSender: Send {
|
||||
fn as_any(&self) -> &dyn Any;
|
||||
fn queued_len(&self) -> usize;
|
||||
@@ -169,7 +228,7 @@ impl<M: Send + 'static> ErasedSender for Sender<M> {
|
||||
|
||||
/// One typed channel of an actor, type-erased. Concretely a `Sender<M>` filed
|
||||
/// under `TypeId::of::<M>()`; `msg_type` is `type_name::<M>()`, kept for
|
||||
/// observers (RFC 014 §4.5) and as the debug cross-check on the downcast.
|
||||
/// observability tooling and as the debug cross-check on the downcast.
|
||||
struct Channel {
|
||||
sender: Box<dyn ErasedSender>,
|
||||
msg_type: &'static str,
|
||||
@@ -185,7 +244,10 @@ struct Mailbox {
|
||||
|
||||
impl Mailbox {
|
||||
fn new(pid: Pid) -> Self {
|
||||
Self { pid, channels: HashMap::new() }
|
||||
Self {
|
||||
pid,
|
||||
channels: HashMap::new(),
|
||||
}
|
||||
}
|
||||
|
||||
/// Clone the `Sender<M>` for this actor, if it has one. Called **under the
|
||||
@@ -204,7 +266,7 @@ impl Mailbox {
|
||||
}
|
||||
}
|
||||
|
||||
/// Per-actor registry view handed to RFC 016 introspection: registered names
|
||||
/// Per-actor registry view handed to runtime introspection: registered names
|
||||
/// and summed mailbox depth, tagged with the mailbox's `pid` so a stale
|
||||
/// incarnation can be filtered against the slab. Covers only *published*
|
||||
/// channels (`register` / `install` / `spawn_addr` / gen_server start); an
|
||||
@@ -219,47 +281,54 @@ pub(crate) struct MailboxInfo {
|
||||
/// The directory. Invariant (held under the registry lock): every value in
|
||||
/// `by_name` is the full [`Pid`] (index *and* generation) of an actor that
|
||||
/// published a [`Mailbox`] into `by_index` at registration time. Stale entries
|
||||
/// (dead holders — including holders whose slot has since been re-tenanted by
|
||||
/// a different actor) violate nothing — they are pruned on contact, and the
|
||||
/// (dead holders, including holders whose slot has since been re-tenanted by
|
||||
/// a different actor) violate nothing: they are pruned on contact, and the
|
||||
/// generation makes "dead" decidable even after slot reuse.
|
||||
pub(crate) struct Registry {
|
||||
/// `pid.index() -> the actor's mailbox`. The handle store.
|
||||
by_index: HashMap<u32, Mailbox>,
|
||||
/// `name -> holder pid`. Several names may map to one actor. The full pid
|
||||
/// (not just the index) is load-bearing: an index alone cannot tell a dead
|
||||
/// holder from the live actor now tenanting its recycled slot, which made
|
||||
/// such a name read as live-held — unresolvable *and* unregisterable — and
|
||||
/// would misdeliver to a same-typed tenant (soak20 signature 2).
|
||||
/// holder from the live actor now tenanting its recycled slot. Comparing
|
||||
/// only the index would make such a name read as live-held (unresolvable
|
||||
/// and unregisterable at once) and could misdeliver to whatever new,
|
||||
/// same-typed actor now sits in that slot.
|
||||
by_name: HashMap<&'static str, Pid>,
|
||||
}
|
||||
|
||||
impl Registry {
|
||||
pub(crate) fn new() -> Self {
|
||||
Self { by_index: HashMap::new(), by_name: HashMap::new() }
|
||||
Self {
|
||||
by_index: HashMap::new(),
|
||||
by_name: HashMap::new(),
|
||||
}
|
||||
}
|
||||
|
||||
/// Drop a dead holder's artifacts: every name bound to it, and its mailbox
|
||||
/// — but only while the mailbox is still *its own*. A recycled slot's
|
||||
/// mailbox belongs to the live tenant (publish replaces it wholesale on
|
||||
/// pid mismatch) and is left untouched.
|
||||
/// Drop a dead holder's artifacts: every name bound to it, and its
|
||||
/// mailbox, but only while the mailbox is still *its own*. A recycled
|
||||
/// slot's mailbox belongs to the live tenant (publish replaces it
|
||||
/// wholesale on pid mismatch) and is left untouched.
|
||||
fn prune_holder(&mut self, holder: Pid) {
|
||||
self.by_name.retain(|_, p| *p != holder);
|
||||
if self.by_index.get(&holder.index()).is_some_and(|mb| mb.pid == holder) {
|
||||
if self
|
||||
.by_index
|
||||
.get(&holder.index())
|
||||
.is_some_and(|mb| mb.pid == holder)
|
||||
{
|
||||
self.by_index.remove(&holder.index());
|
||||
}
|
||||
}
|
||||
|
||||
/// RFC 016 snapshot input: per-slot-index registry view — the actor's
|
||||
/// registered names (inverted from `by_name`) and its mailbox depth (queued
|
||||
/// messages summed across every published typed channel). Built in one pass
|
||||
/// under the registry Leaf; the per-channel `queued_len` takes a Channel
|
||||
/// lock, legal under the Leaf (Leaf → Channel). Carries each mailbox's full
|
||||
/// `pid` so the caller can discard a stale incarnation's entry against the
|
||||
/// slab's live generation. Names are matched to mailboxes by *full pid*, so
|
||||
/// a stale name (dead holder) still annotates the corpse's own mailbox if
|
||||
/// that survives, but never a recycled slot's new tenant; names that attach
|
||||
/// to no mailbox are dropped — they violate no invariant and get pruned on
|
||||
/// next contact.
|
||||
/// Runtime introspection input: per-slot-index registry view, giving the
|
||||
/// actor's registered names (inverted from `by_name`) and its mailbox
|
||||
/// depth (queued messages summed across every published typed channel).
|
||||
/// Carries each mailbox's full `pid` so the caller can discard a stale
|
||||
/// incarnation's entry against the slab's live generation. Names are
|
||||
/// matched to mailboxes by *full pid*, so a stale name (dead holder)
|
||||
/// still annotates the corpse's own mailbox if that survives, but never a
|
||||
/// recycled slot's new tenant; names that attach to no mailbox are
|
||||
/// dropped, since that violates no invariant and they get pruned on next
|
||||
/// contact.
|
||||
pub(crate) fn introspect_map(&self) -> HashMap<u32, MailboxInfo> {
|
||||
let mut names: HashMap<Pid, Vec<&'static str>> = HashMap::new();
|
||||
for (&name, &pid) in &self.by_name {
|
||||
@@ -282,8 +351,9 @@ impl Registry {
|
||||
|
||||
/// Single-actor form of [`introspect_map`](Self::introspect_map): the
|
||||
/// registry view for one slot index, or `None` if no mailbox is published
|
||||
/// there. Used by `actor_info` so its cost stays proportional to the one
|
||||
/// actor rather than locking every channel in the runtime.
|
||||
/// there. Used by the runtime's per-actor introspection so its cost stays
|
||||
/// proportional to the one actor rather than locking every channel in the
|
||||
/// runtime.
|
||||
pub(crate) fn introspect_one(&self, idx: u32) -> Option<MailboxInfo> {
|
||||
let mb = self.by_index.get(&idx)?;
|
||||
let depth: usize = mb.channels.values().map(|c| c.sender.queued_len()).sum();
|
||||
@@ -292,7 +362,11 @@ impl Registry {
|
||||
.iter()
|
||||
.filter_map(|(&n, &p)| (p == mb.pid).then_some(n))
|
||||
.collect();
|
||||
Some(MailboxInfo { pid: mb.pid, names, depth: depth.min(u32::MAX as usize) as u32 })
|
||||
Some(MailboxInfo {
|
||||
pid: mb.pid,
|
||||
names,
|
||||
depth: depth.min(u32::MAX as usize) as u32,
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
@@ -301,14 +375,18 @@ fn live(inner: &crate::runtime::RuntimeInner, pid: Pid) -> bool {
|
||||
inner.slot_at(pid).is_some_and(|s| s.is_live_for(pid))
|
||||
}
|
||||
|
||||
/// Publish the current actor's `Sender<M>` under `name`, capturing the channel
|
||||
/// so the name becomes messageable. Idempotent for the same `(name, type)`;
|
||||
/// registering a *second* type under the same (or another) name on the same
|
||||
/// actor just adds another channel to the actor's mailbox.
|
||||
/// Give the current actor's channel a name, so other actors can find and
|
||||
/// message it by that name instead of needing its [`Pid`].
|
||||
///
|
||||
/// Fails with [`RegisterError::NameTaken`] if the name is held by a *different*
|
||||
/// live actor (a binding to a dead actor is pruned and the name treated as
|
||||
/// free). Panics if called outside `Runtime::run()`.
|
||||
/// Calling this again with the same `(name, type)` from the same actor is
|
||||
/// harmless. Registering a *second* message type under the same (or a
|
||||
/// different) name from the same actor just adds another typed channel to
|
||||
/// that actor's mailbox; it does not replace the first.
|
||||
///
|
||||
/// Fails with [`RegisterError::NameTaken`] if the name is currently held by a
|
||||
/// *different* live actor. A name held by an actor that has since died is not
|
||||
/// considered taken: it is quietly reclaimed and handed to you. Panics if
|
||||
/// called outside [`run`](crate::run).
|
||||
pub fn register<M: Send + 'static>(name: Name<M>, tx: Sender<M>) -> Result<(), RegisterError> {
|
||||
register_with(self_pid(), name.as_str(), tx)
|
||||
}
|
||||
@@ -316,8 +394,8 @@ pub fn register<M: Send + 'static>(name: Name<M>, tx: Sender<M>) -> Result<(), R
|
||||
/// Bind `name` to `pid`'s mailbox and publish `tx` under `M`'s [`TypeId`], for
|
||||
/// an explicit (already-live) actor rather than `self`. The shared core of
|
||||
/// [`register`] (which passes `self_pid()`) and the parent-side server-name
|
||||
/// bind in `gen_server` (which names a freshly spawned server before its body
|
||||
/// has run, so the name resolves the instant `start()` returns). Same collision
|
||||
/// bind in `gen_server`, which names a freshly spawned server before its body
|
||||
/// has run, so the name resolves the instant `start()` returns. Same collision
|
||||
/// rules and lock discipline as `register`.
|
||||
pub(crate) fn register_with<M: Send + 'static>(
|
||||
me: Pid,
|
||||
@@ -325,6 +403,16 @@ pub(crate) fn register_with<M: Send + 'static>(
|
||||
tx: Sender<M>,
|
||||
) -> Result<(), RegisterError> {
|
||||
with_runtime(|inner| {
|
||||
// Stamp-eligibility for the terminal record (soak sig 4): flag the
|
||||
// tenancy BEFORE the binding lands and outside the registry lock (no
|
||||
// nesting), so no successfully-registered actor can die unflagged.
|
||||
// A register that then fails leaves a harmless overshoot; a stale
|
||||
// `me` is screened by the same live() the binding requires below.
|
||||
if live(inner, me) {
|
||||
if let Some(slot) = inner.slot_at(me) {
|
||||
slot.cold.lock().watchable = true;
|
||||
}
|
||||
}
|
||||
let mut reg = inner.registry.lock();
|
||||
if !live(inner, me) {
|
||||
return Err(RegisterError::NoProc);
|
||||
@@ -337,8 +425,8 @@ pub(crate) fn register_with<M: Send + 'static>(
|
||||
} else {
|
||||
// Dead holder: free the name (and its other stale artifacts).
|
||||
// Liveness is judged against the *stored* pid, generation
|
||||
// included — a recycled slot's live tenant no longer makes a
|
||||
// dead name read as taken (soak20 signature 2).
|
||||
// included, so a recycled slot's live tenant no longer makes a
|
||||
// dead name read as taken.
|
||||
reg.prune_holder(holder);
|
||||
}
|
||||
}
|
||||
@@ -355,26 +443,32 @@ pub(crate) fn register_with<M: Send + 'static>(
|
||||
/// index from a dead prior incarnation (pid mismatch) is replaced wholesale.
|
||||
/// Caller holds the registry lock and has established that `me` is live.
|
||||
fn publish_channel<M: Send + 'static>(reg: &mut Registry, me: Pid, tx: Sender<M>) {
|
||||
let mb = reg.by_index.entry(me.index()).or_insert_with(|| Mailbox::new(me));
|
||||
let mb = reg
|
||||
.by_index
|
||||
.entry(me.index())
|
||||
.or_insert_with(|| Mailbox::new(me));
|
||||
if mb.pid != me {
|
||||
*mb = Mailbox::new(me);
|
||||
}
|
||||
mb.channels.insert(
|
||||
TypeId::of::<M>(),
|
||||
Channel { sender: Box::new(tx), msg_type: type_name::<M>() },
|
||||
Channel {
|
||||
sender: Box::new(tx),
|
||||
msg_type: type_name::<M>(),
|
||||
},
|
||||
);
|
||||
}
|
||||
|
||||
/// Publish the current actor's `Sender<A::Msg>` into its mailbox **without**
|
||||
/// binding a name, and hand back the typed [`Pid<A>`] that addresses this
|
||||
/// actor directly. This is the opt-in, lazy install of RFC 014 §5: an actor
|
||||
/// that wants to be reachable by a direct, identity-bound [`Pid<A>`] (rather
|
||||
/// than only via a re-resolving [`Name`]) calls this once with its inbox
|
||||
/// sender, then hands the returned pid out.
|
||||
/// actor directly.
|
||||
///
|
||||
/// Unlike [`register`] there is no name to collide on, and `self` is always a
|
||||
/// live actor inside `run()`, so this is infallible. Panics if called outside
|
||||
/// `Runtime::run()`.
|
||||
/// This is for an actor that wants to be reachable directly by its pid,
|
||||
/// rather than only through a re-resolving [`Name`]: call this once with your
|
||||
/// inbox sender, then hand the returned `Pid<A>` to whoever should be able to
|
||||
/// message you. Unlike [`register`] there is no name to collide on, and the
|
||||
/// current actor is always live while inside `run()`, so this cannot fail.
|
||||
/// Panics if called outside [`run`](crate::run).
|
||||
pub fn install<A: Addressable>(tx: Sender<A::Msg>) -> Pid<A> {
|
||||
let me = self_pid();
|
||||
with_runtime(|inner| {
|
||||
@@ -390,22 +484,27 @@ pub fn install<A: Addressable>(tx: Sender<A::Msg>) -> Pid<A> {
|
||||
/// Publish `tx` into `pid`'s mailbox under `M`'s [`TypeId`], for an explicit
|
||||
/// (freshly minted, already-live) actor rather than `self`. The parent-side
|
||||
/// half of [`spawn_addr`](crate::spawn_addr): the spawner makes the inbox and
|
||||
/// publishes the sender here *before* handing back the `Pid<A>`, so an immediate
|
||||
/// `send_to` on the returned pid always resolves — the address is live the
|
||||
/// instant the caller holds it, with no dependence on the body having run yet.
|
||||
/// publishes the sender here *before* handing back the `Pid<A>`, so an
|
||||
/// immediate `send_to` on the returned pid always resolves. The address is
|
||||
/// live the instant the caller holds it, with no dependence on the spawned
|
||||
/// actor's body having run yet.
|
||||
///
|
||||
/// Caller guarantees `pid` is the just-installed actor (Queued, this exact
|
||||
/// Caller guarantees `pid` is the just-installed actor (queued, this exact
|
||||
/// incarnation); `publish_channel` replaces any stale leftover at the slot.
|
||||
pub(crate) fn install_for<M: Send + 'static>(pid: Pid, tx: Sender<M>) {
|
||||
with_runtime(|inner| {
|
||||
let mut reg = inner.registry.lock();
|
||||
debug_assert!(live(inner, pid), "install_for: pid must be a freshly spawned, live actor");
|
||||
debug_assert!(
|
||||
live(inner, pid),
|
||||
"install_for: pid must be a freshly spawned, live actor"
|
||||
);
|
||||
publish_channel::<M>(&mut reg, pid, tx);
|
||||
});
|
||||
}
|
||||
|
||||
/// The single actor currently registered under `name`, or `None` if unbound or
|
||||
/// no longer live (the stale binding is pruned on the way out).
|
||||
/// Look up which actor currently holds `name`, if any. Returns `None` if the
|
||||
/// name is unbound, or if it was bound to an actor that has since died (the
|
||||
/// stale binding is cleared as a side effect of this call).
|
||||
pub fn whereis(name: &str) -> Option<Pid> {
|
||||
with_runtime(|inner| {
|
||||
let mut reg = inner.registry.lock();
|
||||
@@ -421,67 +520,122 @@ pub fn whereis(name: &str) -> Option<Pid> {
|
||||
})
|
||||
}
|
||||
|
||||
/// Resolve `name` to a *typed* [`Pid<A>`] — the identity-bound counterpart of
|
||||
/// [`whereis`] (RFC 014 §4.4). Recovers the compile-checked
|
||||
/// [`send_to`] path from a durable name: looks the name up,
|
||||
/// then re-types the erased pid as `Pid<A>` via the unchecked
|
||||
/// [`assert_type`](crate::pid::assert_type) primitive. A wrong `A` is not
|
||||
/// unsound — it degrades to [`SendError::NoChannel`] on the next send (routing
|
||||
/// is by message `TypeId`), never a misdelivery. `None` if unbound or dead.
|
||||
/// What a name is bound to, three-valued (bridge soak signature 4).
|
||||
///
|
||||
/// Panics if called outside `Runtime::run()`.
|
||||
/// [`Live`](NameResolution::Live) is [`whereis`]'s `Some`.
|
||||
/// [`Corpse`](NameResolution::Corpse) carries the *stored* holder pid of a
|
||||
/// dead-but-unpruned binding — a state Erlang cannot represent (its name
|
||||
/// death unregisters atomically; smarm's prune is lazy), captured here before
|
||||
/// the prune that `whereis` performs discards it, so the caller can consult
|
||||
/// [`terminal_reason`](crate::monitor::terminal_reason) for the tenancy's
|
||||
/// real down reason. [`Unbound`](NameResolution::Unbound) matches Erlang's
|
||||
/// unregistered name.
|
||||
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
|
||||
pub enum NameResolution {
|
||||
/// The stored holder is live (generation-checked); the binding stands.
|
||||
Live(Pid),
|
||||
/// The stored holder is dead. The binding was pruned on the way out —
|
||||
/// the name heals exactly as `whereis` heals it; only the evidence is
|
||||
/// returned instead of discarded. A second resolve is `Unbound`.
|
||||
Corpse(Pid),
|
||||
/// No binding stored (never registered, or already pruned by any reader).
|
||||
Unbound,
|
||||
}
|
||||
|
||||
/// Resolve `name` like [`whereis`], but keep the corpse: the dead-holder arm
|
||||
/// returns the stored pid it pruned instead of a bare `None`. Same lock
|
||||
/// discipline and pruning behavior as `whereis`; same `Runtime::run()`
|
||||
/// context contract.
|
||||
pub fn resolve_name(name: &str) -> NameResolution {
|
||||
with_runtime(|inner| {
|
||||
let mut reg = inner.registry.lock();
|
||||
let Some(&pid) = reg.by_name.get(name) else {
|
||||
return NameResolution::Unbound;
|
||||
};
|
||||
if live(inner, pid) {
|
||||
NameResolution::Live(pid)
|
||||
} else {
|
||||
reg.prune_holder(pid);
|
||||
NameResolution::Corpse(pid)
|
||||
}
|
||||
})
|
||||
}
|
||||
|
||||
/// Like [`whereis`], but returns a *typed* [`Pid<A>`] instead of a bare
|
||||
/// [`Pid`], so a follow-up [`send_to`] is compile-checked instead of needing
|
||||
/// the untyped [`send_dyn`] escape hatch. `None` if the name is unbound or its
|
||||
/// holder has died.
|
||||
///
|
||||
/// The type `A` is not checked against what the name's holder actually
|
||||
/// published: if you pick the wrong `A`, this still succeeds, but the next
|
||||
/// send against the returned pid degrades to [`SendError::NoChannel`] rather
|
||||
/// than reaching the wrong actor or the wrong channel.
|
||||
///
|
||||
/// Panics if called outside [`run`](crate::run).
|
||||
pub fn lookup_as<A: Addressable>(name: &str) -> Option<Pid<A>> {
|
||||
whereis(name).map(crate::pid::assert_type::<A>)
|
||||
}
|
||||
|
||||
/// Resolve `name` to its actor's pid and a cloned `Sender<M>`, under the Leaf
|
||||
/// lock (clone-under-lock, then release). The crate-internal building block for
|
||||
/// `gen_server`'s by-name addressing: a named server publishes its inbox as a
|
||||
/// `Sender<Envelope<G>>` (via [`register_with`]), and `whereis_server` / `call`
|
||||
/// / `cast` recover that exact typed sender here to rebuild a `GenServerRef<G>`.
|
||||
/// `None` if unbound, dead (pruned on the way out), or holding no `M` channel.
|
||||
/// Resolve `name` to its actor's pid and a cloned `Sender<M>`, all under one
|
||||
/// lock acquisition. The crate-internal building block for `gen_server`'s
|
||||
/// by-name addressing: a named server publishes its inbox as a
|
||||
/// `Sender<Envelope<G>>` (via [`register_with`]), and the server's `call` /
|
||||
/// `cast` / `whereis_server` recover that exact typed sender here to rebuild a
|
||||
/// `GenServerRef<G>`. `None` if unbound, dead (pruned on the way out), or
|
||||
/// holding no `M` channel.
|
||||
pub(crate) fn resolve_named_sender<M: Send + 'static>(name: &str) -> Option<(Pid, Sender<M>)> {
|
||||
with_runtime(|inner| {
|
||||
let mut reg = inner.registry.lock();
|
||||
let pid = *reg.by_name.get(name)?;
|
||||
if !live(inner, pid) {
|
||||
// Stored-pid liveness, generation included: a name whose holder
|
||||
// died is pruned (heals) even if the slot has a new tenant —
|
||||
// previously the tenant's mailbox made the name unresolvable
|
||||
// *without* pruning, wedging it for the tenant's lifetime.
|
||||
// died is pruned (heals) even if the slot has a new tenant.
|
||||
// Otherwise the tenant's mailbox would make the name unresolvable
|
||||
// without pruning, wedging it for the tenant's lifetime.
|
||||
reg.prune_holder(pid);
|
||||
return None;
|
||||
}
|
||||
// A live holder's mailbox is its own (publish replaces wholesale on
|
||||
// pid mismatch, and one live actor per slot), so index lookup is safe.
|
||||
let tx = reg.by_index.get(&pid.index()).and_then(Mailbox::clone_sender::<M>)?;
|
||||
let tx = reg
|
||||
.by_index
|
||||
.get(&pid.index())
|
||||
.and_then(Mailbox::clone_sender::<M>)?;
|
||||
Some((pid, tx))
|
||||
})
|
||||
}
|
||||
|
||||
/// Remove the binding for `name`, returning the actor it pointed at if still
|
||||
/// live. Only the *name* is freed; the actor's mailbox (and any other names for
|
||||
/// it) remain. A binding to a dead actor is reported as `None`.
|
||||
/// Give up a name. Returns the actor it pointed at, if that actor was still
|
||||
/// live. Only the *name* is freed; the actor's mailbox (and any other names
|
||||
/// bound to it) are unaffected. A binding to an already-dead actor reports
|
||||
/// `None`, since there was nothing live to release.
|
||||
pub fn unregister(name: &str) -> Option<Pid> {
|
||||
with_runtime(|inner| {
|
||||
let mut reg = inner.registry.lock();
|
||||
let pid = reg.by_name.remove(name)?;
|
||||
if live(inner, pid) { Some(pid) } else { None }
|
||||
if live(inner, pid) {
|
||||
Some(pid)
|
||||
} else {
|
||||
None
|
||||
}
|
||||
})
|
||||
}
|
||||
|
||||
/// Resolve `name` to its actor's `Sender<M>` and deliver `msg`. The whole point
|
||||
/// of the rework: a name you can *send* to.
|
||||
/// Look `name` up and deliver `msg` to whichever actor currently holds it.
|
||||
/// This is the point of naming an actor: a name you can send a message to
|
||||
/// directly, without a separate lookup step.
|
||||
///
|
||||
/// Errors (message returned in every case): [`SendError::Unresolved`] if no
|
||||
/// live actor holds the name, [`SendError::NoChannel`] if that actor has no
|
||||
/// channel for `M`, [`SendError::Closed`] if its `M` channel's receiver is
|
||||
/// gone. Panics if called outside `Runtime::run()`.
|
||||
/// On failure the message comes back to you, wrapped in the [`SendError`]
|
||||
/// variant that explains why: [`SendError::Unresolved`] if no live actor
|
||||
/// currently holds the name, [`SendError::NoChannel`] if the actor that holds
|
||||
/// it never published a channel for `M`, or [`SendError::Closed`] if it
|
||||
/// published one but has since dropped the receiving end. Panics if called
|
||||
/// outside [`run`](crate::run).
|
||||
pub fn send<M: Send + 'static>(name: Name<M>, msg: M) -> Result<(), SendError<M>> {
|
||||
let key = name.as_str();
|
||||
with_runtime(|inner| {
|
||||
// Resolve + clone the sender under the Leaf lock, then drop the lock
|
||||
// before sending (a send can unpark a receiver).
|
||||
// Resolve + clone the sender under the registry lock, then drop the
|
||||
// lock before sending (a send can unpark a receiver).
|
||||
let tx = {
|
||||
let mut reg = inner.registry.lock();
|
||||
let pid = match reg.by_name.get(key) {
|
||||
@@ -489,43 +643,51 @@ pub fn send<M: Send + 'static>(name: Name<M>, msg: M) -> Result<(), SendError<M>
|
||||
None => return Err(SendError::Unresolved(msg)),
|
||||
};
|
||||
if !live(inner, pid) {
|
||||
// Stored-pid liveness (generation included). Previously this
|
||||
// checked the slot's *current* mailbox pid, so a recycled
|
||||
// slot's live tenant passed — and a same-typed tenant would
|
||||
// have received the message (misdelivery), a differently
|
||||
// typed one a misleading NoChannel.
|
||||
// Stored-pid liveness (generation included), so a recycled
|
||||
// slot's new live tenant is never mistaken for the name's
|
||||
// original (now-dead) holder.
|
||||
reg.prune_holder(pid);
|
||||
return Err(SendError::Unresolved(msg));
|
||||
}
|
||||
match reg.by_index.get(&pid.index()).and_then(Mailbox::clone_sender::<M>) {
|
||||
match reg
|
||||
.by_index
|
||||
.get(&pid.index())
|
||||
.and_then(Mailbox::clone_sender::<M>)
|
||||
{
|
||||
Some(tx) => tx,
|
||||
None => return Err(SendError::NoChannel(msg)),
|
||||
}
|
||||
};
|
||||
tx.send(msg).map_err(|crate::channel::SendError(m)| SendError::Closed(m))
|
||||
tx.send(msg)
|
||||
.map_err(|crate::channel::SendError(m)| SendError::Closed(m))
|
||||
})
|
||||
}
|
||||
|
||||
/// Resolve a *raw* pid to its mailbox and deliver `msg` on the channel for `M`,
|
||||
/// with **no redirect**. The stored mailbox must be this exact incarnation
|
||||
/// (generation included) and still live; otherwise the actor this pid named is
|
||||
/// gone and the result is [`SendError::Dead`] — even when the slot now holds a
|
||||
/// different, live actor (which we leave untouched). Shared by [`send_to`]
|
||||
/// (typed, `M = A::Msg`, channel guaranteed on an installed actor) and
|
||||
/// [`send_dyn`] (explicit `M`, where `NoChannel` is a real outcome).
|
||||
/// (generation included) and still live; otherwise the actor this pid named
|
||||
/// is gone and the result is [`SendError::Dead`], even when the slot now
|
||||
/// holds a different, live actor (which is left untouched). Shared by
|
||||
/// [`send_to`] (typed, `M = A::Msg`, channel guaranteed on an installed
|
||||
/// actor) and [`send_dyn`] (explicit `M`, where `NoChannel` is a real
|
||||
/// outcome).
|
||||
fn send_to_pid<M: Send + 'static>(
|
||||
inner: &crate::runtime::RuntimeInner,
|
||||
pid: Pid,
|
||||
msg: M,
|
||||
) -> Result<(), SendError<M>> {
|
||||
// Resolve + clone the sender under the Leaf lock, then drop the lock before
|
||||
// sending (a send can unpark a receiver) — Leaf -> Channel, as name `send`.
|
||||
// Resolve + clone the sender under the registry lock, then drop the lock
|
||||
// before sending (a send can unpark a receiver), same order as `send`.
|
||||
let tx = {
|
||||
let mut reg = inner.registry.lock();
|
||||
match reg.by_index.get(&pid.index()).map(|m| m.pid) {
|
||||
// Exact incarnation, still alive: its `M` channel, or NoChannel.
|
||||
Some(stored) if stored == pid && live(inner, pid) => {
|
||||
match reg.by_index.get(&pid.index()).and_then(Mailbox::clone_sender::<M>) {
|
||||
match reg
|
||||
.by_index
|
||||
.get(&pid.index())
|
||||
.and_then(Mailbox::clone_sender::<M>)
|
||||
{
|
||||
Some(tx) => tx,
|
||||
None => return Err(SendError::NoChannel(msg)),
|
||||
}
|
||||
@@ -540,35 +702,40 @@ fn send_to_pid<M: Send + 'static>(
|
||||
_ => return Err(SendError::Dead(msg)),
|
||||
}
|
||||
};
|
||||
tx.send(msg).map_err(|crate::channel::SendError(m)| SendError::Closed(m))
|
||||
tx.send(msg)
|
||||
.map_err(|crate::channel::SendError(m)| SendError::Closed(m))
|
||||
}
|
||||
|
||||
/// Deliver `msg` to the exact actor named by `pid` — RFC 014 §4.2's direct,
|
||||
/// identity-bound addressing mode. Unlike name-addressed [`send`] there is **no
|
||||
/// redirect**: if that incarnation has died the message comes back as
|
||||
/// [`SendError::Dead`], even if its slot now holds a different actor.
|
||||
/// Deliver `msg` directly to the exact actor identified by `pid`. Unlike
|
||||
/// name-addressed [`send`], there is **no redirect**: if that specific actor
|
||||
/// has died, the message comes back as [`SendError::Dead`], even if its slot
|
||||
/// has since been taken over by a different, live actor. Use this when you
|
||||
/// already hold a `Pid<A>` and want to talk to that one actor specifically;
|
||||
/// use [`send`] with a [`Name`] when you want whichever actor currently holds
|
||||
/// a name.
|
||||
///
|
||||
/// The message type is the actor's `A::Msg`, so on a live actor that has
|
||||
/// installed its inbox (via [`install`] or [`register`]) the channel is always
|
||||
/// present; [`SendError::NoChannel`] therefore means the actor is live but
|
||||
/// never published a `Pid<A>`-reachable inbox. Panics if called outside
|
||||
/// `Runtime::run()`.
|
||||
/// installed its inbox (via [`install`] or [`register`]) the channel is
|
||||
/// always present; [`SendError::NoChannel`] therefore means the actor is live
|
||||
/// but never published a `Pid<A>`-reachable inbox. Panics if called outside
|
||||
/// [`run`](crate::run).
|
||||
pub fn send_to<A: Addressable>(pid: Pid<A>, msg: A::Msg) -> Result<(), SendError<A::Msg>> {
|
||||
with_runtime(|inner| send_to_pid::<A::Msg>(inner, pid.erase(), msg))
|
||||
}
|
||||
|
||||
/// The explicit bare-pid escape hatch (RFC 014 §4.6): deliver `msg` of type `M`
|
||||
/// to `pid` when all you hold is an untyped [`Pid`] — a pid off a [`Down`], or
|
||||
/// out of a future `members()` — so the typed [`send_to`] is unavailable.
|
||||
/// The escape hatch for sending to a bare, untyped [`Pid`] when the typed
|
||||
/// [`send_to`] is unavailable, for example a pid recovered from a [`Down`]
|
||||
/// notification or a group's `members()` list, where you no longer know the
|
||||
/// actor's message type at compile time.
|
||||
///
|
||||
/// This is the one send whose message type can genuinely be wrong: the actor
|
||||
/// may be live yet expose no channel for `M`, returning [`SendError::NoChannel`]
|
||||
/// (on the typed paths that downcast collapses to a `debug_assert`). It is
|
||||
/// named and documented as the fallible fallback so the typed `Pid<A>` /
|
||||
/// `Name<M>` paths stay the obvious default and an agentic caller reaches for a
|
||||
/// present primitive instead of inventing a workaround. Liveness is identical
|
||||
/// to [`send_to`]: identity-bound, no redirect, [`SendError::Dead`] once the
|
||||
/// addressed incarnation is gone. Panics if called outside `Runtime::run()`.
|
||||
/// Because the message type is not checked at compile time here, this is the
|
||||
/// one send that can genuinely be live-but-wrong: the actor may be alive yet
|
||||
/// expose no channel for `M`, in which case you get [`SendError::NoChannel`]
|
||||
/// back instead of a misdelivery. Liveness and redirect behavior are
|
||||
/// otherwise identical to [`send_to`]: identity-bound, no redirect,
|
||||
/// [`SendError::Dead`] once the addressed incarnation is gone. Prefer
|
||||
/// `send_to` with a typed `Pid<A>` whenever you have one; reach for this only
|
||||
/// when you don't. Panics if called outside [`run`](crate::run).
|
||||
///
|
||||
/// [`Down`]: crate::Down
|
||||
pub fn send_dyn<M: Send + 'static>(pid: Pid, msg: M) -> Result<(), SendError<M>> {
|
||||
|
||||
+28
-5
@@ -222,7 +222,10 @@ impl MpmcRing {
|
||||
if diff == 0 {
|
||||
// Our turn: claim the position.
|
||||
match self.enqueue_pos.0.compare_exchange_weak(
|
||||
pos, pos + 1, Ordering::Relaxed, Ordering::Relaxed,
|
||||
pos,
|
||||
pos + 1,
|
||||
Ordering::Relaxed,
|
||||
Ordering::Relaxed,
|
||||
) {
|
||||
Ok(_) => {
|
||||
// SAFETY: the claim gives us exclusive write access
|
||||
@@ -250,7 +253,10 @@ impl MpmcRing {
|
||||
let diff = seq as isize - (pos + 1) as isize;
|
||||
if diff == 0 {
|
||||
match self.dequeue_pos.0.compare_exchange_weak(
|
||||
pos, pos + 1, Ordering::Relaxed, Ordering::Relaxed,
|
||||
pos,
|
||||
pos + 1,
|
||||
Ordering::Relaxed,
|
||||
Ordering::Relaxed,
|
||||
) {
|
||||
Ok(_) => {
|
||||
// SAFETY: the claim gives us exclusive read access;
|
||||
@@ -464,19 +470,36 @@ mod tests {
|
||||
|
||||
let popped = popped.lock().unwrap();
|
||||
assert_eq!(popped.len(), total, "count mismatch");
|
||||
let set: HashSet<u64> = popped.iter().map(|p| ((p.index() as u64) << 32) | p.generation() as u64).collect();
|
||||
let set: HashSet<u64> = popped
|
||||
.iter()
|
||||
.map(|p| ((p.index() as u64) << 32) | p.generation() as u64)
|
||||
.collect();
|
||||
assert_eq!(set.len(), total, "duplicate or lost element");
|
||||
assert_eq!(pop(&q), None);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn mpmc_exactly_once_contended() {
|
||||
exactly_once(MpmcRing::new(8, 4096), |q, p| q.push(p), |q| q.pop(), 4, 4, 1000);
|
||||
exactly_once(
|
||||
MpmcRing::new(8, 4096),
|
||||
|q, p| q.push(p),
|
||||
|q| q.pop(),
|
||||
4,
|
||||
4,
|
||||
1000,
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn striped_exactly_once_contended() {
|
||||
exactly_once(StripedRing::new(8, 4096), |q, p| q.push(p), |q| q.pop(), 4, 4, 1000);
|
||||
exactly_once(
|
||||
StripedRing::new(8, 4096),
|
||||
|q, p| q.push(p),
|
||||
|q| q.pop(),
|
||||
4,
|
||||
4,
|
||||
1000,
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
|
||||
+740
-286
File diff suppressed because it is too large
Load Diff
+707
-269
File diff suppressed because it is too large
Load Diff
+330
@@ -0,0 +1,330 @@
|
||||
//! RFC 019 §7 — overflow diagnostics.
|
||||
//!
|
||||
//! One process-global SIGSEGV handler, installed once at [`crate::runtime::init`]
|
||||
//! (before any scheduler thread exists, so the PRIOR save is unracing), plus a
|
||||
//! per-scheduler-thread `sigaltstack` registered at `schedule_loop` entry — a
|
||||
//! guard hit means the faulting stack has no room to run anything, so the
|
||||
//! altstack is not optional.
|
||||
//!
|
||||
//! The handler classifies `si_addr` against the *current* actor only, reached
|
||||
//! through `preempt::CURRENT_SLOT` — a const-initialized `Cell<*const Slot>`
|
||||
//! whose access is a plain TLS load (no lazy init, no allocation, no dtor
|
||||
//! registration), and which every scheduler thread has materialized before an
|
||||
//! actor can run on it. The slot's diag atomics (`diag_stack_top` & co) are
|
||||
//! written in `install_actor` before the Release publish and are only consulted
|
||||
//! here while the actor is on-CPU, so they cannot be stale.
|
||||
//!
|
||||
//! Two classification tiers:
|
||||
//! - **In-guard**: definitive. Rust frames probe pages in order
|
||||
//! (`__rust_probestack`), so Rust overflow always lands here; so does any C
|
||||
//! built with `-fstack-clash-protection` (distro-packaged libraries), and —
|
||||
//! with the 1 MiB default guard — nearly every unprobed frame too.
|
||||
//! - **Overshoot**: within [`OVERSHOOT_SLOP`] *below* the guard. An unprobed
|
||||
//! frame (cargo-built C via `cc` almost never enables clash protection)
|
||||
//! large enough to step over the guard in one `sub rsp`. Attribution is
|
||||
//! "probable": the address is in unmapped VA that nothing else owns, an
|
||||
//! actor was on-CPU, and the distance fits a frame — the diagnostic says so.
|
||||
//!
|
||||
//! Classified faults print one line (async-signal-safe: stack buffer +
|
||||
//! `write(2)`, no fmt, no alloc, no locks) and re-raise with default
|
||||
//! disposition — no unwind, no resume, no fail-soft (jarred; UB-adjacent from
|
||||
//! a handler). Unclassified faults reinstate the PRIOR handler and refault, so
|
||||
//! std's own "thread ... has overflowed its stack" diagnostics for OS-thread
|
||||
//! stacks survive our presence. Reinstating deregisters us for good, which is
|
||||
//! fine: the process is dying either way.
|
||||
|
||||
use std::cell::Cell;
|
||||
use std::mem::MaybeUninit;
|
||||
use std::sync::atomic::Ordering;
|
||||
use std::sync::Once;
|
||||
|
||||
/// Tier-2 window below the guard. Matches the guard default (and the kernel's
|
||||
/// `stack_guard_gap`): a frame that out-jumps both the guard and this window
|
||||
/// in one displacement is past what a diagnostic can honestly attribute.
|
||||
pub(crate) const OVERSHOOT_SLOP: usize = 1024 * 1024;
|
||||
|
||||
/// Per-scheduler-thread signal stack. MINSIGSTKSZ is ~11 KiB on AVX-512
|
||||
/// hardware; 64 KiB leaves the formatter room without mattering to anyone.
|
||||
/// One per OS thread, never freed: scheduler threads live for the process in
|
||||
/// practice, and repeated `run()`s on reused threads re-use the registration
|
||||
/// (the TLS flag), so the leak is bounded by the OS thread count.
|
||||
const ALTSTACK_SIZE: usize = 64 * 1024;
|
||||
|
||||
static INSTALL: Once = Once::new();
|
||||
/// The handler that was installed before ours (std's, typically). Written
|
||||
/// exactly once inside INSTALL — which completes in `runtime::init` before
|
||||
/// any scheduler thread (and thus any classifiable fault) can exist — and
|
||||
/// only read from the handler afterwards.
|
||||
static mut PRIOR: MaybeUninit<libc::sigaction> = MaybeUninit::uninit();
|
||||
|
||||
thread_local! {
|
||||
/// Whether this OS thread has registered its altstack.
|
||||
static ALTSTACK_SET: Cell<bool> = const { Cell::new(false) };
|
||||
}
|
||||
|
||||
/// Where a fault landed relative to the current actor's stack.
|
||||
#[derive(Debug, PartialEq, Eq)]
|
||||
pub(crate) enum FaultClass {
|
||||
/// Inside `[top − reserve − guard, top − reserve)`: the guard region.
|
||||
Guard,
|
||||
/// Within `OVERSHOOT_SLOP` below the guard: stepped over it. Payload is
|
||||
/// the distance below `guard_lo`.
|
||||
Overshoot(usize),
|
||||
/// Not ours to explain.
|
||||
Foreign,
|
||||
}
|
||||
|
||||
/// Pure classifier — all edges unit-tested below. `top` is the stack's usable
|
||||
/// top, `reserve`/`guard` its shape; both page-rounded by `Stack::new`.
|
||||
pub(crate) fn classify(addr: usize, top: usize, reserve: usize, guard: usize) -> FaultClass {
|
||||
let guard_hi = top.wrapping_sub(reserve);
|
||||
let guard_lo = guard_hi.wrapping_sub(guard);
|
||||
if addr >= guard_lo && addr < guard_hi {
|
||||
FaultClass::Guard
|
||||
} else if addr < guard_lo && addr >= guard_lo.saturating_sub(OVERSHOOT_SLOP) {
|
||||
FaultClass::Overshoot(guard_lo - addr)
|
||||
} else {
|
||||
FaultClass::Foreign
|
||||
}
|
||||
}
|
||||
|
||||
/// Install the process-global handler. Idempotent; called from
|
||||
/// `runtime::init`.
|
||||
pub(crate) fn install_once() {
|
||||
INSTALL.call_once(|| unsafe {
|
||||
let mut sa: libc::sigaction = std::mem::zeroed();
|
||||
sa.sa_sigaction = handler as *const () as usize;
|
||||
sa.sa_flags = libc::SA_SIGINFO | libc::SA_ONSTACK;
|
||||
libc::sigemptyset(&mut sa.sa_mask);
|
||||
let prior = &mut *std::ptr::addr_of_mut!(PRIOR);
|
||||
libc::sigaction(libc::SIGSEGV, &sa, prior.as_mut_ptr());
|
||||
});
|
||||
}
|
||||
|
||||
/// Register this OS thread's altstack (idempotent per thread). Called at
|
||||
/// `schedule_loop` entry, so every thread that can run an actor has one.
|
||||
pub(crate) fn register_altstack() {
|
||||
ALTSTACK_SET.with(|set| {
|
||||
if set.get() {
|
||||
return;
|
||||
}
|
||||
unsafe {
|
||||
let sp = libc::mmap(
|
||||
std::ptr::null_mut(),
|
||||
ALTSTACK_SIZE,
|
||||
libc::PROT_READ | libc::PROT_WRITE,
|
||||
libc::MAP_PRIVATE | libc::MAP_ANONYMOUS,
|
||||
-1,
|
||||
0,
|
||||
);
|
||||
if sp == libc::MAP_FAILED {
|
||||
// Degrade: no altstack means a guard hit dies without the
|
||||
// message (handler can't run) — the pre-RFC behavior, never
|
||||
// incorrectness.
|
||||
return;
|
||||
}
|
||||
let ss = libc::stack_t {
|
||||
ss_sp: sp,
|
||||
ss_flags: 0,
|
||||
ss_size: ALTSTACK_SIZE,
|
||||
};
|
||||
libc::sigaltstack(&ss, std::ptr::null_mut());
|
||||
}
|
||||
set.set(true);
|
||||
});
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// The handler
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
unsafe extern "C" fn handler(
|
||||
_sig: libc::c_int,
|
||||
info: *mut libc::siginfo_t,
|
||||
_ctx: *mut libc::c_void,
|
||||
) {
|
||||
let slot_ptr = crate::preempt::current_slot_ptr();
|
||||
if !slot_ptr.is_null() {
|
||||
let slot = &*slot_ptr;
|
||||
let top = slot.diag_stack_top.load(Ordering::Relaxed);
|
||||
if top != 0 {
|
||||
let reserve = slot.diag_stack_reserve.load(Ordering::Relaxed);
|
||||
let guard = slot.diag_stack_guard.load(Ordering::Relaxed);
|
||||
let pid = slot.diag_pid.load(Ordering::Relaxed);
|
||||
let addr = (*info).si_addr() as usize;
|
||||
match classify(addr, top, reserve, guard) {
|
||||
FaultClass::Guard => {
|
||||
let mut b = Buf::new();
|
||||
b.s("smarm: actor ");
|
||||
b.pid(pid);
|
||||
b.s(" overflowed its stack: fault in the guard region, depth-at-fault=");
|
||||
b.u(top - addr);
|
||||
b.s(" bytes (reserve=");
|
||||
b.u(reserve);
|
||||
b.s(", guard=");
|
||||
b.u(guard);
|
||||
b.s("). Raise stack_reserve (SpawnOpts or Config).\n");
|
||||
b.emit();
|
||||
die_by_default();
|
||||
return;
|
||||
}
|
||||
FaultClass::Overshoot(below) => {
|
||||
let mut b = Buf::new();
|
||||
b.s("smarm: actor ");
|
||||
b.pid(pid);
|
||||
b.s(" probably overflowed its stack: fault ");
|
||||
b.u(below);
|
||||
b.s(" bytes below the guard - an unprobed (FFI?) frame stepped over it (reserve=");
|
||||
b.u(reserve);
|
||||
b.s(", guard=");
|
||||
b.u(guard);
|
||||
b.s("). Raise stack_guard or stack_reserve.\n");
|
||||
b.emit();
|
||||
die_by_default();
|
||||
return;
|
||||
}
|
||||
FaultClass::Foreign => {}
|
||||
}
|
||||
}
|
||||
}
|
||||
// Not ours: put back whoever was there before us and refault into them.
|
||||
let prior = &*std::ptr::addr_of!(PRIOR);
|
||||
libc::sigaction(libc::SIGSEGV, prior.as_ptr(), std::ptr::null_mut());
|
||||
}
|
||||
|
||||
/// Reset SIGSEGV to default disposition; returning from the handler then
|
||||
/// refaults at the same instruction and the process dies the normal death
|
||||
/// (core-dumpable, correct wait status), exactly as if we were never here —
|
||||
/// but with the message already on stderr.
|
||||
unsafe fn die_by_default() {
|
||||
let mut dfl: libc::sigaction = std::mem::zeroed();
|
||||
dfl.sa_sigaction = libc::SIG_DFL;
|
||||
libc::sigemptyset(&mut dfl.sa_mask);
|
||||
libc::sigaction(libc::SIGSEGV, &dfl, std::ptr::null_mut());
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Async-signal-safe formatting: fixed buffer, decimal itoa, one write(2).
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
struct Buf {
|
||||
b: [u8; 320],
|
||||
len: usize,
|
||||
}
|
||||
|
||||
impl Buf {
|
||||
fn new() -> Self {
|
||||
Buf {
|
||||
b: [0; 320],
|
||||
len: 0,
|
||||
}
|
||||
}
|
||||
fn s(&mut self, s: &str) {
|
||||
for &c in s.as_bytes() {
|
||||
if self.len < self.b.len() {
|
||||
self.b[self.len] = c;
|
||||
self.len += 1;
|
||||
}
|
||||
}
|
||||
}
|
||||
fn u(&mut self, mut n: usize) {
|
||||
let mut tmp = [0u8; 20];
|
||||
let mut i = tmp.len();
|
||||
loop {
|
||||
i -= 1;
|
||||
tmp[i] = b'0' + (n % 10) as u8;
|
||||
n /= 10;
|
||||
if n == 0 {
|
||||
break;
|
||||
}
|
||||
}
|
||||
for &c in &tmp[i..] {
|
||||
if self.len < self.b.len() {
|
||||
self.b[self.len] = c;
|
||||
self.len += 1;
|
||||
}
|
||||
}
|
||||
}
|
||||
/// `idx.gen`, unpacked from the install-time packing.
|
||||
fn pid(&mut self, packed: u64) {
|
||||
self.u((packed >> 32) as usize);
|
||||
self.s(".");
|
||||
self.u((packed & 0xffff_ffff) as usize);
|
||||
}
|
||||
fn emit(&self) {
|
||||
unsafe {
|
||||
libc::write(2, self.b.as_ptr() as *const libc::c_void, self.len);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Classifier units — the arithmetic edges, before anything integrates.
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::{classify, FaultClass, OVERSHOOT_SLOP};
|
||||
|
||||
const PG: usize = 4096;
|
||||
// A synthetic stack far from address-space edges: top at 1 GiB.
|
||||
const TOP: usize = 1 << 30;
|
||||
const RESERVE: usize = 16 * PG;
|
||||
const GUARD: usize = 4 * PG;
|
||||
const GUARD_HI: usize = TOP - RESERVE;
|
||||
const GUARD_LO: usize = GUARD_HI - GUARD;
|
||||
|
||||
#[test]
|
||||
fn inside_guard_both_edges() {
|
||||
assert_eq!(classify(GUARD_LO, TOP, RESERVE, GUARD), FaultClass::Guard);
|
||||
assert_eq!(
|
||||
classify(GUARD_HI - 1, TOP, RESERVE, GUARD),
|
||||
FaultClass::Guard
|
||||
);
|
||||
assert_eq!(
|
||||
classify(GUARD_LO + GUARD / 2, TOP, RESERVE, GUARD),
|
||||
FaultClass::Guard
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn usable_region_is_foreign() {
|
||||
// A fault inside the RW stack itself isn't a guard hit and must not
|
||||
// be explained as one.
|
||||
assert_eq!(classify(GUARD_HI, TOP, RESERVE, GUARD), FaultClass::Foreign);
|
||||
assert_eq!(classify(TOP - 1, TOP, RESERVE, GUARD), FaultClass::Foreign);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn above_top_is_foreign() {
|
||||
assert_eq!(classify(TOP, TOP, RESERVE, GUARD), FaultClass::Foreign);
|
||||
assert_eq!(classify(TOP + PG, TOP, RESERVE, GUARD), FaultClass::Foreign);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn overshoot_window_edges() {
|
||||
assert_eq!(
|
||||
classify(GUARD_LO - 1, TOP, RESERVE, GUARD),
|
||||
FaultClass::Overshoot(1)
|
||||
);
|
||||
assert_eq!(
|
||||
classify(GUARD_LO - OVERSHOOT_SLOP, TOP, RESERVE, GUARD),
|
||||
FaultClass::Overshoot(OVERSHOOT_SLOP)
|
||||
);
|
||||
assert_eq!(
|
||||
classify(GUARD_LO - OVERSHOOT_SLOP - 1, TOP, RESERVE, GUARD),
|
||||
FaultClass::Foreign
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn low_address_stack_saturates_not_wraps() {
|
||||
// A stack mapped so low that the slop window would underflow: the
|
||||
// window clips to 0 instead of wrapping around the address space.
|
||||
let top = RESERVE + GUARD + PG; // guard_lo == PG
|
||||
assert_eq!(classify(0, top, RESERVE, GUARD), FaultClass::Overshoot(PG));
|
||||
// Null-page fault still classified only because it IS within slop
|
||||
// here; with a normal-height stack it is Foreign (covered above by
|
||||
// the window-edge test at realistic addresses).
|
||||
}
|
||||
}
|
||||
+9
-9
@@ -188,8 +188,7 @@ impl StateWord {
|
||||
loop {
|
||||
let w = self.load();
|
||||
debug_assert!(
|
||||
matches!(word_state(w), ST_RUNNING | ST_RUNNING_NOTIFIED)
|
||||
&& word_gen(w) == gen,
|
||||
matches!(word_state(w), ST_RUNNING | ST_RUNNING_NOTIFIED) && word_gen(w) == gen,
|
||||
"yield return from invalid word {w:#x}"
|
||||
);
|
||||
if self
|
||||
@@ -247,8 +246,7 @@ impl StateWord {
|
||||
loop {
|
||||
let w = self.load();
|
||||
debug_assert!(
|
||||
matches!(word_state(w), ST_RUNNING | ST_RUNNING_NOTIFIED)
|
||||
&& word_gen(w) == gen,
|
||||
matches!(word_state(w), ST_RUNNING | ST_RUNNING_NOTIFIED) && word_gen(w) == gen,
|
||||
"begin_wait from invalid word {w:#x}"
|
||||
);
|
||||
let next = word_epoch(w).wrapping_add(1) & EPOCH_MASK;
|
||||
@@ -342,8 +340,7 @@ impl StateWord {
|
||||
loop {
|
||||
let w = self.load();
|
||||
debug_assert!(
|
||||
matches!(word_state(w), ST_RUNNING | ST_RUNNING_NOTIFIED)
|
||||
&& word_gen(w) == gen,
|
||||
matches!(word_state(w), ST_RUNNING | ST_RUNNING_NOTIFIED) && word_gen(w) == gen,
|
||||
"clear_notify from invalid word {w:#x}"
|
||||
);
|
||||
if word_state(w) != ST_RUNNING_NOTIFIED {
|
||||
@@ -372,8 +369,7 @@ impl StateWord {
|
||||
pub(crate) fn set_done(&self, gen: u32) {
|
||||
let prev = self.0.swap(pack(gen, 0, ST_DONE), Ordering::AcqRel);
|
||||
debug_assert!(
|
||||
matches!(word_state(prev), ST_RUNNING | ST_RUNNING_NOTIFIED)
|
||||
&& word_gen(prev) == gen,
|
||||
matches!(word_state(prev), ST_RUNNING | ST_RUNNING_NOTIFIED) && word_gen(prev) == gen,
|
||||
"finalize from invalid word {prev:#x}"
|
||||
);
|
||||
}
|
||||
@@ -538,7 +534,11 @@ mod loom_tests {
|
||||
// not a pending notification.
|
||||
assert!(word.try_claim(0));
|
||||
assert_eq!(word.unpark(0, Some(epoch)), Unpark::Noop);
|
||||
assert_eq!(word_state(word.load()), ST_RUNNING, "stale epoch notified a live run");
|
||||
assert_eq!(
|
||||
word_state(word.load()),
|
||||
ST_RUNNING,
|
||||
"stale epoch notified a live run"
|
||||
);
|
||||
});
|
||||
}
|
||||
|
||||
|
||||
+237
-17
@@ -1,32 +1,45 @@
|
||||
//! mmap-based growable stack with a guard page below.
|
||||
//! mmap-based actor stack with a PROT_NONE guard region below (RFC 019).
|
||||
//!
|
||||
//! Layout (low → high address):
|
||||
//! [ guard page (PROT_NONE) | stack region ]
|
||||
//! ^ top() — initial stack pointer
|
||||
//! [ guard region (PROT_NONE) | stack region ]
|
||||
//! ^ top() — initial stack pointer
|
||||
//!
|
||||
//! Stacks grow downward. Overflow lands in the guard page → SIGSEGV.
|
||||
//! Stacks grow downward. Overflow lands in the guard region → SIGSEGV.
|
||||
//!
|
||||
//! Both the usable reserve and the guard are caller-chosen (page-rounded).
|
||||
//! The reserve is a *virtual* reservation: anonymous mmap is demand-paged,
|
||||
//! so RSS is touched-pages, not reserve × actors. The guard costs address
|
||||
//! space only. A wide guard (the runtime defaults to 64 KiB) exists for
|
||||
//! unprobed FFI frames: Rust frames touch pages in order (probestack), so
|
||||
//! one page catches Rust overflow, but a C frame with a large local can
|
||||
//! step over a single page in one `sub rsp`.
|
||||
|
||||
use std::io;
|
||||
|
||||
pub struct Stack {
|
||||
/// Bottom of the entire mmap'd region (start of guard page).
|
||||
/// Bottom of the entire mmap'd region (start of the guard).
|
||||
base: *mut u8,
|
||||
/// Total mmap'd size: guard_size + stack_size.
|
||||
total_size: usize,
|
||||
/// Usable stack size (excluding guard page).
|
||||
/// Usable stack size (excluding the guard).
|
||||
stack_size: usize,
|
||||
/// PROT_NONE region below the usable stack.
|
||||
guard_size: usize,
|
||||
}
|
||||
|
||||
// Stack owns its memory; safe to send across threads.
|
||||
unsafe impl Send for Stack {}
|
||||
|
||||
impl Stack {
|
||||
/// Allocate a new stack. `stack_size` is the usable region; one page is
|
||||
/// added below as a guard page. Both are rounded up to the page size.
|
||||
pub fn new(stack_size: usize) -> io::Result<Self> {
|
||||
/// Allocate a new stack. `stack_size` is the usable region; `guard_size`
|
||||
/// is mapped PROT_NONE below it. Both are rounded up to the page size
|
||||
/// and must be non-zero.
|
||||
pub fn new(stack_size: usize, guard_size: usize) -> io::Result<Self> {
|
||||
assert!(stack_size > 0, "stack_size must be non-zero");
|
||||
assert!(guard_size > 0, "guard_size must be non-zero");
|
||||
let page = page_size();
|
||||
let stack_size = round_up(stack_size, page);
|
||||
let guard_size = page;
|
||||
let guard_size = round_up(guard_size, page);
|
||||
let total_size = guard_size + stack_size;
|
||||
|
||||
let base = unsafe {
|
||||
@@ -44,16 +57,19 @@ impl Stack {
|
||||
}
|
||||
let base = base as *mut u8;
|
||||
|
||||
let ret = unsafe {
|
||||
libc::mprotect(base as *mut libc::c_void, guard_size, libc::PROT_NONE)
|
||||
};
|
||||
let ret = unsafe { libc::mprotect(base as *mut libc::c_void, guard_size, libc::PROT_NONE) };
|
||||
if ret != 0 {
|
||||
let err = io::Error::last_os_error();
|
||||
unsafe { libc::munmap(base as *mut libc::c_void, total_size) };
|
||||
return Err(err);
|
||||
}
|
||||
|
||||
Ok(Self { base, total_size, stack_size })
|
||||
Ok(Self {
|
||||
base,
|
||||
total_size,
|
||||
stack_size,
|
||||
guard_size,
|
||||
})
|
||||
}
|
||||
|
||||
/// 16-byte-aligned top of the usable region.
|
||||
@@ -62,14 +78,54 @@ impl Stack {
|
||||
(raw_top & !15) as *mut u8
|
||||
}
|
||||
|
||||
/// Pointer to the bottom of the usable region (just above the guard page).
|
||||
/// Pointer to the bottom of the usable region (just above the guard).
|
||||
pub fn usable_base(&self) -> *mut u8 {
|
||||
unsafe { self.base.add(page_size()) }
|
||||
unsafe { self.base.add(self.guard_size) }
|
||||
}
|
||||
|
||||
pub fn stack_size(&self) -> usize {
|
||||
self.stack_size
|
||||
}
|
||||
|
||||
pub fn guard_size(&self) -> usize {
|
||||
self.guard_size
|
||||
}
|
||||
|
||||
/// `(stack_size, guard_size)` after page rounding. The pool rule
|
||||
/// (RFC 019 §1) compares this against the runtime defaults: only
|
||||
/// default-shaped stacks are pooled.
|
||||
pub fn shape(&self) -> (usize, usize) {
|
||||
(self.stack_size, self.guard_size)
|
||||
}
|
||||
|
||||
/// Pool-recycle zap (RFC 019 §6): `MADV_DONTNEED` everything below the
|
||||
/// retained entry end `[top − retain, top)` — the span the next actor's
|
||||
/// shallow frames land in stays resident, the dead spike below it is
|
||||
/// released. The stack is unowned at the call site (its actor is dead),
|
||||
/// so a synchronous eager zap races nothing and the RSS drop is
|
||||
/// immediate — a museum of worst-case spikes is exactly what a pool must
|
||||
/// not be; DONTNEED's ~8× per-page cost vs FREE is irrelevant off the
|
||||
/// hot path. Advisory like the park-path shrink: a failure degrades to
|
||||
/// "the pool keeps RSS", never to incorrectness. No-op (no syscall) when
|
||||
/// `retain` covers the whole usable region — i.e. always, at the 64 KiB
|
||||
/// default reserve.
|
||||
pub(crate) fn recycle_zap(&self, retain: usize) {
|
||||
if let Some((off, len)) = retain_range(self.stack_size, retain, page_size()) {
|
||||
unsafe {
|
||||
libc::madvise(
|
||||
self.usable_base().add(off) as *mut libc::c_void,
|
||||
len,
|
||||
libc::MADV_DONTNEED,
|
||||
);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Round `n` up to whole pages — the same rounding `Stack::new` applies, so
|
||||
/// runtime defaults stored pre-rounded compare exactly against [`Stack::shape`].
|
||||
pub(crate) fn round_to_pages(n: usize) -> usize {
|
||||
round_up(n, page_size())
|
||||
}
|
||||
|
||||
impl Drop for Stack {
|
||||
@@ -80,10 +136,174 @@ impl Drop for Stack {
|
||||
}
|
||||
}
|
||||
|
||||
fn page_size() -> usize {
|
||||
pub(crate) fn page_size() -> usize {
|
||||
unsafe { libc::sysconf(libc::_SC_PAGESIZE) as usize }
|
||||
}
|
||||
|
||||
fn round_up(n: usize, align: usize) -> usize {
|
||||
(n + align - 1) & !(align - 1)
|
||||
}
|
||||
|
||||
/// The whole-page span the park-path shrink may `MADV_FREE` (RFC 019 §3):
|
||||
/// `[page_up(hwm), page_down(sp − redzone))`, or `None` if no full page fits.
|
||||
///
|
||||
/// `hwm` is the sampled high-water (deepest observed `sp`); everything in
|
||||
/// `[hwm, sp)` is below the live frame and dead by definition. One page of
|
||||
/// redzone stays resident under live `sp` — it covers the SysV 128-byte red
|
||||
/// zone plus spill margin with room to spare. Rounding is inward on both
|
||||
/// ends so the result can never touch the redzone, cross `sp`, or dip below
|
||||
/// `hwm`; all arithmetic is checked so adversarial inputs (`sp < redzone`,
|
||||
/// `hwm ≥ sp`, values near the address-space edges) collapse to `None`
|
||||
/// rather than a wild or negative-length range.
|
||||
pub(crate) fn shrink_range(hwm: usize, sp: usize, page: usize) -> Option<(usize, usize)> {
|
||||
debug_assert!(page.is_power_of_two());
|
||||
if hwm >= sp {
|
||||
return None;
|
||||
}
|
||||
let redzone = page;
|
||||
let end = sp.checked_sub(redzone)? & !(page - 1); // page_down(sp − redzone)
|
||||
let start = hwm.checked_add(page - 1)? & !(page - 1); // page_up(hwm)
|
||||
if end > start {
|
||||
Some((start, end - start))
|
||||
} else {
|
||||
None
|
||||
}
|
||||
}
|
||||
|
||||
/// The `(offset_from_usable_base, len)` span the pool recycle DONTNEEDs
|
||||
/// (RFC 019 §6): everything below the retained entry end. "Bottom RETAIN of
|
||||
/// the stack" is read stack-wise (entry frames = highest addresses of a
|
||||
/// downward stack): the retained span is `[top − page_up(retain), top)`, the
|
||||
/// zapped span is the rest — retaining the low-address deep end instead
|
||||
/// would keep the coldest pages and release the ones the next actor faults
|
||||
/// first. `retain` rounds *up* to whole pages (retain more, zap less), so
|
||||
/// with `stack_size` page-rounded by `Stack::new` the result is always
|
||||
/// page-aligned. Checked math: `retain ≥ stack_size` (notably the default
|
||||
/// 64 KiB reserve with the 64 KiB RETAIN) and overflow collapse to `None`.
|
||||
pub(crate) fn retain_range(
|
||||
stack_size: usize,
|
||||
retain: usize,
|
||||
page: usize,
|
||||
) -> Option<(usize, usize)> {
|
||||
debug_assert!(page.is_power_of_two());
|
||||
let retain = retain.checked_add(page - 1)? & !(page - 1); // page_up(retain)
|
||||
let len = stack_size.checked_sub(retain)?;
|
||||
if len == 0 {
|
||||
return None;
|
||||
}
|
||||
Some((0, len))
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::{retain_range, shrink_range};
|
||||
|
||||
const PG: usize = 4096;
|
||||
|
||||
#[test]
|
||||
fn retain_covers_whole_stack_is_a_noop() {
|
||||
// The default config: reserve == RETAIN == 64 KiB. No zap, no syscall.
|
||||
assert_eq!(retain_range(16 * PG, 16 * PG, PG), None);
|
||||
assert_eq!(retain_range(PG, PG, PG), None);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn retain_larger_than_stack_is_a_noop() {
|
||||
assert_eq!(retain_range(16 * PG, 17 * PG, PG), None);
|
||||
assert_eq!(retain_range(PG, usize::MAX, PG), None); // page_up overflows
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn retain_zero_zaps_everything() {
|
||||
assert_eq!(retain_range(16 * PG, 0, PG), Some((0, 16 * PG)));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn retain_rounds_up_zapping_less() {
|
||||
// 1 byte of retain keeps a whole page.
|
||||
assert_eq!(retain_range(16 * PG, 1, PG), Some((0, 15 * PG)));
|
||||
assert_eq!(retain_range(16 * PG, PG + 1, PG), Some((0, 14 * PG)));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn retain_one_page_short_of_stack() {
|
||||
assert_eq!(retain_range(2 * PG, PG, PG), Some((0, PG)));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn retain_range_is_page_aligned() {
|
||||
for size_pg in [1usize, 2, 3, 16, 1024] {
|
||||
for retain in [0usize, 1, PG - 1, PG, PG + 1, 4 * PG, size_pg * PG] {
|
||||
if let Some((off, len)) = retain_range(size_pg * PG, retain, PG) {
|
||||
assert_eq!(off, 0);
|
||||
assert_eq!(len % PG, 0);
|
||||
assert!(len <= size_pg * PG);
|
||||
assert!(len > 0);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn empty_and_inverted_spans_are_none() {
|
||||
assert_eq!(shrink_range(0x8000_0000, 0x8000_0000, PG), None); // hwm == sp
|
||||
assert_eq!(shrink_range(0x8000_1000, 0x8000_0000, PG), None); // hwm > sp
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn span_smaller_than_redzone_plus_page_is_none() {
|
||||
let sp = 0x8000_0000;
|
||||
// Everything within redzone+1 page of sp: no full page clears both
|
||||
// the redzone and the page_up(hwm) rounding.
|
||||
assert_eq!(shrink_range(sp - PG, sp, PG), None);
|
||||
assert_eq!(shrink_range(sp - 2 * PG + 1, sp, PG), None);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn exact_two_pages_frees_one() {
|
||||
let sp = 0x8000_0000;
|
||||
let hwm = sp - 2 * PG;
|
||||
// [hwm, hwm+PG) frees; [sp−PG, sp) is redzone.
|
||||
assert_eq!(shrink_range(hwm, sp, PG), Some((hwm, PG)));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn unaligned_ends_round_inward() {
|
||||
let sp = 0x8000_0123; // live sp mid-page
|
||||
let hwm = 0x7f00_0abc; // high-water mid-page
|
||||
let (start, len) = shrink_range(hwm, sp, PG).unwrap();
|
||||
assert_eq!(start % PG, 0);
|
||||
assert_eq!(len % PG, 0);
|
||||
assert!(start >= hwm); // never below the sampled high-water
|
||||
assert!(start + len <= (sp - PG) & !(PG - 1)); // never into the redzone
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn result_never_crosses_sp() {
|
||||
// Sweep hwm across every offset of the page straddling the boundary.
|
||||
let sp = 0x8000_0000 + 137;
|
||||
for hwm in (sp - 4 * PG)..(sp) {
|
||||
if let Some((start, len)) = shrink_range(hwm, sp, PG) {
|
||||
assert!(start >= hwm);
|
||||
assert!(start + len + PG <= sp + PG); // end ≤ page_down(sp − PG) < sp
|
||||
assert!(len > 0);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn underflow_near_zero_is_none() {
|
||||
assert_eq!(shrink_range(0, PG - 1, PG), None); // sp < redzone
|
||||
assert_eq!(shrink_range(0, 0, PG), None);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn big_span_frees_interior() {
|
||||
let sp = 0x8000_0000;
|
||||
let spike = 4 * 1024 * 1024;
|
||||
let hwm = sp - spike;
|
||||
let (start, len) = shrink_range(hwm, sp, PG).unwrap();
|
||||
assert_eq!(start, hwm); // aligned input: starts exactly at hwm
|
||||
assert_eq!(len, spike - PG); // everything but the redzone page
|
||||
}
|
||||
}
|
||||
|
||||
+197
-74
@@ -123,9 +123,9 @@ pub enum Signal {
|
||||
impl std::fmt::Debug for Signal {
|
||||
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
|
||||
match self {
|
||||
Signal::Exit(pid) => write!(f, "Signal::Exit({:?})", pid),
|
||||
Signal::Exit(pid) => write!(f, "Signal::Exit({:?})", pid),
|
||||
Signal::Panic(pid, _) => write!(f, "Signal::Panic({:?}, ..)", pid),
|
||||
Signal::Stopped(pid) => write!(f, "Signal::Stopped({:?})", pid),
|
||||
Signal::Stopped(pid) => write!(f, "Signal::Stopped({:?})", pid),
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -133,14 +133,15 @@ impl std::fmt::Debug for Signal {
|
||||
impl Signal {
|
||||
pub fn pid(&self) -> Pid {
|
||||
match self {
|
||||
Signal::Exit(p) => *p,
|
||||
Signal::Exit(p) => *p,
|
||||
Signal::Panic(p, _) => *p,
|
||||
Signal::Stopped(p) => *p,
|
||||
Signal::Stopped(p) => *p,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
use crate::channel::channel;
|
||||
use crate::channel::{channel, RecvTimeoutError};
|
||||
use crate::monitor::DownReason;
|
||||
use std::collections::{HashMap, VecDeque};
|
||||
use std::sync::Arc;
|
||||
use std::time::{Duration, Instant};
|
||||
@@ -165,11 +166,52 @@ pub enum Restart {
|
||||
pub struct ChildSpec {
|
||||
start: Arc<dyn Fn() + Send + Sync + 'static>,
|
||||
restart: Restart,
|
||||
shutdown: Shutdown,
|
||||
}
|
||||
|
||||
impl ChildSpec {
|
||||
/// A child with the given restart policy and the default
|
||||
/// [`Shutdown::Timeout`] of 5 seconds.
|
||||
pub fn new(restart: Restart, start: impl Fn() + Send + Sync + 'static) -> Self {
|
||||
Self { start: Arc::new(start), restart }
|
||||
Self {
|
||||
start: Arc::new(start),
|
||||
restart,
|
||||
shutdown: Shutdown::default(),
|
||||
}
|
||||
}
|
||||
|
||||
/// Set how the supervisor stops this child (see [`Shutdown`]). A child
|
||||
/// that is itself a supervisor should use [`Shutdown::Infinity`] so its
|
||||
/// own subtree gets its full grace periods.
|
||||
pub fn shutdown(mut self, shutdown: Shutdown) -> Self {
|
||||
self.shutdown = shutdown;
|
||||
self
|
||||
}
|
||||
}
|
||||
|
||||
/// How a supervisor stops a child it is taking down — the OTP child-spec
|
||||
/// `shutdown` value. Applies to every supervisor-initiated stop: the ordered
|
||||
/// shutdown of the whole set and the sibling cycling of
|
||||
/// [`Strategy::OneForAll`] / [`Strategy::RestForOne`].
|
||||
///
|
||||
/// A graceful stop is a [`request_shutdown`](crate::request_shutdown): a child
|
||||
/// that traps exits receives the request as a message and winds down in its
|
||||
/// own time; one that does not is stopped outright.
|
||||
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
|
||||
pub enum Shutdown {
|
||||
/// `request_stop` immediately; no request, no grace period.
|
||||
BrutalKill,
|
||||
/// `request_shutdown`, wait up to the duration for the child to exit, then
|
||||
/// `request_stop` it. The default, at 5 seconds.
|
||||
Timeout(Duration),
|
||||
/// `request_shutdown` and wait however long the child takes. Use for a
|
||||
/// child supervisor, whose subtree has its own timeouts.
|
||||
Infinity,
|
||||
}
|
||||
|
||||
impl Default for Shutdown {
|
||||
fn default() -> Self {
|
||||
Shutdown::Timeout(Duration::from_secs(5))
|
||||
}
|
||||
}
|
||||
|
||||
@@ -241,23 +283,35 @@ impl OneForOne {
|
||||
}
|
||||
|
||||
/// Run the supervision loop on the current actor. Returns when every child
|
||||
/// has reached a terminal, non-restartable state, or when the restart
|
||||
/// intensity cap is tripped.
|
||||
/// has reached a terminal, non-restartable state, when the restart
|
||||
/// intensity cap is tripped, or when the supervisor is asked to shut down
|
||||
/// (a [`request_shutdown`](crate::request_shutdown) — from its own
|
||||
/// supervisor, or from the app). On every one of those exits the survivors
|
||||
/// are stopped in reverse start order, each per its
|
||||
/// [`Shutdown`] policy, before this returns.
|
||||
///
|
||||
/// The supervisor traps exits for the length of the loop (that is how the
|
||||
/// shutdown request reaches it as a message). Should the supervisor itself
|
||||
/// be hard-stopped with [`request_stop`](crate::request_stop), it unwinds
|
||||
/// without waiting for anything — but a drop guard hard-stops its live
|
||||
/// children on the way out, so the subtree is not orphaned (a child
|
||||
/// supervisor unwinds the same way, recursively).
|
||||
pub fn run(self) {
|
||||
let me = crate::scheduler::self_pid();
|
||||
let (tx, rx) = channel::<Signal>();
|
||||
crate::scheduler::register_supervisor_channel(me, tx);
|
||||
let exits = crate::link::trap_exit();
|
||||
|
||||
// pid -> index into `self.children`, for the children currently alive.
|
||||
let mut by_pid: HashMap<Pid, usize> = HashMap::new();
|
||||
let mut live = Live::default();
|
||||
let mut active: usize = 0;
|
||||
// Sliding window of recent restart instants, for the intensity cap.
|
||||
let mut restarts: Vec<Instant> = Vec::new();
|
||||
|
||||
let start_child = |idx: usize, by_pid: &mut HashMap<Pid, usize>| {
|
||||
let start_child = |idx: usize, live: &mut Live| {
|
||||
let start = self.children[idx].start.clone();
|
||||
let h = crate::scheduler::spawn_under(me, move || (start)());
|
||||
by_pid.insert(h.pid(), idx);
|
||||
live.insert(h.pid(), idx);
|
||||
// We supervise via the signal funnel, not by joining; drop the
|
||||
// handle so the child's slot is reclaimed promptly on death (the
|
||||
// termination Signal is delivered before reclamation regardless).
|
||||
@@ -265,28 +319,105 @@ impl OneForOne {
|
||||
};
|
||||
|
||||
for idx in 0..self.children.len() {
|
||||
start_child(idx, &mut by_pid);
|
||||
start_child(idx, &mut live);
|
||||
active += 1;
|
||||
}
|
||||
|
||||
// A signal that arrives while we are awaiting stop-confirmations (for a
|
||||
// child we are *not* currently stopping) is stashed here and processed
|
||||
// by the main loop before it blocks on `recv` again.
|
||||
// by the main loop before it blocks again.
|
||||
let mut pending: VecDeque<Signal> = VecDeque::new();
|
||||
let next_signal = |pending: &mut VecDeque<Signal>| -> Option<Signal> {
|
||||
if let Some(s) = pending.pop_front() {
|
||||
Some(s)
|
||||
} else {
|
||||
rx.recv().ok()
|
||||
|
||||
// Stop one child per its policy and wait for its termination signal.
|
||||
// Signals for other pids that arrive meanwhile are stashed. Bounded by
|
||||
// construction: `request_stop` (used directly, or as the fallback once
|
||||
// the grace period lapses) always produces a signal.
|
||||
let stop_child = |pid: Pid, idx: usize, pending: &mut VecDeque<Signal>| {
|
||||
let await_one = |deadline: Option<Instant>, pending: &mut VecDeque<Signal>| -> bool {
|
||||
loop {
|
||||
let sig = match pending.iter().position(|s| s.pid() == pid) {
|
||||
Some(i) => pending.remove(i),
|
||||
None => match deadline {
|
||||
None => rx.recv().ok(),
|
||||
Some(dl) => {
|
||||
match rx.recv_timeout(dl.saturating_duration_since(Instant::now()))
|
||||
{
|
||||
Ok(s) => Some(s),
|
||||
Err(RecvTimeoutError::Timeout) => return false,
|
||||
Err(RecvTimeoutError::Disconnected) => None,
|
||||
}
|
||||
}
|
||||
},
|
||||
};
|
||||
match sig {
|
||||
Some(s) if s.pid() == pid => return true,
|
||||
Some(s) => pending.push_back(s),
|
||||
None => return true, // funnel closed: nothing more can arrive
|
||||
}
|
||||
}
|
||||
};
|
||||
match self.children[idx].shutdown {
|
||||
Shutdown::BrutalKill => {
|
||||
crate::scheduler::request_stop(pid);
|
||||
await_one(None, pending);
|
||||
}
|
||||
Shutdown::Timeout(grace) => {
|
||||
crate::scheduler::request_shutdown(pid);
|
||||
if !await_one(Some(Instant::now() + grace), pending) {
|
||||
crate::scheduler::request_stop(pid);
|
||||
await_one(None, pending);
|
||||
}
|
||||
}
|
||||
Shutdown::Infinity => {
|
||||
crate::scheduler::request_shutdown(pid);
|
||||
await_one(None, pending);
|
||||
}
|
||||
}
|
||||
};
|
||||
|
||||
// Stop a set of children in reverse start order, one at a time.
|
||||
let stop_set =
|
||||
|set: &mut Vec<(Pid, usize)>, live: &mut Live, pending: &mut VecDeque<Signal>| {
|
||||
set.sort_unstable_by_key(|x| std::cmp::Reverse(x.1));
|
||||
for (pid, idx) in set.iter() {
|
||||
live.remove(pid);
|
||||
stop_child(*pid, *idx, pending);
|
||||
}
|
||||
};
|
||||
|
||||
// Wait for the next event: a stashed signal, a child signal, or a
|
||||
// shutdown request. `Ok(sig)`, or `Err(())` when we must wind down.
|
||||
let next_event = |pending: &mut VecDeque<Signal>| -> Result<Signal, ()> {
|
||||
loop {
|
||||
if let Some(s) = pending.pop_front() {
|
||||
return Ok(s);
|
||||
}
|
||||
// The trap inbox is arm 0: a shutdown request is noticed even
|
||||
// under a flood of child signals.
|
||||
match crate::channel::select(&[&exits, &rx]) {
|
||||
0 => match exits.try_recv() {
|
||||
Ok(Some(sig)) if sig.reason == DownReason::Shutdown => return Err(()),
|
||||
// Any other exit signal (a linked peer's death — a
|
||||
// supervisor links nothing itself, but may be linked
|
||||
// to) is not ours to act on; a closed trap inbox is
|
||||
// impossible while `exits` is held here.
|
||||
_ => {}
|
||||
},
|
||||
_ => match rx.try_recv() {
|
||||
Ok(Some(s)) => return Ok(s),
|
||||
Ok(None) => {}
|
||||
Err(_) => return Err(()), // funnel closed: nothing left to supervise
|
||||
},
|
||||
}
|
||||
}
|
||||
};
|
||||
|
||||
while active > 0 {
|
||||
let sig = match next_signal(&mut pending) {
|
||||
Some(s) => s,
|
||||
None => break, // mailbox closed: nothing left to supervise
|
||||
let sig = match next_event(&mut pending) {
|
||||
Ok(s) => s,
|
||||
Err(()) => break,
|
||||
};
|
||||
let idx = match by_pid.remove(&sig.pid()) {
|
||||
let idx = match live.remove(&sig.pid()) {
|
||||
Some(i) => i,
|
||||
None => continue, // stray/duplicate signal
|
||||
};
|
||||
@@ -318,78 +449,70 @@ impl OneForOne {
|
||||
restarts.push(now);
|
||||
|
||||
// Which *live* siblings get cycled along with the failed child.
|
||||
// (The failed child is already gone — removed from `by_pid` above.)
|
||||
// (The failed child is already gone — removed from `live` above.)
|
||||
let mut to_stop: Vec<(Pid, usize)> = match self.strategy {
|
||||
Strategy::OneForOne => Vec::new(),
|
||||
Strategy::OneForAll => by_pid.iter().map(|(p, i)| (*p, *i)).collect(),
|
||||
Strategy::RestForOne => by_pid
|
||||
Strategy::OneForAll => live.iter().map(|(p, i)| (*p, *i)).collect(),
|
||||
Strategy::RestForOne => live
|
||||
.iter()
|
||||
.filter(|(_, i)| **i > idx)
|
||||
.map(|(p, i)| (*p, *i))
|
||||
.collect(),
|
||||
};
|
||||
// Stop survivors in reverse start order (highest child index first).
|
||||
to_stop.sort_unstable_by_key(|x| std::cmp::Reverse(x.1));
|
||||
|
||||
// The set we will restart: the failed child plus every sibling we
|
||||
// are about to stop, restarted in start (ascending index) order.
|
||||
let mut restart_set: Vec<usize> = Vec::with_capacity(to_stop.len() + 1);
|
||||
restart_set.push(idx);
|
||||
restart_set.extend(to_stop.iter().map(|(_, i)| *i));
|
||||
|
||||
// Request stops, then await each survivor's termination signal
|
||||
// before restarting. `request_stop` on an already-dead pid is a
|
||||
// no-op; in that case its (already-sent) Exit signal serves as the
|
||||
// confirmation. Any signal for a pid we are *not* awaiting is
|
||||
// stashed for the main loop.
|
||||
let mut awaiting: Vec<Pid> = Vec::with_capacity(to_stop.len());
|
||||
for (pid, cidx) in &to_stop {
|
||||
by_pid.remove(pid);
|
||||
restart_set.push(*cidx);
|
||||
crate::scheduler::request_stop(*pid);
|
||||
awaiting.push(*pid);
|
||||
}
|
||||
while !awaiting.is_empty() {
|
||||
let s = match next_signal(&mut pending) {
|
||||
Some(s) => s,
|
||||
None => break, // mailbox closed mid-await; stop waiting
|
||||
};
|
||||
if let Some(pos) = awaiting.iter().position(|p| *p == s.pid()) {
|
||||
awaiting.swap_remove(pos);
|
||||
} else {
|
||||
pending.push_back(s);
|
||||
}
|
||||
}
|
||||
|
||||
// Restart the whole set in start order. Net effect on `active`:
|
||||
// one child died (idx), `to_stop.len()` were stopped, and
|
||||
// `restart_set.len() == 1 + to_stop.len()` are started — so
|
||||
// Stop the survivors (each per its policy, reverse start order),
|
||||
// then restart the whole set in start order. Net effect on
|
||||
// `active`: one child died (idx), `to_stop.len()` were stopped,
|
||||
// and `restart_set.len() == 1 + to_stop.len()` are started — so
|
||||
// `active` is unchanged and needs no adjustment here.
|
||||
stop_set(&mut to_stop, &mut live, &mut pending);
|
||||
restart_set.sort_unstable();
|
||||
for cidx in restart_set {
|
||||
start_child(cidx, &mut by_pid);
|
||||
start_child(cidx, &mut live);
|
||||
}
|
||||
}
|
||||
|
||||
// Ordered shutdown: stop any survivors in reverse start order and await
|
||||
// their termination. On the normal `active == 0` exit `by_pid` is empty
|
||||
// and this is a no-op; on a cap-trip or mailbox-closed break it tears
|
||||
// the remaining children down deterministically instead of leaking them.
|
||||
let mut survivors: Vec<(Pid, usize)> = by_pid.iter().map(|(p, i)| (*p, *i)).collect();
|
||||
survivors.sort_unstable_by_key(|x| std::cmp::Reverse(x.1));
|
||||
let mut awaiting: Vec<Pid> = Vec::with_capacity(survivors.len());
|
||||
for (pid, _) in &survivors {
|
||||
crate::scheduler::request_stop(*pid);
|
||||
awaiting.push(*pid);
|
||||
}
|
||||
while !awaiting.is_empty() {
|
||||
let s = match next_signal(&mut pending) {
|
||||
Some(s) => s,
|
||||
None => break,
|
||||
};
|
||||
if let Some(pos) = awaiting.iter().position(|p| *p == s.pid()) {
|
||||
awaiting.swap_remove(pos);
|
||||
// Ordered shutdown: stop any survivors in reverse start order, each per
|
||||
// its policy. On the normal `active == 0` exit `live` is empty and this
|
||||
// is a no-op; on a shutdown request, a cap-trip, or a closed funnel it
|
||||
// tears the remaining children down deterministically.
|
||||
let mut survivors: Vec<(Pid, usize)> = live.iter().map(|(p, i)| (*p, *i)).collect();
|
||||
stop_set(&mut survivors, &mut live, &mut pending);
|
||||
}
|
||||
}
|
||||
|
||||
/// The live children of a supervisor, with a drop guard: if the supervisor is
|
||||
/// unwound (a hard `request_stop`, or a panic in the loop) its children are
|
||||
/// hard-stopped rather than orphaned. Fire-and-forget by necessity — a guard
|
||||
/// running mid-unwind cannot park to await anything.
|
||||
#[derive(Default)]
|
||||
struct Live(HashMap<Pid, usize>);
|
||||
|
||||
impl std::ops::Deref for Live {
|
||||
type Target = HashMap<Pid, usize>;
|
||||
fn deref(&self) -> &Self::Target {
|
||||
&self.0
|
||||
}
|
||||
}
|
||||
|
||||
impl std::ops::DerefMut for Live {
|
||||
fn deref_mut(&mut self) -> &mut Self::Target {
|
||||
&mut self.0
|
||||
}
|
||||
}
|
||||
|
||||
impl Drop for Live {
|
||||
fn drop(&mut self) {
|
||||
if std::thread::panicking() {
|
||||
for pid in self.0.keys() {
|
||||
crate::scheduler::request_stop(*pid);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
+11
-2
@@ -6,10 +6,19 @@
|
||||
//! Build the loom models with: `RUSTFLAGS="--cfg loom" cargo test --lib --release`
|
||||
|
||||
#[cfg(loom)]
|
||||
pub(crate) use loom::sync::atomic::{AtomicU64, AtomicUsize, Ordering};
|
||||
pub(crate) use loom::sync::atomic::{fence, AtomicU64, AtomicUsize, Ordering};
|
||||
|
||||
#[cfg(not(loom))]
|
||||
pub(crate) use std::sync::atomic::{AtomicU64, AtomicUsize, Ordering};
|
||||
pub(crate) use std::sync::atomic::{fence, AtomicU64, AtomicUsize, Ordering};
|
||||
|
||||
// park.rs condvar-parker (loom + non-Linux builds only; the Linux non-loom
|
||||
// build parks on a futex and never touches these — gating them identically
|
||||
// keeps the default build free of unused imports).
|
||||
#[cfg(loom)]
|
||||
pub(crate) use loom::sync::{Condvar, Mutex};
|
||||
|
||||
#[cfg(all(not(loom), not(target_os = "linux")))]
|
||||
pub(crate) use std::sync::{Condvar, Mutex};
|
||||
|
||||
/// `UnsafeCell` with loom's `with`/`with_mut` access API; pass-through cost
|
||||
/// is zero in normal builds (`#[inline]`, newtype over std's cell).
|
||||
|
||||
+40
-2
@@ -129,7 +129,9 @@ impl Ord for Entry {
|
||||
// Earlier deadline first; ties broken by insertion order so the
|
||||
// ordering is total. `Reason` and `Pid` deliberately don't
|
||||
// participate.
|
||||
self.deadline.cmp(&other.deadline).then_with(|| self.seq.cmp(&other.seq))
|
||||
self.deadline
|
||||
.cmp(&other.deadline)
|
||||
.then_with(|| self.seq.cmp(&other.seq))
|
||||
}
|
||||
}
|
||||
|
||||
@@ -141,6 +143,14 @@ impl PartialOrd for Entry {
|
||||
|
||||
#[derive(Default)]
|
||||
pub struct Timers {
|
||||
/// RFC 018: the scheduler coordination layer. Attached once at
|
||||
/// `RuntimeInner::new`; every insert notes its deadline (min-maintained
|
||||
/// snapshot for the busy-path due-check + the timekeeper re-arm wake)
|
||||
/// and every pop/clear re-anchors the snapshot to the heap minimum.
|
||||
/// All calls happen under the timers mutex — the serialization the
|
||||
/// coordinator's timer protocol mandates. `None` only in unit tests
|
||||
/// that construct a bare `Timers`.
|
||||
coord: Option<std::sync::Arc<crate::park::Coordinator>>,
|
||||
/// Reverse-wrapped so the smallest deadline is at the top.
|
||||
heap: BinaryHeap<Reverse<Entry>>,
|
||||
/// Monotonic counter for the tiebreaker `seq` field (and the `TimerId` of a
|
||||
@@ -157,7 +167,18 @@ pub struct Timers {
|
||||
|
||||
impl Timers {
|
||||
pub fn new() -> Self {
|
||||
Self { heap: BinaryHeap::new(), next_seq: 0, armed: std::collections::HashSet::new() }
|
||||
Self {
|
||||
coord: None,
|
||||
heap: BinaryHeap::new(),
|
||||
next_seq: 0,
|
||||
armed: std::collections::HashSet::new(),
|
||||
}
|
||||
}
|
||||
|
||||
/// Attach the scheduler coordination layer (RFC 018). Called once, at
|
||||
/// runtime construction, before any scheduler thread exists.
|
||||
pub(crate) fn attach_coordinator(&mut self, c: std::sync::Arc<crate::park::Coordinator>) {
|
||||
self.coord = Some(c);
|
||||
}
|
||||
|
||||
/// Insert a `Sleep` timer. Convenience for the common case.
|
||||
@@ -242,6 +263,13 @@ impl Timers {
|
||||
#[cfg(feature = "smarm-causal")]
|
||||
wall,
|
||||
}));
|
||||
// RFC 018: publish the (possibly new-minimum) deadline to the
|
||||
// busy-path snapshot and wake the timekeeper if it is parked
|
||||
// toward a later one. We hold the timers mutex — the mandated
|
||||
// serialization for both.
|
||||
if let Some(c) = &self.coord {
|
||||
c.note_deadline(deadline);
|
||||
}
|
||||
seq
|
||||
}
|
||||
|
||||
@@ -255,6 +283,9 @@ impl Timers {
|
||||
pub fn clear(&mut self) {
|
||||
self.heap.clear();
|
||||
self.armed.clear();
|
||||
if let Some(c) = &self.coord {
|
||||
c.refresh_deadline(None);
|
||||
}
|
||||
}
|
||||
|
||||
/// Soonest pending deadline, or `None` if the heap is empty.
|
||||
@@ -324,6 +355,13 @@ impl Timers {
|
||||
}
|
||||
out.push(entry);
|
||||
}
|
||||
// RFC 018: re-anchor the busy-path snapshot to the new heap minimum
|
||||
// (still under the timers mutex). A causal-shift re-queue above went
|
||||
// through `heap.push` directly, so this peek is the one place the
|
||||
// snapshot is guaranteed to catch up.
|
||||
if let Some(c) = &self.coord {
|
||||
c.refresh_deadline(self.peek_deadline());
|
||||
}
|
||||
out
|
||||
}
|
||||
}
|
||||
|
||||
+68
-30
@@ -16,13 +16,17 @@
|
||||
#[cfg(feature = "smarm-trace")]
|
||||
#[macro_export]
|
||||
macro_rules! te {
|
||||
($kind:expr) => { $crate::trace::record($kind) };
|
||||
($kind:expr) => {
|
||||
$crate::trace::record($kind)
|
||||
};
|
||||
}
|
||||
|
||||
#[cfg(not(feature = "smarm-trace"))]
|
||||
#[macro_export]
|
||||
macro_rules! te {
|
||||
($kind:expr) => { () };
|
||||
($kind:expr) => {
|
||||
()
|
||||
};
|
||||
}
|
||||
|
||||
#[cfg(feature = "smarm-trace")]
|
||||
@@ -42,17 +46,32 @@ mod inner {
|
||||
#[derive(Clone, Debug)]
|
||||
pub enum Event {
|
||||
// Actor lifecycle
|
||||
Spawn { parent: Pid, child: Pid },
|
||||
Spawn {
|
||||
parent: Pid,
|
||||
child: Pid,
|
||||
},
|
||||
Resume(Pid),
|
||||
Yield(Pid),
|
||||
Park(Pid),
|
||||
Done(Pid),
|
||||
/// Root exit found a live forest root (an actor nobody supervises)
|
||||
/// and delivered `request_shutdown` to it. `trapping` says whether it
|
||||
/// got the chance to drain (true) or was stopped outright (false).
|
||||
/// Every such line is an actor whose lifetime was nobody's business
|
||||
/// but the runtime's — the way to *see* unsupervised leftovers.
|
||||
RootSweep {
|
||||
target: Pid,
|
||||
trapping: bool,
|
||||
},
|
||||
// Wakeup paths
|
||||
UnparkDirect(Pid), // unpark() saw Parked -> re-queued immediately
|
||||
UnparkDeferred(Pid), // unpark() saw Runnable -> set pending_unpark flag
|
||||
UnparkFlagConsumed(Pid), // scheduler saw flag on Park -> re-queued instead
|
||||
// Channel
|
||||
Send { sender: Pid, receiver: Option<Pid> },
|
||||
Send {
|
||||
sender: Pid,
|
||||
receiver: Option<Pid>,
|
||||
},
|
||||
RecvPark(Pid),
|
||||
RecvWake(Pid),
|
||||
// Queue
|
||||
@@ -68,8 +87,8 @@ mod inner {
|
||||
// -----------------------------------------------------------------------
|
||||
|
||||
struct Record {
|
||||
nanos: u64, // ns since open()
|
||||
tid: u64, // OS thread id
|
||||
nanos: u64, // ns since open()
|
||||
tid: u64, // OS thread id
|
||||
event: Event,
|
||||
}
|
||||
|
||||
@@ -84,8 +103,8 @@ mod inner {
|
||||
// -----------------------------------------------------------------------
|
||||
|
||||
struct Global {
|
||||
sender: mpsc::Sender<Msg>,
|
||||
start: Instant,
|
||||
sender: mpsc::Sender<Msg>,
|
||||
start: Instant,
|
||||
}
|
||||
|
||||
static GLOBAL: Mutex<Option<Global>> = Mutex::new(None);
|
||||
@@ -95,7 +114,7 @@ mod inner {
|
||||
// The start Instant is copied alongside it — also one mutex hit per thread.
|
||||
// record() never touches GLOBAL after that.
|
||||
struct LocalState {
|
||||
tx: mpsc::Sender<Msg>,
|
||||
tx: mpsc::Sender<Msg>,
|
||||
start: Instant,
|
||||
}
|
||||
|
||||
@@ -109,8 +128,8 @@ mod inner {
|
||||
// -----------------------------------------------------------------------
|
||||
|
||||
pub fn open() {
|
||||
let path = std::env::var("SMARM_TRACE_FILE")
|
||||
.unwrap_or_else(|_| "smarm_trace.json".to_owned());
|
||||
let path =
|
||||
std::env::var("SMARM_TRACE_FILE").unwrap_or_else(|_| "smarm_trace.json".to_owned());
|
||||
|
||||
let (tx, rx) = mpsc::channel::<Msg>();
|
||||
let start = Instant::now();
|
||||
@@ -164,8 +183,11 @@ mod inner {
|
||||
// which would try to re-acquire inner.shared (already held at many
|
||||
// te!() call sites) -> deadlock. Guard at the very top, before any
|
||||
// allocation-capable call.
|
||||
let was_enabled = crate::preempt::PREEMPTION_ENABLED
|
||||
.with(|e| { let v = e.get(); e.set(false); v });
|
||||
let was_enabled = crate::preempt::PREEMPTION_ENABLED.with(|e| {
|
||||
let v = e.get();
|
||||
e.set(false);
|
||||
v
|
||||
});
|
||||
|
||||
LOCAL_STATE.with(|cell| {
|
||||
let mut opt = cell.borrow_mut();
|
||||
@@ -182,7 +204,7 @@ mod inner {
|
||||
}
|
||||
if let Some(ls) = opt.as_ref() {
|
||||
let nanos = ls.start.elapsed().as_nanos() as u64;
|
||||
let tid = os_tid();
|
||||
let tid = os_tid();
|
||||
let _ = ls.tx.send(Msg::Event(Record { nanos, tid, event }));
|
||||
}
|
||||
});
|
||||
@@ -197,7 +219,10 @@ mod inner {
|
||||
fn drain_thread(rx: mpsc::Receiver<Msg>, path: &str) {
|
||||
let f = match std::fs::File::create(path) {
|
||||
Ok(f) => f,
|
||||
Err(e) => { eprintln!("[smarm-trace] create failed: {}", e); return; }
|
||||
Err(e) => {
|
||||
eprintln!("[smarm-trace] create failed: {}", e);
|
||||
return;
|
||||
}
|
||||
};
|
||||
let mut w = std::io::BufWriter::new(f);
|
||||
let _ = writeln!(w, "{{\"traceEvents\":[");
|
||||
@@ -210,7 +235,9 @@ mod inner {
|
||||
Ok(Msg::Event(r)) => {
|
||||
let (name, actor_idx) = chrome_fields(&r.event);
|
||||
let ts_us = r.nanos as f64 / 1000.0;
|
||||
if !first { let _ = w.write_all(b",\n"); }
|
||||
if !first {
|
||||
let _ = w.write_all(b",\n");
|
||||
}
|
||||
first = false;
|
||||
let _ = write!(w,
|
||||
"{{\"ph\":\"i\",\"ts\":{:.3},\"pid\":{},\"tid\":{},\"name\":{:?},\"s\":\"g\"}}",
|
||||
@@ -234,27 +261,38 @@ mod inner {
|
||||
|
||||
fn chrome_fields(ev: &Event) -> (String, u32) {
|
||||
match ev {
|
||||
Event::Spawn { parent, child } =>
|
||||
(format!("spawn c={}", child.index()), parent.index()),
|
||||
Event::Resume(p) => ("resume".into(), p.index()),
|
||||
Event::Yield(p) => ("yield".into(), p.index()),
|
||||
Event::Park(p) => ("park".into(), p.index()),
|
||||
Event::Done(p) => ("done".into(), p.index()),
|
||||
Event::UnparkDirect(p) => ("unpark_direct".into(), p.index()),
|
||||
Event::UnparkDeferred(p) => ("unpark_deferred".into(), p.index()),
|
||||
Event::Spawn { parent, child } => {
|
||||
(format!("spawn c={}", child.index()), parent.index())
|
||||
}
|
||||
Event::Resume(p) => ("resume".into(), p.index()),
|
||||
Event::Yield(p) => ("yield".into(), p.index()),
|
||||
Event::Park(p) => ("park".into(), p.index()),
|
||||
Event::Done(p) => ("done".into(), p.index()),
|
||||
Event::RootSweep { target, trapping } => (
|
||||
format!(
|
||||
"root_sweep {}",
|
||||
if *trapping { "shutdown" } else { "stopped" }
|
||||
),
|
||||
target.index(),
|
||||
),
|
||||
Event::UnparkDirect(p) => ("unpark_direct".into(), p.index()),
|
||||
Event::UnparkDeferred(p) => ("unpark_deferred".into(), p.index()),
|
||||
Event::UnparkFlagConsumed(p) => ("unpark_flag_consumed".into(), p.index()),
|
||||
Event::Send { sender, receiver } => (
|
||||
format!("send rx={}", receiver
|
||||
.map(|p| p.index().to_string())
|
||||
.unwrap_or_else(|| "none".into())),
|
||||
format!(
|
||||
"send rx={}",
|
||||
receiver
|
||||
.map(|p| p.index().to_string())
|
||||
.unwrap_or_else(|| "none".into())
|
||||
),
|
||||
sender.index(),
|
||||
),
|
||||
Event::RecvPark(p) => ("recv_park".into(), p.index()),
|
||||
Event::RecvWake(p) => ("recv_wake".into(), p.index()),
|
||||
Event::Enqueue(p) => ("enqueue".into(), p.index()),
|
||||
Event::Dequeue(p) => ("dequeue".into(), p.index()),
|
||||
Event::Enqueue(p) => ("enqueue".into(), p.index()),
|
||||
Event::Dequeue(p) => ("dequeue".into(), p.index()),
|
||||
Event::SlotPush(p) => ("slot_push".into(), p.index()),
|
||||
Event::SlotPop(p) => ("slot_pop".into(), p.index()),
|
||||
Event::SlotPop(p) => ("slot_pop".into(), p.index()),
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
+24
-6
@@ -49,8 +49,14 @@ fn looping_actor_on_check_is_stopped() {
|
||||
}
|
||||
let _ = h.join();
|
||||
});
|
||||
assert!(saw_stopped.load(Ordering::SeqCst), "expected DownReason::Stopped");
|
||||
assert!(dropped.load(Ordering::SeqCst), "Drop guard must run during the cancellation unwind");
|
||||
assert!(
|
||||
saw_stopped.load(Ordering::SeqCst),
|
||||
"expected DownReason::Stopped"
|
||||
);
|
||||
assert!(
|
||||
dropped.load(Ordering::SeqCst),
|
||||
"Drop guard must run during the cancellation unwind"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
@@ -79,8 +85,14 @@ fn parked_on_recv_actor_is_stopped() {
|
||||
}
|
||||
let _ = h.join();
|
||||
});
|
||||
assert!(saw_stopped.load(Ordering::SeqCst), "expected DownReason::Stopped");
|
||||
assert!(dropped.load(Ordering::SeqCst), "Drop guard must run on cancellation of a parked actor");
|
||||
assert!(
|
||||
saw_stopped.load(Ordering::SeqCst),
|
||||
"expected DownReason::Stopped"
|
||||
);
|
||||
assert!(
|
||||
dropped.load(Ordering::SeqCst),
|
||||
"Drop guard must run on cancellation of a parked actor"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
@@ -185,6 +197,12 @@ fn stop_flagged_while_queued_lands_at_first_park() {
|
||||
.recv_timeout(Duration::from_secs(10))
|
||||
.expect("runtime deadlocked: stop against a QUEUED actor was lost at its first park");
|
||||
|
||||
assert!(saw_stopped.load(Ordering::SeqCst), "expected DownReason::Stopped");
|
||||
assert!(dropped.load(Ordering::SeqCst), "Drop guard must run during the cancellation unwind");
|
||||
assert!(
|
||||
saw_stopped.load(Ordering::SeqCst),
|
||||
"expected DownReason::Stopped"
|
||||
);
|
||||
assert!(
|
||||
dropped.load(Ordering::SeqCst),
|
||||
"Drop guard must run during the cancellation unwind"
|
||||
);
|
||||
}
|
||||
|
||||
+25
-26
@@ -24,7 +24,11 @@ fn progress_point_counts() {
|
||||
h.join().unwrap();
|
||||
let after = smarm::causal::progress_snapshot();
|
||||
let delta = |name: &str| {
|
||||
after.iter().find(|(n, _)| n == name).map(|(_, c)| *c).unwrap()
|
||||
after
|
||||
.iter()
|
||||
.find(|(n, _)| n == name)
|
||||
.map(|(_, c)| *c)
|
||||
.unwrap()
|
||||
- before
|
||||
.iter()
|
||||
.find(|(n, _)| n == name)
|
||||
@@ -45,21 +49,12 @@ fn site_guard_nesting_restores() {
|
||||
assert_eq!(smarm::causal::current_site_name(), None);
|
||||
{
|
||||
let _outer = smarm::causal_site!("outer");
|
||||
assert_eq!(
|
||||
smarm::causal::current_site_name().as_deref(),
|
||||
Some("outer")
|
||||
);
|
||||
assert_eq!(smarm::causal::current_site_name().as_deref(), Some("outer"));
|
||||
{
|
||||
let _inner = smarm::causal_site!("inner");
|
||||
assert_eq!(
|
||||
smarm::causal::current_site_name().as_deref(),
|
||||
Some("inner")
|
||||
);
|
||||
assert_eq!(smarm::causal::current_site_name().as_deref(), Some("inner"));
|
||||
}
|
||||
assert_eq!(
|
||||
smarm::causal::current_site_name().as_deref(),
|
||||
Some("outer")
|
||||
);
|
||||
assert_eq!(smarm::causal::current_site_name().as_deref(), Some("outer"));
|
||||
}
|
||||
assert_eq!(smarm::causal::current_site_name(), None);
|
||||
});
|
||||
@@ -88,10 +83,7 @@ fn virtual_speedup_ledger() {
|
||||
let bystander = smarm::spawn(move || {
|
||||
while !stop2.load(Ordering::Relaxed) {
|
||||
smarm::check!();
|
||||
out2.store(
|
||||
smarm::causal::my_absorbed_delay_cycles(),
|
||||
Ordering::Relaxed,
|
||||
);
|
||||
out2.store(smarm::causal::my_absorbed_delay_cycles(), Ordering::Relaxed);
|
||||
}
|
||||
});
|
||||
|
||||
@@ -166,10 +158,7 @@ fn runnable_bystander_pays_delay() {
|
||||
while !stop_b.load(Ordering::Relaxed) {
|
||||
iters2.fetch_add(1, Ordering::Relaxed);
|
||||
smarm::check!();
|
||||
absorbed2.store(
|
||||
smarm::causal::my_absorbed_delay_cycles(),
|
||||
Ordering::Relaxed,
|
||||
);
|
||||
absorbed2.store(smarm::causal::my_absorbed_delay_cycles(), Ordering::Relaxed);
|
||||
}
|
||||
});
|
||||
|
||||
@@ -177,8 +166,7 @@ fn runnable_bystander_pays_delay() {
|
||||
let i0 = iters.load(Ordering::Relaxed);
|
||||
let t = std::time::Instant::now();
|
||||
smarm::sleep(Duration::from_millis(150));
|
||||
let rate =
|
||||
(iters.load(Ordering::Relaxed) - i0) as f64 / t.elapsed().as_secs_f64();
|
||||
let rate = (iters.load(Ordering::Relaxed) - i0) as f64 / t.elapsed().as_secs_f64();
|
||||
out.store(rate as u64, Ordering::Relaxed);
|
||||
};
|
||||
|
||||
@@ -448,7 +436,10 @@ fn timer_deadline_shifts_with_injected_delay() {
|
||||
|
||||
// Raw deadline passed, effective deadline not: nothing fires, entry kept.
|
||||
assert!(t.pop_due(now + Duration::from_millis(60)).is_empty());
|
||||
assert!(!t.is_empty(), "shifted entry must be re-queued, not dropped");
|
||||
assert!(
|
||||
!t.is_empty(),
|
||||
"shifted entry must be re-queued, not dropped"
|
||||
);
|
||||
|
||||
// Past raw + injected (with margin) it must fire. Chase in case a
|
||||
// parallel test injected more debt meanwhile.
|
||||
@@ -503,7 +494,11 @@ fn wall_timer_ignores_injected_delay() {
|
||||
// Just past the raw deadline: the wall entry fires, the virtual one is
|
||||
// re-queued at its shifted deadline.
|
||||
let due = t.pop_due(now + Duration::from_millis(60));
|
||||
assert_eq!(due.len(), 1, "exactly the wall entry must fire at raw deadline");
|
||||
assert_eq!(
|
||||
due.len(),
|
||||
1,
|
||||
"exactly the wall entry must fire at raw deadline"
|
||||
);
|
||||
assert_eq!(due[0].pid, Pid::new(0, 0));
|
||||
assert!(!t.is_empty(), "virtual sibling must remain queued, shifted");
|
||||
}
|
||||
@@ -623,7 +618,11 @@ fn wall_send_after_ignores_injected_delay() {
|
||||
|
||||
// Just past the raw deadline: only the wall send pops; run its thunk.
|
||||
let due = t.pop_due(now + Duration::from_millis(60));
|
||||
assert_eq!(due.len(), 1, "exactly the wall send must fire at raw deadline");
|
||||
assert_eq!(
|
||||
due.len(),
|
||||
1,
|
||||
"exactly the wall send must fire at raw deadline"
|
||||
);
|
||||
for e in due {
|
||||
if let smarm::timer::Reason::Send { fire } = e.reason {
|
||||
fire();
|
||||
|
||||
+18
-10
@@ -154,7 +154,10 @@ fn channel_ops_interleaved_with_monitor_churn_multi_thread() {
|
||||
}
|
||||
consumer.join().unwrap();
|
||||
});
|
||||
assert_eq!(total.load(std::sync::atomic::Ordering::Relaxed), (0..32).sum::<i64>());
|
||||
assert_eq!(
|
||||
total.load(std::sync::atomic::Ordering::Relaxed),
|
||||
(0..32).sum::<i64>()
|
||||
);
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
@@ -220,7 +223,10 @@ fn recv_timeout_reports_disconnected_on_close() {
|
||||
fn recv_timeout_zero_duration_is_a_bounded_poll() {
|
||||
run(|| {
|
||||
let (_tx, rx) = channel::<i64>();
|
||||
assert_eq!(rx.recv_timeout(Duration::ZERO), Err(RecvTimeoutError::Timeout));
|
||||
assert_eq!(
|
||||
rx.recv_timeout(Duration::ZERO),
|
||||
Err(RecvTimeoutError::Timeout)
|
||||
);
|
||||
});
|
||||
}
|
||||
|
||||
@@ -262,15 +268,17 @@ fn recv_timeout_many_waiters_multi_thread() {
|
||||
let (tx, rx) = channel::<i64>();
|
||||
let got = got2.clone();
|
||||
let timed_out = timed_out2.clone();
|
||||
handles.push(spawn(move || match rx.recv_timeout(Duration::from_millis(100)) {
|
||||
Ok(v) => {
|
||||
assert_eq!(v, i);
|
||||
got.fetch_add(1, Ordering::Relaxed);
|
||||
handles.push(spawn(move || {
|
||||
match rx.recv_timeout(Duration::from_millis(100)) {
|
||||
Ok(v) => {
|
||||
assert_eq!(v, i);
|
||||
got.fetch_add(1, Ordering::Relaxed);
|
||||
}
|
||||
Err(RecvTimeoutError::Timeout) => {
|
||||
timed_out.fetch_add(1, Ordering::Relaxed);
|
||||
}
|
||||
Err(e) => panic!("unexpected: {e}"),
|
||||
}
|
||||
Err(RecvTimeoutError::Timeout) => {
|
||||
timed_out.fetch_add(1, Ordering::Relaxed);
|
||||
}
|
||||
Err(e) => panic!("unexpected: {e}"),
|
||||
}));
|
||||
if i % 2 == 0 {
|
||||
handles.push(spawn(move || {
|
||||
|
||||
+35
-16
@@ -11,9 +11,15 @@ thread_local! {
|
||||
static LOG: Cell<u64> = const { Cell::new(0) };
|
||||
}
|
||||
|
||||
fn log(v: u64) { LOG.with(|c| c.set(c.get() | v)); }
|
||||
fn get_log() -> u64 { LOG.with(|c| c.get()) }
|
||||
fn reset_log() { LOG.with(|c| c.set(0)); }
|
||||
fn log(v: u64) {
|
||||
LOG.with(|c| c.set(c.get() | v));
|
||||
}
|
||||
fn get_log() -> u64 {
|
||||
LOG.with(|c| c.get())
|
||||
}
|
||||
fn reset_log() {
|
||||
LOG.with(|c| c.set(0));
|
||||
}
|
||||
|
||||
extern "C-unwind" fn actor_simple() {
|
||||
log(0x1);
|
||||
@@ -23,7 +29,7 @@ extern "C-unwind" fn actor_simple() {
|
||||
#[test]
|
||||
fn actor_runs_and_returns_to_scheduler() {
|
||||
reset_log();
|
||||
let stack = Stack::new(64 * 1024).unwrap();
|
||||
let stack = Stack::new(64 * 1024, 4096).unwrap();
|
||||
let sp = init_actor_stack(stack.top(), actor_simple);
|
||||
set_actor_sp(sp);
|
||||
unsafe { switch_to_actor() };
|
||||
@@ -40,7 +46,7 @@ extern "C-unwind" fn actor_two_steps() {
|
||||
#[test]
|
||||
fn actor_yields_and_resumes() {
|
||||
reset_log();
|
||||
let stack = Stack::new(64 * 1024).unwrap();
|
||||
let stack = Stack::new(64 * 1024, 4096).unwrap();
|
||||
let sp = init_actor_stack(stack.top(), actor_two_steps);
|
||||
set_actor_sp(sp);
|
||||
|
||||
@@ -56,7 +62,7 @@ fn actor_yields_and_resumes() {
|
||||
use std::sync::OnceLock;
|
||||
|
||||
static REG_BEFORE: OnceLock<[u64; 4]> = OnceLock::new();
|
||||
static REG_AFTER: OnceLock<[u64; 4]> = OnceLock::new();
|
||||
static REG_AFTER: OnceLock<[u64; 4]> = OnceLock::new();
|
||||
|
||||
extern "C-unwind" fn actor_reg_check() {
|
||||
unsafe {
|
||||
@@ -73,7 +79,10 @@ extern "C-unwind" fn actor_reg_check() {
|
||||
REG_BEFORE.set([s0, s1, s2, s3]).ok();
|
||||
switch_to_scheduler();
|
||||
|
||||
let a0: u64; let a1: u64; let a2: u64; let a3: u64;
|
||||
let a0: u64;
|
||||
let a1: u64;
|
||||
let a2: u64;
|
||||
let a3: u64;
|
||||
core::arch::asm!(
|
||||
"mov {a0}, r12", "mov {a1}, r13", "mov {a2}, r14", "mov {a3}, r15",
|
||||
a0 = out(reg) a0, a1 = out(reg) a1, a2 = out(reg) a2, a3 = out(reg) a3,
|
||||
@@ -85,11 +94,17 @@ extern "C-unwind" fn actor_reg_check() {
|
||||
|
||||
#[test]
|
||||
fn callee_saved_registers_survive_yield() {
|
||||
let stack = Stack::new(64 * 1024).unwrap();
|
||||
let stack = Stack::new(64 * 1024, 4096).unwrap();
|
||||
let sp = init_actor_stack(stack.top(), actor_reg_check);
|
||||
set_actor_sp(sp);
|
||||
unsafe { switch_to_actor(); switch_to_actor(); }
|
||||
assert_eq!(REG_BEFORE.get().copied().unwrap(), REG_AFTER.get().copied().unwrap());
|
||||
unsafe {
|
||||
switch_to_actor();
|
||||
switch_to_actor();
|
||||
}
|
||||
assert_eq!(
|
||||
REG_BEFORE.get().copied().unwrap(),
|
||||
REG_AFTER.get().copied().unwrap()
|
||||
);
|
||||
}
|
||||
|
||||
// Two actors, independent stacks.
|
||||
@@ -117,20 +132,24 @@ extern "C-unwind" fn actor_b() {
|
||||
|
||||
#[test]
|
||||
fn two_actors_dont_corrupt_each_other() {
|
||||
let stack_a = Stack::new(64 * 1024).unwrap();
|
||||
let stack_b = Stack::new(64 * 1024).unwrap();
|
||||
let stack_a = Stack::new(64 * 1024, 4096).unwrap();
|
||||
let stack_b = Stack::new(64 * 1024, 4096).unwrap();
|
||||
|
||||
let sp_a = init_actor_stack(stack_a.top(), actor_a);
|
||||
let sp_b = init_actor_stack(stack_b.top(), actor_b);
|
||||
|
||||
set_actor_sp(sp_a); unsafe { switch_to_actor() };
|
||||
set_actor_sp(sp_a);
|
||||
unsafe { switch_to_actor() };
|
||||
let sp_a = get_actor_sp();
|
||||
|
||||
set_actor_sp(sp_b); unsafe { switch_to_actor() };
|
||||
set_actor_sp(sp_b);
|
||||
unsafe { switch_to_actor() };
|
||||
let sp_b = get_actor_sp();
|
||||
|
||||
set_actor_sp(sp_a); unsafe { switch_to_actor() };
|
||||
set_actor_sp(sp_b); unsafe { switch_to_actor() };
|
||||
set_actor_sp(sp_a);
|
||||
unsafe { switch_to_actor() };
|
||||
set_actor_sp(sp_b);
|
||||
unsafe { switch_to_actor() };
|
||||
|
||||
assert_eq!(A_VAL.with(|c| c.get()), 0xA00D);
|
||||
assert_eq!(B_VAL.with(|c| c.get()), 0xB00D);
|
||||
|
||||
@@ -0,0 +1,125 @@
|
||||
//! Cross-thread wake: a thread that is *not* a smarm scheduler thread must be
|
||||
//! able to wake (and stop) a parked actor.
|
||||
//!
|
||||
//! The gap this pins down: every off-runtime wake primitive (`unpark`,
|
||||
//! `unpark_at`, `request_stop`) reaches the runtime through the `RUNTIME`
|
||||
//! thread-local, which is `None` on any non-scheduler thread — so a wake
|
||||
//! issued from a foreign OS thread is a silent no-op and the parked actor
|
||||
//! sleeps forever. Both failure modes below manifest as `Runtime::run` never
|
||||
//! returning, so each test is wrapped in a watchdog: a timeout is the failure.
|
||||
//!
|
||||
//! The fix mirrors RFC 018's IO backend — the waker reaches the runtime
|
||||
//! through a `Weak<RuntimeInner>` it already holds (the receiver captures one
|
||||
//! when it parks; `Runtime::handle()` hands one to an app thread).
|
||||
|
||||
use std::sync::mpsc;
|
||||
use std::thread;
|
||||
use std::time::Duration;
|
||||
|
||||
const WATCHDOG: Duration = Duration::from_secs(10);
|
||||
/// Give the target actor time to actually park before the foreign thread pokes
|
||||
/// it, so we exercise the *wake* of a parked actor rather than the entry-side
|
||||
/// stop check.
|
||||
const SETTLE: Duration = Duration::from_millis(200);
|
||||
|
||||
fn assert_send_sync<T: Send + Sync>() {}
|
||||
|
||||
/// A cross-thread `send` from a plain OS thread must wake a receiver parked in
|
||||
/// `recv`. Under the thread-local-only wake path the send enqueues the message
|
||||
/// but never wakes the receiver, so `recv` — and therefore `run` — hangs.
|
||||
#[test]
|
||||
fn foreign_thread_send_wakes_parked_receiver() {
|
||||
let (done_tx, done_rx) = mpsc::channel();
|
||||
thread::spawn(move || {
|
||||
let rt = smarm::init(smarm::Config::exact(2));
|
||||
rt.run(|| {
|
||||
let (tx, rx) = smarm::channel::<u32>();
|
||||
// Receiver actor: parks on recv until the foreign thread sends.
|
||||
let h = smarm::spawn(move || {
|
||||
assert_eq!(rx.recv().expect("recv"), 42);
|
||||
});
|
||||
// Foreign (non-scheduler) OS thread owns the Sender and sends
|
||||
// after the receiver has parked.
|
||||
let sender = thread::spawn(move || {
|
||||
thread::sleep(SETTLE);
|
||||
tx.send(42).expect("send");
|
||||
});
|
||||
let _ = h.join();
|
||||
sender.join().expect("sender thread");
|
||||
});
|
||||
let _ = done_tx.send(());
|
||||
});
|
||||
done_rx
|
||||
.recv_timeout(WATCHDOG)
|
||||
.expect("run did not return: a foreign-thread send never woke the parked receiver");
|
||||
}
|
||||
|
||||
/// A cross-thread `request_stop` through a `RuntimeHandle` must wake and stop a
|
||||
/// parked actor. The actor parks on a long sleep (only a stop can end it); the
|
||||
/// handle is grabbed before `run` and driven from a foreign thread.
|
||||
#[test]
|
||||
fn foreign_thread_request_stop_wakes_parked_actor() {
|
||||
assert_send_sync::<smarm::RuntimeHandle>();
|
||||
|
||||
let rt = smarm::init(smarm::Config::exact(2));
|
||||
let handle = rt.handle();
|
||||
|
||||
// Foreign thread: learn the target pid from inside the run, let it park,
|
||||
// then stop it through the handle.
|
||||
let (pid_tx, pid_rx) = mpsc::channel::<smarm::Pid>();
|
||||
let stopper = thread::spawn(move || {
|
||||
let pid = pid_rx.recv().expect("pid");
|
||||
thread::sleep(SETTLE);
|
||||
handle.request_stop(pid);
|
||||
});
|
||||
|
||||
let (done_tx, done_rx) = mpsc::channel();
|
||||
thread::spawn(move || {
|
||||
rt.run(move || {
|
||||
let h = smarm::spawn(|| {
|
||||
// Parks indefinitely; only a cooperative stop unwinds it.
|
||||
smarm::sleep(Duration::from_secs(3600));
|
||||
});
|
||||
pid_tx.send(h.pid()).expect("send pid");
|
||||
let _ = h.join();
|
||||
});
|
||||
let _ = done_tx.send(());
|
||||
});
|
||||
|
||||
done_rx
|
||||
.recv_timeout(WATCHDOG)
|
||||
.expect("run did not return: a foreign-thread request_stop never woke the parked actor");
|
||||
stopper.join().expect("stopper thread");
|
||||
}
|
||||
|
||||
/// A `RuntimeHandle` held across (and beyond) a run must not keep the runtime
|
||||
/// alive or block all-done: `run` still returns, and once the `Runtime` is
|
||||
/// dropped the handle degrades to a harmless no-op (Weak lifecycle) rather than
|
||||
/// panicking or touching freed memory.
|
||||
#[test]
|
||||
fn lingering_handle_does_not_block_all_done() {
|
||||
let rt = smarm::init(smarm::Config::exact(1));
|
||||
let handle = rt.handle(); // outlives the run below
|
||||
|
||||
let (pid_tx, pid_rx) = mpsc::channel::<smarm::Pid>();
|
||||
let (done_tx, done_rx) = mpsc::channel();
|
||||
let runner = thread::spawn(move || {
|
||||
rt.run(move || {
|
||||
let h = smarm::spawn(|| {});
|
||||
pid_tx.send(h.pid()).expect("send pid");
|
||||
let _ = h.join();
|
||||
});
|
||||
// `rt` is dropped here, at the end of this thread.
|
||||
let _ = done_tx.send(());
|
||||
});
|
||||
|
||||
done_rx
|
||||
.recv_timeout(WATCHDOG)
|
||||
.expect("run did not return while a RuntimeHandle was held live");
|
||||
runner.join().expect("runner thread");
|
||||
|
||||
// Runtime is now dropped. A stop through the lingering handle must be a
|
||||
// silent no-op, not a panic or use-after-free.
|
||||
let dead_pid = pid_rx.recv().expect("pid");
|
||||
handle.request_stop(dead_pid);
|
||||
}
|
||||
+22
-7
@@ -11,8 +11,8 @@
|
||||
//! OUTSIDE `run` — an in-actor assertion alone passes vacuously.
|
||||
|
||||
use smarm::{
|
||||
channel, run, select, select_timeout, spawn, try_select, wait_readable,
|
||||
wait_readable_timeout, wait_writable_timeout, yield_now, FdArm,
|
||||
channel, run, select, select_timeout, spawn, try_select, wait_readable, wait_readable_timeout,
|
||||
wait_writable_timeout, yield_now, FdArm,
|
||||
};
|
||||
use std::os::fd::RawFd;
|
||||
use std::sync::atomic::{AtomicBool, AtomicU32, Ordering};
|
||||
@@ -33,7 +33,10 @@ impl Pipe {
|
||||
let mut fds: [libc::c_int; 2] = [0; 2];
|
||||
let r = unsafe { libc::pipe2(fds.as_mut_ptr(), libc::O_CLOEXEC | libc::O_NONBLOCK) };
|
||||
assert_eq!(r, 0, "pipe2 failed");
|
||||
Pipe { read: fds[0], write: fds[1] }
|
||||
Pipe {
|
||||
read: fds[0],
|
||||
write: fds[1],
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -253,12 +256,18 @@ fn wait_readable_timeout_times_out_then_succeeds_with_data() {
|
||||
let (rfd, wfd) = (p.read, p.write);
|
||||
|
||||
let start = Instant::now();
|
||||
assert_eq!(wait_readable_timeout(rfd, Duration::from_millis(30)).unwrap(), false);
|
||||
assert_eq!(
|
||||
wait_readable_timeout(rfd, Duration::from_millis(30)).unwrap(),
|
||||
false
|
||||
);
|
||||
assert!(start.elapsed() >= Duration::from_millis(30));
|
||||
|
||||
// Timed-out wait must leave the fd clean; ready path returns true.
|
||||
assert_eq!(raw_write(wfd, b"d"), 1);
|
||||
assert_eq!(wait_readable_timeout(rfd, Duration::from_secs(5)).unwrap(), true);
|
||||
assert_eq!(
|
||||
wait_readable_timeout(rfd, Duration::from_secs(5)).unwrap(),
|
||||
true
|
||||
);
|
||||
let mut buf = [0u8; 1];
|
||||
assert_eq!(raw_read(rfd, &mut buf), 1);
|
||||
ok2.store(true, Ordering::SeqCst);
|
||||
@@ -274,7 +283,10 @@ fn wait_readable_timeout_wakes_on_late_data() {
|
||||
let p = Pipe::new();
|
||||
let (rfd, wfd) = (p.read, p.write);
|
||||
let h = spawn(move || {
|
||||
assert_eq!(wait_readable_timeout(rfd, Duration::from_secs(5)).unwrap(), true);
|
||||
assert_eq!(
|
||||
wait_readable_timeout(rfd, Duration::from_secs(5)).unwrap(),
|
||||
true
|
||||
);
|
||||
let mut buf = [0u8; 1];
|
||||
assert_eq!(raw_read(rfd, &mut buf), 1);
|
||||
got2.store(buf[0] as u32, Ordering::SeqCst);
|
||||
@@ -292,7 +304,10 @@ fn wait_writable_timeout_ready_now_on_empty_pipe() {
|
||||
run(move || {
|
||||
let p = Pipe::new();
|
||||
// An empty pipe's write end is writable: ready-now path, no park.
|
||||
assert_eq!(wait_writable_timeout(p.write, Duration::from_secs(5)).unwrap(), true);
|
||||
assert_eq!(
|
||||
wait_writable_timeout(p.write, Duration::from_secs(5)).unwrap(),
|
||||
true
|
||||
);
|
||||
ok2.store(true, Ordering::SeqCst);
|
||||
});
|
||||
assert!(ok.load(Ordering::SeqCst));
|
||||
|
||||
+64
-20
@@ -88,7 +88,7 @@ impl GenServer for Lifecycle {
|
||||
}
|
||||
}
|
||||
|
||||
// init -> handle_call -> (drop last ref closes inbox) -> terminate.
|
||||
// init -> handle_call -> shutdown -> terminate.
|
||||
#[test]
|
||||
fn init_and_terminate_run() {
|
||||
let log = Arc::new(Mutex::new(Vec::new()));
|
||||
@@ -96,9 +96,9 @@ fn init_and_terminate_run() {
|
||||
run(move || {
|
||||
let server = start(Lifecycle { log: log2 });
|
||||
server.call(()).unwrap();
|
||||
// Dropping the only ref closes the inbox; the server breaks out of its
|
||||
// recv loop and runs terminate. run() will not return until it has.
|
||||
drop(server);
|
||||
// Refs are addresses: dropping one does not end the server. The
|
||||
// explicit close does, and waits for terminate.
|
||||
server.shutdown();
|
||||
});
|
||||
assert_eq!(*log.lock().unwrap(), vec!["init", "call", "terminate"]);
|
||||
}
|
||||
@@ -403,7 +403,10 @@ fn worker_pool_down_reaches_handle_down() {
|
||||
let got = Arc::new(Mutex::new(Vec::new()));
|
||||
let got2 = got.clone();
|
||||
run(move || {
|
||||
let server = start(Pool { watcher: None, log: Vec::new() });
|
||||
let server = start(Pool {
|
||||
watcher: None,
|
||||
log: Vec::new(),
|
||||
});
|
||||
server.cast(PoolCast::SpawnDoomedWorker).unwrap();
|
||||
let _ = server.call(()).unwrap(); // sync point: cast handled, worker live
|
||||
*got2.lock().unwrap() = server.call(()).unwrap();
|
||||
@@ -421,7 +424,10 @@ fn watch_dead_pid_is_noproc_down() {
|
||||
let h = spawn(|| {});
|
||||
let dead = h.pid();
|
||||
h.join().unwrap();
|
||||
let server = start(Pool { watcher: None, log: Vec::new() });
|
||||
let server = start(Pool {
|
||||
watcher: None,
|
||||
log: Vec::new(),
|
||||
});
|
||||
server.cast(PoolCast::Watch(dead)).unwrap();
|
||||
*got2.lock().unwrap() = server.call(()).unwrap();
|
||||
});
|
||||
@@ -497,7 +503,12 @@ impl GenServer for Timed {
|
||||
}
|
||||
|
||||
fn timed(fired: Arc<Mutex<Vec<u32>>>, cancel_won: Arc<Mutex<Option<bool>>>) -> Timed {
|
||||
Timed { timer: None, fired, cancel_won, last: None }
|
||||
Timed {
|
||||
timer: None,
|
||||
fired,
|
||||
cancel_won,
|
||||
last: None,
|
||||
}
|
||||
}
|
||||
|
||||
// A one-shot armed from a handler fires into handle_timer with its payload.
|
||||
@@ -534,7 +545,11 @@ fn cancel_before_fire_suppresses_it() {
|
||||
let count = server.call(()).unwrap();
|
||||
assert_eq!(count, 0, "cancelled timer must not fire");
|
||||
});
|
||||
assert_eq!(*cancel_won.lock().unwrap(), Some(true), "cancel beat the fire");
|
||||
assert_eq!(
|
||||
*cancel_won.lock().unwrap(),
|
||||
Some(true),
|
||||
"cancel beat the fire"
|
||||
);
|
||||
assert!(fired.lock().unwrap().is_empty());
|
||||
}
|
||||
|
||||
@@ -549,11 +564,16 @@ fn tick_every_rearms_repeatedly() {
|
||||
run(move || {
|
||||
let cw = Arc::new(Mutex::new(None));
|
||||
let server = start(timed(f2, cw));
|
||||
server.cast(TkCast::Tick(Duration::from_millis(20))).unwrap();
|
||||
server
|
||||
.cast(TkCast::Tick(Duration::from_millis(20)))
|
||||
.unwrap();
|
||||
let _ = server.call(()).unwrap(); // sync: periodic armed
|
||||
smarm::sleep(Duration::from_millis(130)); // ~6 periods
|
||||
let count = server.call(()).unwrap();
|
||||
assert!(count >= 3, "periodic should have re-armed several times, got {count}");
|
||||
assert!(
|
||||
count >= 3,
|
||||
"periodic should have re-armed several times, got {count}"
|
||||
);
|
||||
});
|
||||
// Every tick delivered the same payload.
|
||||
assert!(fired.lock().unwrap().iter().all(|&v| v == 9));
|
||||
@@ -568,7 +588,9 @@ fn cancel_stops_a_periodic() {
|
||||
let c2 = cancel_won.clone();
|
||||
run(move || {
|
||||
let server = start(timed(f2, c2));
|
||||
server.cast(TkCast::Tick(Duration::from_millis(20))).unwrap();
|
||||
server
|
||||
.cast(TkCast::Tick(Duration::from_millis(20)))
|
||||
.unwrap();
|
||||
let _ = server.call(()).unwrap();
|
||||
smarm::sleep(Duration::from_millis(70)); // a few ticks
|
||||
server.cast(TkCast::CancelLast).unwrap();
|
||||
@@ -616,11 +638,17 @@ fn idle_fires_repeatedly_on_quiet() {
|
||||
let idles = Arc::new(Mutex::new(0));
|
||||
let i2 = idles.clone();
|
||||
run(move || {
|
||||
let server = start(Idler { window: Duration::from_millis(25), idles: i2 });
|
||||
let server = start(Idler {
|
||||
window: Duration::from_millis(25),
|
||||
idles: i2,
|
||||
});
|
||||
smarm::sleep(Duration::from_millis(130)); // quiet ⇒ ~5 windows
|
||||
drop(server); // keep the server alive across the quiet span
|
||||
});
|
||||
assert!(*idles.lock().unwrap() >= 2, "idle should re-arm and fire several times");
|
||||
assert!(
|
||||
*idles.lock().unwrap() >= 2,
|
||||
"idle should re-arm and fire several times"
|
||||
);
|
||||
}
|
||||
|
||||
// Traffic within the window keeps idle from firing; only once the inbox goes
|
||||
@@ -632,7 +660,10 @@ fn traffic_resets_the_idle_window() {
|
||||
let before_quiet = Arc::new(Mutex::new(u32::MAX));
|
||||
let bq = before_quiet.clone();
|
||||
run(move || {
|
||||
let server = start(Idler { window: Duration::from_millis(60), idles: i2 });
|
||||
let server = start(Idler {
|
||||
window: Duration::from_millis(60),
|
||||
idles: i2,
|
||||
});
|
||||
// Poke every 25ms (< 60ms window) for ~100ms: each cast resets the
|
||||
// window before it can elapse.
|
||||
for _ in 0..4 {
|
||||
@@ -643,8 +674,15 @@ fn traffic_resets_the_idle_window() {
|
||||
smarm::sleep(Duration::from_millis(140)); // now genuinely quiet
|
||||
drop(server);
|
||||
});
|
||||
assert_eq!(*before_quiet.lock().unwrap(), 0, "steady traffic must suppress idle");
|
||||
assert!(*idles.lock().unwrap() >= 1, "idle fires once the inbox falls quiet");
|
||||
assert_eq!(
|
||||
*before_quiet.lock().unwrap(),
|
||||
0,
|
||||
"steady traffic must suppress idle"
|
||||
);
|
||||
assert!(
|
||||
*idles.lock().unwrap() >= 1,
|
||||
"idle fires once the inbox falls quiet"
|
||||
);
|
||||
}
|
||||
|
||||
// RFC 015 §4.7 — no armed timer survives loop exit. A server with a live
|
||||
@@ -658,15 +696,21 @@ fn no_timer_survives_exit() {
|
||||
let f_read = fired.clone();
|
||||
run(move || {
|
||||
let server = start(timed(f_server, Arc::new(Mutex::new(None))));
|
||||
server.cast(TkCast::Tick(Duration::from_millis(15))).unwrap();
|
||||
server
|
||||
.cast(TkCast::Tick(Duration::from_millis(15)))
|
||||
.unwrap();
|
||||
let _ = server.call(()).unwrap(); // sync: periodic armed
|
||||
smarm::sleep(Duration::from_millis(45)); // a couple of ticks
|
||||
let mon = smarm::monitor(server.pid());
|
||||
drop(server); // inbox closes → loop exits → guard drains timers
|
||||
// Clean Down ⇒ the loop returned without the no-leak assert aborting.
|
||||
server.shutdown(); // loop exits → guard drains timers
|
||||
// Clean Down ⇒ the loop returned without the no-leak assert aborting.
|
||||
assert!(mon.rx.recv().is_ok());
|
||||
let at_exit = f_read.lock().unwrap().len();
|
||||
smarm::sleep(Duration::from_millis(90)); // would be several more ticks
|
||||
assert_eq!(f_read.lock().unwrap().len(), at_exit, "no tick may fire after exit");
|
||||
assert_eq!(
|
||||
f_read.lock().unwrap().len(),
|
||||
at_exit,
|
||||
"no tick may fire after exit"
|
||||
);
|
||||
});
|
||||
}
|
||||
|
||||
@@ -0,0 +1,243 @@
|
||||
//! gen_server lifetime is the actor's, not its refs' (OTP: a pid is an
|
||||
//! address, a process lives until it stops, is shut down, or is killed).
|
||||
//!
|
||||
//! - Dropping the last `GenServerRef` does NOT end the server. It ends via
|
||||
//! `StopHandle::stop`, `request_shutdown` / `GenServerRef::shutdown`,
|
||||
//! `request_stop`, or a handler panic.
|
||||
//! - `GenServerBuilder::named(N).run()` runs the loop inline as the *current*
|
||||
//! actor, so a server is a direct `ChildSpec` child: the supervisor's
|
||||
//! shutdown reaches it as `handle_shutdown`, a restart re-binds the name,
|
||||
//! and by-name `call`/`cast` reach whichever incarnation is live.
|
||||
|
||||
use smarm::gen_server::{
|
||||
self, GenServer, GenServerBuilder, GenServerCtx, GenServerName, ShutdownAction, StopHandle,
|
||||
};
|
||||
use smarm::registry::RegisterError;
|
||||
use smarm::supervisor::{ChildSpec, OneForOne, Restart, Shutdown};
|
||||
use smarm::{monitor, request_shutdown, run, sleep, spawn, DownReason};
|
||||
use std::sync::atomic::{AtomicUsize, Ordering};
|
||||
use std::sync::{Arc, Mutex};
|
||||
use std::time::Duration;
|
||||
|
||||
#[derive(Default, Clone)]
|
||||
struct Log(Arc<Mutex<Vec<String>>>);
|
||||
impl Log {
|
||||
fn push(&self, s: impl Into<String>) {
|
||||
self.0.lock().unwrap().push(s.into());
|
||||
}
|
||||
fn get(&self) -> Vec<String> {
|
||||
self.0.lock().unwrap().clone()
|
||||
}
|
||||
}
|
||||
|
||||
struct Counter {
|
||||
log: Log,
|
||||
n: u64,
|
||||
trap: bool,
|
||||
stop: Option<StopHandle<Counter>>,
|
||||
}
|
||||
|
||||
enum Call {
|
||||
Get,
|
||||
}
|
||||
enum Cast {
|
||||
Inc,
|
||||
Stop,
|
||||
}
|
||||
|
||||
impl GenServer for Counter {
|
||||
type Call = Call;
|
||||
type Reply = u64;
|
||||
type Cast = Cast;
|
||||
type Info = ();
|
||||
type Timer = ();
|
||||
|
||||
fn init(&mut self, ctx: &GenServerCtx<Self>) {
|
||||
if self.trap {
|
||||
ctx.trap_exit();
|
||||
}
|
||||
self.stop = Some(ctx.stop_handle());
|
||||
self.log.push("init");
|
||||
}
|
||||
fn handle_call(&mut self, Call::Get: Call) -> u64 {
|
||||
self.n
|
||||
}
|
||||
fn handle_cast(&mut self, c: Cast) {
|
||||
match c {
|
||||
Cast::Inc => self.n += 1,
|
||||
Cast::Stop => self.stop.as_ref().unwrap().stop(),
|
||||
}
|
||||
}
|
||||
fn handle_shutdown(&mut self) -> ShutdownAction {
|
||||
self.log.push("handle_shutdown");
|
||||
ShutdownAction::Exit
|
||||
}
|
||||
fn terminate(&mut self) {
|
||||
self.log.push("terminate");
|
||||
}
|
||||
}
|
||||
|
||||
fn counter(log: &Log, trap: bool) -> Counter {
|
||||
Counter {
|
||||
log: log.clone(),
|
||||
n: 0,
|
||||
trap,
|
||||
stop: None,
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Refs are addresses: dropping the last one does not end the server.
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
#[test]
|
||||
fn dropping_last_ref_does_not_end_server() {
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let srv = gen_server::start(counter(&l, true));
|
||||
let pid = srv.pid();
|
||||
srv.cast(Cast::Inc).unwrap();
|
||||
assert_eq!(srv.call(Call::Get).unwrap(), 1);
|
||||
let mon = monitor(pid);
|
||||
drop(srv);
|
||||
sleep(Duration::from_millis(30));
|
||||
assert!(
|
||||
mon.rx.try_recv().unwrap().is_none(),
|
||||
"server must outlive its last ref"
|
||||
);
|
||||
assert_eq!(l.get(), vec!["init"], "terminate must not have run");
|
||||
// Explicit teardown still works, and is what ends it.
|
||||
request_shutdown(pid);
|
||||
let down = mon.rx.recv().unwrap();
|
||||
assert_eq!(down.reason, DownReason::Exit);
|
||||
});
|
||||
assert_eq!(log.get(), vec!["init", "handle_shutdown", "terminate"]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn ref_shutdown_is_the_explicit_close() {
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let srv = gen_server::start(counter(&l, true));
|
||||
srv.call(Call::Get).unwrap(); // sync: init (and trap_exit) has run
|
||||
srv.shutdown(); // graceful, waits
|
||||
assert_eq!(l.get(), vec!["init", "handle_shutdown", "terminate"]);
|
||||
});
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn forgotten_server_is_shut_down_at_root_exit() {
|
||||
// A ref-less server is not a hung run: root exit shuts it down.
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let srv = gen_server::start(counter(&l, false));
|
||||
drop(srv);
|
||||
});
|
||||
assert_eq!(log.get(), vec!["init", "terminate"]);
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Inline run: a gen_server as a direct ChildSpec child.
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
const COUNTER: GenServerName<Counter> = GenServerName::new("lifetime-counter");
|
||||
|
||||
#[test]
|
||||
fn named_run_is_a_direct_supervised_child_and_gets_shutdown() {
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let l2 = l.clone();
|
||||
let sup = spawn(move || {
|
||||
let l3 = l2.clone();
|
||||
OneForOne::new()
|
||||
.child(
|
||||
ChildSpec::new(Restart::Permanent, move || {
|
||||
GenServerBuilder::new(counter(&l3, true))
|
||||
.named(COUNTER)
|
||||
.run()
|
||||
.expect("name free");
|
||||
})
|
||||
.shutdown(Shutdown::Infinity),
|
||||
)
|
||||
.run();
|
||||
});
|
||||
sleep(Duration::from_millis(10));
|
||||
gen_server::cast(COUNTER, Cast::Inc).unwrap();
|
||||
assert_eq!(gen_server::call(COUNTER, Call::Get).unwrap(), 1);
|
||||
request_shutdown(sup.pid());
|
||||
sup.join()
|
||||
.expect("ordered shutdown, supervisor returns normally");
|
||||
assert_eq!(l.get(), vec!["init", "handle_shutdown", "terminate"]);
|
||||
assert!(gen_server::whereis_server(COUNTER).is_none());
|
||||
});
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn named_run_child_restarts_and_rebinds_name() {
|
||||
let log = Log::default();
|
||||
let inits = Arc::new(AtomicUsize::new(0));
|
||||
let l = log.clone();
|
||||
let i = inits.clone();
|
||||
run(move || {
|
||||
let l2 = l.clone();
|
||||
let i2 = i.clone();
|
||||
let sup = spawn(move || {
|
||||
let l3 = l2.clone();
|
||||
let i3 = i2.clone();
|
||||
OneForOne::new()
|
||||
.child(ChildSpec::new(Restart::Permanent, move || {
|
||||
i3.fetch_add(1, Ordering::SeqCst);
|
||||
GenServerBuilder::new(counter(&l3, false))
|
||||
.named(COUNTER)
|
||||
.run()
|
||||
.expect("name free on (re)start");
|
||||
}))
|
||||
.run();
|
||||
});
|
||||
sleep(Duration::from_millis(10));
|
||||
gen_server::cast(COUNTER, Cast::Inc).unwrap();
|
||||
assert_eq!(gen_server::call(COUNTER, Call::Get).unwrap(), 1);
|
||||
// Normal self-exit → Permanent restarts it, fresh state, same name.
|
||||
gen_server::cast(COUNTER, Cast::Stop).unwrap();
|
||||
sleep(Duration::from_millis(30));
|
||||
assert_eq!(i.load(Ordering::SeqCst), 2, "restarted once");
|
||||
assert_eq!(gen_server::call(COUNTER, Call::Get).unwrap(), 0);
|
||||
request_shutdown(sup.pid());
|
||||
sup.join().unwrap();
|
||||
});
|
||||
assert_eq!(log.get(), vec!["init", "terminate", "init", "terminate"]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn named_run_name_clash_fails_before_init() {
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let first = GenServerBuilder::new(counter(&l, false))
|
||||
.named(COUNTER)
|
||||
.start()
|
||||
.unwrap();
|
||||
let l2 = l.clone();
|
||||
let res = Arc::new(Mutex::new(None));
|
||||
let r2 = res.clone();
|
||||
let first_pid = first.pid();
|
||||
spawn(move || {
|
||||
let r = GenServerBuilder::new(counter(&l2, false))
|
||||
.named(COUNTER)
|
||||
.run();
|
||||
*r2.lock().unwrap() = Some(r);
|
||||
})
|
||||
.join()
|
||||
.unwrap();
|
||||
assert_eq!(
|
||||
*res.lock().unwrap(),
|
||||
Some(Err(RegisterError::NameTaken { holder: first_pid }))
|
||||
);
|
||||
assert_eq!(l.get(), vec!["init"], "clashing server never ran init");
|
||||
first.shutdown();
|
||||
});
|
||||
}
|
||||
@@ -0,0 +1,222 @@
|
||||
//! gen_server graceful shutdown.
|
||||
//!
|
||||
//! - A server that does not opt in (`ctx.trap_exit()` in `init`) is stopped
|
||||
//! outright by `request_shutdown`, exactly as by `request_stop`.
|
||||
//! - A trapping server receives the request as `handle_shutdown`. The default
|
||||
//! returns `ShutdownAction::Exit`: the loop breaks and `terminate` runs on
|
||||
//! the normal (non-unwind) path, so it may block. `Continue` keeps the loop
|
||||
//! dispatching; the state later ends itself with a `StopHandle` — the only
|
||||
//! way for a gen_server to exit *normally* on its own (`request_stop` on
|
||||
//! self is an abnormal `Stopped`, which `Transient` restarts).
|
||||
//! - Other exit signals (linked peers dying) reach a trapping server via
|
||||
//! `handle_exit`.
|
||||
|
||||
use smarm::gen_server::{
|
||||
start, GenServer, GenServerBuilder, GenServerCtx, GenServerRef, ShutdownAction, StopHandle,
|
||||
};
|
||||
use smarm::supervisor::{ChildSpec, OneForOne, Restart};
|
||||
use smarm::{link, monitor, request_shutdown, run, self_pid, sleep, spawn, DownReason, ExitSignal};
|
||||
use std::sync::atomic::{AtomicBool, AtomicUsize, Ordering};
|
||||
use std::sync::{Arc, Mutex};
|
||||
use std::time::Duration;
|
||||
|
||||
#[derive(Default, Clone)]
|
||||
struct Log {
|
||||
events: Arc<Mutex<Vec<&'static str>>>,
|
||||
}
|
||||
impl Log {
|
||||
fn push(&self, e: &'static str) {
|
||||
self.events.lock().unwrap().push(e);
|
||||
}
|
||||
fn get(&self) -> Vec<&'static str> {
|
||||
self.events.lock().unwrap().clone()
|
||||
}
|
||||
}
|
||||
|
||||
/// A server with configurable shutdown behaviour.
|
||||
struct Srv {
|
||||
log: Log,
|
||||
trap: bool,
|
||||
action: ShutdownAction,
|
||||
stop: Option<StopHandle<Srv>>,
|
||||
exits: Arc<Mutex<Vec<ExitSignal>>>,
|
||||
}
|
||||
|
||||
impl Srv {
|
||||
fn new(log: &Log, trap: bool, action: ShutdownAction) -> Self {
|
||||
Srv {
|
||||
log: log.clone(),
|
||||
trap,
|
||||
action,
|
||||
stop: None,
|
||||
exits: Default::default(),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
enum Cast {
|
||||
Note(&'static str),
|
||||
StopNow,
|
||||
}
|
||||
|
||||
impl GenServer for Srv {
|
||||
type Call = ();
|
||||
type Reply = ();
|
||||
type Cast = Cast;
|
||||
type Info = ();
|
||||
type Timer = ();
|
||||
|
||||
fn init(&mut self, ctx: &GenServerCtx<Self>) {
|
||||
if self.trap {
|
||||
ctx.trap_exit();
|
||||
}
|
||||
self.stop = Some(ctx.stop_handle());
|
||||
}
|
||||
fn handle_call(&mut self, _: ()) {}
|
||||
fn handle_cast(&mut self, c: Cast) {
|
||||
match c {
|
||||
Cast::Note(s) => self.log.push(s),
|
||||
Cast::StopNow => self.stop.as_ref().unwrap().stop(),
|
||||
}
|
||||
}
|
||||
fn handle_shutdown(&mut self) -> ShutdownAction {
|
||||
self.log.push("handle_shutdown");
|
||||
self.action
|
||||
}
|
||||
fn handle_exit(&mut self, sig: ExitSignal) {
|
||||
self.log.push("handle_exit");
|
||||
self.exits.lock().unwrap().push(sig);
|
||||
}
|
||||
fn terminate(&mut self) {
|
||||
// Allowed to block on the graceful path.
|
||||
if self.trap {
|
||||
sleep(Duration::from_millis(10));
|
||||
}
|
||||
self.log.push("terminate");
|
||||
}
|
||||
}
|
||||
|
||||
fn spawn_settled<G: GenServer>(state: G) -> GenServerRef<G> {
|
||||
let r = start(state);
|
||||
sleep(Duration::from_millis(20)); // let init (trap_exit) run
|
||||
r
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn non_trapping_server_is_stopped_outright() {
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let r = spawn_settled(Srv::new(&l, false, ShutdownAction::Exit));
|
||||
let mon = monitor(r.pid());
|
||||
request_shutdown(r.pid());
|
||||
let d = mon.rx.recv().unwrap();
|
||||
assert_eq!(d.reason, DownReason::Stopped);
|
||||
});
|
||||
assert_eq!(log.get(), vec!["terminate"]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn trapping_server_exits_normally_via_handle_shutdown() {
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let r = spawn_settled(Srv::new(&l, true, ShutdownAction::Exit));
|
||||
let mon = monitor(r.pid());
|
||||
request_shutdown(r.pid());
|
||||
let d = mon.rx.recv().unwrap();
|
||||
assert_eq!(d.reason, DownReason::Exit);
|
||||
});
|
||||
assert_eq!(log.get(), vec!["handle_shutdown", "terminate"]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn continue_keeps_dispatching_until_stop_handle() {
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let r = spawn_settled(Srv::new(&l, true, ShutdownAction::Continue));
|
||||
let mon = monitor(r.pid());
|
||||
request_shutdown(r.pid());
|
||||
sleep(Duration::from_millis(20));
|
||||
r.cast(Cast::Note("after-shutdown-request")).unwrap();
|
||||
r.cast(Cast::StopNow).unwrap();
|
||||
let d = mon.rx.recv().unwrap();
|
||||
assert_eq!(d.reason, DownReason::Exit);
|
||||
});
|
||||
assert_eq!(
|
||||
log.get(),
|
||||
vec!["handle_shutdown", "after-shutdown-request", "terminate"]
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn stop_handle_is_a_normal_exit_that_transient_does_not_restart() {
|
||||
let starts = Arc::new(AtomicUsize::new(0));
|
||||
let s = starts.clone();
|
||||
run(move || {
|
||||
let s2 = s.clone();
|
||||
let sup = spawn(move || {
|
||||
let s3 = s2.clone();
|
||||
OneForOne::new()
|
||||
.child(ChildSpec::new(Restart::Transient, move || {
|
||||
s3.fetch_add(1, Ordering::SeqCst);
|
||||
let log = Log::default();
|
||||
let r = GenServerBuilder::new(Srv::new(&log, false, ShutdownAction::Exit))
|
||||
.under(self_pid())
|
||||
.start();
|
||||
r.cast(Cast::StopNow).unwrap();
|
||||
// Block until the server is gone; a bare spawn parent
|
||||
// returning would not itself end the server.
|
||||
let mon = monitor(r.pid());
|
||||
let _ = mon.rx.recv();
|
||||
}))
|
||||
.run();
|
||||
});
|
||||
sup.join().unwrap(); // returns only if the child was not restarted forever
|
||||
});
|
||||
assert_eq!(starts.load(Ordering::SeqCst), 1);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn linked_peer_death_reaches_handle_exit() {
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
let alive = Arc::new(AtomicBool::new(false));
|
||||
let a = alive.clone();
|
||||
run(move || {
|
||||
let r = spawn_settled(Srv::new(&l, true, ShutdownAction::Exit));
|
||||
let pid = r.pid();
|
||||
let peer = spawn(move || {
|
||||
link(pid);
|
||||
panic!("peer dies");
|
||||
});
|
||||
let _ = peer.join();
|
||||
sleep(Duration::from_millis(20));
|
||||
r.cast(Cast::Note("still-serving")).unwrap();
|
||||
sleep(Duration::from_millis(20));
|
||||
a.store(true, Ordering::SeqCst);
|
||||
r.shutdown(); // graceful; waits for terminate
|
||||
});
|
||||
assert!(alive.load(Ordering::SeqCst));
|
||||
assert_eq!(
|
||||
log.get(),
|
||||
vec![
|
||||
"handle_exit",
|
||||
"still-serving",
|
||||
"handle_shutdown",
|
||||
"terminate"
|
||||
]
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn gen_server_ref_shutdown_is_graceful_for_a_trapping_server() {
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let r = spawn_settled(Srv::new(&l, true, ShutdownAction::Exit));
|
||||
r.shutdown();
|
||||
});
|
||||
assert_eq!(log.get(), vec!["handle_shutdown", "terminate"]);
|
||||
}
|
||||
+91
-15
@@ -100,7 +100,15 @@ fn state_timeout_fires() {
|
||||
let got = Arc::new(Mutex::new(0u32));
|
||||
let got2 = got.clone();
|
||||
run(move || {
|
||||
let m = TimerSm::start(T::Idle, TData { enters: 0, st_fires: 0, named_fires: 0, st_window: 5 });
|
||||
let m = TimerSm::start(
|
||||
T::Idle,
|
||||
TData {
|
||||
enters: 0,
|
||||
st_fires: 0,
|
||||
named_fires: 0,
|
||||
st_window: 5,
|
||||
},
|
||||
);
|
||||
m.send(Ev2::Cast(TCast::Arm)).unwrap(); // -> Armed, arms 5ms state-timeout
|
||||
smarm::sleep(Duration::from_millis(40)); // let it fire
|
||||
*got2.lock().unwrap() = m.call(|r| Ev2::Call(TCall::StFires(r))).unwrap();
|
||||
@@ -116,13 +124,25 @@ fn state_timeout_auto_resets_on_transition() {
|
||||
let got2 = got.clone();
|
||||
run(move || {
|
||||
// Long window so the explicit Disarm beats it comfortably.
|
||||
let m = TimerSm::start(T::Idle, TData { enters: 0, st_fires: 0, named_fires: 0, st_window: 50 });
|
||||
let m = TimerSm::start(
|
||||
T::Idle,
|
||||
TData {
|
||||
enters: 0,
|
||||
st_fires: 0,
|
||||
named_fires: 0,
|
||||
st_window: 50,
|
||||
},
|
||||
);
|
||||
m.send(Ev2::Cast(TCast::Arm)).unwrap(); // -> Armed, arms 50ms state-timeout
|
||||
m.send(Ev2::Cast(TCast::Disarm)).unwrap(); // -> Idle, auto-resets it
|
||||
smarm::sleep(Duration::from_millis(80)); // past the original window
|
||||
*got2.lock().unwrap() = m.call(|r| Ev2::Call(TCall::StFires(r))).unwrap();
|
||||
});
|
||||
assert_eq!(*got.lock().unwrap(), 0, "auto-reset cancelled the pending state-timeout");
|
||||
assert_eq!(
|
||||
*got.lock().unwrap(),
|
||||
0,
|
||||
"auto-reset cancelled the pending state-timeout"
|
||||
);
|
||||
}
|
||||
|
||||
// A named timeout survives a state change: armed in Idle, it still fires after
|
||||
@@ -133,14 +153,26 @@ fn named_timeout_survives_transition() {
|
||||
let got2 = got.clone();
|
||||
run(move || {
|
||||
// Armed's own state-timeout is long so it doesn't interfere.
|
||||
let m = TimerSm::start(T::Idle, TData { enters: 0, st_fires: 0, named_fires: 0, st_window: 200 });
|
||||
let m = TimerSm::start(
|
||||
T::Idle,
|
||||
TData {
|
||||
enters: 0,
|
||||
st_fires: 0,
|
||||
named_fires: 0,
|
||||
st_window: 200,
|
||||
},
|
||||
);
|
||||
m.send(Ev2::Cast(TCast::Ping(20))).unwrap(); // arm "ping" for 20ms (in Idle)
|
||||
m.send(Ev2::Cast(TCast::Arm)).unwrap(); // -> Armed (ping must survive this)
|
||||
m.send(Ev2::Cast(TCast::Disarm)).unwrap(); // -> Idle (and this)
|
||||
smarm::sleep(Duration::from_millis(60)); // let "ping" fire
|
||||
*got2.lock().unwrap() = m.call(|r| Ev2::Call(TCall::NamedFires(r))).unwrap();
|
||||
});
|
||||
assert_eq!(*got.lock().unwrap(), 1, "named timeout fired across the transitions");
|
||||
assert_eq!(
|
||||
*got.lock().unwrap(),
|
||||
1,
|
||||
"named timeout fired across the transitions"
|
||||
);
|
||||
}
|
||||
|
||||
// Cancelling a named timeout before its window prevents the fire.
|
||||
@@ -149,13 +181,25 @@ fn named_timeout_cancel() {
|
||||
let got = Arc::new(Mutex::new(99u32));
|
||||
let got2 = got.clone();
|
||||
run(move || {
|
||||
let m = TimerSm::start(T::Idle, TData { enters: 0, st_fires: 0, named_fires: 0, st_window: 200 });
|
||||
let m = TimerSm::start(
|
||||
T::Idle,
|
||||
TData {
|
||||
enters: 0,
|
||||
st_fires: 0,
|
||||
named_fires: 0,
|
||||
st_window: 200,
|
||||
},
|
||||
);
|
||||
m.send(Ev2::Cast(TCast::Ping(30))).unwrap(); // arm "ping" for 30ms
|
||||
m.send(Ev2::Cast(TCast::CancelPing)).unwrap(); // cancel before it fires
|
||||
smarm::sleep(Duration::from_millis(60)); // past the original window
|
||||
*got2.lock().unwrap() = m.call(|r| Ev2::Call(TCall::NamedFires(r))).unwrap();
|
||||
});
|
||||
assert_eq!(*got.lock().unwrap(), 0, "cancel prevented the named-timeout fire");
|
||||
assert_eq!(
|
||||
*got.lock().unwrap(),
|
||||
0,
|
||||
"cancel prevented the named-timeout fire"
|
||||
);
|
||||
}
|
||||
|
||||
// ===========================================================================
|
||||
@@ -170,15 +214,27 @@ fn cast_then_call_roundtrip() {
|
||||
let got2 = got.clone();
|
||||
run(move || {
|
||||
// Long state-timeout window so it never fires during the test.
|
||||
let m = TimerSm::start(T::Idle, TData { enters: 0, st_fires: 0, named_fires: 0, st_window: 10_000 });
|
||||
m.send(Ev2::Cast(TCast::Arm)).unwrap(); // Idle -> Armed (enter)
|
||||
let m = TimerSm::start(
|
||||
T::Idle,
|
||||
TData {
|
||||
enters: 0,
|
||||
st_fires: 0,
|
||||
named_fires: 0,
|
||||
st_window: 10_000,
|
||||
},
|
||||
);
|
||||
m.send(Ev2::Cast(TCast::Arm)).unwrap(); // Idle -> Armed (enter)
|
||||
m.send(Ev2::Cast(TCast::Disarm)).unwrap(); // Armed -> Idle (enter)
|
||||
m.send(Ev2::Cast(TCast::Arm)).unwrap(); // Idle -> Armed (enter)
|
||||
m.send(Ev2::Cast(TCast::Arm)).unwrap(); // Idle -> Armed (enter)
|
||||
m.send(Ev2::Cast(TCast::Disarm)).unwrap(); // Armed -> Idle (enter)
|
||||
// enters = 1 (start) + 4 transitions = 5.
|
||||
// enters = 1 (start) + 4 transitions = 5.
|
||||
*got2.lock().unwrap() = m.call(|r| Ev2::Call(TCall::Enters(r))).unwrap();
|
||||
});
|
||||
assert_eq!(*got.lock().unwrap(), 5, "one enter on start, one per real transition");
|
||||
assert_eq!(
|
||||
*got.lock().unwrap(),
|
||||
5,
|
||||
"one enter on start, one per real transition"
|
||||
);
|
||||
}
|
||||
|
||||
// `enter` fires once on start and once per *real* transition; a stay (a call
|
||||
@@ -188,7 +244,15 @@ fn enter_on_start_and_each_transition_but_not_stay() {
|
||||
let got = Arc::new(Mutex::new((0u32, 0u32, 0u32)));
|
||||
let got2 = got.clone();
|
||||
run(move || {
|
||||
let m = TimerSm::start(T::Idle, TData { enters: 0, st_fires: 0, named_fires: 0, st_window: 10_000 }); // enter -> 1
|
||||
let m = TimerSm::start(
|
||||
T::Idle,
|
||||
TData {
|
||||
enters: 0,
|
||||
st_fires: 0,
|
||||
named_fires: 0,
|
||||
st_window: 10_000,
|
||||
},
|
||||
); // enter -> 1
|
||||
let after_start = m.call(|r| Ev2::Call(TCall::Enters(r))).unwrap();
|
||||
// A stay (a counter read returns `prev`) must not bump enters.
|
||||
let _ = m.call(|r| Ev2::Call(TCall::StFires(r))).unwrap();
|
||||
@@ -208,7 +272,15 @@ fn call_to_panicking_handler_is_down() {
|
||||
let got = Arc::new(Mutex::new(None::<Result<u32, CallError>>));
|
||||
let got2 = got.clone();
|
||||
run(move || {
|
||||
let m = TimerSm::start(T::Idle, TData { enters: 0, st_fires: 0, named_fires: 0, st_window: 10_000 });
|
||||
let m = TimerSm::start(
|
||||
T::Idle,
|
||||
TData {
|
||||
enters: 0,
|
||||
st_fires: 0,
|
||||
named_fires: 0,
|
||||
st_window: 10_000,
|
||||
},
|
||||
);
|
||||
let r = m.call(|rep| Ev2::Call(TCall::Boom(rep)));
|
||||
*got2.lock().unwrap() = Some(r);
|
||||
});
|
||||
@@ -299,7 +371,11 @@ fn postponed_call_answered_after_transition() {
|
||||
smarm::sleep(Duration::from_millis(20)); // let the child wake with its reply
|
||||
*g2.lock().unwrap() = *taken.lock().unwrap();
|
||||
});
|
||||
assert_eq!(*got.lock().unwrap(), Some(42), "postponed call answered by the Filled state");
|
||||
assert_eq!(
|
||||
*got.lock().unwrap(),
|
||||
Some(42),
|
||||
"postponed call answered by the Filled state"
|
||||
);
|
||||
}
|
||||
|
||||
#[derive(Clone, Copy, PartialEq, Eq, Debug)]
|
||||
|
||||
@@ -0,0 +1,128 @@
|
||||
//! gen_statem lifetime parity with gen_server: a machine lives until it
|
||||
//! stops, is shut down, or is killed — its refs are addresses. And
|
||||
//! `gen_statem::run_named` runs a machine inline as the current actor, so it
|
||||
//! is a direct `ChildSpec` child addressed by name.
|
||||
|
||||
use smarm::gen_statem::{self, GenStatemName, Reply};
|
||||
use smarm::supervisor::{ChildSpec, OneForOne, Restart, Shutdown};
|
||||
use smarm::{monitor, request_shutdown, run, sleep, spawn, DownReason};
|
||||
use std::sync::{Arc, Mutex};
|
||||
use std::time::Duration;
|
||||
|
||||
#[derive(Default, Clone)]
|
||||
struct Log(Arc<Mutex<Vec<&'static str>>>);
|
||||
impl Log {
|
||||
fn push(&self, e: &'static str) {
|
||||
self.0.lock().unwrap().push(e);
|
||||
}
|
||||
fn get(&self) -> Vec<&'static str> {
|
||||
self.0.lock().unwrap().clone()
|
||||
}
|
||||
}
|
||||
|
||||
#[derive(Clone, Copy, PartialEq, Eq, Debug)]
|
||||
enum S {
|
||||
On,
|
||||
}
|
||||
|
||||
struct D {
|
||||
log: Log,
|
||||
trap: bool,
|
||||
n: u64,
|
||||
}
|
||||
|
||||
enum Cast {
|
||||
Inc,
|
||||
StopNow,
|
||||
}
|
||||
enum Call {
|
||||
Get(Reply<u64>),
|
||||
}
|
||||
|
||||
smarm::gen_statem! {
|
||||
machine: Sm { state: S, data: D };
|
||||
event: Ev { cast: Cast, call: Call, info: () };
|
||||
context(data, prev, cx);
|
||||
|
||||
enter {
|
||||
S::On => { data.log.push("enter"); if data.trap { cx.trap_exit() } },
|
||||
}
|
||||
|
||||
on S::On => {
|
||||
cast Cast::Inc => { data.n += 1; prev },
|
||||
cast Cast::StopNow => stop,
|
||||
call Call::Get(r) => { r.reply(data.n); prev },
|
||||
shutdown => { data.log.push("shutdown"); cx.stop(); prev },
|
||||
state_timeout => unhandled,
|
||||
timeout _ => unhandled,
|
||||
}
|
||||
|
||||
terminate { data.log.push("terminate"); }
|
||||
}
|
||||
|
||||
fn d(log: &Log, trap: bool) -> D {
|
||||
D {
|
||||
log: log.clone(),
|
||||
trap,
|
||||
n: 0,
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn dropping_last_ref_does_not_end_machine() {
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let m = Sm::start(S::On, d(&l, true));
|
||||
let pid = m.pid();
|
||||
m.send(Ev::Cast(Cast::Inc)).unwrap();
|
||||
assert_eq!(m.call(|r| Ev::Call(Call::Get(r))).unwrap(), 1);
|
||||
let mon = monitor(pid);
|
||||
drop(m);
|
||||
sleep(Duration::from_millis(30));
|
||||
assert!(
|
||||
mon.rx.try_recv().unwrap().is_none(),
|
||||
"machine must outlive its refs"
|
||||
);
|
||||
assert_eq!(l.get(), vec!["enter"]);
|
||||
request_shutdown(pid);
|
||||
assert_eq!(mon.rx.recv().unwrap().reason, DownReason::Exit);
|
||||
});
|
||||
assert_eq!(log.get(), vec!["enter", "shutdown", "terminate"]);
|
||||
}
|
||||
|
||||
const SM: GenStatemName<Sm> = GenStatemName::new("lifetime-sm");
|
||||
|
||||
#[test]
|
||||
fn run_named_is_a_direct_supervised_child() {
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let l2 = l.clone();
|
||||
let sup = spawn(move || {
|
||||
let l3 = l2.clone();
|
||||
OneForOne::new()
|
||||
.child(
|
||||
ChildSpec::new(Restart::Permanent, move || {
|
||||
gen_statem::run_named(SM, Sm::new(S::On, d(&l3, true))).expect("name free");
|
||||
})
|
||||
.shutdown(Shutdown::Infinity),
|
||||
)
|
||||
.run();
|
||||
});
|
||||
sleep(Duration::from_millis(10));
|
||||
gen_statem::send(SM, Ev::Cast(Cast::Inc)).unwrap();
|
||||
assert_eq!(gen_statem::call(SM, |r| Ev::Call(Call::Get(r))).unwrap(), 1);
|
||||
// Normal self-exit → Permanent restart → fresh data, same name.
|
||||
gen_statem::send(SM, Ev::Cast(Cast::StopNow)).unwrap();
|
||||
sleep(Duration::from_millis(30));
|
||||
assert_eq!(gen_statem::call(SM, |r| Ev::Call(Call::Get(r))).unwrap(), 0);
|
||||
request_shutdown(sup.pid());
|
||||
sup.join().unwrap();
|
||||
assert!(gen_statem::whereis_machine(SM).is_none());
|
||||
});
|
||||
assert_eq!(
|
||||
log.get(),
|
||||
vec!["enter", "terminate", "enter", "shutdown", "terminate"]
|
||||
);
|
||||
}
|
||||
@@ -0,0 +1,219 @@
|
||||
//! gen_statem graceful shutdown — the gen_server surface, in state-machine
|
||||
//! clothes. Where gen_server routes a shutdown request to a `handle_shutdown`
|
||||
//! method, a gen_statem gets it as an **event** so it can be routed by state:
|
||||
//!
|
||||
//! - A machine that does not opt in (`cx.trap_exit()` in the initial `enter`)
|
||||
//! is stopped outright by `request_shutdown`, exactly as by `request_stop`.
|
||||
//! - A trapping machine sees the request as a `shutdown` row (a unit event
|
||||
//! like `state_timeout`). The macro's default, when a state writes no
|
||||
//! `shutdown` row, is `stop` — the loop breaks and `terminate` runs on the
|
||||
//! normal path. A row may instead transition (e.g. into a Draining state)
|
||||
//! and `stop` later from any row via the `stop` tail keyword.
|
||||
//! - Linked-peer deaths reach a trapping machine as `exit <pat>` rows; an
|
||||
//! unmatched exit is silently dropped, like an unmatched info.
|
||||
//! - `terminate { … }` is an optional macro block, run on every exit path.
|
||||
|
||||
use smarm::gen_statem;
|
||||
use smarm::gen_statem::{GenStatemRef, Reply};
|
||||
use smarm::{link, monitor, request_shutdown, run, sleep, spawn, DownReason, ExitSignal};
|
||||
use std::sync::{Arc, Mutex};
|
||||
use std::time::Duration;
|
||||
|
||||
#[derive(Default, Clone)]
|
||||
struct Log(Arc<Mutex<Vec<&'static str>>>);
|
||||
impl Log {
|
||||
fn push(&self, e: &'static str) {
|
||||
self.0.lock().unwrap().push(e);
|
||||
}
|
||||
fn get(&self) -> Vec<&'static str> {
|
||||
self.0.lock().unwrap().clone()
|
||||
}
|
||||
}
|
||||
|
||||
#[derive(Clone, Copy, PartialEq, Eq, Debug)]
|
||||
enum S {
|
||||
Idle,
|
||||
Draining,
|
||||
}
|
||||
|
||||
struct D {
|
||||
log: Log,
|
||||
trap: bool,
|
||||
exits: Vec<ExitSignal>,
|
||||
}
|
||||
|
||||
enum Cast {
|
||||
Note(&'static str),
|
||||
StopNow,
|
||||
}
|
||||
enum Call {
|
||||
Exits(Reply<usize>),
|
||||
}
|
||||
|
||||
gen_statem! {
|
||||
machine: Sm { state: S, data: D };
|
||||
event: Ev { cast: Cast, call: Call, info: () };
|
||||
context(data, prev, cx);
|
||||
|
||||
enter {
|
||||
S::Idle => if data.trap { cx.trap_exit() },
|
||||
S::Draining => { data.log.push("draining"); cx.state_timeout(Duration::from_millis(30)); },
|
||||
}
|
||||
|
||||
on S::Idle => {
|
||||
// Shutdown in Idle: go drain first, stop later.
|
||||
shutdown => S::Draining,
|
||||
cast Cast::StopNow => stop,
|
||||
state_timeout => unhandled,
|
||||
}
|
||||
|
||||
on S::Draining => {
|
||||
// Drained: end the machine normally.
|
||||
state_timeout => { data.log.push("drained"); cx.stop(); prev },
|
||||
// A second request while draining is ignored.
|
||||
shutdown => unhandled,
|
||||
cast Cast::StopNow => stop,
|
||||
}
|
||||
|
||||
on _ => {
|
||||
cast Cast::Note(s) => { data.log.push(s); prev },
|
||||
call Call::Exits(r) => { r.reply(data.exits.len()); prev },
|
||||
exit sig => { data.log.push("exit"); data.exits.push(sig); prev },
|
||||
timeout _ => unhandled,
|
||||
}
|
||||
|
||||
terminate {
|
||||
data.log.push("terminate");
|
||||
}
|
||||
}
|
||||
|
||||
/// A machine with no `shutdown` rows at all: the macro default applies.
|
||||
#[derive(Clone, Copy, PartialEq, Eq, Debug)]
|
||||
enum P {
|
||||
On,
|
||||
}
|
||||
struct PD {
|
||||
log: Log,
|
||||
}
|
||||
enum PCast {}
|
||||
enum PCall {}
|
||||
|
||||
gen_statem! {
|
||||
machine: Plain { state: P, data: PD };
|
||||
event: PEv { cast: PCast, call: PCall, info: () };
|
||||
context(data, prev, cx);
|
||||
enter { P::On => cx.trap_exit(), }
|
||||
on P::On => {
|
||||
cast _ => unhandled,
|
||||
call _ => unhandled,
|
||||
state_timeout => unhandled,
|
||||
timeout _ => unhandled,
|
||||
}
|
||||
terminate { data.log.push("terminate"); }
|
||||
}
|
||||
|
||||
fn settled(log: &Log, trap: bool) -> GenStatemRef<Sm> {
|
||||
let r = Sm::start(
|
||||
S::Idle,
|
||||
D {
|
||||
log: log.clone(),
|
||||
trap,
|
||||
exits: Vec::new(),
|
||||
},
|
||||
);
|
||||
sleep(Duration::from_millis(20)); // let on_start (trap_exit) run
|
||||
r
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn non_trapping_machine_is_stopped_outright() {
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let r = settled(&l, false);
|
||||
let mon = monitor(r.pid());
|
||||
request_shutdown(r.pid());
|
||||
assert_eq!(mon.rx.recv().unwrap().reason, DownReason::Stopped);
|
||||
});
|
||||
assert_eq!(log.get(), vec!["terminate"]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn shutdown_row_routes_by_state_and_stop_tail_exits_normally() {
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let r = settled(&l, true);
|
||||
let mon = monitor(r.pid());
|
||||
request_shutdown(r.pid());
|
||||
// The second request lands in Draining and is `unhandled` (ignored).
|
||||
sleep(Duration::from_millis(5));
|
||||
request_shutdown(r.pid());
|
||||
assert_eq!(mon.rx.recv().unwrap().reason, DownReason::Exit);
|
||||
});
|
||||
assert_eq!(log.get(), vec!["draining", "drained", "terminate"]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn default_shutdown_is_stop() {
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let r = Plain::start(P::On, PD { log: l });
|
||||
sleep(Duration::from_millis(20));
|
||||
let mon = monitor(r.pid());
|
||||
request_shutdown(r.pid());
|
||||
assert_eq!(mon.rx.recv().unwrap().reason, DownReason::Exit);
|
||||
});
|
||||
assert_eq!(log.get(), vec!["terminate"]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn stop_tail_from_a_cast_is_a_normal_exit() {
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let r = settled(&l, false);
|
||||
let mon = monitor(r.pid());
|
||||
r.send(Ev::Cast(Cast::Note("a"))).unwrap();
|
||||
r.send(Ev::Cast(Cast::StopNow)).unwrap();
|
||||
r.send(Ev::Cast(Cast::Note("after-stop"))).unwrap(); // never dispatched
|
||||
assert_eq!(mon.rx.recv().unwrap().reason, DownReason::Exit);
|
||||
});
|
||||
assert_eq!(log.get(), vec!["a", "terminate"]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn linked_peer_death_reaches_exit_row() {
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let r = settled(&l, true);
|
||||
let pid = r.pid();
|
||||
let peer = spawn(move || {
|
||||
link(pid);
|
||||
panic!("peer dies");
|
||||
});
|
||||
let _ = peer.join();
|
||||
sleep(Duration::from_millis(20));
|
||||
r.send(Ev::Cast(Cast::Note("still-running"))).unwrap();
|
||||
assert_eq!(r.call(|r| Ev::Call(Call::Exits(r))).unwrap(), 1);
|
||||
r.shutdown();
|
||||
});
|
||||
assert_eq!(
|
||||
log.get(),
|
||||
vec!["exit", "still-running", "draining", "drained", "terminate"]
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn ref_shutdown_is_graceful_and_waits() {
|
||||
let log = Log::default();
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let r = settled(&l, true);
|
||||
r.shutdown();
|
||||
// terminate has run by the time shutdown() returns.
|
||||
assert_eq!(l.get(), vec!["draining", "drained", "terminate"]);
|
||||
});
|
||||
}
|
||||
+162
-5
@@ -56,7 +56,11 @@ fn snapshot_lists_actors_with_parent_edge() {
|
||||
|
||||
// The root itself is on-CPU (it's running this code) and rooted under
|
||||
// the forest sentinel.
|
||||
let root = snap.actors.iter().find(|a| a.pid == me).expect("root present");
|
||||
let root = snap
|
||||
.actors
|
||||
.iter()
|
||||
.find(|a| a.pid == me)
|
||||
.expect("root present");
|
||||
assert_eq!(root.state, ActorState::Running);
|
||||
assert_eq!(root.supervisor, smarm::Pid::new(u32::MAX, u32::MAX));
|
||||
|
||||
@@ -200,7 +204,11 @@ fn tree_places_child_under_its_spawner() {
|
||||
|
||||
// The root is parented at the forest sentinel, so it's a genuine root,
|
||||
// and the worker it spawned hangs beneath it.
|
||||
let root = t.roots.iter().find(|n| n.info.pid == me).expect("root in forest");
|
||||
let root = t
|
||||
.roots
|
||||
.iter()
|
||||
.find(|n| n.info.pid == me)
|
||||
.expect("root in forest");
|
||||
assert!(!root.orphaned);
|
||||
assert!(
|
||||
root.children.iter().any(|c| c.info.pid == h.pid()),
|
||||
@@ -237,6 +245,13 @@ fn tree_from_nests_children_and_reroots_orphans() {
|
||||
overruns: 0,
|
||||
messages_received: 0,
|
||||
budget_cycles: 0,
|
||||
stack: smarm::StackInfo {
|
||||
reserve: 0,
|
||||
guard: 0,
|
||||
depth_high_water: 0,
|
||||
parks_since_shrink: 0,
|
||||
shrinks: 0,
|
||||
},
|
||||
};
|
||||
|
||||
let snap = RuntimeSnapshot {
|
||||
@@ -251,14 +266,25 @@ fn tree_from_nests_children_and_reroots_orphans() {
|
||||
let t = tree_from(snap);
|
||||
assert_eq!(t.roots.len(), 2);
|
||||
|
||||
let root = t.roots.iter().find(|n| n.info.pid == root_pid).expect("root present");
|
||||
let root = t
|
||||
.roots
|
||||
.iter()
|
||||
.find(|n| n.info.pid == root_pid)
|
||||
.expect("root present");
|
||||
assert!(!root.orphaned);
|
||||
assert_eq!(root.children.len(), 1);
|
||||
assert_eq!(root.children[0].info.pid, child);
|
||||
assert!(!root.children[0].orphaned);
|
||||
|
||||
let o = t.roots.iter().find(|n| n.info.pid == orphan).expect("orphan re-rooted");
|
||||
assert!(o.orphaned, "an actor whose parent is absent must be flagged orphaned");
|
||||
let o = t
|
||||
.roots
|
||||
.iter()
|
||||
.find(|n| n.info.pid == orphan)
|
||||
.expect("orphan re-rooted");
|
||||
assert!(
|
||||
o.orphaned,
|
||||
"an actor whose parent is absent must be flagged orphaned"
|
||||
);
|
||||
assert!(o.children.is_empty());
|
||||
}
|
||||
|
||||
@@ -352,3 +378,134 @@ fn budget_cycles_accumulate_when_enabled() {
|
||||
h.join().unwrap();
|
||||
});
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// RFC 019 §8 — the stack introspection surface.
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/// Burn ~`frames` × 4 KiB of stack with a yield at max depth, so the context
|
||||
/// save samples the high-water there (RFC 019 §2: hwm is SAMPLED at
|
||||
/// deschedule, not tracked continuously).
|
||||
#[inline(never)]
|
||||
fn burn_stack_yielding(frames: usize) -> u64 {
|
||||
let mut local = [0u8; 4096];
|
||||
local[0] = frames as u8;
|
||||
let below = if frames == 0 {
|
||||
smarm::yield_now();
|
||||
0
|
||||
} else {
|
||||
burn_stack_yielding(frames - 1)
|
||||
};
|
||||
std::hint::black_box(&mut local);
|
||||
below.wrapping_add(local[0] as u64)
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn stack_info_reports_defaults_and_sampled_depth() {
|
||||
run(|| {
|
||||
let (ready_tx, ready_rx) = channel::<()>();
|
||||
let (gate_tx, gate_rx) = channel::<()>();
|
||||
|
||||
let h = spawn(move || {
|
||||
// ~32 KiB deep with a yield at the bottom: the sample point.
|
||||
std::hint::black_box(burn_stack_yielding(8));
|
||||
ready_tx.send(()).unwrap();
|
||||
gate_rx.recv().unwrap();
|
||||
});
|
||||
ready_rx.recv().unwrap();
|
||||
|
||||
let info = spin_until(h.pid(), |a| a.state == ActorState::Parked);
|
||||
let s = info.stack;
|
||||
assert_eq!(s.reserve, 64 * 1024, "default reserve");
|
||||
assert_eq!(
|
||||
s.guard,
|
||||
1024 * 1024,
|
||||
"default guard (kernel stack_guard_gap convention)"
|
||||
);
|
||||
assert!(
|
||||
s.depth_high_water >= 8 * 4096,
|
||||
"hwm sampled at the deep yield: expected ≥ 32 KiB, got {}",
|
||||
s.depth_high_water
|
||||
);
|
||||
assert!(
|
||||
s.depth_high_water < s.reserve,
|
||||
"depth {} cannot exceed the reserve {}",
|
||||
s.depth_high_water,
|
||||
s.reserve
|
||||
);
|
||||
// Parked at the gate right now, never shrunk (64 KiB reserve cannot
|
||||
// cross the shrink threshold).
|
||||
assert!(s.parks_since_shrink >= 1, "the gate park must be counted");
|
||||
assert_eq!(s.shrinks, 0);
|
||||
|
||||
gate_tx.send(()).unwrap();
|
||||
h.join().unwrap();
|
||||
});
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn stack_info_shrink_counters_are_live() {
|
||||
use smarm::runtime::{Config, SHRINK_COOLDOWN, SHRINK_THRESHOLD};
|
||||
use smarm::{spawn_with, SpawnOpts};
|
||||
|
||||
let rt = smarm::runtime::init(Config::exact(1));
|
||||
rt.run(|| {
|
||||
let (park_tx, park_rx) = channel::<()>();
|
||||
|
||||
let spike = 768 * 4096;
|
||||
assert!(spike > SHRINK_THRESHOLD);
|
||||
let worker = spawn_with(
|
||||
SpawnOpts {
|
||||
stack_reserve: Some(8 * 1024 * 1024),
|
||||
..SpawnOpts::default()
|
||||
},
|
||||
move || {
|
||||
std::hint::black_box(burn_stack_yielding(768));
|
||||
for _ in 0..(SHRINK_COOLDOWN + 8) {
|
||||
park_rx.recv().unwrap();
|
||||
}
|
||||
},
|
||||
);
|
||||
|
||||
let wpid = worker.pid();
|
||||
// Before any parks complete: the spike depth is visible.
|
||||
let info = spin_until(wpid, |a| a.state == ActorState::Parked);
|
||||
assert!(
|
||||
info.stack.depth_high_water >= spike,
|
||||
"spike should be sampled: {} < {spike}",
|
||||
info.stack.depth_high_water
|
||||
);
|
||||
|
||||
// Cross the cooldown, then read the counters live while the worker
|
||||
// is parked waiting for the remaining rounds (post-join the slot is
|
||||
// reclaimed and the generation check correctly hides it).
|
||||
for _ in 0..(SHRINK_COOLDOWN + 2) {
|
||||
spin_until(wpid, |a| a.state == ActorState::Parked);
|
||||
park_tx.send(()).unwrap();
|
||||
}
|
||||
let info = spin_until(wpid, |a| {
|
||||
a.state == ActorState::Parked && a.stack.shrinks >= 1
|
||||
});
|
||||
let s = info.stack;
|
||||
assert!(
|
||||
s.shrinks >= 1,
|
||||
"cooldown was crossed with a spike above threshold"
|
||||
);
|
||||
assert!(
|
||||
s.parks_since_shrink < SHRINK_COOLDOWN,
|
||||
"counter must reset at shrink: {}",
|
||||
s.parks_since_shrink
|
||||
);
|
||||
assert!(
|
||||
s.depth_high_water < spike,
|
||||
"hwm resets to the shallow park sp at shrink; got {}",
|
||||
s.depth_high_water
|
||||
);
|
||||
|
||||
for _ in 0..6 {
|
||||
spin_until(wpid, |a| a.state == ActorState::Parked);
|
||||
park_tx.send(()).unwrap();
|
||||
}
|
||||
worker.join().unwrap();
|
||||
});
|
||||
}
|
||||
|
||||
+13
-3
@@ -56,8 +56,16 @@ fn other_actors_run_while_block_on_io_is_in_flight() {
|
||||
let pos_2 = v.iter().position(|&x| x == 2).unwrap();
|
||||
let pos_3 = v.iter().position(|&x| x == 3).unwrap();
|
||||
let pos_4 = v.iter().position(|&x| x == 4).unwrap();
|
||||
assert!(pos_2 < pos_4, "B's first step ran after A resumed: {:?}", *v);
|
||||
assert!(pos_3 < pos_4, "B's second step ran after A resumed: {:?}", *v);
|
||||
assert!(
|
||||
pos_2 < pos_4,
|
||||
"B's first step ran after A resumed: {:?}",
|
||||
*v
|
||||
);
|
||||
assert!(
|
||||
pos_3 < pos_4,
|
||||
"B's second step ran after A resumed: {:?}",
|
||||
*v
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
@@ -76,7 +84,9 @@ fn many_concurrent_block_on_io_calls_all_complete() {
|
||||
cc.fetch_add(n, Ordering::SeqCst);
|
||||
}));
|
||||
}
|
||||
for h in handles { h.join().unwrap(); }
|
||||
for h in handles {
|
||||
h.join().unwrap();
|
||||
}
|
||||
});
|
||||
assert_eq!(counter.load(Ordering::SeqCst), 10);
|
||||
}
|
||||
|
||||
+11
-4
@@ -144,8 +144,7 @@ fn write_sugar_sends_bytes_to_pipe() {
|
||||
// Pipe is empty + has buffer space, so this returns immediately
|
||||
// after wait_writable wakes (which happens fast because the
|
||||
// kernel marks an empty pipe as immediately writable).
|
||||
let n = smarm::scheduler::write(p_writer.write, b"smarm")
|
||||
.expect("write failed");
|
||||
let n = smarm::scheduler::write(p_writer.write, b"smarm").expect("write failed");
|
||||
assert_eq!(n, 5);
|
||||
c.fetch_add(1, Ordering::SeqCst);
|
||||
});
|
||||
@@ -209,10 +208,18 @@ fn other_actors_run_while_one_is_parked_on_wait_readable() {
|
||||
let pos_lit_a = v.iter().position(|&c| c == b'a').unwrap();
|
||||
let big_b_count = v.iter().filter(|&&c| c == b'B').count();
|
||||
assert_eq!(big_b_count, 3, "B should have made 3 steps: {:?}", *v);
|
||||
assert!(pos_big_a < pos_lit_a, "A pre-park before A post-park: {:?}", *v);
|
||||
assert!(
|
||||
pos_big_a < pos_lit_a,
|
||||
"A pre-park before A post-park: {:?}",
|
||||
*v
|
||||
);
|
||||
// At least the last B step should be before A resumes.
|
||||
let last_big_b = v.iter().rposition(|&c| c == b'B').unwrap();
|
||||
assert!(last_big_b < pos_lit_a, "B should finish before A resumes: {:?}", *v);
|
||||
assert!(
|
||||
last_big_b < pos_lit_a,
|
||||
"B should finish before A resumes: {:?}",
|
||||
*v
|
||||
);
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
+10
-4
@@ -57,7 +57,10 @@ fn linked_pair_one_panics_other_is_stopped() {
|
||||
panic!("boom");
|
||||
});
|
||||
|
||||
let dn = down_b.rx.recv().expect("monitor channel closed before Down");
|
||||
let dn = down_b
|
||||
.rx
|
||||
.recv()
|
||||
.expect("monitor channel closed before Down");
|
||||
assert_eq!(dn.pid, b, "Down reported the wrong pid");
|
||||
if matches!(dn.reason, DownReason::Stopped) {
|
||||
s.store(true, Ordering::SeqCst);
|
||||
@@ -117,7 +120,7 @@ fn normal_exit_does_not_propagate() {
|
||||
let a = ha.pid();
|
||||
link(a);
|
||||
yield_now(); // let A run to completion and finalize
|
||||
// A exited normally: nothing should have landed on the inbox.
|
||||
// A exited normally: nothing should have landed on the inbox.
|
||||
if let Ok(None) = inbox.try_recv() {
|
||||
e.store(true, Ordering::SeqCst);
|
||||
}
|
||||
@@ -152,7 +155,10 @@ fn link_to_dead_pid_stops_a_nontrapping_caller() {
|
||||
});
|
||||
let b = hb.pid();
|
||||
let down_b = monitor(b);
|
||||
let dn = down_b.rx.recv().expect("monitor channel closed before Down");
|
||||
let dn = down_b
|
||||
.rx
|
||||
.recv()
|
||||
.expect("monitor channel closed before Down");
|
||||
if matches!(dn.reason, DownReason::Stopped) {
|
||||
s.store(true, Ordering::SeqCst);
|
||||
}
|
||||
@@ -208,7 +214,7 @@ fn unlink_prevents_propagation() {
|
||||
panic!("boom"); // abnormal, but the link is gone
|
||||
});
|
||||
yield_now(); // let A link, unlink, and panic
|
||||
// Unlinked before death → no ExitSignal should have arrived.
|
||||
// Unlinked before death → no ExitSignal should have arrived.
|
||||
if let Ok(None) = inbox.try_recv() {
|
||||
sv.store(true, Ordering::SeqCst);
|
||||
}
|
||||
|
||||
+27
-6
@@ -67,7 +67,10 @@ fn monitor_already_dead_target_is_noproc() {
|
||||
// and its generation bumped, so `pid` is now stale.
|
||||
h.join().unwrap();
|
||||
let down = monitor(pid);
|
||||
let d = down.rx.recv().expect("NoProc Down should be delivered immediately");
|
||||
let d = down
|
||||
.rx
|
||||
.recv()
|
||||
.expect("NoProc Down should be delivered immediately");
|
||||
assert_eq!(d.pid, pid);
|
||||
if matches!(d.reason, DownReason::NoProc) {
|
||||
o.store(true, Ordering::SeqCst);
|
||||
@@ -91,7 +94,11 @@ fn multiple_monitors_all_notified() {
|
||||
}
|
||||
}
|
||||
});
|
||||
assert_eq!(count.load(Ordering::SeqCst), 3, "every monitor should see the Down");
|
||||
assert_eq!(
|
||||
count.load(Ordering::SeqCst),
|
||||
3,
|
||||
"every monitor should see the Down"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
@@ -103,8 +110,15 @@ fn demonitor_stops_delivery() {
|
||||
let h = spawn(|| {});
|
||||
let pid = h.pid();
|
||||
let m = monitor(pid);
|
||||
assert_eq!(demonitor(&m), Some(m.id), "live registration should be removed");
|
||||
assert!(m.rx.recv().is_err(), "no Down should arrive after demonitor");
|
||||
assert_eq!(
|
||||
demonitor(&m),
|
||||
Some(m.id),
|
||||
"live registration should be removed"
|
||||
);
|
||||
assert!(
|
||||
m.rx.recv().is_err(),
|
||||
"no Down should arrive after demonitor"
|
||||
);
|
||||
let _ = h.join();
|
||||
});
|
||||
}
|
||||
@@ -122,7 +136,10 @@ fn demonitor_one_of_many() {
|
||||
let _ = h.join();
|
||||
assert!(matches!(ms[0].rx.recv().unwrap().reason, DownReason::Exit));
|
||||
assert!(matches!(ms[2].rx.recv().unwrap().reason, DownReason::Exit));
|
||||
assert!(ms[1].rx.recv().is_err(), "demonitored channel should be closed");
|
||||
assert!(
|
||||
ms[1].rx.recv().is_err(),
|
||||
"demonitored channel should be closed"
|
||||
);
|
||||
});
|
||||
}
|
||||
|
||||
@@ -136,7 +153,11 @@ fn demonitor_after_fire_is_none() {
|
||||
let m = monitor(pid);
|
||||
let d = m.rx.recv().expect("Down before close");
|
||||
assert!(matches!(d.reason, DownReason::Exit));
|
||||
assert_eq!(demonitor(&m), None, "already-fired monitor has nothing to remove");
|
||||
assert_eq!(
|
||||
demonitor(&m),
|
||||
None,
|
||||
"already-fired monitor has nothing to remove"
|
||||
);
|
||||
let _ = h.join();
|
||||
});
|
||||
}
|
||||
|
||||
+16
-4
@@ -3,9 +3,9 @@
|
||||
//! needs to be able to park.
|
||||
|
||||
use smarm::{run, spawn, yield_now, LockTimeout, Mutex};
|
||||
use std::sync::atomic::{AtomicU32, Ordering};
|
||||
use std::sync::Arc;
|
||||
use std::sync::Mutex as StdMutex;
|
||||
use std::sync::atomic::{AtomicU32, Ordering};
|
||||
use std::time::{Duration, Instant};
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
@@ -111,8 +111,16 @@ fn contended_lock_parks_until_holder_releases() {
|
||||
let pos_b_locked = v.iter().position(|s| *s == "B_locked").unwrap();
|
||||
|
||||
assert!(pos_a_locked < pos_b_try, "log: {:?}", *v);
|
||||
assert!(pos_b_try < pos_a_dropped, "B should attempt before A drops: {:?}", *v);
|
||||
assert!(pos_a_dropped < pos_b_locked, "B should lock only after A drops: {:?}", *v);
|
||||
assert!(
|
||||
pos_b_try < pos_a_dropped,
|
||||
"B should attempt before A drops: {:?}",
|
||||
*v
|
||||
);
|
||||
assert!(
|
||||
pos_a_dropped < pos_b_locked,
|
||||
"B should lock only after A drops: {:?}",
|
||||
*v
|
||||
);
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
@@ -209,7 +217,11 @@ fn waiters_are_granted_the_lock_in_fifo_order() {
|
||||
});
|
||||
|
||||
let v = order.lock().unwrap().clone();
|
||||
assert_eq!(v, vec![1, 2, 3, 4], "waiters should acquire in arrival order");
|
||||
assert_eq!(
|
||||
v,
|
||||
vec![1, 2, 3, 4],
|
||||
"waiters should acquire in arrival order"
|
||||
);
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
+5
-3
@@ -77,8 +77,7 @@ fn observer_reports_none_for_a_forged_pid() {
|
||||
// An index that is not in the slab at all — the verb relays the
|
||||
// primitive's `None` faithfully.
|
||||
let forged = smarm::Pid::new(u32::MAX - 1, 0);
|
||||
let ObserverReply::ActorInfo(none) =
|
||||
obs.call(ObserverRequest::ActorInfo(forged)).unwrap()
|
||||
let ObserverReply::ActorInfo(none) = obs.call(ObserverRequest::ActorInfo(forged)).unwrap()
|
||||
else {
|
||||
panic!("ActorInfo verb must reply ActorInfo");
|
||||
};
|
||||
@@ -113,7 +112,10 @@ fn observer_sees_a_parked_actor_as_parked() {
|
||||
}
|
||||
smarm::yield_now();
|
||||
}
|
||||
assert!(parked, "observer should eventually report the worker as Parked");
|
||||
assert!(
|
||||
parked,
|
||||
"observer should eventually report the worker as Parked"
|
||||
);
|
||||
|
||||
gate_tx.send(()).unwrap();
|
||||
worker.join().unwrap();
|
||||
|
||||
@@ -0,0 +1,69 @@
|
||||
//! RFC 018 scheduler park/wake — observable-behavior guards.
|
||||
//!
|
||||
//! These pin the two timer-latency properties the park/wake swap must
|
||||
//! preserve or introduce:
|
||||
//!
|
||||
//! - `sleep_fires_under_saturation`: due timers fire even when every
|
||||
//! scheduler is busy (nobody parked ⇒ no timekeeper) — the busy-path
|
||||
//! due-check, ratified design point (a). The old drain phase gave this
|
||||
//! for free (timers drained every loop iteration); the new design must
|
||||
//! not lose it.
|
||||
//! - `submillisecond_sleep_is_prompt`: a sub-ms sleep completes promptly.
|
||||
//! Under the old wake pipe, `poll_wake`'s `as_millis` truncation turned
|
||||
//! sub-ms deadlines into 0ms busy-polls (correct wall time, pathological
|
||||
//! CPU); under park/wake the futex timespec carries full nanosecond
|
||||
//! precision.
|
||||
|
||||
use std::sync::atomic::{AtomicBool, Ordering};
|
||||
use std::sync::Arc;
|
||||
use std::time::{Duration, Instant};
|
||||
|
||||
#[test]
|
||||
fn sleep_fires_under_saturation() {
|
||||
let rt = smarm::runtime::init(smarm::runtime::Config::exact(4));
|
||||
rt.run(|| {
|
||||
let stop = Arc::new(AtomicBool::new(false));
|
||||
let mut spinners = Vec::new();
|
||||
// 8 spinners over 4 schedulers: the run queue never empties, so no
|
||||
// scheduler ever parks and no timekeeper exists. Only the busy-path
|
||||
// due-check can fire the sleeper's timer before the spinners quit.
|
||||
for _ in 0..8 {
|
||||
let stop = stop.clone();
|
||||
spinners.push(smarm::spawn(move || {
|
||||
let t0 = Instant::now();
|
||||
while !stop.load(Ordering::Relaxed) && t0.elapsed() < Duration::from_secs(5) {
|
||||
smarm::yield_now();
|
||||
}
|
||||
}));
|
||||
}
|
||||
let t0 = Instant::now();
|
||||
smarm::sleep(Duration::from_millis(10));
|
||||
let dt = t0.elapsed();
|
||||
stop.store(true, Ordering::Relaxed);
|
||||
for s in spinners {
|
||||
let _ = s.join();
|
||||
}
|
||||
assert!(
|
||||
dt < Duration::from_millis(500),
|
||||
"10ms sleep took {dt:?} under scheduler saturation — busy-path \
|
||||
timer firing is broken (timekeeper-only firing stalls under load)"
|
||||
);
|
||||
});
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn submillisecond_sleep_is_prompt() {
|
||||
let rt = smarm::runtime::init(smarm::runtime::Config::exact(2));
|
||||
rt.run(|| {
|
||||
// Warm one iteration, then measure.
|
||||
smarm::sleep(Duration::from_micros(500));
|
||||
let t0 = Instant::now();
|
||||
smarm::sleep(Duration::from_micros(500));
|
||||
let dt = t0.elapsed();
|
||||
assert!(dt >= Duration::from_micros(400), "woke early: {dt:?}");
|
||||
assert!(
|
||||
dt < Duration::from_millis(100),
|
||||
"500µs sleep took {dt:?} — sub-ms deadline handling is broken"
|
||||
);
|
||||
});
|
||||
}
|
||||
+13
-3
@@ -44,7 +44,10 @@ fn a_dead_actor_vanishes_from_every_group_it_joined() {
|
||||
// Drain-on-contact: touching g1 detects the death and sweeps the pid
|
||||
// out of every group (g2 included), not just g1.
|
||||
assert!(members("g1").is_empty(), "evicted from the touched group");
|
||||
assert!(members("g2").is_empty(), "and swept from the untouched group");
|
||||
assert!(
|
||||
members("g2").is_empty(),
|
||||
"and swept from the untouched group"
|
||||
);
|
||||
assert_eq!(pick("g1"), None);
|
||||
});
|
||||
}
|
||||
@@ -83,7 +86,11 @@ fn live_members_survive_a_peers_death() {
|
||||
tx_a.send(()).unwrap();
|
||||
a.join().unwrap();
|
||||
|
||||
assert_eq!(members("svc"), vec![b.pid()], "only the dead peer is reaped");
|
||||
assert_eq!(
|
||||
members("svc"),
|
||||
vec![b.pid()],
|
||||
"only the dead peer is reaped"
|
||||
);
|
||||
assert_eq!(pick("svc"), Some(b.pid()));
|
||||
|
||||
tx_b.send(()).unwrap();
|
||||
@@ -125,7 +132,10 @@ fn joining_an_already_dead_pid_is_evicted_on_next_contact() {
|
||||
// monitor() on a gone pid queues a NoProc Down immediately, so the
|
||||
// membership is reaped the next time the group is touched.
|
||||
join("late", pid);
|
||||
assert!(members("late").is_empty(), "dead-at-join member is reaped on read");
|
||||
assert!(
|
||||
members("late").is_empty(),
|
||||
"dead-at-join member is reaped on read"
|
||||
);
|
||||
assert_eq!(pick("late"), None);
|
||||
});
|
||||
}
|
||||
|
||||
+10
-2
@@ -41,7 +41,11 @@ fn stop_storm_does_not_poison_runtime() {
|
||||
}
|
||||
c.fetch_add(1, Ordering::SeqCst);
|
||||
});
|
||||
assert_eq!(completed.load(Ordering::SeqCst), 1, "root completed cleanly");
|
||||
assert_eq!(
|
||||
completed.load(Ordering::SeqCst),
|
||||
1,
|
||||
"root completed cleanly"
|
||||
);
|
||||
}
|
||||
|
||||
/// The sharper repro: a stop-flagged actor whose *next allocation* is the
|
||||
@@ -85,5 +89,9 @@ fn self_stop_during_spawn_does_not_poison_shared_mutex() {
|
||||
}
|
||||
c.fetch_add(1, Ordering::SeqCst);
|
||||
});
|
||||
assert_eq!(completed.load(Ordering::SeqCst), 1, "root completed cleanly");
|
||||
assert_eq!(
|
||||
completed.load(Ordering::SeqCst),
|
||||
1,
|
||||
"root completed cleanly"
|
||||
);
|
||||
}
|
||||
|
||||
+15
-4
@@ -43,10 +43,21 @@ fn check_yields_when_timeslice_expired() {
|
||||
let pos_big_b = v.iter().position(|&c| c == b'B').unwrap();
|
||||
let pos_lit_a = v.iter().position(|&c| c == b'a').unwrap();
|
||||
let pos_lit_b = v.iter().position(|&c| c == b'b').unwrap();
|
||||
assert!(pos_big_a < pos_lit_a, "A's tail ran before B's head: {:?}", *v);
|
||||
assert!(pos_big_b < pos_lit_b, "B's tail ran before A's head: {:?}", *v);
|
||||
assert!(pos_big_a.max(pos_big_b) < pos_lit_a.min(pos_lit_b),
|
||||
"preemption didn't interleave: {:?}", *v);
|
||||
assert!(
|
||||
pos_big_a < pos_lit_a,
|
||||
"A's tail ran before B's head: {:?}",
|
||||
*v
|
||||
);
|
||||
assert!(
|
||||
pos_big_b < pos_lit_b,
|
||||
"B's tail ran before A's head: {:?}",
|
||||
*v
|
||||
);
|
||||
assert!(
|
||||
pos_big_a.max(pos_big_b) < pos_lit_a.min(pos_lit_b),
|
||||
"preemption didn't interleave: {:?}",
|
||||
*v
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
|
||||
+13
-4
@@ -65,7 +65,10 @@ fn name_held_by_live_actor_is_taken() {
|
||||
ready_rx.recv().unwrap();
|
||||
// Root tries to claim a live actor's name for itself -> NameTaken.
|
||||
let (tx_b, _rx_b) = channel::<u64>();
|
||||
assert_eq!(register(SVC, tx_b), Err(RegisterError::NameTaken { holder: a.pid() }));
|
||||
assert_eq!(
|
||||
register(SVC, tx_b),
|
||||
Err(RegisterError::NameTaken { holder: a.pid() })
|
||||
);
|
||||
send(SVC, 0).unwrap(); // release a (delivers to the holder, a)
|
||||
a.join().unwrap();
|
||||
});
|
||||
@@ -105,7 +108,10 @@ fn dead_holder_is_pruned_and_name_taken_over() {
|
||||
fn send_errors_unresolved_and_no_channel() {
|
||||
run(|| {
|
||||
// No actor at all.
|
||||
assert!(matches!(send(Name::<u64>::new("ghost"), 1u64), Err(SendError::Unresolved(_))));
|
||||
assert!(matches!(
|
||||
send(Name::<u64>::new("ghost"), 1u64),
|
||||
Err(SendError::Unresolved(_))
|
||||
));
|
||||
|
||||
let (ready_tx, ready_rx) = channel::<()>();
|
||||
let (tx, rx) = channel::<u64>();
|
||||
@@ -227,8 +233,11 @@ fn send_dyn_delivers_and_reports_wrong_type() {
|
||||
ready_rx.recv().unwrap();
|
||||
let p = h.pid(); // a bare Pid<Erased>, as if recovered off a Down
|
||||
send_dyn::<u64>(p, 3u64).unwrap(); // right type: delivered
|
||||
// Live actor, but it has no channel for &str — the genuinely-fallible case.
|
||||
assert!(matches!(send_dyn::<&'static str>(p, "nope"), Err(SendError::NoChannel(_))));
|
||||
// Live actor, but it has no channel for &str — the genuinely-fallible case.
|
||||
assert!(matches!(
|
||||
send_dyn::<&'static str>(p, "nope"),
|
||||
Err(SendError::NoChannel(_))
|
||||
));
|
||||
done_tx.send(()).unwrap();
|
||||
h.join().unwrap();
|
||||
});
|
||||
|
||||
@@ -0,0 +1,288 @@
|
||||
//! Root exit — the run's initial actor returning means "the program is done".
|
||||
//!
|
||||
//! When the root finalizes, the runtime delivers `request_shutdown` to every
|
||||
//! **forest root**: each live actor whose parent is the run itself (a plain
|
||||
//! `spawn` from the root closure) or is already dead. Nothing below a live
|
||||
//! parent is touched directly — a supervisor gets one request and runs its
|
||||
//! own ordered shutdown per child `Shutdown` policy.
|
||||
//!
|
||||
//! - Non-trapping actors are stopped outright, exactly as by
|
||||
//! `request_shutdown` — a `spawn(|| { sleep(..); work() })` the root did
|
||||
//! not `join` does NOT get to finish. Join it, supervise it, or trap.
|
||||
//! - Trapping actors get `handle_shutdown` / an `ExitSignal{Shutdown}` and
|
||||
//! may keep running (`Continue`, drain, then stop themselves) — timers and
|
||||
//! all; the run ends when they do. There is no second, forcing sweep.
|
||||
//! - A periodic-timer daemon (the classic wedge) never blocks `run()`.
|
||||
|
||||
use smarm::gen_server::{start, GenServer, GenServerCtx, ShutdownAction, StopHandle, TimerHandle};
|
||||
use smarm::supervisor::{ChildSpec, OneForOne, Restart, Shutdown};
|
||||
use smarm::{run, sleep, spawn, trap_exit, DownReason};
|
||||
use std::sync::atomic::{AtomicBool, AtomicUsize, Ordering};
|
||||
use std::sync::{Arc, Mutex};
|
||||
use std::time::{Duration, Instant};
|
||||
|
||||
fn assert_prompt(start: Instant, what: &str) {
|
||||
assert!(
|
||||
start.elapsed() < Duration::from_secs(2),
|
||||
"{what}: run() took {:?}",
|
||||
start.elapsed()
|
||||
);
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Bare (non-gen_server) actors
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/// A non-trapping sleeper the root did not join is stopped, not waited for.
|
||||
#[test]
|
||||
fn unjoined_non_trapping_sleeper_is_stopped() {
|
||||
let finished = Arc::new(AtomicBool::new(false));
|
||||
let f = finished.clone();
|
||||
let t = Instant::now();
|
||||
run(move || {
|
||||
spawn(move || {
|
||||
sleep(Duration::from_secs(5));
|
||||
f.store(true, Ordering::SeqCst);
|
||||
});
|
||||
});
|
||||
assert_prompt(t, "sleeper");
|
||||
assert!(
|
||||
!finished.load(Ordering::SeqCst),
|
||||
"sleeper should have been stopped"
|
||||
);
|
||||
}
|
||||
|
||||
/// A trapping bare actor sees `Shutdown` from the root's exit and may keep
|
||||
/// working — here it sleeps (a timer!) after the signal, then returns. The
|
||||
/// run waits for it: no forcing sweep.
|
||||
#[test]
|
||||
fn trapping_actor_may_finish_after_shutdown_signal() {
|
||||
let finished = Arc::new(AtomicBool::new(false));
|
||||
let f = finished.clone();
|
||||
run(move || {
|
||||
spawn(move || {
|
||||
let inbox = trap_exit();
|
||||
let sig = inbox.recv().expect("shutdown signal");
|
||||
assert_eq!(sig.reason, DownReason::Shutdown);
|
||||
sleep(Duration::from_millis(100));
|
||||
f.store(true, Ordering::SeqCst);
|
||||
});
|
||||
// Trapping is a runtime opt-in: give the actor a chance to run
|
||||
// `trap_exit()`; a not-yet-run actor is non-trapping and is stopped.
|
||||
sleep(Duration::from_millis(20));
|
||||
});
|
||||
assert!(
|
||||
finished.load(Ordering::SeqCst),
|
||||
"trapping actor must be allowed to finish"
|
||||
);
|
||||
}
|
||||
|
||||
/// The classic wedge: a lazily spawned daemon that never returns on its own.
|
||||
#[test]
|
||||
fn parked_forever_daemon_does_not_block_run() {
|
||||
let t = Instant::now();
|
||||
run(|| {
|
||||
let (_tx, rx) = smarm::channel::<()>();
|
||||
spawn(move || {
|
||||
let _ = rx.recv(); // parked forever: sender is held by the root, which returns
|
||||
});
|
||||
// Leak the sender into the daemon's own scope so nothing else drops it.
|
||||
std::mem::forget(_tx);
|
||||
});
|
||||
assert_prompt(t, "daemon");
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// gen_servers
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/// A non-trapping ticker with a periodic timer: the timer wheel is never empty,
|
||||
/// and root exit must still end the run.
|
||||
struct Ticker {
|
||||
ticks: Arc<AtomicUsize>,
|
||||
}
|
||||
impl GenServer for Ticker {
|
||||
type Call = ();
|
||||
type Reply = ();
|
||||
type Cast = ();
|
||||
type Info = ();
|
||||
type Timer = ();
|
||||
fn init(&mut self, ctx: &GenServerCtx<Self>) {
|
||||
ctx.timer().tick_every(Duration::from_millis(5), ());
|
||||
}
|
||||
fn handle_call(&mut self, _: ()) {}
|
||||
fn handle_cast(&mut self, _: ()) {}
|
||||
fn handle_timer(&mut self, _: ()) {
|
||||
self.ticks.fetch_add(1, Ordering::SeqCst);
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn periodic_timer_daemon_does_not_block_run() {
|
||||
let ticks = Arc::new(AtomicUsize::new(0));
|
||||
let tk = ticks.clone();
|
||||
let t = Instant::now();
|
||||
run(move || {
|
||||
let _r = start(Ticker { ticks: tk });
|
||||
sleep(Duration::from_millis(50));
|
||||
});
|
||||
assert_prompt(t, "ticker");
|
||||
assert!(
|
||||
ticks.load(Ordering::SeqCst) >= 3,
|
||||
"ticker should have ticked"
|
||||
);
|
||||
}
|
||||
|
||||
/// A trapping server that answers `Continue`, keeps ticking on its own timer
|
||||
/// (draining), and stops itself later. Root exit must not cut it short.
|
||||
struct Drainer {
|
||||
log: Arc<Mutex<Vec<&'static str>>>,
|
||||
shutdowns: Arc<AtomicUsize>,
|
||||
ticks_after_shutdown: usize,
|
||||
stop: Option<StopHandle<Drainer>>,
|
||||
timer: Option<TimerHandle<Drainer>>,
|
||||
draining: bool,
|
||||
}
|
||||
impl GenServer for Drainer {
|
||||
type Call = ();
|
||||
type Reply = ();
|
||||
type Cast = ();
|
||||
type Info = ();
|
||||
type Timer = ();
|
||||
fn init(&mut self, ctx: &GenServerCtx<Self>) {
|
||||
ctx.trap_exit();
|
||||
self.stop = Some(ctx.stop_handle());
|
||||
self.timer = Some(ctx.timer());
|
||||
}
|
||||
fn handle_call(&mut self, _: ()) {}
|
||||
fn handle_cast(&mut self, _: ()) {}
|
||||
fn handle_shutdown(&mut self) -> ShutdownAction {
|
||||
self.shutdowns.fetch_add(1, Ordering::SeqCst);
|
||||
self.log.lock().unwrap().push("handle_shutdown");
|
||||
self.draining = true;
|
||||
self.timer
|
||||
.as_ref()
|
||||
.unwrap()
|
||||
.tick_every(Duration::from_millis(10), ());
|
||||
ShutdownAction::Continue
|
||||
}
|
||||
fn handle_timer(&mut self, _: ()) {
|
||||
if !self.draining {
|
||||
return;
|
||||
}
|
||||
self.ticks_after_shutdown += 1;
|
||||
if self.ticks_after_shutdown == 3 {
|
||||
self.log.lock().unwrap().push("drained");
|
||||
self.stop.as_ref().unwrap().stop();
|
||||
}
|
||||
}
|
||||
fn terminate(&mut self) {
|
||||
self.log.lock().unwrap().push("terminate");
|
||||
}
|
||||
}
|
||||
|
||||
fn drainer(log: &Arc<Mutex<Vec<&'static str>>>, shutdowns: &Arc<AtomicUsize>) -> Drainer {
|
||||
Drainer {
|
||||
log: log.clone(),
|
||||
shutdowns: shutdowns.clone(),
|
||||
ticks_after_shutdown: 0,
|
||||
stop: None,
|
||||
timer: None,
|
||||
draining: false,
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn trapping_server_drains_with_timers_after_root_exit() {
|
||||
let log = Arc::new(Mutex::new(Vec::new()));
|
||||
let shutdowns = Arc::new(AtomicUsize::new(0));
|
||||
let (l, s) = (log.clone(), shutdowns.clone());
|
||||
run(move || {
|
||||
let r = start(drainer(&l, &s));
|
||||
// A gen_server's lifetime is governed by its refs: dropping the last
|
||||
// one closes the inbox and ends the loop cleanly, which would cut the
|
||||
// drain short for a reason unrelated to root exit. Pin it the way a
|
||||
// registered name would.
|
||||
std::mem::forget(r);
|
||||
sleep(Duration::from_millis(20)); // let init (trap_exit) run
|
||||
});
|
||||
assert_eq!(
|
||||
*log.lock().unwrap(),
|
||||
vec!["handle_shutdown", "drained", "terminate"]
|
||||
);
|
||||
assert_eq!(shutdowns.load(Ordering::SeqCst), 1);
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Supervision trees: only forest roots are addressed
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/// The supervisor gets ONE request and runs its ordered shutdown; a trapping
|
||||
/// child under it sees exactly one `Shutdown` — from the supervisor, not a
|
||||
/// second one from the runtime — and is allowed to finish its drain (a sleep,
|
||||
/// i.e. a timer) under `Shutdown::Infinity`.
|
||||
#[test]
|
||||
fn supervised_children_are_shut_down_only_via_their_supervisor() {
|
||||
let signals = Arc::new(AtomicUsize::new(0));
|
||||
let drained = Arc::new(AtomicBool::new(false));
|
||||
let sup_returned = Arc::new(AtomicBool::new(false));
|
||||
let (sg, dr, sr) = (signals.clone(), drained.clone(), sup_returned.clone());
|
||||
run(move || {
|
||||
spawn(move || {
|
||||
OneForOne::new()
|
||||
.child(
|
||||
ChildSpec::new(Restart::Permanent, move || {
|
||||
let inbox = trap_exit();
|
||||
while let Ok(sig) = inbox.recv() {
|
||||
if sig.reason == DownReason::Shutdown {
|
||||
sg.fetch_add(1, Ordering::SeqCst);
|
||||
sleep(Duration::from_millis(100));
|
||||
// A late second signal would land here.
|
||||
while let Ok(Some(sig)) = inbox.try_recv() {
|
||||
if sig.reason == DownReason::Shutdown {
|
||||
sg.fetch_add(1, Ordering::SeqCst);
|
||||
}
|
||||
}
|
||||
dr.store(true, Ordering::SeqCst);
|
||||
return;
|
||||
}
|
||||
}
|
||||
})
|
||||
.shutdown(Shutdown::Infinity),
|
||||
)
|
||||
.run();
|
||||
sr.store(true, Ordering::SeqCst);
|
||||
});
|
||||
sleep(Duration::from_millis(30)); // let the tree settle
|
||||
});
|
||||
assert!(
|
||||
sup_returned.load(Ordering::SeqCst),
|
||||
"supervisor should return normally"
|
||||
);
|
||||
assert!(
|
||||
drained.load(Ordering::SeqCst),
|
||||
"child should finish its drain"
|
||||
);
|
||||
assert_eq!(signals.load(Ordering::SeqCst), 1);
|
||||
}
|
||||
|
||||
/// A supervised non-trapping child under `Shutdown::Timeout` is stopped by
|
||||
/// the supervisor's policy, and the run ends promptly.
|
||||
#[test]
|
||||
fn supervisor_tree_is_torn_down_promptly_on_root_exit() {
|
||||
let t = Instant::now();
|
||||
run(|| {
|
||||
spawn(|| {
|
||||
OneForOne::new()
|
||||
.child(
|
||||
ChildSpec::new(Restart::Permanent, || loop {
|
||||
sleep(Duration::from_millis(5));
|
||||
})
|
||||
.shutdown(Shutdown::Timeout(Duration::from_millis(50))),
|
||||
)
|
||||
.run();
|
||||
});
|
||||
sleep(Duration::from_millis(30));
|
||||
});
|
||||
assert_prompt(t, "tree");
|
||||
}
|
||||
@@ -0,0 +1,36 @@
|
||||
//! Under `smarm-trace`, every actor the root-exit sweep reaches is recorded
|
||||
//! as a `root_sweep` event — the way to *see* unsupervised leftovers. One test
|
||||
//! per binary: the trace file is process-global.
|
||||
#![cfg(feature = "smarm-trace")]
|
||||
|
||||
use smarm::gen_server::{self, GenServer, GenServerCtx};
|
||||
use smarm::run;
|
||||
|
||||
struct Quiet;
|
||||
impl GenServer for Quiet {
|
||||
type Call = ();
|
||||
type Reply = ();
|
||||
type Cast = ();
|
||||
type Info = ();
|
||||
type Timer = ();
|
||||
fn init(&mut self, _: &GenServerCtx<Self>) {}
|
||||
fn handle_call(&mut self, _: ()) {}
|
||||
fn handle_cast(&mut self, _: ()) {}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn forgotten_server_shows_up_as_root_sweep() {
|
||||
let path = std::env::temp_dir().join(format!("smarm_root_sweep_{}.json", std::process::id()));
|
||||
std::env::set_var("SMARM_TRACE_FILE", &path);
|
||||
run(|| {
|
||||
let srv = gen_server::start(Quiet);
|
||||
srv.call(()).unwrap();
|
||||
drop(srv); // forgotten: nobody supervises it, nobody holds it
|
||||
});
|
||||
let trace = std::fs::read_to_string(&path).expect("trace file written");
|
||||
let _ = std::fs::remove_file(&path);
|
||||
assert!(
|
||||
trace.contains("root_sweep stopped"),
|
||||
"expected a root_sweep line for the non-trapping leftover; got:\n{trace}"
|
||||
);
|
||||
}
|
||||
+110
-25
@@ -14,10 +14,17 @@
|
||||
//! - No slot leaks under high spawn/join churn
|
||||
//! - Panic on one scheduler thread doesn't kill others
|
||||
|
||||
use smarm::{channel, runtime::{Config, Runtime}, spawn, yield_now, JoinHandle};
|
||||
use std::sync::{atomic::{AtomicBool, AtomicU64, Ordering}, Arc};
|
||||
use std::time::Duration;
|
||||
use smarm::{
|
||||
channel,
|
||||
runtime::{Config, Runtime},
|
||||
spawn, yield_now, JoinHandle,
|
||||
};
|
||||
use std::collections::HashSet;
|
||||
use std::sync::{
|
||||
atomic::{AtomicBool, AtomicU64, Ordering},
|
||||
Arc,
|
||||
};
|
||||
use std::time::Duration;
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Helpers
|
||||
@@ -29,7 +36,9 @@ fn rt(n: usize) -> Runtime {
|
||||
}
|
||||
|
||||
/// Convenient single-threaded runtime (regression guard).
|
||||
fn rt1() -> Runtime { rt(1) }
|
||||
fn rt1() -> Runtime {
|
||||
rt(1)
|
||||
}
|
||||
|
||||
/// Multi-threaded runtime using all available parallelism.
|
||||
fn rt_par() -> Runtime {
|
||||
@@ -79,7 +88,9 @@ fn config_min_1_max_1_is_single_threaded() {
|
||||
fn runtime_run_executes_closure() {
|
||||
let flag = Arc::new(AtomicBool::new(false));
|
||||
let f = flag.clone();
|
||||
rt(1).run(move || { f.store(true, Ordering::SeqCst); });
|
||||
rt(1).run(move || {
|
||||
f.store(true, Ordering::SeqCst);
|
||||
});
|
||||
assert!(flag.load(Ordering::SeqCst));
|
||||
}
|
||||
|
||||
@@ -111,8 +122,12 @@ fn runtime_can_be_used_multiple_times_sequentially() {
|
||||
let b = Arc::new(AtomicU64::new(0));
|
||||
let ac = a.clone();
|
||||
let bc = b.clone();
|
||||
r.run(move || { ac.fetch_add(1, Ordering::SeqCst); });
|
||||
r.run(move || { bc.fetch_add(1, Ordering::SeqCst); });
|
||||
r.run(move || {
|
||||
ac.fetch_add(1, Ordering::SeqCst);
|
||||
});
|
||||
r.run(move || {
|
||||
bc.fetch_add(1, Ordering::SeqCst);
|
||||
});
|
||||
assert_eq!(a.load(Ordering::SeqCst), 1);
|
||||
assert_eq!(b.load(Ordering::SeqCst), 1);
|
||||
}
|
||||
@@ -126,7 +141,9 @@ fn exact_1_spawn_join_works() {
|
||||
let v = Arc::new(AtomicU64::new(0));
|
||||
let vc = v.clone();
|
||||
rt1().run(move || {
|
||||
let h = spawn(move || { vc.store(42, Ordering::SeqCst); });
|
||||
let h = spawn(move || {
|
||||
vc.store(42, Ordering::SeqCst);
|
||||
});
|
||||
h.join().unwrap();
|
||||
});
|
||||
assert_eq!(v.load(Ordering::SeqCst), 42);
|
||||
@@ -155,7 +172,9 @@ fn exact_1_panic_captured() {
|
||||
let s = saw_err.clone();
|
||||
rt1().run(move || {
|
||||
let h = spawn(|| panic!("oops"));
|
||||
if h.join().is_err() { s.store(true, Ordering::SeqCst); }
|
||||
if h.join().is_err() {
|
||||
s.store(true, Ordering::SeqCst);
|
||||
}
|
||||
});
|
||||
assert!(saw_err.load(Ordering::SeqCst));
|
||||
}
|
||||
@@ -176,7 +195,9 @@ fn multi_thread_all_actors_complete() {
|
||||
cc.fetch_add(1, Ordering::SeqCst);
|
||||
}));
|
||||
}
|
||||
for h in handles { h.join().unwrap(); }
|
||||
for h in handles {
|
||||
h.join().unwrap();
|
||||
}
|
||||
});
|
||||
assert_eq!(counter.load(Ordering::SeqCst), 100);
|
||||
}
|
||||
@@ -221,7 +242,9 @@ fn multi_thread_many_channels_no_lost_wakeups() {
|
||||
tx.send(1).unwrap();
|
||||
}));
|
||||
}
|
||||
for h in handles { h.join().unwrap(); }
|
||||
for h in handles {
|
||||
h.join().unwrap();
|
||||
}
|
||||
});
|
||||
assert_eq!(count.load(Ordering::SeqCst), PAIRS as u64);
|
||||
}
|
||||
@@ -247,7 +270,9 @@ fn multi_thread_mutex_contention_no_deadlock() {
|
||||
}
|
||||
}));
|
||||
}
|
||||
for h in handles { h.join().unwrap(); }
|
||||
for h in handles {
|
||||
h.join().unwrap();
|
||||
}
|
||||
let g = m.lock_timeout(Duration::from_secs(1)).unwrap();
|
||||
t.store(*g, Ordering::SeqCst);
|
||||
});
|
||||
@@ -262,7 +287,9 @@ fn multi_thread_join_across_threads() {
|
||||
rt_par().run(move || {
|
||||
let h = spawn(move || {
|
||||
// Do some work to make scheduling interesting.
|
||||
for _ in 0..10 { yield_now(); }
|
||||
for _ in 0..10 {
|
||||
yield_now();
|
||||
}
|
||||
vc.store(1, Ordering::SeqCst);
|
||||
});
|
||||
h.join().unwrap();
|
||||
@@ -279,8 +306,7 @@ fn multi_thread_join_across_threads() {
|
||||
|
||||
#[test]
|
||||
fn actors_run_on_multiple_os_threads() {
|
||||
let thread_ids: Arc<smarm::Mutex<HashSet<u64>>> =
|
||||
Arc::new(smarm::Mutex::new(HashSet::new()));
|
||||
let thread_ids: Arc<smarm::Mutex<HashSet<u64>>> = Arc::new(smarm::Mutex::new(HashSet::new()));
|
||||
|
||||
rt_par().run({
|
||||
let ids = thread_ids.clone();
|
||||
@@ -294,11 +320,15 @@ fn actors_run_on_multiple_os_threads() {
|
||||
g.insert(tid);
|
||||
}));
|
||||
}
|
||||
for h in handles { h.join().unwrap(); }
|
||||
for h in handles {
|
||||
h.join().unwrap();
|
||||
}
|
||||
}
|
||||
});
|
||||
|
||||
let n = std::thread::available_parallelism().map(|n| n.get()).unwrap_or(1);
|
||||
let n = std::thread::available_parallelism()
|
||||
.map(|n| n.get())
|
||||
.unwrap_or(1);
|
||||
|
||||
let ids = thread_ids.lock_timeout(Duration::from_secs(1)).unwrap();
|
||||
// If we have >1 scheduler threads, we expect >1 OS thread IDs.
|
||||
@@ -326,11 +356,17 @@ fn scheduler_stats_run_queue_len_is_observable() {
|
||||
// run() completes (queue len == 0 at quiescence).
|
||||
let r = rt_par();
|
||||
r.run(|| {
|
||||
for _ in 0..10 { spawn(|| {}); }
|
||||
for _ in 0..10 {
|
||||
spawn(|| {});
|
||||
}
|
||||
// Don't join — let them drain naturally.
|
||||
});
|
||||
let stats = r.stats();
|
||||
assert_eq!(stats.total_run_queue_len(), 0, "queue should be empty after run()");
|
||||
assert_eq!(
|
||||
stats.total_run_queue_len(),
|
||||
0,
|
||||
"queue should be empty after run()"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
@@ -359,7 +395,9 @@ fn panic_in_actor_does_not_kill_runtime() {
|
||||
}));
|
||||
}
|
||||
let _ = bad.join(); // expect Err
|
||||
for h in good_handles { h.join().unwrap(); }
|
||||
for h in good_handles {
|
||||
h.join().unwrap();
|
||||
}
|
||||
});
|
||||
assert_eq!(completed.load(Ordering::SeqCst), 10);
|
||||
}
|
||||
@@ -379,9 +417,11 @@ fn no_slot_leak_under_churn() {
|
||||
rt_par().run(move || {
|
||||
for _ in 0..500 {
|
||||
let cc = c.clone();
|
||||
spawn(move || { cc.fetch_add(1, Ordering::SeqCst); })
|
||||
.join()
|
||||
.unwrap();
|
||||
spawn(move || {
|
||||
cc.fetch_add(1, Ordering::SeqCst);
|
||||
})
|
||||
.join()
|
||||
.unwrap();
|
||||
}
|
||||
});
|
||||
assert_eq!(counter.load(Ordering::SeqCst), 500);
|
||||
@@ -474,7 +514,11 @@ fn multi_thread_timer_only_no_pipe_contention() {
|
||||
}
|
||||
});
|
||||
|
||||
assert_eq!(count.load(Ordering::SeqCst), ACTORS as u64, "not all actors completed");
|
||||
assert_eq!(
|
||||
count.load(Ordering::SeqCst),
|
||||
ACTORS as u64,
|
||||
"not all actors completed"
|
||||
);
|
||||
|
||||
let elapsed = start.elapsed();
|
||||
assert!(
|
||||
@@ -515,5 +559,46 @@ fn runtime_reusable_after_root_panic() {
|
||||
let ran = Arc::new(AtomicBool::new(false));
|
||||
let ran_t = ran.clone();
|
||||
r.run(move || ran_t.store(true, Ordering::Relaxed));
|
||||
assert!(ran.load(Ordering::Relaxed), "runtime unusable after root panic");
|
||||
assert!(
|
||||
ran.load(Ordering::Relaxed),
|
||||
"runtime unusable after root panic"
|
||||
);
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// RFC 019 — Config stack knobs
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/// Burn ~`frames` × 4 KiB of stack; probestack touches pages in order so
|
||||
/// exceeding the reserve would hit the guard and SIGSEGV the process.
|
||||
#[inline(never)]
|
||||
fn burn_stack(frames: usize) -> u64 {
|
||||
let mut local = [0u8; 4096];
|
||||
local[0] = frames as u8;
|
||||
let below = if frames == 0 {
|
||||
0
|
||||
} else {
|
||||
burn_stack(frames - 1)
|
||||
};
|
||||
std::hint::black_box(&mut local);
|
||||
below.wrapping_add(local[0] as u64)
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn config_stack_reserve_permits_deep_recursion() {
|
||||
// ~256 KiB of frames: four times the old fixed 64 KiB reserve. With
|
||||
// Config::stack_reserve raised this must complete; before RFC 019 it
|
||||
// could only segfault.
|
||||
let rt = smarm::runtime::init(Config::exact(1).stack_reserve(1024 * 1024));
|
||||
let done = Arc::new(AtomicBool::new(false));
|
||||
let done2 = done.clone();
|
||||
rt.run(move || {
|
||||
spawn(move || {
|
||||
std::hint::black_box(burn_stack(64));
|
||||
done2.store(true, Ordering::SeqCst);
|
||||
})
|
||||
.join()
|
||||
.unwrap();
|
||||
});
|
||||
assert!(done.load(Ordering::SeqCst));
|
||||
}
|
||||
|
||||
+7
-4
@@ -14,7 +14,9 @@ use std::sync::Arc;
|
||||
fn root_actor_runs() {
|
||||
let captured = Arc::new(AtomicI64::new(0));
|
||||
let c = captured.clone();
|
||||
run(move || { c.store(99, Ordering::SeqCst); });
|
||||
run(move || {
|
||||
c.store(99, Ordering::SeqCst);
|
||||
});
|
||||
assert_eq!(captured.load(Ordering::SeqCst), 99);
|
||||
}
|
||||
|
||||
@@ -27,7 +29,9 @@ fn spawn_and_join_returns_exit() {
|
||||
let captured = Arc::new(AtomicI64::new(0));
|
||||
let c = captured.clone();
|
||||
run(move || {
|
||||
let h = spawn(move || { c.store(7, Ordering::SeqCst); });
|
||||
let h = spawn(move || {
|
||||
c.store(7, Ordering::SeqCst);
|
||||
});
|
||||
let res = h.join();
|
||||
assert!(res.is_ok(), "join returned {:?}", res);
|
||||
});
|
||||
@@ -68,8 +72,7 @@ fn yield_now_interleaves_actors() {
|
||||
|
||||
#[test]
|
||||
fn self_pid_is_stable_within_an_actor() {
|
||||
let pid_cell: Arc<std::sync::Mutex<Option<smarm::Pid>>> =
|
||||
Arc::new(std::sync::Mutex::new(None));
|
||||
let pid_cell: Arc<std::sync::Mutex<Option<smarm::Pid>>> = Arc::new(std::sync::Mutex::new(None));
|
||||
let p2 = pid_cell.clone();
|
||||
run(move || {
|
||||
let h = spawn(move || {
|
||||
|
||||
+14
-3
@@ -19,7 +19,12 @@ fn ready_arm_returns_immediately_without_parking() {
|
||||
txa.send(42).unwrap();
|
||||
let i = select(&[&rxb, &rxa]);
|
||||
assert_eq!(i, 1);
|
||||
out2.store(rxa.try_recv().unwrap().expect("ready arm must hold a message"), Ordering::SeqCst);
|
||||
out2.store(
|
||||
rxa.try_recv()
|
||||
.unwrap()
|
||||
.expect("ready arm must hold a message"),
|
||||
Ordering::SeqCst,
|
||||
);
|
||||
});
|
||||
assert_eq!(out.load(Ordering::SeqCst), 42);
|
||||
}
|
||||
@@ -276,7 +281,10 @@ fn select_timeout_ready_arm_wins_without_arming_a_timer() {
|
||||
let (txa, rxa) = channel::<i64>();
|
||||
let (_keep_b, rxb) = channel::<i64>();
|
||||
txa.send(5).unwrap();
|
||||
assert_eq!(select_timeout(&[&rxb, &rxa], Duration::from_millis(500)), Some(1));
|
||||
assert_eq!(
|
||||
select_timeout(&[&rxb, &rxa], Duration::from_millis(500)),
|
||||
Some(1)
|
||||
);
|
||||
assert_eq!(rxa.try_recv().unwrap(), Some(5));
|
||||
});
|
||||
}
|
||||
@@ -342,7 +350,10 @@ fn select_timeout_closed_arm_is_ready_not_a_timeout() {
|
||||
let (_keep_a, rxa) = channel::<i64>();
|
||||
let (txb, rxb) = channel::<i64>();
|
||||
drop(txb);
|
||||
assert_eq!(select_timeout(&[&rxa, &rxb], Duration::from_millis(200)), Some(1));
|
||||
assert_eq!(
|
||||
select_timeout(&[&rxa, &rxb], Duration::from_millis(200)),
|
||||
Some(1)
|
||||
);
|
||||
assert!(rxb.try_recv().is_err());
|
||||
});
|
||||
}
|
||||
|
||||
@@ -0,0 +1,122 @@
|
||||
//! Graceful shutdown — `request_shutdown` (OTP `exit(Pid, shutdown)`).
|
||||
//!
|
||||
//! `request_stop` is `exit(Pid, kill)`: an uncatchable unwind at the target's
|
||||
//! next observation point. `request_shutdown` is the polite form:
|
||||
//! - a target that is NOT trapping exits is stopped exactly as by
|
||||
//! `request_stop` (OTP's rule: don't trap, you die);
|
||||
//! - a target that IS trapping receives an `ExitSignal { reason: Shutdown }`
|
||||
//! on its trap inbox and keeps running — it is expected to wind down and
|
||||
//! exit normally on its own.
|
||||
|
||||
use smarm::{monitor, request_shutdown, run, self_pid, sleep, spawn, trap_exit, DownReason, Pid};
|
||||
use std::sync::atomic::{AtomicBool, Ordering};
|
||||
use std::sync::{mpsc, Arc};
|
||||
use std::thread;
|
||||
use std::time::Duration;
|
||||
|
||||
const WATCHDOG: Duration = Duration::from_secs(10);
|
||||
|
||||
#[test]
|
||||
fn request_shutdown_stops_a_non_trapping_actor() {
|
||||
run(|| {
|
||||
let h = spawn(|| sleep(Duration::from_secs(3600)));
|
||||
let mon = monitor(h.pid());
|
||||
request_shutdown(h.pid());
|
||||
let down = mon.rx.recv().expect("down");
|
||||
assert_eq!(down.reason, DownReason::Stopped);
|
||||
});
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn request_shutdown_is_a_message_to_a_trapping_actor() {
|
||||
let unwound = Arc::new(AtomicBool::new(false));
|
||||
let u = unwound.clone();
|
||||
run(move || {
|
||||
struct Unwound(Arc<AtomicBool>);
|
||||
impl Drop for Unwound {
|
||||
fn drop(&mut self) {
|
||||
if std::thread::panicking() {
|
||||
self.0.store(true, Ordering::SeqCst);
|
||||
}
|
||||
}
|
||||
}
|
||||
let (tx, rx) = smarm::channel::<(Pid, DownReason)>();
|
||||
let (ready_tx, ready_rx) = smarm::channel::<()>();
|
||||
let h = spawn(move || {
|
||||
let _g = Unwound(u);
|
||||
let inbox = trap_exit();
|
||||
let _ = ready_tx.send(());
|
||||
let sig = inbox.recv().expect("exit signal");
|
||||
let _ = tx.send((sig.from, sig.reason));
|
||||
// Keep doing work after the request: shutdown is advisory.
|
||||
sleep(Duration::from_millis(20));
|
||||
});
|
||||
// Trapping is set by the target itself; a request that beats it is a
|
||||
// plain stop (same window as OTP's exit-before-process_flag).
|
||||
ready_rx.recv().expect("ready");
|
||||
let me = self_pid();
|
||||
let mon = monitor(h.pid());
|
||||
request_shutdown(h.pid());
|
||||
let (from, reason) = rx.recv().expect("relayed");
|
||||
assert_eq!(from, me);
|
||||
assert_eq!(reason, DownReason::Shutdown);
|
||||
let down = mon.rx.recv().expect("down");
|
||||
assert_eq!(
|
||||
down.reason,
|
||||
DownReason::Exit,
|
||||
"target exited normally, not stopped"
|
||||
);
|
||||
});
|
||||
assert!(
|
||||
!unwound.load(Ordering::SeqCst),
|
||||
"trapping target must not be unwound"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn request_shutdown_on_dead_pid_is_a_no_op() {
|
||||
run(|| {
|
||||
let h = spawn(|| {});
|
||||
let pid = h.pid();
|
||||
let _ = h.join();
|
||||
request_shutdown(pid); // must not panic
|
||||
});
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn handle_request_shutdown_from_foreign_thread() {
|
||||
let rt = smarm::init(smarm::Config::exact(2));
|
||||
let handle = rt.handle();
|
||||
|
||||
let (pid_tx, pid_rx) = mpsc::channel::<Pid>();
|
||||
let requester = thread::spawn(move || {
|
||||
let pid = pid_rx.recv().expect("pid");
|
||||
thread::sleep(Duration::from_millis(50));
|
||||
handle.request_shutdown(pid);
|
||||
});
|
||||
|
||||
let (done_tx, done_rx) = mpsc::channel();
|
||||
thread::spawn(move || {
|
||||
rt.run(move || {
|
||||
let (tx, rx) = smarm::channel::<DownReason>();
|
||||
let (ready_tx, ready_rx) = smarm::channel::<()>();
|
||||
let h = spawn(move || {
|
||||
let inbox = trap_exit();
|
||||
let _ = ready_tx.send(());
|
||||
let sig = inbox.recv().expect("exit signal");
|
||||
let _ = tx.send(sig.reason);
|
||||
});
|
||||
ready_rx.recv().expect("ready");
|
||||
pid_tx.send(h.pid()).expect("send pid");
|
||||
let reason = rx.recv().expect("relayed");
|
||||
assert_eq!(reason, DownReason::Shutdown);
|
||||
let _ = h.join();
|
||||
});
|
||||
let _ = done_tx.send(());
|
||||
});
|
||||
|
||||
done_rx
|
||||
.recv_timeout(WATCHDOG)
|
||||
.expect("run did not return: foreign-thread request_shutdown never reached the target");
|
||||
requester.join().expect("requester thread");
|
||||
}
|
||||
@@ -0,0 +1,235 @@
|
||||
//! RFC 019 commit 2 — the `SpawnOpts` surface.
|
||||
//!
|
||||
//! Covers: per-spawn stack shape overrides on every spawn surface, the
|
||||
//! `None ⇒ Config default` resolution, the pool rule from the outside
|
||||
//! (obligation 4: a custom-shaped stack never enters the pool), and that a
|
||||
//! big reserve behaviorally takes effect (deep recursion completes).
|
||||
|
||||
use smarm::runtime::{Config, DEFAULT_STACK_GUARD, DEFAULT_STACK_RESERVE};
|
||||
use smarm::{self_pid, spawn, spawn_under_with, spawn_with, GenServerBuilder, SpawnOpts};
|
||||
use std::sync::atomic::{AtomicBool, Ordering};
|
||||
use std::sync::Arc;
|
||||
|
||||
fn rt1() -> smarm::runtime::Runtime {
|
||||
smarm::runtime::init(Config::exact(1))
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn default_spawn_has_default_shape() {
|
||||
rt1().run(|| {
|
||||
let h = spawn(|| {
|
||||
let shape = smarm::introspect::stack_shape(self_pid()).unwrap();
|
||||
assert_eq!(shape, (DEFAULT_STACK_RESERVE, DEFAULT_STACK_GUARD));
|
||||
});
|
||||
h.join().unwrap();
|
||||
});
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn spawn_with_overrides_reserve_and_guard() {
|
||||
rt1().run(|| {
|
||||
let opts = SpawnOpts {
|
||||
stack_reserve: Some(1024 * 1024),
|
||||
guard_size: Some(256 * 1024),
|
||||
};
|
||||
let h = spawn_with(opts, || {
|
||||
let shape = smarm::introspect::stack_shape(self_pid()).unwrap();
|
||||
assert_eq!(shape, (1024 * 1024, 256 * 1024));
|
||||
});
|
||||
h.join().unwrap();
|
||||
});
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn spawn_with_partial_override_keeps_config_default_for_the_rest() {
|
||||
rt1().run(|| {
|
||||
let opts = SpawnOpts {
|
||||
stack_reserve: Some(1024 * 1024),
|
||||
..SpawnOpts::default()
|
||||
};
|
||||
let h = spawn_with(opts, || {
|
||||
let shape = smarm::introspect::stack_shape(self_pid()).unwrap();
|
||||
assert_eq!(shape, (1024 * 1024, DEFAULT_STACK_GUARD));
|
||||
});
|
||||
h.join().unwrap();
|
||||
});
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn spawn_with_rounds_to_pages() {
|
||||
rt1().run(|| {
|
||||
let opts = SpawnOpts {
|
||||
stack_reserve: Some(64 * 1024 + 1),
|
||||
guard_size: Some(4097),
|
||||
};
|
||||
let h = spawn_with(opts, || {
|
||||
let (reserve, guard) = smarm::introspect::stack_shape(self_pid()).unwrap();
|
||||
assert_eq!(reserve % 4096, 0);
|
||||
assert_eq!(guard % 4096, 0);
|
||||
assert!(reserve >= 64 * 1024 + 1);
|
||||
assert!(guard >= 4097);
|
||||
});
|
||||
h.join().unwrap();
|
||||
});
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn spawn_under_with_takes_opts() {
|
||||
rt1().run(|| {
|
||||
let me = self_pid();
|
||||
let opts = SpawnOpts {
|
||||
stack_reserve: Some(128 * 1024),
|
||||
..SpawnOpts::default()
|
||||
};
|
||||
let h = spawn_under_with(me, opts, || {
|
||||
let (reserve, _) = smarm::introspect::stack_shape(self_pid()).unwrap();
|
||||
assert_eq!(reserve, 128 * 1024);
|
||||
});
|
||||
h.join().unwrap();
|
||||
});
|
||||
}
|
||||
|
||||
/// Obligation 4, from the outside: a dead custom stack must not be handed to
|
||||
/// the next default spawn. The pool is LIFO, so if the custom stack had been
|
||||
/// (wrongly) pushed at death, the very next default-shaped spawn on this
|
||||
/// single-threaded runtime would pop it and report a custom shape.
|
||||
#[test]
|
||||
fn custom_stack_never_enters_the_pool() {
|
||||
rt1().run(|| {
|
||||
spawn_with(
|
||||
SpawnOpts {
|
||||
stack_reserve: Some(512 * 1024),
|
||||
guard_size: Some(128 * 1024),
|
||||
},
|
||||
|| {},
|
||||
)
|
||||
.join()
|
||||
.unwrap();
|
||||
let h = spawn(|| {
|
||||
let shape = smarm::introspect::stack_shape(self_pid()).unwrap();
|
||||
assert_eq!(shape, (DEFAULT_STACK_RESERVE, DEFAULT_STACK_GUARD));
|
||||
});
|
||||
h.join().unwrap();
|
||||
});
|
||||
}
|
||||
|
||||
/// The reverse direction of the pool rule: a default-shaped stack IS pooled
|
||||
/// and reused (cap = threads × 4 ≥ 1 here, pool empty at start).
|
||||
#[test]
|
||||
fn default_stack_is_recycled() {
|
||||
rt1().run(|| {
|
||||
spawn(|| {}).join().unwrap();
|
||||
let h = spawn(|| {
|
||||
let shape = smarm::introspect::stack_shape(self_pid()).unwrap();
|
||||
assert_eq!(shape, (DEFAULT_STACK_RESERVE, DEFAULT_STACK_GUARD));
|
||||
});
|
||||
h.join().unwrap();
|
||||
});
|
||||
}
|
||||
|
||||
/// Burn ~`frames` × 4 KiB of stack (see tests/runtime.rs twin).
|
||||
#[inline(never)]
|
||||
fn burn_stack(frames: usize) -> u64 {
|
||||
let mut local = [0u8; 4096];
|
||||
local[0] = frames as u8;
|
||||
let below = if frames == 0 {
|
||||
0
|
||||
} else {
|
||||
burn_stack(frames - 1)
|
||||
};
|
||||
std::hint::black_box(&mut local);
|
||||
below.wrapping_add(local[0] as u64)
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn big_reserve_behaviorally_takes_effect() {
|
||||
// ~1 MiB deep on an 8 MiB per-spawn reserve, runtime default untouched.
|
||||
rt1().run(|| {
|
||||
let done = Arc::new(AtomicBool::new(false));
|
||||
let done2 = done.clone();
|
||||
spawn_with(
|
||||
SpawnOpts {
|
||||
stack_reserve: Some(8 * 1024 * 1024),
|
||||
..SpawnOpts::default()
|
||||
},
|
||||
move || {
|
||||
std::hint::black_box(burn_stack(256));
|
||||
done2.store(true, Ordering::SeqCst);
|
||||
},
|
||||
)
|
||||
.join()
|
||||
.unwrap();
|
||||
assert!(done.load(Ordering::SeqCst));
|
||||
});
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Builder surfaces
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
struct Echo;
|
||||
impl smarm::GenServer for Echo {
|
||||
type Call = ();
|
||||
type Reply = (usize, usize);
|
||||
type Cast = ();
|
||||
type Info = ();
|
||||
type Timer = ();
|
||||
fn handle_call(&mut self, _c: ()) -> (usize, usize) {
|
||||
smarm::introspect::stack_shape(self_pid()).unwrap()
|
||||
}
|
||||
fn handle_cast(&mut self, _c: ()) {}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn gen_server_builder_stack_opts() {
|
||||
rt1().run(|| {
|
||||
let server = GenServerBuilder::new(Echo)
|
||||
.stack_opts(SpawnOpts {
|
||||
stack_reserve: Some(256 * 1024),
|
||||
..SpawnOpts::default()
|
||||
})
|
||||
.start();
|
||||
let (reserve, guard) = server.call(()).unwrap();
|
||||
assert_eq!(reserve, 256 * 1024);
|
||||
assert_eq!(guard, DEFAULT_STACK_GUARD);
|
||||
server.shutdown();
|
||||
});
|
||||
}
|
||||
|
||||
struct Probe;
|
||||
impl smarm::Machine for Probe {
|
||||
type Ev = smarm::channel::Sender<(usize, usize)>;
|
||||
fn state_timeout_ev() -> Self::Ev {
|
||||
unreachable!("no timers in this test")
|
||||
}
|
||||
fn timeout_ev(_name: &'static str) -> Self::Ev {
|
||||
unreachable!("no timers in this test")
|
||||
}
|
||||
fn on_start(&mut self, _cx: &mut smarm::Cx<Self::Ev>) {}
|
||||
fn handle(
|
||||
&mut self,
|
||||
ev: Self::Ev,
|
||||
_cx: &mut smarm::Cx<Self::Ev>,
|
||||
) -> smarm::gen_statem::Step<Self::Ev> {
|
||||
let _ = ev.send(smarm::introspect::stack_shape(self_pid()).unwrap());
|
||||
smarm::gen_statem::Step::Stayed
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn gen_statem_spawn_with_stack_opts() {
|
||||
rt1().run(|| {
|
||||
let m = smarm::gen_statem::spawn_with(
|
||||
SpawnOpts {
|
||||
stack_reserve: Some(256 * 1024),
|
||||
..SpawnOpts::default()
|
||||
},
|
||||
Probe,
|
||||
);
|
||||
let (tx, rx) = smarm::channel::channel();
|
||||
m.send(tx).unwrap();
|
||||
let (reserve, guard) = rx.recv().unwrap();
|
||||
assert_eq!(reserve, 256 * 1024);
|
||||
assert_eq!(guard, DEFAULT_STACK_GUARD);
|
||||
});
|
||||
}
|
||||
+103
-11
@@ -7,13 +7,13 @@ use smarm::stack::Stack;
|
||||
|
||||
#[test]
|
||||
fn top_is_16_byte_aligned() {
|
||||
let s = Stack::new(64 * 1024).unwrap();
|
||||
let s = Stack::new(64 * 1024, 4096).unwrap();
|
||||
assert_eq!(s.top() as usize % 16, 0);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn top_is_within_allocation() {
|
||||
let s = Stack::new(64 * 1024).unwrap();
|
||||
let s = Stack::new(64 * 1024, 4096).unwrap();
|
||||
let top = s.top() as usize;
|
||||
let base = s.usable_base() as usize;
|
||||
assert!(top > base);
|
||||
@@ -22,7 +22,7 @@ fn top_is_within_allocation() {
|
||||
|
||||
#[test]
|
||||
fn write_and_read_top_of_stack() {
|
||||
let s = Stack::new(64 * 1024).unwrap();
|
||||
let s = Stack::new(64 * 1024, 4096).unwrap();
|
||||
let sentinel: u64 = 0xDEAD_BEEF_CAFE_1234;
|
||||
unsafe {
|
||||
let ptr = s.top().sub(8) as *mut u64;
|
||||
@@ -33,7 +33,7 @@ fn write_and_read_top_of_stack() {
|
||||
|
||||
#[test]
|
||||
fn write_and_read_bottom_of_usable_region() {
|
||||
let s = Stack::new(64 * 1024).unwrap();
|
||||
let s = Stack::new(64 * 1024, 4096).unwrap();
|
||||
let sentinel: u64 = 0x0102_0304_0506_0708;
|
||||
unsafe {
|
||||
let ptr = s.usable_base() as *mut u64;
|
||||
@@ -44,17 +44,17 @@ fn write_and_read_bottom_of_usable_region() {
|
||||
|
||||
#[test]
|
||||
fn small_stack_allocates() {
|
||||
assert!(Stack::new(4096).is_ok());
|
||||
assert!(Stack::new(4096, 4096).is_ok());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn large_stack_allocates() {
|
||||
assert!(Stack::new(8 * 1024 * 1024).is_ok());
|
||||
assert!(Stack::new(8 * 1024 * 1024, 4096).is_ok());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn stack_size_at_least_requested() {
|
||||
let s = Stack::new(64 * 1024).unwrap();
|
||||
let s = Stack::new(64 * 1024, 4096).unwrap();
|
||||
assert!(s.stack_size() >= 64 * 1024);
|
||||
}
|
||||
|
||||
@@ -68,15 +68,32 @@ use std::process::Command;
|
||||
fn run_as_child_if_requested() {
|
||||
match env::var("SMARM_SUBTEST").as_deref() {
|
||||
Ok("guard_page_direct") => {
|
||||
let s = Stack::new(64 * 1024).unwrap();
|
||||
let s = Stack::new(64 * 1024, 4096).unwrap();
|
||||
unsafe {
|
||||
let guard_ptr = s.usable_base().sub(1);
|
||||
guard_ptr.write_volatile(0xAB);
|
||||
}
|
||||
std::process::exit(0);
|
||||
}
|
||||
Ok("wide_guard_top") => {
|
||||
// One byte below the usable region, 64 KiB guard: must fault.
|
||||
let s = Stack::new(64 * 1024, 64 * 1024).unwrap();
|
||||
unsafe {
|
||||
s.usable_base().sub(1).write_volatile(0xAB);
|
||||
}
|
||||
std::process::exit(0);
|
||||
}
|
||||
Ok("wide_guard_bottom") => {
|
||||
// The very bottom page of a 64 KiB guard: an unprobed C-style
|
||||
// leap over a small guard lands here — must still fault.
|
||||
let s = Stack::new(64 * 1024, 64 * 1024).unwrap();
|
||||
unsafe {
|
||||
s.usable_base().sub(64 * 1024).write_volatile(0xAB);
|
||||
}
|
||||
std::process::exit(0);
|
||||
}
|
||||
Ok("stack_overflow") => {
|
||||
let s = Stack::new(64 * 1024).unwrap();
|
||||
let s = Stack::new(64 * 1024, 4096).unwrap();
|
||||
unsafe {
|
||||
let mut ptr = s.top().sub(1);
|
||||
let stop = s.usable_base().sub(1);
|
||||
@@ -107,7 +124,12 @@ fn guard_page_causes_sigsegv() {
|
||||
#[cfg(unix)]
|
||||
{
|
||||
use std::os::unix::process::ExitStatusExt;
|
||||
assert_eq!(status.signal(), Some(11), "expected SIGSEGV, got: {:?}", status);
|
||||
assert_eq!(
|
||||
status.signal(),
|
||||
Some(11),
|
||||
"expected SIGSEGV, got: {:?}",
|
||||
status
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -118,6 +140,76 @@ fn stack_overflow_causes_sigsegv() {
|
||||
#[cfg(unix)]
|
||||
{
|
||||
use std::os::unix::process::ExitStatusExt;
|
||||
assert_eq!(status.signal(), Some(11), "expected SIGSEGV, got: {:?}", status);
|
||||
assert_eq!(
|
||||
status.signal(),
|
||||
Some(11),
|
||||
"expected SIGSEGV, got: {:?}",
|
||||
status
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// RFC 019 — explicit shape: rounding, guard accessor, wide-guard coverage.
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
#[test]
|
||||
fn sizes_round_up_to_page() {
|
||||
let s = Stack::new(64 * 1024 + 1, 4096 + 1).unwrap();
|
||||
assert_eq!(s.stack_size() % 4096, 0);
|
||||
assert_eq!(s.guard_size() % 4096, 0);
|
||||
assert!(s.stack_size() >= 64 * 1024 + 1);
|
||||
assert!(s.guard_size() >= 4096 + 1);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn shape_reports_rounded_sizes() {
|
||||
let s = Stack::new(64 * 1024, 64 * 1024).unwrap();
|
||||
assert_eq!(s.shape(), (64 * 1024, 64 * 1024));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn usable_base_sits_above_guard() {
|
||||
let s = Stack::new(64 * 1024, 64 * 1024).unwrap();
|
||||
// The usable region must start exactly guard_size above the mapping
|
||||
// base: a write at usable_base is legal, one byte below is not (the
|
||||
// subprocess tests below prove the "not").
|
||||
let sentinel: u64 = 0x1111_2222_3333_4444;
|
||||
unsafe {
|
||||
let ptr = s.usable_base() as *mut u64;
|
||||
ptr.write_volatile(sentinel);
|
||||
assert_eq!(ptr.read_volatile(), sentinel);
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn wide_guard_faults_at_top() {
|
||||
run_as_child_if_requested();
|
||||
let status = spawn_subtest("wide_guard_top");
|
||||
#[cfg(unix)]
|
||||
{
|
||||
use std::os::unix::process::ExitStatusExt;
|
||||
assert_eq!(
|
||||
status.signal(),
|
||||
Some(11),
|
||||
"expected SIGSEGV, got: {:?}",
|
||||
status
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn wide_guard_faults_at_bottom() {
|
||||
run_as_child_if_requested();
|
||||
let status = spawn_subtest("wide_guard_bottom");
|
||||
#[cfg(unix)]
|
||||
{
|
||||
use std::os::unix::process::ExitStatusExt;
|
||||
assert_eq!(
|
||||
status.signal(),
|
||||
Some(11),
|
||||
"expected SIGSEGV, got: {:?}",
|
||||
status
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,153 @@
|
||||
//! RFC 019 §7 — overflow diagnostics, observed from outside via subprocess
|
||||
//! (mirrors tests/stack.rs's harness, plus stderr capture).
|
||||
//!
|
||||
//! Four cases:
|
||||
//! - Rust recursion at defaults: probed frames walk into the guard →
|
||||
//! tier-1 definitive message, death by SIGSEGV.
|
||||
//! - FFI canary (96 KiB unprobed C local) at defaults: first touch lands
|
||||
//! inside the 1 MiB guard → tier-1 message.
|
||||
//! - FFI canary with the guard shrunk to 4 KiB: the frame steps over it
|
||||
//! into unmapped VA below → tier-2 "stepped over" message. This is the
|
||||
//! RFC's motivating incident (cargo-vendored gz build) reproduced.
|
||||
//! - FFI canary with reserve raised to 256 KiB: fits, runs clean, exits 0 —
|
||||
//! the §1 knob is the fix, proven by the same frame.
|
||||
|
||||
use std::env;
|
||||
use std::process::Command;
|
||||
|
||||
unsafe extern "C" {
|
||||
fn smarm_canary_burn();
|
||||
}
|
||||
|
||||
/// Unbounded probed recursion; each frame dirties 4 KiB. black_box defeats
|
||||
/// tail-call elision so the walk is real.
|
||||
#[inline(never)]
|
||||
#[allow(unconditional_recursion)]
|
||||
fn recurse_forever(depth: u64) -> u64 {
|
||||
let mut local = [0u8; 4096];
|
||||
local[0] = depth as u8;
|
||||
std::hint::black_box(&mut local);
|
||||
recurse_forever(depth + 1).wrapping_add(local[0] as u64)
|
||||
}
|
||||
|
||||
fn run_as_child_if_requested() {
|
||||
let mode = match env::var("SMARM_DIAG_SUBTEST") {
|
||||
Ok(m) => m,
|
||||
Err(_) => return,
|
||||
};
|
||||
use smarm::runtime::Config;
|
||||
use smarm::{spawn_with, SpawnOpts};
|
||||
let rt = smarm::runtime::init(Config::exact(1));
|
||||
rt.run(move || {
|
||||
let opts = match mode.as_str() {
|
||||
"rust_overflow" | "ffi_tier1" => SpawnOpts::default(),
|
||||
// Small guard: the canary's 96 KiB displacement clears it.
|
||||
"ffi_tier2" => SpawnOpts {
|
||||
guard_size: Some(4096),
|
||||
..SpawnOpts::default()
|
||||
},
|
||||
// Enough reserve: the same frame simply fits.
|
||||
"ffi_clean" => SpawnOpts {
|
||||
stack_reserve: Some(256 * 1024),
|
||||
..SpawnOpts::default()
|
||||
},
|
||||
other => panic!("unknown subtest {other}"),
|
||||
};
|
||||
let is_rust = mode == "rust_overflow";
|
||||
spawn_with(opts, move || {
|
||||
if is_rust {
|
||||
std::hint::black_box(recurse_forever(0));
|
||||
} else {
|
||||
unsafe { smarm_canary_burn() };
|
||||
}
|
||||
})
|
||||
.join()
|
||||
.unwrap();
|
||||
});
|
||||
std::process::exit(0);
|
||||
}
|
||||
|
||||
fn spawn_subtest(name: &str) -> std::process::Output {
|
||||
let exe = env::current_exe().unwrap();
|
||||
Command::new(exe)
|
||||
.env("SMARM_DIAG_SUBTEST", name)
|
||||
.args(["--test-threads=1", "--quiet"])
|
||||
.output()
|
||||
.expect("failed to spawn subprocess")
|
||||
}
|
||||
|
||||
#[cfg(unix)]
|
||||
fn assert_died_sigsegv(out: &std::process::Output) {
|
||||
use std::os::unix::process::ExitStatusExt;
|
||||
assert_eq!(
|
||||
out.status.signal(),
|
||||
Some(11),
|
||||
"expected death by SIGSEGV, got {:?}; stderr:\n{}",
|
||||
out.status,
|
||||
String::from_utf8_lossy(&out.stderr)
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn rust_overflow_dies_with_tier1_message() {
|
||||
run_as_child_if_requested();
|
||||
let out = spawn_subtest("rust_overflow");
|
||||
assert_died_sigsegv(&out);
|
||||
let err = String::from_utf8_lossy(&out.stderr);
|
||||
assert!(
|
||||
err.contains("overflowed its stack") && err.contains("in the guard region"),
|
||||
"missing tier-1 diagnostic; stderr:\n{err}"
|
||||
);
|
||||
assert!(
|
||||
err.contains("reserve=65536"),
|
||||
"wrong reserve in message:\n{err}"
|
||||
);
|
||||
assert!(
|
||||
err.contains("guard=1048576"),
|
||||
"wrong guard in message:\n{err}"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn ffi_canary_at_defaults_dies_with_tier1_message() {
|
||||
run_as_child_if_requested();
|
||||
let out = spawn_subtest("ffi_tier1");
|
||||
assert_died_sigsegv(&out);
|
||||
let err = String::from_utf8_lossy(&out.stderr);
|
||||
// 96 KiB displacement from a 64 KiB reserve lands ~32 KiB into the
|
||||
// 1 MiB guard: definitively classified.
|
||||
assert!(
|
||||
err.contains("in the guard region"),
|
||||
"wide guard should catch the unprobed frame in tier 1; stderr:\n{err}"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn ffi_canary_over_small_guard_dies_with_tier2_message() {
|
||||
run_as_child_if_requested();
|
||||
let out = spawn_subtest("ffi_tier2");
|
||||
assert_died_sigsegv(&out);
|
||||
let err = String::from_utf8_lossy(&out.stderr);
|
||||
assert!(
|
||||
err.contains("stepped over it") && err.contains("below the guard"),
|
||||
"expected tier-2 overshoot attribution; stderr:\n{err}"
|
||||
);
|
||||
assert!(err.contains("guard=4096"), "wrong guard in message:\n{err}");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn ffi_canary_with_enough_reserve_runs_clean() {
|
||||
run_as_child_if_requested();
|
||||
let out = spawn_subtest("ffi_clean");
|
||||
assert!(
|
||||
out.status.success(),
|
||||
"canary should fit in 256 KiB reserve, got {:?}; stderr:\n{}",
|
||||
out.status,
|
||||
String::from_utf8_lossy(&out.stderr)
|
||||
);
|
||||
let err = String::from_utf8_lossy(&out.stderr);
|
||||
assert!(
|
||||
!err.contains("smarm: actor"),
|
||||
"no diagnostic expected on the clean path; stderr:\n{err}"
|
||||
);
|
||||
}
|
||||
@@ -0,0 +1,133 @@
|
||||
//! RFC 019 commit 5 — pool recycle zaps a dead stack down to its retained
|
||||
//! entry end, observed from the outside.
|
||||
//!
|
||||
//! A default-shaped stack that spiked deep and then died must not carry its
|
||||
//! spike into the pool as resident RSS: `recycle_stack` DONTNEEDs everything
|
||||
//! below the top `RECYCLE_RETAIN` bytes before pushing. The zap is
|
||||
//! synchronous on the death path, so the drop is immediate — but the death
|
||||
//! path itself races the observer's `join` return, hence the brief poll.
|
||||
//!
|
||||
//! Residency is measured with `mincore`, not smaps: a neighboring rw anon
|
||||
//! mapping can land flush against the stack top and the kernel merges the
|
||||
//! VMAs (observed under the full test run), so per-mapping smaps fields
|
||||
//! over-count. The PROT_NONE guard below can never merge, so the usable
|
||||
//! base is exactly the anchor VMA's start, and `mincore` counts pages
|
||||
//! within [usable_base, usable_base + reserve) regardless of merging.
|
||||
|
||||
use smarm::runtime::{Config, RECYCLE_RETAIN};
|
||||
use smarm::{channel, spawn, yield_now};
|
||||
|
||||
const RESERVE: usize = 4 * 1024 * 1024;
|
||||
|
||||
/// Burn ~`frames` × 4 KiB of stack, dirtying every frame.
|
||||
#[inline(never)]
|
||||
fn burn_stack(frames: usize) -> u64 {
|
||||
let mut local = [0u8; 4096];
|
||||
local[0] = frames as u8;
|
||||
let below = if frames == 0 {
|
||||
0
|
||||
} else {
|
||||
burn_stack(frames - 1)
|
||||
};
|
||||
std::hint::black_box(&mut local);
|
||||
below.wrapping_add(local[0] as u64)
|
||||
}
|
||||
|
||||
/// Resident-page count over [lo, lo + len) via mincore (len page-aligned).
|
||||
fn resident_pages(lo: usize, len: usize) -> usize {
|
||||
let page = 4096;
|
||||
let mut vec = vec![0u8; len / page];
|
||||
let ret = unsafe { libc::mincore(lo as *mut libc::c_void, len, vec.as_mut_ptr()) };
|
||||
assert_eq!(
|
||||
ret,
|
||||
0,
|
||||
"mincore failed: {}",
|
||||
std::io::Error::last_os_error()
|
||||
);
|
||||
vec.iter().filter(|&&b| b & 1 != 0).count()
|
||||
}
|
||||
|
||||
/// The [start, end) of the VMA containing `addr`.
|
||||
fn vma_containing(addr: usize) -> (usize, usize) {
|
||||
let maps = std::fs::read_to_string("/proc/self/maps").unwrap();
|
||||
for line in maps.lines() {
|
||||
if let Some((range, _)) = line.split_once(' ') {
|
||||
if let Some((a, b)) = range.split_once('-') {
|
||||
if let (Ok(start), Ok(end)) =
|
||||
(usize::from_str_radix(a, 16), usize::from_str_radix(b, 16))
|
||||
{
|
||||
if start <= addr && addr < end {
|
||||
return (start, end);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
panic!("no VMA contains {addr:#x}");
|
||||
}
|
||||
|
||||
fn vma_exists(addr: usize) -> bool {
|
||||
let maps = std::fs::read_to_string("/proc/self/maps").unwrap();
|
||||
for line in maps.lines() {
|
||||
if let Some((range, _)) = line.split_once(' ') {
|
||||
if let Some((a, b)) = range.split_once('-') {
|
||||
if let (Ok(start), Ok(end)) =
|
||||
(usize::from_str_radix(a, 16), usize::from_str_radix(b, 16))
|
||||
{
|
||||
if start <= addr && addr < end {
|
||||
return true;
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
false
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn recycle_zaps_dead_stack_down_to_retain() {
|
||||
// Default reserve raised so the pool holds big stacks (default-shaped ⇒
|
||||
// pooled) and the zap has something to bite; single scheduler.
|
||||
let rt = smarm::runtime::init(Config::exact(1).stack_reserve(RESERVE));
|
||||
rt.run(|| {
|
||||
let (tx, rx) = channel::<usize>();
|
||||
|
||||
let h = spawn(move || {
|
||||
let probe = 0u8;
|
||||
let anchor = &probe as *const u8 as usize;
|
||||
// The guard below is PROT_NONE and can never merge with the
|
||||
// usable region, so the anchor VMA's start IS the usable base.
|
||||
let (vlo, _) = vma_containing(anchor);
|
||||
// Dirty ~3 MiB of the 4 MiB reserve, then die.
|
||||
std::hint::black_box(burn_stack(768));
|
||||
tx.send(vlo).unwrap();
|
||||
});
|
||||
|
||||
let usable_base = rx.recv().unwrap();
|
||||
h.join().unwrap();
|
||||
|
||||
// The zap span is everything below the retained entry end. DONTNEED
|
||||
// on private anon discards synchronously and unconditionally, so
|
||||
// this must go to exactly zero resident pages; the poll only covers
|
||||
// the death path racing join's return.
|
||||
let zap_len = RESERVE - RECYCLE_RETAIN;
|
||||
let mut resident = usize::MAX;
|
||||
for _ in 0..10_000 {
|
||||
resident = resident_pages(usable_base, zap_len);
|
||||
if resident == 0 {
|
||||
break;
|
||||
}
|
||||
yield_now();
|
||||
}
|
||||
assert_eq!(
|
||||
resident, 0,
|
||||
"recycled stack's zap span still resident: {resident} pages in \
|
||||
[{usable_base:#x}, +{zap_len:#x})"
|
||||
);
|
||||
// Pooled, not munmapped: the mapping must still be there.
|
||||
assert!(
|
||||
vma_exists(usable_base),
|
||||
"default-shaped stack was unmapped instead of pooled"
|
||||
);
|
||||
});
|
||||
}
|
||||
@@ -0,0 +1,161 @@
|
||||
//! RFC 019 commit 3 — park-path stack shrink, observed from the outside.
|
||||
//!
|
||||
//! The one integration-level claim of the shrink machinery: an actor that
|
||||
//! spikes deep, returns shallow, and then parks past the cooldown gets its
|
||||
//! dead span MADV_FREE'd — visible as `LazyFree` in `/proc/self/smaps`
|
||||
//! within the stack's address range — while everything live survives.
|
||||
//!
|
||||
//! The high-water mark is *sampled* at context-save, so the spike yields
|
||||
//! once at max depth to guarantee a sample there (in production, preemption
|
||||
//! provides the quasi-random samples; a test must not rely on luck).
|
||||
|
||||
use smarm::runtime::{Config, SHRINK_COOLDOWN, SHRINK_THRESHOLD};
|
||||
use smarm::{actor_info, channel, spawn, spawn_with, yield_now, ActorState, SpawnOpts};
|
||||
|
||||
/// Burn ~`frames` × 4 KiB of stack, yielding once at the bottom so the
|
||||
/// context-save samples `sp` at max depth.
|
||||
#[inline(never)]
|
||||
fn burn_stack_yielding(frames: usize) -> u64 {
|
||||
let mut local = [0u8; 4096];
|
||||
local[0] = frames as u8;
|
||||
let below = if frames == 0 {
|
||||
yield_now();
|
||||
0
|
||||
} else {
|
||||
burn_stack_yielding(frames - 1)
|
||||
};
|
||||
std::hint::black_box(&mut local);
|
||||
below.wrapping_add(local[0] as u64)
|
||||
}
|
||||
|
||||
/// Sum the `LazyFree:` kB of every smaps mapping intersecting [lo, hi).
|
||||
fn lazy_free_bytes_in(lo: usize, hi: usize) -> usize {
|
||||
let smaps = std::fs::read_to_string("/proc/self/smaps").unwrap();
|
||||
let mut total_kb = 0usize;
|
||||
let mut in_range = false;
|
||||
for line in smaps.lines() {
|
||||
if let Some((range, _)) = line.split_once(' ') {
|
||||
if let Some((a, b)) = range.split_once('-') {
|
||||
if let (Ok(start), Ok(end)) =
|
||||
(usize::from_str_radix(a, 16), usize::from_str_radix(b, 16))
|
||||
{
|
||||
in_range = start < hi && end > lo;
|
||||
continue;
|
||||
}
|
||||
}
|
||||
}
|
||||
if in_range {
|
||||
if let Some(rest) = line.strip_prefix("LazyFree:") {
|
||||
let kb: usize = rest.trim().trim_end_matches(" kB").trim().parse().unwrap();
|
||||
total_kb += kb;
|
||||
}
|
||||
}
|
||||
}
|
||||
total_kb * 1024
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn spike_then_parks_marks_lazyfree_and_keeps_live_data() {
|
||||
// Single scheduler: the controller can gate on the worker being Parked.
|
||||
let rt = smarm::runtime::init(Config::exact(1));
|
||||
rt.run(|| {
|
||||
let (park_tx, park_rx) = channel::<()>();
|
||||
let (done_tx, done_rx) = channel::<(usize, u64)>();
|
||||
|
||||
let spike = 768 * 4096; // ~3 MiB, well past SHRINK_THRESHOLD
|
||||
assert!(spike > SHRINK_THRESHOLD);
|
||||
|
||||
let worker = spawn_with(
|
||||
SpawnOpts {
|
||||
stack_reserve: Some(8 * 1024 * 1024),
|
||||
..SpawnOpts::default()
|
||||
},
|
||||
move || {
|
||||
// Live data that must survive the shrink, and an anchor
|
||||
// address inside the stack for the smaps scan.
|
||||
let live = [0xA5u8; 64];
|
||||
let anchor = live.as_ptr() as usize;
|
||||
|
||||
// Spike: ~3 MiB deep, sampled at the bottom, unwound.
|
||||
std::hint::black_box(burn_stack_yielding(768));
|
||||
|
||||
// Park past the cooldown. Each recv on the drained inbox is
|
||||
// one park; the controller sends only when it sees us Parked.
|
||||
for _ in 0..(SHRINK_COOLDOWN + 8) {
|
||||
park_rx.recv().unwrap();
|
||||
}
|
||||
|
||||
// Measure from inside: the stack spans ≤ 8 MiB below anchor.
|
||||
let lazy = lazy_free_bytes_in(anchor - 8 * 1024 * 1024, anchor + 4096);
|
||||
let checksum = live.iter().map(|&b| b as u64).sum();
|
||||
done_tx.send((lazy, checksum)).unwrap();
|
||||
},
|
||||
);
|
||||
|
||||
let wpid = worker.pid();
|
||||
for _ in 0..(SHRINK_COOLDOWN + 8) {
|
||||
// Gate: send only once the worker is genuinely parked so every
|
||||
// round is a real park-on-empty-mailbox.
|
||||
loop {
|
||||
match actor_info(wpid) {
|
||||
Some(info) if info.state == ActorState::Parked => break,
|
||||
Some(_) => yield_now(),
|
||||
None => panic!("worker died early"),
|
||||
}
|
||||
}
|
||||
park_tx.send(()).unwrap();
|
||||
}
|
||||
|
||||
let (lazy, checksum) = done_rx.recv().unwrap();
|
||||
// The spike was ~3 MiB; demand at least 2 MiB marked to leave slack
|
||||
// for the redzone, rounding, and pages the unwind re-dirtied.
|
||||
assert!(
|
||||
lazy >= 2 * 1024 * 1024,
|
||||
"expected ≥ 2 MiB LazyFree in the stack range, got {} bytes",
|
||||
lazy
|
||||
);
|
||||
assert_eq!(
|
||||
checksum,
|
||||
64 * 0xA5u64,
|
||||
"live stack data corrupted by shrink"
|
||||
);
|
||||
worker.join().unwrap();
|
||||
});
|
||||
}
|
||||
|
||||
/// Steady-state actors must never pay the syscall: an actor that parks a lot
|
||||
/// but never spikes past the threshold ends with zero LazyFree in its stack.
|
||||
#[test]
|
||||
fn shallow_actor_never_shrinks() {
|
||||
let rt = smarm::runtime::init(Config::exact(1));
|
||||
rt.run(|| {
|
||||
let (park_tx, park_rx) = channel::<()>();
|
||||
let (done_tx, done_rx) = channel::<usize>();
|
||||
|
||||
let worker = spawn(move || {
|
||||
let probe = 0u8;
|
||||
let anchor = &probe as *const u8 as usize;
|
||||
for _ in 0..(SHRINK_COOLDOWN + 8) {
|
||||
park_rx.recv().unwrap();
|
||||
}
|
||||
done_tx
|
||||
.send(lazy_free_bytes_in(anchor - 64 * 1024, anchor + 4096))
|
||||
.unwrap();
|
||||
});
|
||||
|
||||
let wpid = worker.pid();
|
||||
for _ in 0..(SHRINK_COOLDOWN + 8) {
|
||||
loop {
|
||||
match actor_info(wpid) {
|
||||
Some(info) if info.state == ActorState::Parked => break,
|
||||
Some(_) => yield_now(),
|
||||
None => panic!("worker died early"),
|
||||
}
|
||||
}
|
||||
park_tx.send(()).unwrap();
|
||||
}
|
||||
|
||||
assert_eq!(done_rx.recv().unwrap(), 0, "steady-state actor was shrunk");
|
||||
worker.join().unwrap();
|
||||
});
|
||||
}
|
||||
@@ -18,8 +18,8 @@
|
||||
//! registry entry guarantees for every named server.
|
||||
|
||||
use smarm::{
|
||||
call, channel, init, request_stop, spawn, Config, GenServer, GenServerBuilder, GenServerName,
|
||||
CallError, Receiver, RecvTimeoutError,
|
||||
call, channel, init, request_stop, spawn, CallError, Config, GenServer, GenServerBuilder,
|
||||
GenServerName, Receiver, RecvTimeoutError,
|
||||
};
|
||||
use std::sync::{Arc, Mutex};
|
||||
use std::time::Duration;
|
||||
@@ -73,10 +73,12 @@ fn named_server_request_stop_releases_queued_caller_with_server_down() {
|
||||
let (res_tx, res_rx) = channel::<Result<(), CallError>>();
|
||||
|
||||
// 1. Start the named server and keep its ref alive.
|
||||
let server = GenServerBuilder::new(Blocker { gate: Some(gate_rx) })
|
||||
.named(BLOCKER)
|
||||
.start()
|
||||
.expect("name should be free");
|
||||
let server = GenServerBuilder::new(Blocker {
|
||||
gate: Some(gate_rx),
|
||||
})
|
||||
.named(BLOCKER)
|
||||
.start()
|
||||
.expect("name should be free");
|
||||
let spid = server.pid();
|
||||
|
||||
// 2. Send the cast and let the server dequeue it and park on the gate.
|
||||
|
||||
+16
-8
@@ -10,7 +10,11 @@
|
||||
//! out rather than produce a false pass — run with `cargo test -- --timeout`
|
||||
//! or under a CI timeout.
|
||||
|
||||
use smarm::{channel, runtime::{Config, Runtime}, spawn, yield_now, JoinHandle};
|
||||
use smarm::{
|
||||
channel,
|
||||
runtime::{Config, Runtime},
|
||||
spawn, yield_now, JoinHandle,
|
||||
};
|
||||
use std::sync::{
|
||||
atomic::{AtomicU64, AtomicUsize, Ordering},
|
||||
Arc,
|
||||
@@ -199,7 +203,9 @@ fn thundering_herd_all_wake() {
|
||||
}
|
||||
|
||||
// Let all receivers park before we send.
|
||||
for _ in 0..4 { yield_now(); }
|
||||
for _ in 0..4 {
|
||||
yield_now();
|
||||
}
|
||||
|
||||
// Coordinator blasts all channels.
|
||||
handles.push(spawn(move || {
|
||||
@@ -240,8 +246,7 @@ fn concurrent_spawn_join_churn() {
|
||||
for _ in 0..PARENTS {
|
||||
let tc = t.clone();
|
||||
parent_handles.push(spawn(move || {
|
||||
let mut child_handles: Vec<JoinHandle> =
|
||||
Vec::with_capacity(CHILDREN_PER_PARENT);
|
||||
let mut child_handles: Vec<JoinHandle> = Vec::with_capacity(CHILDREN_PER_PARENT);
|
||||
|
||||
for _ in 0..CHILDREN_PER_PARENT {
|
||||
let tcc = tc.clone();
|
||||
@@ -292,7 +297,9 @@ fn join_race_child_finishes_first() {
|
||||
}
|
||||
|
||||
// Yield enough to let children run to completion before we join.
|
||||
for _ in 0..8 { yield_now(); }
|
||||
for _ in 0..8 {
|
||||
yield_now();
|
||||
}
|
||||
|
||||
for h in handles {
|
||||
// If child already finished, join must return immediately with Ok.
|
||||
@@ -374,8 +381,7 @@ fn panic_storm_does_not_corrupt_scheduler() {
|
||||
fn pid_generation_increments_on_reuse() {
|
||||
use smarm::self_pid;
|
||||
|
||||
let pids: Arc<smarm::Mutex<Vec<smarm::Pid>>> =
|
||||
Arc::new(smarm::Mutex::new(Vec::new()));
|
||||
let pids: Arc<smarm::Mutex<Vec<smarm::Pid>>> = Arc::new(smarm::Mutex::new(Vec::new()));
|
||||
|
||||
let p = pids.clone();
|
||||
rt(1).run(move || {
|
||||
@@ -392,7 +398,9 @@ fn pid_generation_increments_on_reuse() {
|
||||
}
|
||||
});
|
||||
|
||||
let g = pids.lock_timeout(std::time::Duration::from_secs(1)).unwrap();
|
||||
let g = pids
|
||||
.lock_timeout(std::time::Duration::from_secs(1))
|
||||
.unwrap();
|
||||
// Any two PIDs that share an index must have different generations.
|
||||
for i in 0..g.len() {
|
||||
for j in (i + 1)..g.len() {
|
||||
|
||||
+10
-2
@@ -51,7 +51,11 @@ fn transient_child_is_restarted_on_panic_then_settles() {
|
||||
});
|
||||
sup.join().unwrap();
|
||||
});
|
||||
assert_eq!(runs.load(Ordering::SeqCst), 3, "two restarts then a clean exit");
|
||||
assert_eq!(
|
||||
runs.load(Ordering::SeqCst),
|
||||
3,
|
||||
"two restarts then a clean exit"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
@@ -167,7 +171,11 @@ fn one_for_all_restarts_a_normally_exited_sibling() {
|
||||
sup.join().unwrap();
|
||||
});
|
||||
assert_eq!(a.load(Ordering::SeqCst), 2, "A: crash then clean run");
|
||||
assert_eq!(b.load(Ordering::SeqCst), 2, "B cycled with the group despite a clean exit");
|
||||
assert_eq!(
|
||||
b.load(Ordering::SeqCst),
|
||||
2,
|
||||
"B cycled with the group despite a clean exit"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
|
||||
@@ -0,0 +1,303 @@
|
||||
//! Supervisor shutdown — the OTP child-spec `shutdown` policy.
|
||||
//!
|
||||
//! A supervisor traps exits. A `request_shutdown` reaching it (from its parent
|
||||
//! supervisor, or from the app via `request_shutdown`/`RuntimeHandle`) runs
|
||||
//! the ordered shutdown: children are stopped in reverse start order, each
|
||||
//! per its `Shutdown` policy — `request_shutdown`, wait up to the timeout for
|
||||
//! its termination signal, `request_stop` if it overstays — and then `run()`
|
||||
//! returns normally. Every supervisor-initiated child stop (ordered shutdown,
|
||||
//! OneForAll/RestForOne sibling cycling) goes through the same policy.
|
||||
//!
|
||||
//! A *hard* `request_stop` on a supervisor unwinds it; a drop guard then
|
||||
//! hard-stops its live children so the subtree is never orphaned.
|
||||
|
||||
use smarm::supervisor::{ChildSpec, OneForOne, Restart, Shutdown, Strategy};
|
||||
use smarm::{
|
||||
monitor, request_shutdown, request_stop, run, sleep, spawn, trap_exit, DownReason, JoinHandle,
|
||||
};
|
||||
use std::sync::atomic::{AtomicBool, AtomicUsize, Ordering};
|
||||
use std::sync::{Arc, Mutex};
|
||||
use std::time::{Duration, Instant};
|
||||
|
||||
/// A child that traps exits, records the order it was shut down in, and exits
|
||||
/// normally on the request (after `delay`). Ignores the request if `comply`
|
||||
/// is false — a straggler that must be hard-stopped.
|
||||
fn polite_child(
|
||||
tag: usize,
|
||||
log: &Arc<Mutex<Vec<usize>>>,
|
||||
delay: Duration,
|
||||
comply: bool,
|
||||
) -> impl Fn() + Send + Sync + 'static {
|
||||
let log = log.clone();
|
||||
move || {
|
||||
let inbox = trap_exit();
|
||||
loop {
|
||||
let sig = match inbox.recv() {
|
||||
Ok(s) => s,
|
||||
Err(_) => return,
|
||||
};
|
||||
if sig.reason == DownReason::Shutdown {
|
||||
log.lock().unwrap().push(tag);
|
||||
if comply {
|
||||
sleep(delay);
|
||||
return;
|
||||
}
|
||||
// Not complying: keep running until hard-stopped.
|
||||
loop {
|
||||
sleep(Duration::from_millis(5));
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Spawn `sup`, let its children reach `trap_exit`, return the handle.
|
||||
fn spawn_settled(sup: OneForOne) -> JoinHandle {
|
||||
let h = spawn(move || sup.run());
|
||||
sleep(Duration::from_millis(30));
|
||||
h
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn shutdown_stops_children_in_reverse_order_and_returns_normally() {
|
||||
let log = Arc::new(Mutex::new(Vec::new()));
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let sup = OneForOne::new()
|
||||
.child(ChildSpec::new(
|
||||
Restart::Permanent,
|
||||
polite_child(1, &l, Duration::ZERO, true),
|
||||
))
|
||||
.child(ChildSpec::new(
|
||||
Restart::Permanent,
|
||||
polite_child(2, &l, Duration::ZERO, true),
|
||||
))
|
||||
.child(ChildSpec::new(
|
||||
Restart::Permanent,
|
||||
polite_child(3, &l, Duration::ZERO, true),
|
||||
));
|
||||
let h = spawn_settled(sup);
|
||||
let mon = monitor(h.pid());
|
||||
request_shutdown(h.pid());
|
||||
let down = mon.rx.recv().expect("down");
|
||||
assert_eq!(
|
||||
down.reason,
|
||||
DownReason::Exit,
|
||||
"supervisor exits normally after shutdown"
|
||||
);
|
||||
});
|
||||
assert_eq!(*log.lock().unwrap(), vec![3, 2, 1]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn non_trapping_child_is_simply_stopped() {
|
||||
let dropped = Arc::new(AtomicBool::new(false));
|
||||
let d = dropped.clone();
|
||||
run(move || {
|
||||
struct G(Arc<AtomicBool>);
|
||||
impl Drop for G {
|
||||
fn drop(&mut self) {
|
||||
self.0.store(true, Ordering::SeqCst);
|
||||
}
|
||||
}
|
||||
let sup = OneForOne::new().child(ChildSpec::new(Restart::Permanent, move || {
|
||||
let _g = G(d.clone());
|
||||
loop {
|
||||
sleep(Duration::from_millis(5));
|
||||
}
|
||||
}));
|
||||
let h = spawn_settled(sup);
|
||||
request_shutdown(h.pid());
|
||||
h.join().expect("sup");
|
||||
});
|
||||
assert!(dropped.load(Ordering::SeqCst));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn straggler_is_hard_stopped_after_timeout() {
|
||||
let log = Arc::new(Mutex::new(Vec::new()));
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let sup = OneForOne::new().child(
|
||||
ChildSpec::new(
|
||||
Restart::Permanent,
|
||||
polite_child(1, &l, Duration::ZERO, false),
|
||||
)
|
||||
.shutdown(Shutdown::Timeout(Duration::from_millis(50))),
|
||||
);
|
||||
let h = spawn_settled(sup);
|
||||
let t0 = Instant::now();
|
||||
request_shutdown(h.pid());
|
||||
h.join().expect("sup");
|
||||
let took = t0.elapsed();
|
||||
assert!(
|
||||
took >= Duration::from_millis(50),
|
||||
"returned before the grace period: {took:?}"
|
||||
);
|
||||
assert!(
|
||||
took < Duration::from_secs(2),
|
||||
"did not fall back to a hard stop: {took:?}"
|
||||
);
|
||||
});
|
||||
assert_eq!(
|
||||
*log.lock().unwrap(),
|
||||
vec![1],
|
||||
"the straggler did receive the request"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn infinity_waits_for_a_slow_but_compliant_child() {
|
||||
let log = Arc::new(Mutex::new(Vec::new()));
|
||||
let l = log.clone();
|
||||
let finished = Arc::new(AtomicBool::new(false));
|
||||
let f = finished.clone();
|
||||
run(move || {
|
||||
let f2 = f.clone();
|
||||
let l2 = l.clone();
|
||||
let sup = OneForOne::new().child(
|
||||
ChildSpec::new(Restart::Permanent, move || {
|
||||
let inbox = trap_exit();
|
||||
let _ = inbox.recv();
|
||||
l2.lock().unwrap().push(1);
|
||||
sleep(Duration::from_millis(150));
|
||||
f2.store(true, Ordering::SeqCst); // only reached if not hard-stopped
|
||||
})
|
||||
.shutdown(Shutdown::Infinity),
|
||||
);
|
||||
let h = spawn_settled(sup);
|
||||
request_shutdown(h.pid());
|
||||
h.join().expect("sup");
|
||||
});
|
||||
assert!(
|
||||
finished.load(Ordering::SeqCst),
|
||||
"Infinity must not hard-stop a compliant child"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn brutal_kill_skips_the_request() {
|
||||
let log = Arc::new(Mutex::new(Vec::new()));
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let sup = OneForOne::new().child(
|
||||
ChildSpec::new(
|
||||
Restart::Permanent,
|
||||
polite_child(1, &l, Duration::ZERO, true),
|
||||
)
|
||||
.shutdown(Shutdown::BrutalKill),
|
||||
);
|
||||
let h = spawn_settled(sup);
|
||||
request_shutdown(h.pid());
|
||||
h.join().expect("sup");
|
||||
});
|
||||
assert!(
|
||||
log.lock().unwrap().is_empty(),
|
||||
"a BrutalKill child never sees the request"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn hard_stop_of_supervisor_does_not_orphan_children() {
|
||||
let alive = Arc::new(AtomicUsize::new(0));
|
||||
let a = alive.clone();
|
||||
run(move || {
|
||||
struct Alive(Arc<AtomicUsize>);
|
||||
impl Drop for Alive {
|
||||
fn drop(&mut self) {
|
||||
self.0.fetch_sub(1, Ordering::SeqCst);
|
||||
}
|
||||
}
|
||||
let mk = |a: Arc<AtomicUsize>| {
|
||||
move || {
|
||||
a.fetch_add(1, Ordering::SeqCst);
|
||||
let _g = Alive(a.clone());
|
||||
loop {
|
||||
sleep(Duration::from_millis(5));
|
||||
}
|
||||
}
|
||||
};
|
||||
let sup = OneForOne::new()
|
||||
.child(ChildSpec::new(Restart::Permanent, mk(a.clone())))
|
||||
.child(ChildSpec::new(Restart::Permanent, mk(a.clone())));
|
||||
let h = spawn_settled(sup);
|
||||
assert_eq!(a.load(Ordering::SeqCst), 2);
|
||||
let mon = monitor(h.pid());
|
||||
request_stop(h.pid());
|
||||
let _ = mon.rx.recv();
|
||||
sleep(Duration::from_millis(50));
|
||||
assert_eq!(
|
||||
a.load(Ordering::SeqCst),
|
||||
0,
|
||||
"children orphaned by a hard supervisor stop"
|
||||
);
|
||||
});
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn nested_shutdown_reaches_grandchildren() {
|
||||
let log = Arc::new(Mutex::new(Vec::new()));
|
||||
let l = log.clone();
|
||||
run(move || {
|
||||
let l_inner = l.clone();
|
||||
let inner = move || {
|
||||
OneForOne::new()
|
||||
.child(ChildSpec::new(
|
||||
Restart::Permanent,
|
||||
polite_child(10, &l_inner, Duration::ZERO, true),
|
||||
))
|
||||
.child(ChildSpec::new(
|
||||
Restart::Permanent,
|
||||
polite_child(11, &l_inner, Duration::ZERO, true),
|
||||
))
|
||||
.run()
|
||||
};
|
||||
let sup = OneForOne::new()
|
||||
.child(ChildSpec::new(
|
||||
Restart::Permanent,
|
||||
polite_child(1, &l, Duration::ZERO, true),
|
||||
))
|
||||
.child(ChildSpec::new(Restart::Permanent, inner).shutdown(Shutdown::Infinity));
|
||||
let h = spawn_settled(sup);
|
||||
request_shutdown(h.pid());
|
||||
h.join().expect("sup");
|
||||
});
|
||||
assert_eq!(*log.lock().unwrap(), vec![11, 10, 1]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn sibling_cycling_uses_graceful_shutdown() {
|
||||
// OneForAll: when child A dies, sibling B (trapping) must receive a
|
||||
// Shutdown request rather than a bare stop.
|
||||
let log = Arc::new(Mutex::new(Vec::new()));
|
||||
let l = log.clone();
|
||||
let a_runs = Arc::new(AtomicUsize::new(0));
|
||||
let ar = a_runs.clone();
|
||||
run(move || {
|
||||
let ar2 = ar.clone();
|
||||
let sup = OneForOne::new()
|
||||
.strategy(Strategy::OneForAll)
|
||||
.intensity(5, Duration::from_secs(60))
|
||||
.child(ChildSpec::new(Restart::Transient, move || {
|
||||
let n = ar2.fetch_add(1, Ordering::SeqCst) + 1;
|
||||
sleep(Duration::from_millis(30));
|
||||
if n == 1 {
|
||||
panic!("first run dies");
|
||||
}
|
||||
// Second run: park until shut down.
|
||||
let inbox = trap_exit();
|
||||
let _ = inbox.recv();
|
||||
}))
|
||||
.child(ChildSpec::new(
|
||||
Restart::Permanent,
|
||||
polite_child(2, &l, Duration::ZERO, true),
|
||||
));
|
||||
let h = spawn(move || sup.run());
|
||||
sleep(Duration::from_millis(150));
|
||||
request_shutdown(h.pid());
|
||||
h.join().expect("sup");
|
||||
});
|
||||
// B was shut down once by the cycle and once by the final shutdown.
|
||||
assert_eq!(*log.lock().unwrap(), vec![2, 2]);
|
||||
assert_eq!(a_runs.load(Ordering::SeqCst), 2);
|
||||
}
|
||||
@@ -0,0 +1,276 @@
|
||||
//! The terminal-record contract (bridge soak signature 4): a watch installed
|
||||
//! *after* its target's death — the async-install race the bridge's proxies
|
||||
//! live with — must be able to recover the real down reason instead of a
|
||||
//! blanket `NoProc`. Two primitives carry it:
|
||||
//!
|
||||
//! - `finalize_actor` stamps the slot with `(generation, DownReason)`; the
|
||||
//! record survives reclaim, registry pruning, and the next tenant's
|
||||
//! install, and is overwritten only by the slot's next death.
|
||||
//! [`terminal_reason`] reads it generation-matched.
|
||||
//! - [`resolve_name`] is `whereis` with the corpse kept: the dead-holder arm
|
||||
//! returns the stored pid it prunes ([`NameResolution::Corpse`]) instead
|
||||
//! of discarding the only evidence of *who* died. `Unbound` stays the
|
||||
//! Erlang-shaped `noproc` for names that were never (or are no longer)
|
||||
//! bound.
|
||||
//!
|
||||
//! `monitor()` of a stale pid still queues plain `NoProc` — the upgrade is a
|
||||
//! caller's deliberate act, not a semantics change.
|
||||
|
||||
use smarm::{
|
||||
init, mark_watchable, request_stop, resolve_name, terminal_reason, CallError, Config,
|
||||
DownReason, GenServer, GenServerBuilder, GenServerName, NameResolution,
|
||||
};
|
||||
use std::sync::{Arc, Mutex};
|
||||
use std::time::Duration;
|
||||
|
||||
const TARGET: GenServerName<Target> = GenServerName::new("terminal_target");
|
||||
|
||||
/// Named server that panics on cast — the sig-4 death.
|
||||
struct Target;
|
||||
|
||||
impl GenServer for Target {
|
||||
type Call = ();
|
||||
type Reply = ();
|
||||
type Cast = ();
|
||||
type Info = ();
|
||||
type Timer = ();
|
||||
|
||||
fn handle_call(&mut self, _req: ()) {}
|
||||
fn handle_cast(&mut self, _op: ()) {
|
||||
panic!("terminal_target: induced panic");
|
||||
}
|
||||
}
|
||||
|
||||
/// Slot filler for the re-tenancy phase (distinct type, held alive).
|
||||
struct Filler;
|
||||
|
||||
impl GenServer for Filler {
|
||||
type Call = ();
|
||||
type Reply = ();
|
||||
type Cast = ();
|
||||
type Info = ();
|
||||
type Timer = ();
|
||||
|
||||
fn handle_call(&mut self, _req: ()) {}
|
||||
fn handle_cast(&mut self, _op: ()) {}
|
||||
}
|
||||
|
||||
#[derive(Debug)]
|
||||
struct Observed {
|
||||
exit_reason: Option<DownReason>,
|
||||
anon_reason: Option<DownReason>,
|
||||
/// Anonymous but export-marked while alive — must stamp (sig 5).
|
||||
marked_reason: Option<DownReason>,
|
||||
/// Marked only after death — must remain unknowable.
|
||||
marked_late_reason: Option<DownReason>,
|
||||
panic_reason: Option<DownReason>,
|
||||
stopped_reason: Option<DownReason>,
|
||||
live_reason: Option<DownReason>,
|
||||
live_resolution_is_live: bool,
|
||||
unknown_resolution: NameResolution,
|
||||
/// First resolve after the named target's panic — must be Corpse(old pid).
|
||||
corpse_resolution_matches: bool,
|
||||
/// Second resolve — the Corpse arm pruned, so the name has healed.
|
||||
resolution_after_prune: NameResolution,
|
||||
/// Read AFTER the prune above: the record is slot-side, not registry-side.
|
||||
corpse_reason_after_prune: Option<DownReason>,
|
||||
/// Record survives the slot being re-tenanted (new tenant still alive).
|
||||
corpse_reason_after_reuse: Option<DownReason>,
|
||||
/// ... and dies with the next tenancy's death (overwritten).
|
||||
corpse_reason_after_tenant_death: Option<DownReason>,
|
||||
tenant_reason: Option<DownReason>,
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn terminal_record_recovers_the_reason_a_raced_watch_lost() {
|
||||
let out: Arc<Mutex<Option<Observed>>> = Arc::new(Mutex::new(None));
|
||||
let out_w = out.clone();
|
||||
|
||||
// Tiny slab: prompt slot recycling for the re-tenancy phase.
|
||||
init(Config::exact(2).max_actors(32)).run(move || {
|
||||
// --- Registered plain actors: one record per way of dying. The
|
||||
// record is named-tenancy-only, so each actor self-registers a
|
||||
// throwaway channel before dying; the anonymous control below pins
|
||||
// the complement.
|
||||
let h = smarm::spawn(|| {
|
||||
let (tx, _rx) = smarm::channel::<()>();
|
||||
let _ = smarm::register(smarm::Name::<()>::new("terminal_probe_exit"), tx);
|
||||
});
|
||||
let pid_exit = h.pid();
|
||||
let _ = h.join();
|
||||
let exit_reason = terminal_reason(pid_exit);
|
||||
|
||||
let h = smarm::spawn(|| {
|
||||
let (tx, _rx) = smarm::channel::<()>();
|
||||
let _ = smarm::register(smarm::Name::<()>::new("terminal_probe_panic"), tx);
|
||||
panic!("induced");
|
||||
});
|
||||
let pid_panic = h.pid();
|
||||
let _ = h.join();
|
||||
let panic_reason = terminal_reason(pid_panic);
|
||||
|
||||
let h = smarm::spawn(|| {
|
||||
let (tx, _rx) = smarm::channel::<()>();
|
||||
let _ = smarm::register(smarm::Name::<()>::new("terminal_probe_stop"), tx);
|
||||
loop {
|
||||
smarm::sleep(Duration::from_millis(2));
|
||||
}
|
||||
});
|
||||
let pid_stop = h.pid();
|
||||
request_stop(pid_stop);
|
||||
let _ = h.join();
|
||||
let stopped_reason = terminal_reason(pid_stop);
|
||||
|
||||
// --- Anonymous control: an unregistered death must NOT stamp (nor
|
||||
// evict) — the free list is LIFO, so green-thread churn would
|
||||
// otherwise overwrite a watchable record faster than any race
|
||||
// window this exists to cover.
|
||||
let h = smarm::spawn(|| panic!("anonymous"));
|
||||
let pid_anon = h.pid();
|
||||
let _ = h.join();
|
||||
let anon_reason = terminal_reason(pid_anon);
|
||||
|
||||
// --- mark_watchable: the bridge's export-seam eligibility (sig 5).
|
||||
// An anonymous actor marked while alive stamps like a named one ...
|
||||
let h = smarm::spawn(|| loop {
|
||||
smarm::sleep(Duration::from_millis(2));
|
||||
});
|
||||
let pid_marked = h.pid();
|
||||
mark_watchable(pid_marked);
|
||||
request_stop(pid_marked);
|
||||
let _ = h.join();
|
||||
let marked_reason = terminal_reason(pid_marked);
|
||||
|
||||
// ... while marking a pid whose tenancy already ended is a no-op:
|
||||
// the history is honestly unknowable, not retroactively invented.
|
||||
mark_watchable(pid_anon);
|
||||
let marked_late_reason = terminal_reason(pid_anon);
|
||||
|
||||
// --- The named target: live readings first. -----------------------
|
||||
let target = GenServerBuilder::new(Target)
|
||||
.named(TARGET)
|
||||
.start()
|
||||
.expect("name free at test start");
|
||||
let old_pid = target.pid();
|
||||
let live_reason = terminal_reason(old_pid);
|
||||
let live_resolution_is_live =
|
||||
resolve_name(TARGET.as_str()) == NameResolution::Live(old_pid.erase());
|
||||
let unknown_resolution = resolve_name("terminal_never_bound");
|
||||
|
||||
// --- Kill it by panic; confirm death via the ref, NEVER the name
|
||||
// (any name reader would take the prune arm and destroy the corpse
|
||||
// precondition — the same trap stale_name_slot_reuse.rs documents).
|
||||
let _ = target.cast(());
|
||||
loop {
|
||||
match target.call(()) {
|
||||
Err(CallError::ServerDown) => break,
|
||||
Ok(()) => smarm::sleep(Duration::from_millis(2)),
|
||||
}
|
||||
}
|
||||
|
||||
let corpse_resolution_matches =
|
||||
resolve_name(TARGET.as_str()) == NameResolution::Corpse(old_pid.erase());
|
||||
let resolution_after_prune = resolve_name(TARGET.as_str());
|
||||
let corpse_reason_after_prune = terminal_reason(old_pid);
|
||||
|
||||
// --- Re-tenant the freed slot; the record must outlive the install
|
||||
// and die only with the next tenancy's death.
|
||||
let mut fillers = Vec::new();
|
||||
let mut tenant = None;
|
||||
for i in 0..24 {
|
||||
let name: &'static str = Box::leak(format!("terminal_filler_{i}").into_boxed_str());
|
||||
let f = GenServerBuilder::new(Filler)
|
||||
.named(GenServerName::<Filler>::new(name))
|
||||
.start()
|
||||
.expect("filler names are fresh");
|
||||
let fp = f.pid();
|
||||
let landed = fp.index() == old_pid.index();
|
||||
fillers.push(f);
|
||||
if landed {
|
||||
tenant = Some((fillers.len() - 1, fp));
|
||||
break;
|
||||
}
|
||||
}
|
||||
let (tenant_at, tenant_pid) = tenant.expect(
|
||||
"precondition: the freed slot must be re-tenanted within the tiny slab \
|
||||
(slots are recycled; every filler is held alive)",
|
||||
);
|
||||
let corpse_reason_after_reuse = terminal_reason(old_pid);
|
||||
|
||||
request_stop(tenant_pid);
|
||||
loop {
|
||||
match fillers[tenant_at].call(()) {
|
||||
Err(CallError::ServerDown) => break,
|
||||
Ok(()) => smarm::sleep(Duration::from_millis(2)),
|
||||
}
|
||||
}
|
||||
let corpse_reason_after_tenant_death = terminal_reason(old_pid);
|
||||
let tenant_reason = terminal_reason(tenant_pid);
|
||||
|
||||
*out_w.lock().unwrap() = Some(Observed {
|
||||
exit_reason,
|
||||
anon_reason,
|
||||
panic_reason,
|
||||
stopped_reason,
|
||||
live_reason,
|
||||
live_resolution_is_live,
|
||||
unknown_resolution,
|
||||
corpse_resolution_matches,
|
||||
resolution_after_prune,
|
||||
marked_reason,
|
||||
marked_late_reason,
|
||||
corpse_reason_after_prune,
|
||||
corpse_reason_after_reuse,
|
||||
corpse_reason_after_tenant_death,
|
||||
tenant_reason,
|
||||
});
|
||||
});
|
||||
|
||||
let o = out.lock().unwrap().take().expect("runtime body completed");
|
||||
assert_eq!(o.exit_reason, Some(DownReason::Exit), "{o:?}");
|
||||
assert_eq!(
|
||||
o.anon_reason, None,
|
||||
"anonymous deaths must not stamp: {o:?}"
|
||||
);
|
||||
assert_eq!(o.panic_reason, Some(DownReason::Panic), "{o:?}");
|
||||
assert_eq!(
|
||||
o.marked_reason,
|
||||
Some(DownReason::Stopped),
|
||||
"mark_watchable while alive must make the death stamp: {o:?}"
|
||||
);
|
||||
assert_eq!(
|
||||
o.marked_late_reason, None,
|
||||
"marking a dead tenancy must not invent history: {o:?}"
|
||||
);
|
||||
assert_eq!(o.stopped_reason, Some(DownReason::Stopped), "{o:?}");
|
||||
assert_eq!(
|
||||
o.live_reason, None,
|
||||
"live tenancy must have no record: {o:?}"
|
||||
);
|
||||
assert!(o.live_resolution_is_live, "{o:?}");
|
||||
assert_eq!(o.unknown_resolution, NameResolution::Unbound, "{o:?}");
|
||||
assert!(
|
||||
o.corpse_resolution_matches,
|
||||
"first post-death resolve must carry the corpse: {o:?}"
|
||||
);
|
||||
assert_eq!(
|
||||
o.resolution_after_prune,
|
||||
NameResolution::Unbound,
|
||||
"the Corpse arm prunes — the name heals: {o:?}"
|
||||
);
|
||||
assert_eq!(
|
||||
o.corpse_reason_after_prune,
|
||||
Some(DownReason::Panic),
|
||||
"the record is slot-side; registry pruning must not touch it: {o:?}"
|
||||
);
|
||||
assert_eq!(
|
||||
o.corpse_reason_after_reuse,
|
||||
Some(DownReason::Panic),
|
||||
"a new tenant's install must leave the previous tenancy's record: {o:?}"
|
||||
);
|
||||
assert_eq!(
|
||||
o.corpse_reason_after_tenant_death, None,
|
||||
"the next death overwrites — the old generation no longer matches: {o:?}"
|
||||
);
|
||||
assert_eq!(o.tenant_reason, Some(DownReason::Stopped), "{o:?}");
|
||||
}
|
||||
@@ -35,7 +35,10 @@ impl PipePair {
|
||||
let mut fds: [libc::c_int; 2] = [0; 2];
|
||||
let r = unsafe { libc::pipe2(fds.as_mut_ptr(), libc::O_CLOEXEC | libc::O_NONBLOCK) };
|
||||
assert_eq!(r, 0, "pipe2 failed");
|
||||
PipePair { read: fds[0], write: fds[1] }
|
||||
PipePair {
|
||||
read: fds[0],
|
||||
write: fds[1],
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -67,9 +70,9 @@ fn run_with_watchdog(limit: Duration, body: impl FnOnce() + Send + 'static) {
|
||||
rt.run(body);
|
||||
let _ = done_tx.send(());
|
||||
});
|
||||
done_rx
|
||||
.recv_timeout(limit)
|
||||
.expect("Runtime::run did not return: idle scheduler thread was never woken at termination");
|
||||
done_rx.recv_timeout(limit).expect(
|
||||
"Runtime::run did not return: idle scheduler thread was never woken at termination",
|
||||
);
|
||||
}
|
||||
|
||||
/// Permanent-hang variant: sibling blocked in `poll_wake(wake_fd, None)`
|
||||
|
||||
+30
-10
@@ -166,14 +166,19 @@ fn timers_only_pop_entries_whose_deadline_has_passed() {
|
||||
#[test]
|
||||
fn timers_mix_sleep_and_wait_timeout_reasons() {
|
||||
let mut t = Timers::new();
|
||||
let target = Arc::new(RecordingTarget { calls: Mutex::new(Vec::new()) });
|
||||
let target = Arc::new(RecordingTarget {
|
||||
calls: Mutex::new(Vec::new()),
|
||||
});
|
||||
let now = Instant::now();
|
||||
|
||||
t.insert_sleep(now + Duration::from_millis(5), Pid::new(0, 0), 1);
|
||||
t.insert(
|
||||
now + Duration::from_millis(10),
|
||||
Pid::new(1, 0),
|
||||
Reason::WaitTimeout { target: target.clone(), epoch: 42 },
|
||||
Reason::WaitTimeout {
|
||||
target: target.clone(),
|
||||
epoch: 42,
|
||||
},
|
||||
);
|
||||
|
||||
let due = t.pop_due(now + Duration::from_millis(20));
|
||||
@@ -238,7 +243,10 @@ fn armed_send_timer_is_returned_and_fires() {
|
||||
|
||||
let mut due = t.pop_due(now + Duration::from_millis(20));
|
||||
assert_eq!(due.len(), 1, "an armed send timer should pop when due");
|
||||
assert!(!fired.load(Ordering::SeqCst), "pop must not fire on its own");
|
||||
assert!(
|
||||
!fired.load(Ordering::SeqCst),
|
||||
"pop must not fire on its own"
|
||||
);
|
||||
run_fire(due.pop().unwrap());
|
||||
assert!(fired.load(Ordering::SeqCst), "running the thunk delivers");
|
||||
assert!(t.is_empty());
|
||||
@@ -282,7 +290,11 @@ fn cancel_after_fire_returns_false() {
|
||||
fn cancel_unknown_id_returns_false() {
|
||||
let mut t = Timers::new();
|
||||
let now = Instant::now();
|
||||
let id = t.insert_send(now + Duration::from_millis(5), Pid::new(0, 0), Box::new(|| {}));
|
||||
let id = t.insert_send(
|
||||
now + Duration::from_millis(5),
|
||||
Pid::new(0, 0),
|
||||
Box::new(|| {}),
|
||||
);
|
||||
assert!(t.cancel(id));
|
||||
// Second cancel of the same id: already gone.
|
||||
assert!(!t.cancel(id));
|
||||
@@ -293,7 +305,11 @@ fn send_timers_interleave_with_sleep_in_deadline_order() {
|
||||
let mut t = Timers::new();
|
||||
let now = Instant::now();
|
||||
t.insert_sleep(now + Duration::from_millis(30), Pid::new(0, 0), 1);
|
||||
let _id = t.insert_send(now + Duration::from_millis(10), Pid::new(1, 0), Box::new(|| {}));
|
||||
let _id = t.insert_send(
|
||||
now + Duration::from_millis(10),
|
||||
Pid::new(1, 0),
|
||||
Box::new(|| {}),
|
||||
);
|
||||
t.insert_sleep(now + Duration::from_millis(20), Pid::new(2, 0), 1);
|
||||
|
||||
let due = t.pop_due(now + Duration::from_millis(50));
|
||||
@@ -308,7 +324,11 @@ fn send_timers_interleave_with_sleep_in_deadline_order() {
|
||||
fn clear_drops_armed_send_timers() {
|
||||
let mut t = Timers::new();
|
||||
let now = Instant::now();
|
||||
let id = t.insert_send(now + Duration::from_millis(10), Pid::new(0, 0), Box::new(|| {}));
|
||||
let id = t.insert_send(
|
||||
now + Duration::from_millis(10),
|
||||
Pid::new(0, 0),
|
||||
Box::new(|| {}),
|
||||
);
|
||||
t.clear();
|
||||
assert!(t.is_empty());
|
||||
// The arm record is gone too: cancelling reports nothing to cancel.
|
||||
@@ -356,7 +376,7 @@ fn send_after_to_unresolved_name_is_silent() {
|
||||
// Nobody registered NOPE; firing resolves to nothing and is dropped.
|
||||
let _id = send_after_named(Duration::from_millis(10), NOPE, 1);
|
||||
sleep(Duration::from_millis(40)); // let it fire and no-op
|
||||
// Reaching here without a panic is the assertion.
|
||||
// Reaching here without a panic is the assertion.
|
||||
});
|
||||
}
|
||||
|
||||
@@ -401,9 +421,9 @@ fn send_after_to_dead_typed_pid_is_silent() {
|
||||
assert_eq!(report_rx.recv().unwrap(), 1); // sink has now exited
|
||||
let _id = send_after(Duration::from_millis(15), sink, 2);
|
||||
sleep(Duration::from_millis(45)); // let it fire against the dead pid
|
||||
// No panic; the sink is gone, so its report sender dropped with it —
|
||||
// closed+empty is Err (documented), which also proves nothing
|
||||
// further was delivered.
|
||||
// No panic; the sink is gone, so its report sender dropped with it —
|
||||
// closed+empty is Err (documented), which also proves nothing
|
||||
// further was delivered.
|
||||
assert!(report_rx.try_recv().is_err(), "nothing further delivered");
|
||||
});
|
||||
}
|
||||
|
||||
@@ -0,0 +1,185 @@
|
||||
//! Non-panicking spawn at slab capacity (`try_spawn`).
|
||||
//!
|
||||
//! Covers: parity with `spawn` when slots are free; `Err(AtCapacity)` instead
|
||||
//! of a panic on a full slab (the spawning actor survives — the crash-loop
|
||||
//! from the motivating slowloris incident cannot start); self-heal (a freed
|
||||
//! slot makes the next `try_spawn` succeed); and exact claim-or-report
|
||||
//! accounting under a multi-thread race for the last slots (no TOCTOU
|
||||
//! overshoot, no panic).
|
||||
|
||||
use smarm::runtime::Config;
|
||||
use smarm::{spawn, try_spawn, try_spawn_under_with, yield_now, SpawnError, SpawnOpts};
|
||||
use std::sync::atomic::{AtomicBool, AtomicUsize, Ordering};
|
||||
use std::sync::Arc;
|
||||
|
||||
/// A child that holds its slot until `release` flips, without parking
|
||||
/// machinery: busy-yield keeps the scheduler moving and the slot occupied.
|
||||
fn holder(release: Arc<AtomicBool>) -> impl FnOnce() + Send + 'static {
|
||||
move || {
|
||||
while !release.load(Ordering::Acquire) {
|
||||
yield_now();
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn try_spawn_is_spawn_when_slots_free() {
|
||||
smarm::runtime::init(Config::exact(1)).run(|| {
|
||||
let ran = Arc::new(AtomicBool::new(false));
|
||||
let flag = ran.clone();
|
||||
let h = try_spawn(move || flag.store(true, Ordering::Release))
|
||||
.expect("slots free — must behave exactly like spawn");
|
||||
h.join().unwrap();
|
||||
assert!(ran.load(Ordering::Acquire));
|
||||
});
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn at_capacity_is_err_not_panic_and_accounting_is_exact() {
|
||||
const MAX: usize = 8;
|
||||
smarm::runtime::init(Config::exact(1).max_actors(MAX)).run(|| {
|
||||
let release = Arc::new(AtomicBool::new(false));
|
||||
// Fill the slab from the initial actor: slots are claimed at spawn
|
||||
// time, so children need not have run yet. Count until refusal.
|
||||
let mut held = Vec::new();
|
||||
loop {
|
||||
match try_spawn(holder(release.clone())) {
|
||||
Ok(h) => held.push(h),
|
||||
Err(e) => {
|
||||
assert_eq!(e, SpawnError::AtCapacity);
|
||||
break;
|
||||
}
|
||||
}
|
||||
}
|
||||
// Initial actor occupies one slot; the rest were spawnable.
|
||||
assert_eq!(held.len(), MAX - 1, "slab accounting must be exact");
|
||||
// Still refusing (and still not panicking) on repeat.
|
||||
assert!(matches!(try_spawn(|| ()), Err(SpawnError::AtCapacity)));
|
||||
// The `_with` surface refuses identically — a custom shape must not
|
||||
// reach stack allocation when there is no slot for it.
|
||||
let opts = SpawnOpts {
|
||||
stack_reserve: Some(1024 * 1024),
|
||||
..SpawnOpts::default()
|
||||
};
|
||||
assert!(matches!(
|
||||
try_spawn_under_with(smarm::self_pid(), opts, || ()),
|
||||
Err(SpawnError::AtCapacity)
|
||||
));
|
||||
|
||||
// Self-heal: free the slots, join, and the next try_spawn succeeds.
|
||||
release.store(true, Ordering::Release);
|
||||
for h in held {
|
||||
h.join().unwrap();
|
||||
}
|
||||
let h = try_spawn(|| ()).expect("slots freed — must succeed again");
|
||||
h.join().unwrap();
|
||||
});
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn plain_spawn_still_panics_at_capacity() {
|
||||
// The existing invariant-check semantics of `spawn` are untouched: at a
|
||||
// full slab it panics, the panic is caught at the actor isolation
|
||||
// boundary, and it surfaces as a join error — exactly as before. The
|
||||
// bomb actor is spawned into the LAST slot (so the slab is full only
|
||||
// once the bomb itself is live) and the panic lands inside the bomb,
|
||||
// not the initial actor.
|
||||
const MAX: usize = 6;
|
||||
smarm::runtime::init(Config::exact(1).max_actors(MAX)).run(|| {
|
||||
let release = Arc::new(AtomicBool::new(false));
|
||||
let mut held = Vec::new();
|
||||
for _ in 0..MAX - 2 {
|
||||
held.push(spawn(holder(release.clone())));
|
||||
}
|
||||
let armed = Arc::new(AtomicBool::new(false));
|
||||
let armed2 = armed.clone();
|
||||
let bomb = spawn(move || {
|
||||
armed2.store(true, Ordering::Release);
|
||||
// Slab is now full (initial + MAX−2 holders + this actor); the
|
||||
// plain spawn must panic this actor.
|
||||
let _ = spawn(|| ());
|
||||
unreachable!("allocate_slot must have panicked");
|
||||
});
|
||||
let err = bomb
|
||||
.join()
|
||||
.expect_err("bomb must die by panic, not run through");
|
||||
assert!(armed.load(Ordering::Acquire), "bomb must have actually run");
|
||||
// The panic message is a formatted String (panic! with args).
|
||||
let msg = err
|
||||
.payload
|
||||
.downcast_ref::<String>()
|
||||
.cloned()
|
||||
.unwrap_or_else(|| "<non-string payload>".into());
|
||||
assert!(
|
||||
msg.contains("slot table exhausted"),
|
||||
"panic must be the slab-exhaustion invariant message, got: {msg}"
|
||||
);
|
||||
release.store(true, Ordering::Release);
|
||||
for h in held {
|
||||
h.join().unwrap();
|
||||
}
|
||||
});
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn racing_try_spawns_claim_exactly_the_free_slots() {
|
||||
// 4 scheduler threads, 4 spawner actors hammering try_spawn for a small
|
||||
// pool of remaining slots. Claim-or-report must hand out exactly the
|
||||
// free slots across all racers — no overshoot (TOCTOU), no panic.
|
||||
const MAX: usize = 32;
|
||||
const SPAWNERS: usize = 4;
|
||||
smarm::runtime::init(Config::exact(4).max_actors(MAX)).run(|| {
|
||||
let release = Arc::new(AtomicBool::new(false));
|
||||
let won = Arc::new(AtomicUsize::new(0));
|
||||
let done = Arc::new(AtomicUsize::new(0));
|
||||
|
||||
// Occupy some slots up front so the racers fight over a remainder.
|
||||
let mut pre = Vec::new();
|
||||
for _ in 0..8 {
|
||||
pre.push(spawn(holder(release.clone())));
|
||||
}
|
||||
// Free slots now: MAX − 1 (initial) − 8 (pre) − SPAWNERS.
|
||||
let up_for_grabs = MAX - 1 - 8 - SPAWNERS;
|
||||
|
||||
let mut spawners = Vec::new();
|
||||
for _ in 0..SPAWNERS {
|
||||
let release = release.clone();
|
||||
let won = won.clone();
|
||||
let done = done.clone();
|
||||
spawners.push(spawn(move || {
|
||||
loop {
|
||||
match try_spawn(holder(release.clone())) {
|
||||
Ok(h) => {
|
||||
won.fetch_add(1, Ordering::AcqRel);
|
||||
drop(h); // detached; slot held by the holder
|
||||
}
|
||||
Err(SpawnError::AtCapacity) => break,
|
||||
Err(_) => unreachable!("non_exhaustive future-proofing"),
|
||||
}
|
||||
}
|
||||
done.fetch_add(1, Ordering::AcqRel);
|
||||
}));
|
||||
}
|
||||
// Wait for every racer to hit AtCapacity.
|
||||
while done.load(Ordering::Acquire) < SPAWNERS {
|
||||
yield_now();
|
||||
}
|
||||
assert_eq!(won.load(Ordering::Acquire), up_for_grabs);
|
||||
|
||||
release.store(true, Ordering::Release);
|
||||
for h in pre.into_iter().chain(spawners) {
|
||||
h.join().unwrap();
|
||||
}
|
||||
});
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn spawn_error_is_a_real_error() {
|
||||
let e = SpawnError::AtCapacity;
|
||||
let msg = format!("{e}");
|
||||
assert!(
|
||||
msg.contains("capacity"),
|
||||
"Display should name the condition: {msg}"
|
||||
);
|
||||
let _: &dyn std::error::Error = &e;
|
||||
}
|
||||
Reference in New Issue
Block a user