Files
smarm/tests
claude 9c8f59ca53 feat(scheduler,supervisor,gen_server): graceful shutdown — request_shutdown, child Shutdown policy, handle_shutdown
Lift OTP's `exit(Pid, shutdown)` + child-spec `shutdown` wholesale.

scheduler / runtime
- `request_shutdown(pid)`: the polite stop. A target trapping exits gets an
  `ExitSignal { reason: DownReason::Shutdown }` on its trap inbox and keeps
  running; a non-trapping target is stopped as by `request_stop`, which is
  now documented as the hard stop (`exit(Pid, kill)`). Dead pid: no-op.
- `RuntimeHandle::request_shutdown` for the off-runtime (signal thread) path;
  `from == ROOT_PID` there.
- `DownReason::Shutdown` — appears only in ExitSignal, never in Down (a
  complying target exits *normally*).

supervisor
- `ChildSpec::shutdown(Shutdown::{BrutalKill, Timeout(d), Infinity})`,
  default Timeout(5s). Every supervisor-initiated stop (ordered shutdown and
  OneForAll/RestForOne sibling cycling) is: request_shutdown → await the
  child's Signal up to the grace → request_stop → await. Sequential, reverse
  start order.
- The supervisor traps exits; a Shutdown ExitSignal runs the ordered
  shutdown and `run()` returns normally, so `request_shutdown(root_sup)`
  tears a whole tree down top-down with each child's grace period.
- FIX: a hard `request_stop` on a supervisor previously orphaned its
  children (the ordered shutdown lived after the loop, and the unwind
  skipped it). `Live` (the by_pid map) now carries a drop guard that
  fire-and-forget hard-stops live children when unwinding.

gen_server
- `GenServerCtx::trap_exit()` opt-in in `init`; the trap inbox becomes arm 0
  of the loop's select. Shutdown ExitSignal → `handle_shutdown() ->
  ShutdownAction::{Exit, Continue}` (default Exit: loop breaks, `terminate`
  runs on the normal path and may block). Other ExitSignals →
  `handle_exit(sig)`.
- `GenServerCtx::stop_handle() -> StopHandle`, `stop()` ends the server
  after the current message with a *normal* exit — the missing
  `{stop, normal, State}`; `request_stop(self_pid())` was the only self-exit
  and it is abnormal (Transient restarts it).
- `GenServerRef::shutdown()` / `gen_server::shutdown(name)` now go through
  `request_shutdown`.

Tests: tests/shutdown.rs, tests/supervisor_shutdown.rs,
tests/gen_server_shutdown.rs. Full suite green; fmt + clippy --lib clean.
2026-08-19 06:34:48 +00:00
..
2026-05-23 16:09:35 +00:00

Tests

Integration tests for the runtime. Each file owns one feature area or one class of bug. Everything here runs under plain cargo test; the loom model tests are the exception — they live in the library (src/slot_state.rs, src/run_queue.rs), not in this directory, because loom must compile the production code with shimmed atomics (see "Loom" below).

Running

cargo test                # debug build — RUN THIS ONE: all invariant asserts live
cargo test --release      # what users actually execute (LTO, no debug_asserts)

Debug builds are not just "slower tests": the runtime self-checks its invariants only there — every StateWord transition asserts its precondition, enqueue asserts the exact (gen, Queued) word, RawMutex enforces the never-two-cold-locks leaf rule with a per-thread held-count, live_actors checks for double-finalize underflow. A green release run with a red debug run means an invariant broke without (yet) corrupting behavior — treat it as a real failure.

The queue-variant matrix

The run queue is compile-time selected; the suite must pass under all three (features are additive, so drop the default first):

cargo test                                              # rq-mutex (default)
cargo test --no-default-features --features rq-mpmc
cargo test --no-default-features --features rq-striped

Loom (model checking)

RUSTFLAGS="--cfg loom" cargo test --lib --release

Exhaustively explores interleavings of the slot state machine (src/slot_state.rs: lost-wakeup, at-most-once-enqueue, the stale-pid ABA theorem, unpark-vs-claim) and the ring queues (src/run_queue.rs: exactly-once through lap wraparound, push/pop races). Models run the production transitions through src/sync_shim.rs — std atomics normally, loom::sync under --cfg loom. RawMutex is deliberately not modeled: futexes can't be, and it's the textbook Drepper mutex3 with stress and unwind-safety tests of its own.

Trace feature

cargo test --features smarm-trace exists mainly to catch bit-rot in the te!() call sites; run it after touching scheduler paths.

Before a runtime-core PR

The full matrix, in rough order of bug-finding power per minute:

  1. cargo test (debug, default variant)
  2. debug under rq-mpmc and rq-striped
  3. cargo test --release
  4. loom
  5. cargo build --features smarm-trace

Catalog

Low-level units (no scheduler)

file covers
context.rs init_actor_stack + the naked-asm context-switch shims, poked directly
stack.rs the mmap'd stack allocator
pid.rs pid packing/equality

Feature areas (run under a real runtime)

file covers
runtime.rs Config, Runtime::run, re-running a runtime, correctness under genuine parallelism
scheduler.rs spawn / join / panic delivery / yield_now / self_pid
channel.rs send/recv (recv parks, so these need the runtime)
selective_recv.rs recv_match / try_recv_match
mutex.rs the actor-blocking Mutex<T> (lock parks)
timer.rs sleep ordering — time-sensitive, generous tolerances by design
io.rs block_on_io: blocking closures on the pool while the actor parks
io_epoll.rs wait_readable / wait_writable + the read/write sugar
preempt.rs explicit preemption via smarm::check!()
cancel.rs cooperative cancellation (request_stop) — the keystone semantics
monitor.rs monitor delivers exactly one Down; demonitor
link.rs bidirectional links + trap_exit
supervisor.rs one-for-one supervision
gen_server.rs call/cast round-trips, lifecycle callbacks, server-down detection

Regression & stress

file covers
stress.rs lost wakeups, pid-table pressure, thundering herds, panic isolation under concurrency. Where the phase-2 RefCell-migration bug was caught.
poison_stop.rs request_stop racing an alloc-under-lock must not poison/abort. See its header for the full story.
many_timers_multi_thread.rs multi-thread sleep-timer lost-wakeup regression

Conventions

  • Each test owns its runtime. init(Config::exact(N)) + rt.run(...); never share a Runtime between tests. Oversubscription (exact(4) on one core) is deliberate — forced interleaving at yield points is how single-core CI finds races at all.
  • Regression tests must be validated against the bug. A regression test that passes with the bug reintroduced is documentation, not a test. Reintroduce the fix's inverse locally and watch it fail before trusting it (poison_stop.rs went through exactly this: its first version never fired the sentinel under a lock, and was rewritten until it SIGABRT'd pre-fix).
  • Stochastic tests get the odds stacked. Use Config::alloc_interval(1) to make every allocation an observation point, many actors, and both phases of any every-other-allocation cadence (see poison_stop::self_stop_during_spawn...).
  • Time-based assertions use ordering, not durations. Assert "didn't return instantly" / "A woke before B", with generous tolerances; CI machines are slow and noisy.
  • New invariants added to the runtime should come with the assert at the point of reliance (debug_assert on hot paths) and, where the invariant is a protocol, a loom model in the owning module — that combination is what made phases 2–5 land without a single post-merge race so far.