allocate_slot() panics on a full slab; for a load-shedding caller (an accept loop spawning one actor per connection) that panic lands in the spawning actor, which then crash-loops under Restart::Transient into the still-full slab until its restart budget is spent — and the service stops accepting entirely. Observed live (urus slowloris scaling, 2026-08-10). A full slab is a routine overload condition for such callers, not an invariant violation. - RuntimeInner::try_allocate_slot() -> Option<u32>: the non-panicking core; a single pop under the free-list lock, so the claim is atomic (claim-or-report — no check-then-spawn TOCTOU, no headroom margin). allocate_slot() is now a thin panicking wrapper over it. - scheduler::try_spawn / try_spawn_under_with -> Result<JoinHandle, SpawnError>: parity with spawn/spawn_under_with except a full slab returns Err(SpawnError::AtCapacity) instead of panicking. Minimal surface per the agreed strategy; the remaining _with/_addr mirrors are trivial wrappers if ever needed. - Slot-first ordering on the try path (reverse of spawn's stack-first): under overload Err is the hot path, and a rejection costs one mutex pop — no mmap/pool-pop + init + recycle per shed unit of work. A drop-guard returns the claimed slot if stack allocation panics in the claim-to-install window (would otherwise leak and trip run()'s teardown slot-leak debug_assert). - SpawnError: non_exhaustive, Display + std::error::Error. - spawn and every existing call site untouched: the panic remains the correct loud invariant check at internal/bounded spawn sites. tests/try_spawn.rs: parity when slots free; exact slab accounting at capacity (Err, no panic, repeatable); custom-shape try refuses before stack allocation; self-heal after slots free; plain spawn still panics (surfaced via JoinError payload); 4-thread race for the last slots claims exactly the free count; SpawnError impl checks. Design doc: smarm-suggestion-try-spawn.md. Downstream consumer change (canned 503 on AtCapacity in urus's accept loop) is urus scope, not smarm. (cherry picked from commit 36de4b36aeaa72b2a5f9f3797b9854652656dcf6)
Tests
Integration tests for the runtime. Each file owns one feature area or one
class of bug. Everything here runs under plain cargo test; the loom model
tests are the exception — they live in the library (src/slot_state.rs,
src/run_queue.rs), not in this directory, because loom must compile the
production code with shimmed atomics (see "Loom" below).
Running
cargo test # debug build — RUN THIS ONE: all invariant asserts live
cargo test --release # what users actually execute (LTO, no debug_asserts)
Debug builds are not just "slower tests": the runtime self-checks its
invariants only there — every StateWord transition asserts its
precondition, enqueue asserts the exact (gen, Queued) word, RawMutex
enforces the never-two-cold-locks leaf rule with a per-thread held-count,
live_actors checks for double-finalize underflow. A green release run with
a red debug run means an invariant broke without (yet) corrupting behavior —
treat it as a real failure.
The queue-variant matrix
The run queue is compile-time selected; the suite must pass under all three (features are additive, so drop the default first):
cargo test # rq-mutex (default)
cargo test --no-default-features --features rq-mpmc
cargo test --no-default-features --features rq-striped
Loom (model checking)
RUSTFLAGS="--cfg loom" cargo test --lib --release
Exhaustively explores interleavings of the slot state machine
(src/slot_state.rs: lost-wakeup, at-most-once-enqueue, the stale-pid ABA
theorem, unpark-vs-claim) and the ring queues (src/run_queue.rs:
exactly-once through lap wraparound, push/pop races). Models run the
production transitions through src/sync_shim.rs — std atomics normally,
loom::sync under --cfg loom. RawMutex is deliberately not modeled:
futexes can't be, and it's the textbook Drepper mutex3 with stress and
unwind-safety tests of its own.
Trace feature
cargo test --features smarm-trace exists mainly to catch bit-rot in the
te!() call sites; run it after touching scheduler paths.
Before a runtime-core PR
The full matrix, in rough order of bug-finding power per minute:
cargo test(debug, default variant)- debug under
rq-mpmcandrq-striped cargo test --release- loom
cargo build --features smarm-trace
Catalog
Low-level units (no scheduler)
| file | covers |
|---|---|
context.rs |
init_actor_stack + the naked-asm context-switch shims, poked directly |
stack.rs |
the mmap'd stack allocator |
pid.rs |
pid packing/equality |
Feature areas (run under a real runtime)
| file | covers |
|---|---|
runtime.rs |
Config, Runtime::run, re-running a runtime, correctness under genuine parallelism |
scheduler.rs |
spawn / join / panic delivery / yield_now / self_pid |
channel.rs |
send/recv (recv parks, so these need the runtime) |
selective_recv.rs |
recv_match / try_recv_match |
mutex.rs |
the actor-blocking Mutex<T> (lock parks) |
timer.rs |
sleep ordering — time-sensitive, generous tolerances by design |
io.rs |
block_on_io: blocking closures on the pool while the actor parks |
io_epoll.rs |
wait_readable / wait_writable + the read/write sugar |
preempt.rs |
explicit preemption via smarm::check!() |
cancel.rs |
cooperative cancellation (request_stop) — the keystone semantics |
monitor.rs |
monitor delivers exactly one Down; demonitor |
link.rs |
bidirectional links + trap_exit |
supervisor.rs |
one-for-one supervision |
gen_server.rs |
call/cast round-trips, lifecycle callbacks, server-down detection |
Regression & stress
| file | covers |
|---|---|
stress.rs |
lost wakeups, pid-table pressure, thundering herds, panic isolation under concurrency. Where the phase-2 RefCell-migration bug was caught. |
poison_stop.rs |
request_stop racing an alloc-under-lock must not poison/abort. See its header for the full story. |
many_timers_multi_thread.rs |
multi-thread sleep-timer lost-wakeup regression |
Conventions
- Each test owns its runtime.
init(Config::exact(N))+rt.run(...); never share aRuntimebetween tests. Oversubscription (exact(4)on one core) is deliberate — forced interleaving at yield points is how single-core CI finds races at all. - Regression tests must be validated against the bug. A regression test
that passes with the bug reintroduced is documentation, not a test.
Reintroduce the fix's inverse locally and watch it fail before trusting it
(
poison_stop.rswent through exactly this: its first version never fired the sentinel under a lock, and was rewritten until it SIGABRT'd pre-fix). - Stochastic tests get the odds stacked. Use
Config::alloc_interval(1)to make every allocation an observation point, many actors, and both phases of any every-other-allocation cadence (seepoison_stop::self_stop_during_spawn...). - Time-based assertions use ordering, not durations. Assert "didn't return instantly" / "A woke before B", with generous tolerances; CI machines are slow and noisy.
- New invariants added to the runtime should come with the assert at the point of reliance (debug_assert on hot paths) and, where the invariant is a protocol, a loom model in the owning module — that combination is what made phases 2–5 land without a single post-merge race so far.