Files
smarm/tests
Claude (sandbox)andClaude (sandbox) ca1c98336e feat(scheduler,runtime): non-panicking try_spawn for at-capacity load shedding
allocate_slot() panics on a full slab; for a load-shedding caller (an
accept loop spawning one actor per connection) that panic lands in the
spawning actor, which then crash-loops under Restart::Transient into the
still-full slab until its restart budget is spent — and the service stops
accepting entirely. Observed live (urus slowloris scaling, 2026-08-10).
A full slab is a routine overload condition for such callers, not an
invariant violation.

- RuntimeInner::try_allocate_slot() -> Option<u32>: the non-panicking
  core; a single pop under the free-list lock, so the claim is atomic
  (claim-or-report — no check-then-spawn TOCTOU, no headroom margin).
  allocate_slot() is now a thin panicking wrapper over it.
- scheduler::try_spawn / try_spawn_under_with -> Result<JoinHandle,
  SpawnError>: parity with spawn/spawn_under_with except a full slab
  returns Err(SpawnError::AtCapacity) instead of panicking. Minimal
  surface per the agreed strategy; the remaining _with/_addr mirrors are
  trivial wrappers if ever needed.
- Slot-first ordering on the try path (reverse of spawn's stack-first):
  under overload Err is the hot path, and a rejection costs one mutex
  pop — no mmap/pool-pop + init + recycle per shed unit of work. A
  drop-guard returns the claimed slot if stack allocation panics in the
  claim-to-install window (would otherwise leak and trip run()'s
  teardown slot-leak debug_assert).
- SpawnError: non_exhaustive, Display + std::error::Error.
- spawn and every existing call site untouched: the panic remains the
  correct loud invariant check at internal/bounded spawn sites.

tests/try_spawn.rs: parity when slots free; exact slab accounting at
capacity (Err, no panic, repeatable); custom-shape try refuses before
stack allocation; self-heal after slots free; plain spawn still panics
(surfaced via JoinError payload); 4-thread race for the last slots
claims exactly the free count; SpawnError impl checks.

Design doc: smarm-suggestion-try-spawn.md. Downstream consumer change
(canned 503 on AtCapacity in urus's accept loop) is urus scope, not
smarm.

(cherry picked from commit 36de4b36aeaa72b2a5f9f3797b9854652656dcf6)
2026-08-13 15:03:16 +02:00
..
2026-05-23 16:09:35 +00:00

Tests

Integration tests for the runtime. Each file owns one feature area or one class of bug. Everything here runs under plain cargo test; the loom model tests are the exception — they live in the library (src/slot_state.rs, src/run_queue.rs), not in this directory, because loom must compile the production code with shimmed atomics (see "Loom" below).

Running

cargo test                # debug build — RUN THIS ONE: all invariant asserts live
cargo test --release      # what users actually execute (LTO, no debug_asserts)

Debug builds are not just "slower tests": the runtime self-checks its invariants only there — every StateWord transition asserts its precondition, enqueue asserts the exact (gen, Queued) word, RawMutex enforces the never-two-cold-locks leaf rule with a per-thread held-count, live_actors checks for double-finalize underflow. A green release run with a red debug run means an invariant broke without (yet) corrupting behavior — treat it as a real failure.

The queue-variant matrix

The run queue is compile-time selected; the suite must pass under all three (features are additive, so drop the default first):

cargo test                                              # rq-mutex (default)
cargo test --no-default-features --features rq-mpmc
cargo test --no-default-features --features rq-striped

Loom (model checking)

RUSTFLAGS="--cfg loom" cargo test --lib --release

Exhaustively explores interleavings of the slot state machine (src/slot_state.rs: lost-wakeup, at-most-once-enqueue, the stale-pid ABA theorem, unpark-vs-claim) and the ring queues (src/run_queue.rs: exactly-once through lap wraparound, push/pop races). Models run the production transitions through src/sync_shim.rs — std atomics normally, loom::sync under --cfg loom. RawMutex is deliberately not modeled: futexes can't be, and it's the textbook Drepper mutex3 with stress and unwind-safety tests of its own.

Trace feature

cargo test --features smarm-trace exists mainly to catch bit-rot in the te!() call sites; run it after touching scheduler paths.

Before a runtime-core PR

The full matrix, in rough order of bug-finding power per minute:

  1. cargo test (debug, default variant)
  2. debug under rq-mpmc and rq-striped
  3. cargo test --release
  4. loom
  5. cargo build --features smarm-trace

Catalog

Low-level units (no scheduler)

file covers
context.rs init_actor_stack + the naked-asm context-switch shims, poked directly
stack.rs the mmap'd stack allocator
pid.rs pid packing/equality

Feature areas (run under a real runtime)

file covers
runtime.rs Config, Runtime::run, re-running a runtime, correctness under genuine parallelism
scheduler.rs spawn / join / panic delivery / yield_now / self_pid
channel.rs send/recv (recv parks, so these need the runtime)
selective_recv.rs recv_match / try_recv_match
mutex.rs the actor-blocking Mutex<T> (lock parks)
timer.rs sleep ordering — time-sensitive, generous tolerances by design
io.rs block_on_io: blocking closures on the pool while the actor parks
io_epoll.rs wait_readable / wait_writable + the read/write sugar
preempt.rs explicit preemption via smarm::check!()
cancel.rs cooperative cancellation (request_stop) — the keystone semantics
monitor.rs monitor delivers exactly one Down; demonitor
link.rs bidirectional links + trap_exit
supervisor.rs one-for-one supervision
gen_server.rs call/cast round-trips, lifecycle callbacks, server-down detection

Regression & stress

file covers
stress.rs lost wakeups, pid-table pressure, thundering herds, panic isolation under concurrency. Where the phase-2 RefCell-migration bug was caught.
poison_stop.rs request_stop racing an alloc-under-lock must not poison/abort. See its header for the full story.
many_timers_multi_thread.rs multi-thread sleep-timer lost-wakeup regression

Conventions

  • Each test owns its runtime. init(Config::exact(N)) + rt.run(...); never share a Runtime between tests. Oversubscription (exact(4) on one core) is deliberate — forced interleaving at yield points is how single-core CI finds races at all.
  • Regression tests must be validated against the bug. A regression test that passes with the bug reintroduced is documentation, not a test. Reintroduce the fix's inverse locally and watch it fail before trusting it (poison_stop.rs went through exactly this: its first version never fired the sentinel under a lock, and was rewritten until it SIGABRT'd pre-fix).
  • Stochastic tests get the odds stacked. Use Config::alloc_interval(1) to make every allocation an observation point, many actors, and both phases of any every-other-allocation cadence (see poison_stop::self_stop_during_spawn...).
  • Time-based assertions use ordering, not durations. Assert "didn't return instantly" / "A woke before B", with generous tolerances; CI machines are slow and noisy.
  • New invariants added to the runtime should come with the assert at the point of reliance (debug_assert on hot paths) and, where the invariant is a protocol, a loom model in the owning module — that combination is what made phases 2–5 land without a single post-merge race so far.