- src/signal.rs: process-global SA_SIGINFO|SA_ONSTACK handler installed once
at runtime::init (before any scheduler thread -> unracing PRIOR save);
per-scheduler-thread 64 KiB sigaltstack registered at schedule_loop entry
(a guard hit leaves no stack to handle on). Async-signal-safe throughout:
classification is plain loads (const-init TLS Cell + slot atomics), print
is fixed-buffer itoa + one write(2), death is SIG_DFL + refault at the
same instruction (core-dumpable, correct wait status).
- Two-tier classification (agreed): in-guard = definitive; OVERSHOOT window
below the guard = 'unprobed (FFI?) frame stepped over it' probable
attribution -- the RFC's motivating incident (cargo-vendored gz, not
SQLite as the RFC text says) faults there under a small guard. Pure
classify() fn, 5 adversarial units incl. saturation at low addresses.
- DEFAULT_STACK_GUARD 64 KiB -> 1 MiB (agreed): kernel stack_guard_gap
anchor post-Stack-Clash; PROT_NONE is VA-only (no RSS, no page tables,
no overcommit charge) so width is free at any actor count.
- Unclassified faults reinstate the PRIOR sigaction and refault (agreed):
std's own OS-thread overflow diagnostics survive our presence.
- Slot: diag_{stack_top,stack_reserve,stack_guard,pid} atomics written in
install_actor pre-publish; readable without the cold lock (Stack lives
under it); only consulted while CURRENT_SLOT points at the slot, so
never stale where read. preempt::current_slot_ptr ungated from
smarm-causal (now also the classifier's anchor).
- build.rs + cc (agreed Q3): canary/canary.c, 96 KiB local touched low-end
first, -fno-stack-clash-protection pinned so hardened toolchains don't
probe the canary into uselessness.
- tests/stack_diag.rs: subprocess x4 -- Rust recursion tier-1; FFI canary
tier-1 at defaults (1 MiB guard catches the jump); tier-2 at guard=4 KiB
('stepped over', reproduces the incident); clean at reserve=256 KiB
(the §1 knob is the fix, same frame).
FLAGGED (Claude-solo calls):
- OVERSHOOT_SLOP = 1 MiB (matches guard default/kernel gap; beyond it
attribution would be dishonest).
- Altstack 64 KiB, mmap'd once per OS thread, never freed (bounded by
thread count; reused across run()s via TLS flag).
- Foreign-fault reinstate permanently deregisters our handler; accepted --
the process is dying either way.
- Diag geometry as 4 slot atomics (install-time cost only) over a per-switch
TLS snapshot (hot-path stores).
Tests
Integration tests for the runtime. Each file owns one feature area or one
class of bug. Everything here runs under plain cargo test; the loom model
tests are the exception — they live in the library (src/slot_state.rs,
src/run_queue.rs), not in this directory, because loom must compile the
production code with shimmed atomics (see "Loom" below).
Running
cargo test # debug build — RUN THIS ONE: all invariant asserts live
cargo test --release # what users actually execute (LTO, no debug_asserts)
Debug builds are not just "slower tests": the runtime self-checks its
invariants only there — every StateWord transition asserts its
precondition, enqueue asserts the exact (gen, Queued) word, RawMutex
enforces the never-two-cold-locks leaf rule with a per-thread held-count,
live_actors checks for double-finalize underflow. A green release run with
a red debug run means an invariant broke without (yet) corrupting behavior —
treat it as a real failure.
The queue-variant matrix
The run queue is compile-time selected; the suite must pass under all three (features are additive, so drop the default first):
cargo test # rq-mutex (default)
cargo test --no-default-features --features rq-mpmc
cargo test --no-default-features --features rq-striped
Loom (model checking)
RUSTFLAGS="--cfg loom" cargo test --lib --release
Exhaustively explores interleavings of the slot state machine
(src/slot_state.rs: lost-wakeup, at-most-once-enqueue, the stale-pid ABA
theorem, unpark-vs-claim) and the ring queues (src/run_queue.rs:
exactly-once through lap wraparound, push/pop races). Models run the
production transitions through src/sync_shim.rs — std atomics normally,
loom::sync under --cfg loom. RawMutex is deliberately not modeled:
futexes can't be, and it's the textbook Drepper mutex3 with stress and
unwind-safety tests of its own.
Trace feature
cargo test --features smarm-trace exists mainly to catch bit-rot in the
te!() call sites; run it after touching scheduler paths.
Before a runtime-core PR
The full matrix, in rough order of bug-finding power per minute:
cargo test(debug, default variant)- debug under
rq-mpmcandrq-striped cargo test --release- loom
cargo build --features smarm-trace
Catalog
Low-level units (no scheduler)
| file | covers |
|---|---|
context.rs |
init_actor_stack + the naked-asm context-switch shims, poked directly |
stack.rs |
the mmap'd stack allocator |
pid.rs |
pid packing/equality |
Feature areas (run under a real runtime)
| file | covers |
|---|---|
runtime.rs |
Config, Runtime::run, re-running a runtime, correctness under genuine parallelism |
scheduler.rs |
spawn / join / panic delivery / yield_now / self_pid |
channel.rs |
send/recv (recv parks, so these need the runtime) |
selective_recv.rs |
recv_match / try_recv_match |
mutex.rs |
the actor-blocking Mutex<T> (lock parks) |
timer.rs |
sleep ordering — time-sensitive, generous tolerances by design |
io.rs |
block_on_io: blocking closures on the pool while the actor parks |
io_epoll.rs |
wait_readable / wait_writable + the read/write sugar |
preempt.rs |
explicit preemption via smarm::check!() |
cancel.rs |
cooperative cancellation (request_stop) — the keystone semantics |
monitor.rs |
monitor delivers exactly one Down; demonitor |
link.rs |
bidirectional links + trap_exit |
supervisor.rs |
one-for-one supervision |
gen_server.rs |
call/cast round-trips, lifecycle callbacks, server-down detection |
Regression & stress
| file | covers |
|---|---|
stress.rs |
lost wakeups, pid-table pressure, thundering herds, panic isolation under concurrency. Where the phase-2 RefCell-migration bug was caught. |
poison_stop.rs |
request_stop racing an alloc-under-lock must not poison/abort. See its header for the full story. |
many_timers_multi_thread.rs |
multi-thread sleep-timer lost-wakeup regression |
Conventions
- Each test owns its runtime.
init(Config::exact(N))+rt.run(...); never share aRuntimebetween tests. Oversubscription (exact(4)on one core) is deliberate — forced interleaving at yield points is how single-core CI finds races at all. - Regression tests must be validated against the bug. A regression test
that passes with the bug reintroduced is documentation, not a test.
Reintroduce the fix's inverse locally and watch it fail before trusting it
(
poison_stop.rswent through exactly this: its first version never fired the sentinel under a lock, and was rewritten until it SIGABRT'd pre-fix). - Stochastic tests get the odds stacked. Use
Config::alloc_interval(1)to make every allocation an observation point, many actors, and both phases of any every-other-allocation cadence (seepoison_stop::self_stop_during_spawn...). - Time-based assertions use ordering, not durations. Assert "didn't return instantly" / "A woke before B", with generous tolerances; CI machines are slow and noisy.
- New invariants added to the runtime should come with the assert at the point of reliance (debug_assert on hot paths) and, where the invariant is a protocol, a loom model in the owning module — that combination is what made phases 2–5 land without a single post-merge race so far.