Nothing local is remotely reachable by default (RFC §4). expose(Name<M>)
marks a name remotely addressable and registers M's decoder under
type_hash::<M>(); expose_type::<M>() registers only the decoder (the
reply-to path). exposed_names() is the auditable remote surface.
D3's watchable fold, resolved against the code as it stands (stated in the
module docs): register() ALREADY stamps every named holder watchable ('no
successfully-registered actor can die unflagged', registry.rs), so an
exposed name's holder needs no extra mark — and re-registration after a
holder's death re-stamps the new holder for free, which a per-tenancy mark
taken at expose time could not do. The cluster's own mark_watchable
set-site is therefore the pid crossing the wire (frame serialization, c10)
— the exact analog of the membrane crossing. c8 adds only the name/type
state neither the registry nor slot bits can carry. No new pid registry;
RFC §4 honored.
One-viable calls, flagged:
- State lives on RuntimeInner (the pg pattern: leaf RawMutex field,
cfg-gated behind cluster, zero-cost-when-off per c1) — c9's inbound
decode consults it per frame; manager-held state would serialize every
remote delivery through one gen_server.
- type_hash = FNV-1a 64 (fixed seed: the offset basis) over TypeId: a
constant of the binary — stable across runs of the same build (the scope
the build-hash handshake reduces the mesh to), deliberately not across
builds. Collisions degrade to decode error / refused channel, never a
misroute (the NoChannel guarantee, RFC §3).
- Decoder = decode-and-deliver-to-pid Arc closure capturing M (the one
typed site): decode_payload then send_dyn. Wire-name → pid resolution
stays OUTSIDE — that is c9's single seam, which calls decode_deliver.
Arc so the call happens with the exposure lock RELEASED: send_dyn takes
the registry lock, a mutual Leaf (the runtime asserts on nesting — caught
live by the first test run).
- expose is a name-level fact, valid for an unregistered name (names
late-bind; c9 resolves per delivery).
tests/cluster_expose.rs 5/0 stable x5, purely local per roadmap:
exposed/unexposed lookup + audit listing; decoder registration and the
delivery contract (happy path into a registered String channel; unknown
hash; corrupt bytes; wrong channel refused — never misrouted); distinct
types distinct hashes; expose/bridge-crossing agreement via the shared
watchable observable (terminal_reason after holder death); hash stability
across runs in the same binary via a c4-harness re-exec. Payload types are
std types — the crate's serde is derive-less by design, user crates bring
their own derive.
All cluster suites regression-clean (envelope 15, handshake 11, transport
11, lifecycle 1, liveness 3, connect 9, two_node 3, membership 4, mesh 2);
clippy --lib green both configs; fmt clean; default build compiles.
Tests
Integration tests for the runtime. Each file owns one feature area or one
class of bug. Everything here runs under plain cargo test; the loom model
tests are the exception — they live in the library (src/slot_state.rs,
src/run_queue.rs), not in this directory, because loom must compile the
production code with shimmed atomics (see "Loom" below).
Running
cargo test # debug build — RUN THIS ONE: all invariant asserts live
cargo test --release # what users actually execute (LTO, no debug_asserts)
Debug builds are not just "slower tests": the runtime self-checks its
invariants only there — every StateWord transition asserts its
precondition, enqueue asserts the exact (gen, Queued) word, RawMutex
enforces the never-two-cold-locks leaf rule with a per-thread held-count,
live_actors checks for double-finalize underflow. A green release run with
a red debug run means an invariant broke without (yet) corrupting behavior —
treat it as a real failure.
The queue-variant matrix
The run queue is compile-time selected; the suite must pass under all three (features are additive, so drop the default first):
cargo test # rq-mutex (default)
cargo test --no-default-features --features rq-mpmc
cargo test --no-default-features --features rq-striped
Loom (model checking)
RUSTFLAGS="--cfg loom" cargo test --lib --release
Exhaustively explores interleavings of the slot state machine
(src/slot_state.rs: lost-wakeup, at-most-once-enqueue, the stale-pid ABA
theorem, unpark-vs-claim) and the ring queues (src/run_queue.rs:
exactly-once through lap wraparound, push/pop races). Models run the
production transitions through src/sync_shim.rs — std atomics normally,
loom::sync under --cfg loom. RawMutex is deliberately not modeled:
futexes can't be, and it's the textbook Drepper mutex3 with stress and
unwind-safety tests of its own.
Trace feature
cargo test --features smarm-trace exists mainly to catch bit-rot in the
te!() call sites; run it after touching scheduler paths.
Before a runtime-core PR
The full matrix, in rough order of bug-finding power per minute:
cargo test(debug, default variant)- debug under
rq-mpmcandrq-striped cargo test --release- loom
cargo build --features smarm-trace
Catalog
Low-level units (no scheduler)
| file | covers |
|---|---|
context.rs |
init_actor_stack + the naked-asm context-switch shims, poked directly |
stack.rs |
the mmap'd stack allocator |
pid.rs |
pid packing/equality |
Feature areas (run under a real runtime)
| file | covers |
|---|---|
runtime.rs |
Config, Runtime::run, re-running a runtime, correctness under genuine parallelism |
scheduler.rs |
spawn / join / panic delivery / yield_now / self_pid |
channel.rs |
send/recv (recv parks, so these need the runtime) |
selective_recv.rs |
recv_match / try_recv_match |
mutex.rs |
the actor-blocking Mutex<T> (lock parks) |
timer.rs |
sleep ordering — time-sensitive, generous tolerances by design |
io.rs |
block_on_io: blocking closures on the pool while the actor parks |
io_epoll.rs |
wait_readable / wait_writable + the read/write sugar |
preempt.rs |
explicit preemption via smarm::check!() |
cancel.rs |
cooperative cancellation (request_stop) — the keystone semantics |
monitor.rs |
monitor delivers exactly one Down; demonitor |
link.rs |
bidirectional links + trap_exit |
supervisor.rs |
one-for-one supervision |
gen_server.rs |
call/cast round-trips, lifecycle callbacks, server-down detection |
Regression & stress
| file | covers |
|---|---|
stress.rs |
lost wakeups, pid-table pressure, thundering herds, panic isolation under concurrency. Where the phase-2 RefCell-migration bug was caught. |
poison_stop.rs |
request_stop racing an alloc-under-lock must not poison/abort. See its header for the full story. |
many_timers_multi_thread.rs |
multi-thread sleep-timer lost-wakeup regression |
Conventions
- Each test owns its runtime.
init(Config::exact(N))+rt.run(...); never share aRuntimebetween tests. Oversubscription (exact(4)on one core) is deliberate — forced interleaving at yield points is how single-core CI finds races at all. - Regression tests must be validated against the bug. A regression test
that passes with the bug reintroduced is documentation, not a test.
Reintroduce the fix's inverse locally and watch it fail before trusting it
(
poison_stop.rswent through exactly this: its first version never fired the sentinel under a lock, and was rewritten until it SIGABRT'd pre-fix). - Stochastic tests get the odds stacked. Use
Config::alloc_interval(1)to make every allocation an observation point, many actors, and both phases of any every-other-allocation cadence (seepoison_stop::self_stop_during_spawn...). - Time-based assertions use ordering, not durations. Assert "didn't return instantly" / "A woke before B", with generous tolerances; CI machines are slow and noisy.
- New invariants added to the runtime should come with the assert at the point of reliance (debug_assert on hot paths) and, where the invariant is a protocol, a loom model in the owning module — that combination is what made phases 2–5 land without a single post-merge race so far.