8 Commits
Author SHA1 Message Date
Claude (sandbox) 301e3463e3 chore(release): v0.6.0 — RFC 019: actor stack reserve & shrink
Per-actor stack shapes on every spawn surface (SpawnOpts stack_reserve/
guard_size, Config defaults, pool rule: only default-shaped recycle);
sampled stack high-water + MADV_FREE shrink at actor-park (THRESHOLD
256 KiB, COOLDOWN 64 parks, redzone 1 page); pool-recycle MADV_DONTNEED
above the retained 64 KiB entry end; SIGSEGV overflow diagnostics
(two-tier: in-guard definitive / 1 MiB overshoot 'stepped over', prior
handler chained for foreign faults) with per-scheduler sigaltstack; and
the per-actor introspection surface (ActorInfo.stack: reserve, guard,
sampled depth, parks_since_shrink, shrinks).

Amendments ratified during implementation, for the RFC changelog:
- DEFAULT_STACK_GUARD 64 KiB -> 1 MiB, following the kernel's post-Stack-
  Clash stack_guard_gap convention; PROT_NONE width is VA-only and free.
- §7's motivating segfault was a cargo-vendored gz build, not SQLite as
  the RFC text says (cc-built C lacks -fstack-clash-protection; distro
  libraries have it — the risky class is vendored builds).
- §4 hibernate() deferred to the jar (bolt-on: force-flag on the §3
  shrink path, ~10 lines when wanted).

Gates (jobrunner box, 2026-08-08): reclaim gate PASS at c3 and again at
tip (3.0 MiB LazyFree -> kernel reclaim -> Rss to one live page ->
re-spike bit-identical, live data intact; MADV_PAGEOUT stands in for
memcg — cgroup2 is RO in the job container — driving the same reclaim
path). E1 interleaved A/B vs v0.5.0: every ka cell (the E1 subject)
within +0.3..+2.9% at tip; close-mode control cells within noise except
t8-c4 close, which is bistable (~40-44k vs ~46-49k modes for BOTH
variants, base self-disagrees by 11% across rounds); 6 rounds across two
runs are inconclusive there and a 10-round focused run is noted in the
handoff as deferred follow-up, accepted for this release.

No breaking API changes since v0.5.0: SpawnOpts fields and ActorInfo
gained members (exhaustive-construction downstream will need the new
ActorInfo.stack field; urus does not construct it).
2026-08-08 19:44:55 +00:00
Claude (sandbox) 410ba33d82 feat(introspect,runtime): per-actor stack surface on ActorInfo (RFC 019 §8)
- introspect::StackInfo { reserve, guard, depth_high_water,
  parks_since_shrink, shrinks } as ActorInfo.stack; re-exported at crate
  root beside ActorInfo.
- All reads lock-free: geometry from the c6 diag slot atomics, depth =
  top - hwm (the §2 sampled high-water; doc spells out sampled-not-exact
  and that 0 means never-descheduled-at-depth), counters straight off the
  §3 atomics. Coherence for the incarnation rides read_slot's existing
  generation check, same as overruns/messages_received.
- Slot::stack_introspect(): one pub(crate) tuple accessor beside the other
  counter accessors.
- Exact RSS deliberately absent per RFC (mincore = debug tooling only,
  never a runtime path); stack_shape(pid) untouched (cold-lock exact
  variant from c2).
- tests/introspect.rs: defaults surface (64 KiB reserve / 1 MiB guard /
  sampled ~32 KiB depth / gate park counted / zero shrinks) + live shrink
  counters (spike visible pre-shrink; shrinks>=1, cooldown counter reset,
  hwm reset after crossing COOLDOWN) read mid-run -- post-join the slot
  reclaim correctly hides the incarnation, which the first draft of the
  test learned the hard way.

FLAGGED (Claude-solo calls):
- Nested StackInfo struct over five flat ActorInfo fields (grain break;
  the five fields are one concern and ActorInfo is already 12 fields).
- Field names reserve/guard/shrinks (RFC says stack_reserve/stack_guard/
  shrink count; the stack_ prefix is redundant inside StackInfo).
2026-08-08 19:12:46 +00:00
Claude (sandbox) 5fd8aecf55 feat(signal,runtime,stack): SIGSEGV overflow diagnostics + 1 MiB guard default (RFC 019 §7)
- src/signal.rs: process-global SA_SIGINFO|SA_ONSTACK handler installed once
  at runtime::init (before any scheduler thread -> unracing PRIOR save);
  per-scheduler-thread 64 KiB sigaltstack registered at schedule_loop entry
  (a guard hit leaves no stack to handle on). Async-signal-safe throughout:
  classification is plain loads (const-init TLS Cell + slot atomics), print
  is fixed-buffer itoa + one write(2), death is SIG_DFL + refault at the
  same instruction (core-dumpable, correct wait status).
- Two-tier classification (agreed): in-guard = definitive; OVERSHOOT window
  below the guard = 'unprobed (FFI?) frame stepped over it' probable
  attribution -- the RFC's motivating incident (cargo-vendored gz, not
  SQLite as the RFC text says) faults there under a small guard. Pure
  classify() fn, 5 adversarial units incl. saturation at low addresses.
- DEFAULT_STACK_GUARD 64 KiB -> 1 MiB (agreed): kernel stack_guard_gap
  anchor post-Stack-Clash; PROT_NONE is VA-only (no RSS, no page tables,
  no overcommit charge) so width is free at any actor count.
- Unclassified faults reinstate the PRIOR sigaction and refault (agreed):
  std's own OS-thread overflow diagnostics survive our presence.
- Slot: diag_{stack_top,stack_reserve,stack_guard,pid} atomics written in
  install_actor pre-publish; readable without the cold lock (Stack lives
  under it); only consulted while CURRENT_SLOT points at the slot, so
  never stale where read. preempt::current_slot_ptr ungated from
  smarm-causal (now also the classifier's anchor).
- build.rs + cc (agreed Q3): canary/canary.c, 96 KiB local touched low-end
  first, -fno-stack-clash-protection pinned so hardened toolchains don't
  probe the canary into uselessness.
- tests/stack_diag.rs: subprocess x4 -- Rust recursion tier-1; FFI canary
  tier-1 at defaults (1 MiB guard catches the jump); tier-2 at guard=4 KiB
  ('stepped over', reproduces the incident); clean at reserve=256 KiB
  (the §1 knob is the fix, same frame).

FLAGGED (Claude-solo calls):
- OVERSHOOT_SLOP = 1 MiB (matches guard default/kernel gap; beyond it
  attribution would be dishonest).
- Altstack 64 KiB, mmap'd once per OS thread, never freed (bounded by
  thread count; reused across run()s via TLS flag).
- Foreign-fault reinstate permanently deregisters our handler; accepted --
  the process is dying either way.
- Diag geometry as 4 slot atomics (install-time cost only) over a per-switch
  TLS snapshot (hot-path stores).
2026-08-08 18:58:30 +00:00
Claude (sandbox) 7d8b9e0310 feat(stack,runtime): pool-recycle DONTNEED above the retained entry end (RFC 019 §6)
- stack::retain_range: pure checked span fn (retain page-up = zap less;
  None when retain covers the reserve, so the 64 KiB default config never
  pays a syscall) + 6 adversarial units mirroring shrink_range's.
- Stack::recycle_zap: advisory MADV_DONTNEED of [usable_base, top-RETAIN);
  stack is unowned at the call site, synchronous eager zap races nothing.
- recycle_stack: zap OFF-LOCK before pool admission (acquire_stack's
  no-syscall-under-the-pool-lock invariant); rare cap-overflow pays a
  wasted zap ahead of munmap, accepted over a second lock round-trip.
- pub const RECYCLE_RETAIN = 64 KiB beside the shrink knobs, ratified-as-
  constant rationale in doc.
- tests/stack_recycle.rs: mincore-based exact-zero-resident assert over
  the zap span. smaps was tried first and over-counts: a neighboring rw
  anon VMA can merge flush against the stack top (observed once under the
  full-suite run); the PROT_NONE guard pins the usable base exactly.

FLAGGED (Claude-solo calls):
- RFC §6 'above the bottom RETAIN' is direction-ambiguous in address
  terms; implemented as retain the ENTRY end (highest addresses, the
  pages the next actor faults first), zap the cold deep span below.
- Const named RECYCLE_RETAIN (RFC says RETAIN) to sit beside SHRINK_*.
2026-08-08 16:13:53 +00:00
Claude (sandbox) 8225716b11 feat(runtime,stack): sampled stack high-water + MADV_FREE shrink at actor-park (RFC 019 §§2–3)
hwm: AtomicUsize lands beside sp on the slot: the single context-save
site min-updates it (one branch + at most one Relaxed store into the
line the sp store just dirtied), install resets it to the fresh top.
Advisory by construction — correctness never depends on it. The mod-doc
ordering chain gains a line: hwm piggybacks the existing
Relaxed-store-before-Release pattern and adds no edges.

Shrink hook in the YieldIntent::Park arm only, before the park_return
Release transition — the owned window (obligation 1's assert-comment at
the site): after the sp store, before Parked is published, scheduler on
its own stack, actor saved and unstealable. It runs on both arms of the
park_return race (a consumed unpark flag means one wasted-but-harmless
madvise). The preempt/yield path deliberately never checks: §4's
bounded, self-healing leak under saturation, when syscalls are least
affordable.

SHRINK_THRESHOLD = 256 KiB and SHRINK_COOLDOWN = 64 parks are pub
constants with the ratified doc rationale, not Config fields. The freed
span is shrink_range(hwm, sp, page): whole pages of [hwm, sp − 1-page
redzone), rounded inward, checked arithmetic — adversarial inputs
collapse to None (obligation 2). MADV_FREE marks lazily; the kernel's
reclaim-under-pressure IS the hysteresis, cancel-on-write is the safety
net. parks_since_shrink + shrink_count ride the slot for the cooldown
and the future introspect surface.

Tests: 7 adversarial shrink_range units (inverted/empty spans, redzone
underflow, unaligned ends, sp-crossing sweep); integration — 8 MiB
reserve, ~3 MiB spike sampled via yield-at-depth, parks gated on
introspected Parked state past the cooldown, then ≥ 2 MiB LazyFree
asserted inside the stack's smaps range with live data intact; and the
inverse guard — a shallow never-spiking actor ends at exactly 0
LazyFree (also proves the parser isn't vacuously zero via the first
test).
2026-08-08 14:30:32 +00:00
Claude (sandbox) 3cb64eefc2 feat(scheduler,gen_server,gen_statem,introspect): SpawnOpts — per-actor stack shape on every spawn surface (RFC 019 §1)
SpawnOpts { stack_reserve, guard_size } with Option<usize> fields, None
resolving to the Config defaults at spawn time — a deliberate deviation
from the RFC's plain-usize struct so struct-update syntax works without
a runtime handle in scope. Threaded across the five surfaces:
spawn_with, spawn_under_with, spawn_addr_with,
GenServerBuilder::stack_opts (mirrored on NamedGenServerBuilder), and
gen_statem::spawn_with (gen_statem has no builder, so the opts ride a
_with variant — Claude-solo surface call, flagged for review). Existing
spawns forward defaults; no call-site churn.

introspect::stack_shape(pid) pulled forward (agreed) as the first slice
of the RFC 019 introspection surface, giving tests an observable.

Tests (tests/spawn_opts.rs): override/partial-override/rounding on each
surface; obligation 4 from the outside — a dead custom stack is never
handed to the next default spawn (LIFO pool would expose it), and the
reverse (default stacks ARE recycled); 8 MiB reserve behaviorally
permits ~1 MiB recursion. Also: silence unused-Result in the c1
runtime test (join now unwrapped).
2026-08-08 14:22:38 +00:00
Claude (sandbox) 0fe052bc7e feat(stack,runtime): per-shape actor stacks — Stack::new(reserve, guard), Config knobs, pool rule (RFC 019 §1)
Stack takes an explicit (reserve, guard) shape, both page-rounded and
stored; usable_base derives from the stored guard. Guard default raised
4 KiB -> 64 KiB (DEFAULT_STACK_GUARD): probestack makes one page enough
for Rust frames, but an unprobed C frame can leap a page in one sub rsp
— the motivating SQLite segfault. Reserve default stays 64 KiB
(DEFAULT_STACK_RESERVE); ACTOR_STACK_SIZE retired.

Config::{stack_reserve, stack_guard} thread the runtime defaults into
RuntimeInner pre-rounded. All acquisition/recycling now goes through
acquire_stack/recycle_stack carrying the pool rule: only default-shaped
stacks are pooled (pooled ⇒ default-shaped by induction); custom shapes
mmap fresh and munmap at death. Pool lock still dropped before any mmap.

No public spawn API change (SpawnOpts is the next commit).

Tests: shape rounding + accessors, wide-guard faults at both ends
(subprocess), Config::stack_reserve permits >64 KiB recursion that
previously could only segfault.
2026-08-08 14:18:27 +00:00
smarm a03a7ca01e chore(release): v0.5.0
Breaking API rename since v0.4.0: gen_server's ServerRef/ServerBuilder/
ServerCtx -> GenServerRef/GenServerBuilder/GenServerCtx, Watcher<G> is
now generic over its GenServer, and GenServer gained a required
associated Timer type for timer-fire payloads (arm_after/handle_timer).
Downstream consumers (urus) have been ported.
2026-08-08 11:44:35 +02:00
21 changed files with 1894 additions and 63 deletions
+4 -1
View File
@@ -1,6 +1,6 @@
[package]
name = "smarm"
version = "0.4.0"
version = "0.6.0"
edition = "2021"
rust-version = "1.95"
@@ -39,6 +39,9 @@ rq-mutex = []
rq-mpmc = []
rq-striped = []
[build-dependencies]
cc = "1"
[dependencies]
libc = "0.2"
+11
View File
@@ -0,0 +1,11 @@
fn main() {
// RFC 019 §7 test canary (agreed Q3): compiled without stack-clash
// protection so its 96 KiB local is a genuine one-displacement guard
// jumper; distro-hardened compilers would otherwise probe it page-wise
// and defeat the test's purpose.
cc::Build::new()
.file("canary/canary.c")
.flag_if_supported("-fno-stack-clash-protection")
.compile("smarm_canary");
println!("cargo:rerun-if-changed=canary/canary.c");
}
+14
View File
@@ -0,0 +1,14 @@
/* RFC 019 §7 FFI canary: an honest unprobed C frame with a 96 KiB local,
* touched from its LOW end first — the exact "one sub rsp steps over a small
* guard" pattern the RFC's motivating incident hit (a cargo-vendored gz
* build; cc-invoked builds do not enable -fstack-clash-protection, and this
* file pins that off explicitly so the canary stays a canary even on
* hardened-default toolchains). */
void smarm_canary_burn(void) {
volatile char buf[96 * 1024];
buf[0] = 1; /* deepest address first */
for (unsigned i = 0; i < sizeof buf; i += 4096) {
buf[i] = (char)i;
}
buf[sizeof buf - 1] = 1;
}
+1 -1
View File
@@ -75,7 +75,7 @@ genuine advantage over tokio's task abort model.
### Spawn-heavy workloads (19–70×)
Every smarm actor `mmap`s a 64 KiB stack with a guard page. This is
Every smarm actor `mmap`s a 64 KiB stack reserve with a 64 KiB PROT_NONE guard below (both per-actor configurable since RFC 019; the reserve is demand-paged). This is
a syscall. Tokio tasks are heap-allocated state machines — no stack,
no syscall, ~100 bytes each. For workloads that spawn thousands of
short-lived actors per second, this is a structural disadvantage.
+31 -5
View File
@@ -182,7 +182,7 @@ use crate::channel::{channel, select, select_timeout, Receiver, RecvTimeoutError
use crate::monitor::{demonitor, monitor, Down, Monitor};
use crate::pid::Pid;
use crate::registry::{register_with, resolve_named_sender, RegisterError};
use crate::scheduler::{cancel_timer, request_stop, send_after_to, spawn, spawn_under};
use crate::scheduler::{cancel_timer, request_stop, send_after_to};
use crate::timer::TimerId;
use std::cell::Cell;
use std::collections::HashMap;
@@ -643,11 +643,17 @@ pub struct GenServerBuilder<G: GenServer> {
state: G,
infos: Vec<Receiver<G::Info>>,
supervisor: Option<Pid>,
stack_opts: crate::scheduler::SpawnOpts,
}
impl<G: GenServer> GenServerBuilder<G> {
pub fn new(state: G) -> Self {
GenServerBuilder { state, infos: Vec::new(), supervisor: None }
GenServerBuilder {
state,
infos: Vec::new(),
supervisor: None,
stack_opts: crate::scheduler::SpawnOpts::default(),
}
}
/// Add an out-of-band channel; messages arriving on it are dispatched to
@@ -665,6 +671,14 @@ impl<G: GenServer> GenServerBuilder<G> {
self
}
/// Stack shape for the server actor (RFC 019) — see
/// [`SpawnOpts`](crate::SpawnOpts). Useful for servers that recurse
/// deeply or call into FFI with large C frames.
pub fn stack_opts(mut self, opts: crate::scheduler::SpawnOpts) -> Self {
self.stack_opts = opts;
self
}
/// Spawn the server actor and hand back its [`GenServerRef`]. The server's
/// lifetime is governed by its refs, not by joining, so the backing join
/// handle is dropped.
@@ -686,10 +700,16 @@ impl<G: GenServer> GenServerBuilder<G> {
/// under the name before returning.
fn spawn_server(self) -> GenServerRef<G> {
let (tx, rx) = channel::<Envelope<G>>();
let GenServerBuilder { state, infos, supervisor } = self;
let GenServerBuilder { state, infos, supervisor, stack_opts } = self;
let handle = match supervisor {
Some(sup) => spawn_under(sup, move || server_loop::<G>(rx, state, infos)),
None => spawn(move || server_loop::<G>(rx, state, infos)),
Some(sup) => {
crate::scheduler::spawn_under_with(sup, stack_opts, move || {
server_loop::<G>(rx, state, infos)
})
}
None => crate::scheduler::spawn_with(stack_opts, move || {
server_loop::<G>(rx, state, infos)
}),
};
GenServerRef { tx, pid: handle.pid() }
}
@@ -758,6 +778,12 @@ impl<G: GenServer> NamedGenServerBuilder<G> {
self
}
/// Stack shape for the server actor (see [`GenServerBuilder::stack_opts`]).
pub fn stack_opts(mut self, opts: crate::scheduler::SpawnOpts) -> Self {
self.builder = self.builder.stack_opts(opts);
self
}
/// Spawn the server and bind its name in one step. Fallible: returns
/// [`RegisterError::NameTaken`] if the name is already held by a different
/// live server.
+15 -2
View File
@@ -71,7 +71,7 @@
use crate::channel::{channel, select, Receiver, Sender};
use crate::pid::Pid;
use crate::scheduler::{cancel_timer, send_after_to, spawn as spawn_actor};
use crate::scheduler::{cancel_timer, send_after_to};
use crate::timer::TimerId;
use std::collections::{HashMap, VecDeque};
use std::marker::PhantomData;
@@ -434,8 +434,21 @@ impl<M: Machine> GenStatemRef<M> {
///
/// Panics if called outside `Runtime::run()`.
pub fn spawn<M: Machine>(machine: M) -> GenStatemRef<M> {
spawn_with(crate::scheduler::SpawnOpts::default(), machine)
}
/// [`spawn`] with per-actor stack shape overrides (RFC 019) for the machine's
/// actor — see [`SpawnOpts`](crate::SpawnOpts). gen_statem has no builder
/// (its one-shot `spawn(machine)` shape predates RFC 019), so the opts ride
/// a `_with` variant like the scheduler's own spawns.
///
/// Panics if called outside `Runtime::run()`.
pub fn spawn_with<M: Machine>(
opts: crate::scheduler::SpawnOpts,
machine: M,
) -> GenStatemRef<M> {
let (tx, rx) = channel::<M::Ev>();
let handle = spawn_actor(move || statem_loop(rx, machine));
let handle = crate::scheduler::spawn_with(opts, move || statem_loop(rx, machine));
GenStatemRef { tx, pid: handle.pid() }
}
+53
View File
@@ -179,6 +179,34 @@ pub struct ActorInfo {
/// `budget-accounting` feature is enabled, since measuring it costs a
/// timestamp read on every resume.
pub budget_cycles: u64,
/// RFC 019 §8 — this actor's stack, as the runtime sees it. All fields
/// are lock-free atomic reads, coherent for this incarnation via the
/// same generation check as the counters above. Exact RSS is
/// deliberately absent: `mincore` is debug tooling, never a runtime
/// path.
pub stack: StackInfo,
}
/// RFC 019 §8 — per-actor stack introspection. Sizes are page-rounded, as
/// [`Stack::new`](crate::stack::Stack::new) rounds them.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub struct StackInfo {
/// Usable stack size ([`SpawnOpts::stack_reserve`]
/// (crate::SpawnOpts::stack_reserve) or the Config/default).
pub reserve: usize,
/// PROT_NONE guard below the usable region.
pub guard: usize,
/// Sampled high-water depth in bytes: `top − lowest saved sp`. Sampled,
/// not exact — the context save at yields/parks/preemptions is the
/// sampler (RFC 019 §2), so a spike the actor never yielded inside is
/// invisible. 0 depth means "never descheduled at any depth", not
/// "never ran".
pub depth_high_water: usize,
/// Parks on this incarnation since its last shrink (or since install if
/// it has never shrunk) — the §3 cooldown counter, live.
pub parks_since_shrink: u32,
/// §3 shrinks performed on this incarnation.
pub shrinks: u32,
}
/// A snapshot of every actor in the runtime at (approximately) one moment.
@@ -220,6 +248,22 @@ pub fn snapshot() -> RuntimeSnapshot {
/// slot was reused by another), out of range, or was never a real pid at
/// all. Unlike [`snapshot`], every field of the result describes the same
/// instant, since there is only one actor to read.
/// The stack shape `(reserve, guard)` of a live actor, page-rounded — the
/// RFC 019 introspection surface's first field (depth sampling and shrink
/// counters land with the shrink machinery). `None` if `pid` no longer names
/// a live actor. Takes the actor's cold lock briefly; debugging/assertion
/// use, not a hot-path call.
pub fn stack_shape(pid: Pid) -> Option<(usize, usize)> {
with_runtime(|inner| {
let slot = inner.slot_at(pid)?;
let cold = slot.cold.lock();
if slot.generation() != pid.generation() {
return None;
}
cold.actor.as_ref().map(|a| a.stack.shape())
})
}
pub fn actor_info(pid: Pid) -> Option<ActorInfo> {
with_runtime(|inner| {
let slot = inner.slot_at(pid)?;
@@ -265,6 +309,14 @@ fn read_slot(slot: &Slot, idx: u32, mail: Option<&MailboxInfo>) -> Option<ActorI
drop(cold);
// Counters are plain atomics, read lock-free.
let (reserve, guard, top, hwm, parks_since_shrink, shrinks) = slot.stack_introspect();
let stack = StackInfo {
reserve,
guard,
depth_high_water: top.saturating_sub(hwm),
parks_since_shrink,
shrinks,
};
let overruns = slot.overruns();
let messages_received = slot.messages_received();
let budget_cycles = slot.budget_cycles();
@@ -290,6 +342,7 @@ fn read_slot(slot: &Slot, idx: u32, mail: Option<&MailboxInfo>) -> Option<ActorI
overruns,
messages_received,
budget_cycles,
stack,
})
}
+5 -3
View File
@@ -12,6 +12,7 @@
//! See `LOOM.md` for the design intent and the deferred-for-later list.
pub mod stack;
pub(crate) mod signal;
pub mod context;
pub mod preempt;
pub mod pid;
@@ -64,7 +65,7 @@ pub use gen_statem::{
CallError as GenStatemCallError, Cx, Machine, Reply, Resolution, SendError as GenStatemSendError,
GenStatemRef,
};
pub use introspect::{
pub use introspect::{StackInfo,
actor_info, snapshot, tree, tree_from, ActorInfo, ActorState, RuntimeSnapshot, RuntimeTree,
TreeNode, SNAPSHOT_FORMAT_VERSION,
};
@@ -83,8 +84,9 @@ pub use runtime::{init, Config, Runtime};
pub use scheduler::{
block_on_io, cancel_timer, request_stop, run, self_pid, send_after, send_after_named,
send_after_named_wall, send_after_wall, sleep, sleep_wall,
spawn, spawn_addr, spawn_under, wait_readable, wait_readable_timeout, wait_writable,
wait_writable_timeout, yield_now, FdArm, JoinError, JoinHandle,
spawn, spawn_addr, spawn_addr_with, spawn_under, spawn_under_with, spawn_with,
wait_readable, wait_readable_timeout, wait_writable,
wait_writable_timeout, yield_now, FdArm, JoinError, JoinHandle, SpawnOpts,
};
pub use supervisor::{ChildSpec, OneForOne, Restart, Signal, Strategy};
pub use timer::TimerId;
+7 -4
View File
@@ -98,10 +98,13 @@ pub(crate) fn clear_current_slot() {
CURRENT_SLOT.with(|c| c.set(std::ptr::null()));
}
/// RFC 007 (`smarm-causal`) — raw pointer to the on-CPU actor's slot, null on
/// the scheduler's own stack. Same lifetime argument as `note_overrun`: the
/// slot is never reclaimed while its actor is on-CPU.
#[cfg(feature = "smarm-causal")]
/// Raw pointer to the on-CPU actor's slot, null on the scheduler's own
/// stack. Same lifetime argument as `note_overrun`: the slot is never
/// reclaimed while its actor is on-CPU. Consumers: the `smarm-causal`
/// profiler (RFC 007) and — unconditionally — the SIGSEGV classifier
/// (RFC 019 §7), which additionally relies on this being a plain load of a
/// const-initialized TLS Cell (no lazy init, no allocation, no dtor): safe
/// from a signal handler.
#[inline]
pub(crate) fn current_slot_ptr() -> *const crate::runtime::Slot {
CURRENT_SLOT.with(|c| c.get())
+280 -10
View File
@@ -65,6 +65,8 @@
//! word stores are `Release`, loads are `Acquire`. The chain that matters:
//! the park path stores `sp` (Relaxed) *before* its Release transition; any
//! later Acquire transition/load of the word therefore observes that `sp`.
//! RFC 019's `hwm` (and the shrink that reads it) piggybacks this exact
//! pattern in the same pre-Release window and adds no edges.
//! The run-queue mutex independently provides the same edges today; the
//! word's own ordering is what phase 3's lock-free queue will rely on.
//!
@@ -160,6 +162,8 @@ pub struct Config {
alloc_interval: u32,
timeslice_cycles: u64,
stack_pool_cap: usize,
stack_reserve: usize,
stack_guard: usize,
max_actors: usize,
wake_slot: bool,
node_id: crate::pg::NodeId,
@@ -175,6 +179,8 @@ impl Config {
alloc_interval: crate::preempt::DEFAULT_ALLOC_INTERVAL,
timeslice_cycles: crate::preempt::DEFAULT_TIMESLICE_CYCLES,
stack_pool_cap: n * 4,
stack_reserve: DEFAULT_STACK_RESERVE,
stack_guard: DEFAULT_STACK_GUARD,
max_actors: DEFAULT_MAX_ACTORS,
wake_slot: false,
node_id: crate::pg::DEFAULT_NODE_ID,
@@ -194,6 +200,8 @@ impl Config {
alloc_interval: crate::preempt::DEFAULT_ALLOC_INTERVAL,
timeslice_cycles: crate::preempt::DEFAULT_TIMESLICE_CYCLES,
stack_pool_cap: max * 4,
stack_reserve: DEFAULT_STACK_RESERVE,
stack_guard: DEFAULT_STACK_GUARD,
max_actors: DEFAULT_MAX_ACTORS,
wake_slot: false,
node_id: crate::pg::DEFAULT_NODE_ID,
@@ -226,6 +234,31 @@ impl Config {
self
}
/// Default per-actor stack reserve (RFC 019). A *virtual* reservation —
/// anonymous mmap is demand-paged, so RSS follows touched pages, not
/// this number — but overflowing it hits the guard and dies. Page-rounded.
/// Per-actor override: `SpawnOpts::stack_reserve`.
/// Default: [`DEFAULT_STACK_RESERVE`] (64 KiB) — the million-cheap-actors
/// story is unchanged; big stacks are opt-in.
pub fn stack_reserve(mut self, n: usize) -> Self {
assert!(n > 0, "stack_reserve must be non-zero");
self.stack_reserve = n;
self
}
/// Default PROT_NONE guard below each stack (RFC 019). Address space
/// only. Page-rounded. Rust overflow is caught by any single page
/// (probestack touches pages in order); the wide default exists for
/// unprobed FFI frames, which can step over a small guard in one
/// `sub rsp`. Per-actor override: `SpawnOpts::guard_size`.
/// Default: [`DEFAULT_STACK_GUARD`] (1 MiB — the kernel's
/// `stack_guard_gap` convention; see its doc for why width is free).
pub fn stack_guard(mut self, n: usize) -> Self {
assert!(n > 0, "stack_guard must be non-zero");
self.stack_guard = n;
self
}
/// Capacity of the actor slot table — the maximum number of
/// **simultaneously live** actors (total spawned over a run is unbounded;
/// slots are recycled). The table is a fixed slab allocated once at
@@ -293,6 +326,8 @@ impl Default for Config {
alloc_interval: crate::preempt::DEFAULT_ALLOC_INTERVAL,
timeslice_cycles: crate::preempt::DEFAULT_TIMESLICE_CYCLES,
stack_pool_cap: avail * 4,
stack_reserve: DEFAULT_STACK_RESERVE,
stack_guard: DEFAULT_STACK_GUARD,
max_actors: DEFAULT_MAX_ACTORS,
wake_slot: false,
node_id: crate::pg::DEFAULT_NODE_ID,
@@ -383,7 +418,48 @@ impl RuntimeStats {
// Slot — packed state word + hot atomics + cold lifecycle data
// ---------------------------------------------------------------------------
pub(crate) const ACTOR_STACK_SIZE: usize = 64 * 1024;
/// Default usable stack reserve per actor (RFC 019). See [`Config::stack_reserve`].
pub const DEFAULT_STACK_RESERVE: usize = 64 * 1024;
/// Default PROT_NONE guard below each actor stack (RFC 019). Raised from one
/// page so unprobed C frames cannot leap it. See [`Config::stack_guard`].
///
/// 1 MiB, following the kernel's own answer to the same problem: after Stack
/// Clash (2017) the main-thread guard gap became `stack_guard_gap` = 256
/// pages, because 4 KiB was jumpable by one honest `sub rsp` and no small
/// constant was defensible. Guard pages are PROT_NONE: virtual address space
/// only — zero RSS, zero page-table entries, no overcommit charge — so the
/// wide default is free at any actor count (1 M actors ≈ 1 TiB of VA against
/// a 128 TiB budget). A frame that jumps even this lands in the tier-2
/// overshoot window of the SIGSEGV diagnostic (`signal.rs`) instead of
/// silence.
pub const DEFAULT_STACK_GUARD: usize = 1024 * 1024;
/// RFC 019 §3: minimum releasable span (`sp − hwm` at park) before the
/// park-path shrink spends a syscall. A constant, not a `Config` field
/// (ratified): nobody tunes this well and the measured stakes are low — a
/// threshold-sized `MADV_FREE` costs ~3 µs against a ~100 ns park, paid
/// only on spike-recovery parks, which are rare by construction and *were*
/// the spike. Steady-state actors never reach the syscall: their check is
/// two Relaxed loads and a compare on a line the context-save just wrote.
pub const SHRINK_THRESHOLD: usize = 256 * 1024;
/// RFC 019 §3: parks between shrinks of one actor. Guards a few-µs cost, so
/// it can be coarse (parks, not wall time); the kernel's
/// reclaim-under-pressure-only handling of `MADV_FREE` is the real release
/// hysteresis — re-touched-before-pressure pages cost a 0.24 µs/page
/// cancel-write and no fault. A constant, not `Config` (ratified, same
/// rationale as [`SHRINK_THRESHOLD`]).
pub const SHRINK_COOLDOWN: u32 = 64;
/// RFC 019 §6: the entry-end span (highest addresses — the frames the next
/// actor faults first) a recycled stack keeps resident; everything below it
/// is `MADV_DONTNEED`ed before the stack re-enters the pool. Ratified as a
/// constant, not Config, alongside the shrink knobs; the 64 KiB value was a
/// flagged Claude-solo call at ratification — it equals the default reserve,
/// so with an unraised Config the zap is a no-op and only Configs that raise
/// the default reserve pay it.
pub const RECYCLE_RETAIN: usize = 64 * 1024;
pub(crate) type Closure = Box<dyn FnOnce() + Send>;
@@ -426,6 +502,35 @@ pub(crate) struct Slot {
/// Release transition out of Running; read after the Acquire transition
/// Queued→Running. Relaxed is sufficient — ordering rides on `word`.
sp: AtomicUsize,
/// RFC 019: sampled stack high-water — the minimum `sp` ever stored above,
/// i.e. the deepest excursion *observed at a switch point*. Advisory:
/// correctness never depends on it; its one job is "is a shrink worth a
/// syscall?". Declared adjacent to `sp` so the min-update dirties the
/// line the context-save just wrote. Same single-writer Relaxed
/// discipline as `sp`; reset to the fresh `sp` at install.
hwm: AtomicUsize,
/// RFC 019: parks since the last shrink (or install). Counted on every
/// pass through the Park arm by the owning scheduler thread; the shrink
/// fires only once this clears [`SHRINK_COOLDOWN`] *and* the releasable
/// span clears [`SHRINK_THRESHOLD`]. Single-writer Relaxed.
parks_since_shrink: AtomicU32,
/// RFC 019: shrinks performed on this incarnation (introspection lands
/// with the RFC's introspect surface; the counter exists from birth so
/// tests can rely on install resetting it). Single-writer Relaxed.
shrink_count: AtomicU32,
/// RFC 019 §7 — stack geometry for the SIGSEGV classifier, readable
/// without the cold lock (the `Stack` itself lives under it). Written in
/// `install_actor` before the Release publish; consulted by the handler
/// only while `preempt::CURRENT_SLOT` points here, i.e. while this actor
/// is on-CPU, so the values are never stale where they are read. 0 =
/// never installed. Usable top of the stack.
pub(crate) diag_stack_top: AtomicUsize,
/// See `diag_stack_top`: the reserve (usable) size.
pub(crate) diag_stack_reserve: AtomicUsize,
/// See `diag_stack_top`: the guard size.
pub(crate) diag_stack_guard: AtomicUsize,
/// See `diag_stack_top`: `(idx << 32) | generation`, for the message.
pub(crate) diag_pid: AtomicU64,
/// Pointer into the actor's `Arc<AtomicBool>` stop flag. Set at spawn,
/// nulled at finalize. The box outlives every read: it is only ever read
/// on the resume path while the actor cannot be finalized (it is on-CPU).
@@ -497,6 +602,13 @@ impl Slot {
Self {
word: StateWord::new(),
sp: AtomicUsize::new(0),
hwm: AtomicUsize::new(0),
parks_since_shrink: AtomicU32::new(0),
shrink_count: AtomicU32::new(0),
diag_stack_top: AtomicUsize::new(0),
diag_stack_reserve: AtomicUsize::new(0),
diag_stack_guard: AtomicUsize::new(0),
diag_pid: AtomicU64::new(0),
stop_ptr: AtomicPtr::new(std::ptr::null_mut()),
closure: AtomicPtr::new(std::ptr::null_mut()),
overruns: AtomicU64::new(0),
@@ -550,6 +662,23 @@ impl Slot {
/// Read the overrun tally (Relaxed; the snapshot reads cross-thread).
#[inline]
/// RFC 019 §8 — the stack introspection tuple, all lock-free:
/// `(reserve, guard, top, hwm, parks_since_shrink, shrink_count)`.
/// Geometry from the c6 diag atomics (install-time, gen-coherent under
/// `read_slot`'s gen check exactly like the other counters); `hwm` is the
/// §2 sampled high-water (lowest saved sp). All zeros before first
/// install.
pub(crate) fn stack_introspect(&self) -> (usize, usize, usize, usize, u32, u32) {
(
self.diag_stack_reserve.load(Ordering::Relaxed),
self.diag_stack_guard.load(Ordering::Relaxed),
self.diag_stack_top.load(Ordering::Relaxed),
self.hwm.load(Ordering::Relaxed),
self.parks_since_shrink.load(Ordering::Relaxed),
self.shrink_count.load(Ordering::Relaxed),
)
}
pub(crate) fn overruns(&self) -> u64 {
self.overruns.load(Ordering::Relaxed)
}
@@ -802,17 +931,23 @@ pub(crate) struct RuntimeInner {
pub(crate) stack_pool: RawMutex<Vec<crate::stack::Stack>>,
/// Maximum number of stacks to retain in the pool.
pub(crate) stack_pool_cap: usize,
/// Default stack shape (RFC 019), pre-page-rounded so it compares exactly
/// against `Stack::shape()`. Only stacks of exactly this shape are pooled.
pub(crate) stack_reserve: usize,
pub(crate) stack_guard: usize,
}
impl RuntimeInner {
// Private constructor taking the parsed Config fields one-for-one; a params
// struct would only move the same 8 values across the call boundary.
// struct would only move the same 10 values across the call boundary.
#[allow(clippy::too_many_arguments)]
fn new(
thread_count: usize,
alloc_interval: u32,
timeslice_cycles: u64,
stack_pool_cap: usize,
stack_reserve: usize,
stack_guard: usize,
max_actors: usize,
wake_slot: bool,
node_id: crate::pg::NodeId,
@@ -854,6 +989,8 @@ impl RuntimeInner {
process_groups: RawMutex::new(crate::pg::ProcessGroups::new()),
stack_pool: RawMutex::new(Vec::new()),
stack_pool_cap,
stack_reserve: crate::stack::round_to_pages(stack_reserve),
stack_guard: crate::stack::round_to_pages(stack_guard),
})
}
@@ -1045,6 +1182,9 @@ pub struct Runtime {
/// Initialise the runtime with the given config. Returns a reusable handle.
pub fn init(config: Config) -> Runtime {
// RFC 019 §7: one process-global SIGSEGV handler, installed before any
// scheduler thread (and so before any classifiable fault) can exist.
crate::signal::install_once();
let n = config.resolved_thread_count();
Runtime {
inner: RuntimeInner::new(
@@ -1052,6 +1192,8 @@ pub fn init(config: Config) -> Runtime {
config.alloc_interval,
config.timeslice_cycles,
config.stack_pool_cap,
config.stack_reserve,
config.stack_guard,
config.max_actors,
config.wake_slot,
config.node_id,
@@ -1309,6 +1451,99 @@ pub const ROOT_PID: Pid = Pid::new(u32::MAX, u32::MAX);
// Spawn-side slot installation
// ---------------------------------------------------------------------------
// ---------------------------------------------------------------------------
// Stack shrink — RFC 019 §3 (park path only)
// ---------------------------------------------------------------------------
/// The per-park shrink check. Called from the `YieldIntent::Park` arm inside
/// the owned window (see the assert-comment at the call site). Fast path —
/// no spike since the last shrink — is two Relaxed loads, a compare, and the
/// park counter bump, all on the slot line the context-save just wrote.
///
/// On a shrink: `MADV_FREE` the whole pages of `[hwm, sp − redzone)` (the
/// inward-rounded range from [`crate::stack::shrink_range`]), then reset
/// `hwm = sp` and the park counter. MADV_FREE only *marks*: the kernel
/// reclaims under pressure, skips re-dirtied pages, and refaults zero pages
/// for writes after reclaim — so an over-eager mark costs a cancel-write,
/// never data.
fn maybe_shrink_stack(slot: &Slot) {
let parks = slot.parks_since_shrink.load(Ordering::Relaxed).saturating_add(1);
slot.parks_since_shrink.store(parks, Ordering::Relaxed);
let sp = slot.sp.load(Ordering::Relaxed);
let hwm = slot.hwm.load(Ordering::Relaxed);
if sp.wrapping_sub(hwm) < SHRINK_THRESHOLD || sp < hwm {
return; // common case: nothing worth a syscall
}
if parks < SHRINK_COOLDOWN {
return;
}
let page = crate::stack::page_size();
if let Some((addr, len)) = crate::stack::shrink_range(hwm, sp, page) {
// Advisory: on the (kernel-config) chance MADV_FREE is unsupported,
// failing silently degrades to "never shrinks", which is correct.
unsafe {
libc::madvise(addr as *mut libc::c_void, len, libc::MADV_FREE);
}
slot.hwm.store(sp, Ordering::Relaxed);
slot.parks_since_shrink.store(0, Ordering::Relaxed);
slot.shrink_count.store(
slot.shrink_count.load(Ordering::Relaxed).saturating_add(1),
Ordering::Relaxed,
);
}
}
// ---------------------------------------------------------------------------
// Stack acquisition / recycling — RFC 019 pool rule
// ---------------------------------------------------------------------------
/// Get a stack of the shape `opts` requests (`None` fields ⇒ the runtime
/// defaults).
///
/// Pool rule (RFC 019 §1): the pool is a uniform `Vec<Stack>` of
/// default-shaped stacks and stays that way. Default-shaped requests try the
/// pool first; custom shapes always mmap fresh (and `recycle_stack` never
/// admits them, so a pooled stack is default-shaped by induction). The pool
/// lock is dropped before any mmap: no syscall ever stalls another spawner.
pub(crate) fn acquire_stack(
inner: &RuntimeInner,
opts: crate::scheduler::SpawnOpts,
) -> crate::stack::Stack {
let reserve = opts.stack_reserve.unwrap_or(inner.stack_reserve);
let guard = opts.guard_size.unwrap_or(inner.stack_guard);
let default_shaped = crate::stack::round_to_pages(reserve) == inner.stack_reserve
&& crate::stack::round_to_pages(guard) == inner.stack_guard;
if default_shaped {
if let Some(stack) = inner.stack_pool.lock().pop() {
return stack;
}
}
match crate::stack::Stack::new(reserve, guard) {
Ok(stack) => stack,
Err(e) => panic!("stack allocation failed: {e}"),
}
}
/// Return a dead actor's stack: pooled if default-shaped and under cap,
/// otherwise dropped here → munmap (custom shapes and cap overflow alike).
pub(crate) fn recycle_stack(inner: &RuntimeInner, stack: crate::stack::Stack) {
if stack.shape() == (inner.stack_reserve, inner.stack_guard) {
// RFC 019 §6: zap the dead spike before pooling, BEFORE taking the
// pool lock — acquire_stack's invariant is that no syscall ever
// stalls another spawner under it. On the rare cap-overflow the zap
// is wasted work ahead of the munmap; harmless, and cheaper than a
// second lock round-trip to find out.
stack.recycle_zap(RECYCLE_RETAIN);
let mut pool = inner.stack_pool.lock();
if pool.len() < inner.stack_pool_cap {
pool.push(stack);
}
// else: fall through — drop → munmap.
}
// Custom-shaped (or cap overflow): `stack` drops here → munmap.
}
/// Install a freshly spawned actor into the slot `idx` (which must have come
/// from `allocate_slot`) and publish it as Queued. Returns the new `Pid`.
/// Called by `scheduler::spawn_under`; lives here next to its inverse
@@ -1326,6 +1561,11 @@ pub(crate) fn install_actor(
let pid = Pid::new(idx, gen);
let stop = Arc::new(AtomicBool::new(false));
// RFC 019 §7: geometry for the SIGSEGV classifier, captured before the
// Stack moves under the cold lock. Ordered before readers by the
// publish below.
let (diag_reserve, diag_guard) = stack.shape();
let diag_top = stack.top() as usize;
slot.stop_ptr.store(Arc::as_ptr(&stop) as *mut _, Ordering::Release);
{
let mut cold = slot.cold.lock();
@@ -1337,6 +1577,15 @@ pub(crate) fn install_actor(
cold.pending_io_result = None;
}
slot.sp.store(sp, Ordering::Relaxed);
// RFC 019: a fresh incarnation starts with its high-water at the fresh
// top-of-stack `sp` and its shrink bookkeeping zeroed.
slot.hwm.store(sp, Ordering::Relaxed);
slot.parks_since_shrink.store(0, Ordering::Relaxed);
slot.shrink_count.store(0, Ordering::Relaxed);
slot.diag_stack_top.store(diag_top, Ordering::Relaxed);
slot.diag_stack_reserve.store(diag_reserve, Ordering::Relaxed);
slot.diag_stack_guard.store(diag_guard, Ordering::Relaxed);
slot.diag_pid.store(((idx as u64) << 32) | gen as u64, Ordering::Relaxed);
slot.store_closure(closure);
slot.reset_counters();
inner.live_actors.fetch_add(1, Ordering::Relaxed);
@@ -1437,13 +1686,7 @@ fn finalize_actor(inner: &Arc<RuntimeInner>, pid: Pid, outcome: Outcome) {
// (the trap sender can unpark its receiver — keep that outside too).
let supervisor_pid = actor.supervisor;
let Actor { stack, .. } = actor;
{
let mut pool = inner.stack_pool.lock();
if pool.len() < inner.stack_pool_cap {
pool.push(stack);
}
// else: drop here → munmap, same as before
}
recycle_stack(inner, stack);
// Deliver to supervisor. ROOT_PID resolves to no slot → silently absorbed.
let sender = inner.slot_at(supervisor_pid).and_then(|sup| {
@@ -1594,6 +1837,8 @@ fn fire_due_timers(inner: &Arc<RuntimeInner>, try_only: bool) {
// ---------------------------------------------------------------------------
fn schedule_loop(inner: &Arc<RuntimeInner>, slot_idx: usize) {
// RFC 019 §7: a guard hit leaves no stack to handle the signal on.
crate::signal::register_altstack();
crate::preempt::configure_preempt(inner.alloc_interval, inner.timeslice_cycles);
let stats = &inner.stats[slot_idx];
@@ -1841,7 +2086,15 @@ fn schedule_loop(inner: &Arc<RuntimeInner>, slot_idx: usize) {
crate::preempt::clear_current_slot();
let intent = YIELD_INTENT.with(|c| c.get());
slot.sp.store(get_actor_sp(), Ordering::Relaxed);
let saved_sp = get_actor_sp();
slot.sp.store(saved_sp, Ordering::Relaxed);
// RFC 019 §2: sampled high-water — one branch + at most one store
// into the line the store above just dirtied. Relaxed and advisory;
// it piggybacks the existing Relaxed-store-before-Release pattern
// (mod docs, "Memory ordering") and adds no edges.
if saved_sp < slot.hwm.load(Ordering::Relaxed) {
slot.hwm.store(saved_sp, Ordering::Relaxed);
}
if is_actor_done() {
crate::te!(crate::trace::Event::Done(pid));
@@ -1863,6 +2116,23 @@ fn schedule_loop(inner: &Arc<RuntimeInner>, slot_idx: usize) {
inner.enqueue(pid);
}
YieldIntent::Park => {
// RFC 019 §3 shrink window (correctness obligation 1):
// this site sits after the `sp` store above and before
// the `park_return` Release transition below publishes
// Parked — the scheduler is on its own stack and the
// actor is saved but not yet stealable, so the madvise
// races nothing (belt). MADV_FREE's cancel-on-write is
// the suspenders: even a racing writer could lose
// nothing written after the mark, and everything below
// live `sp` is dead by definition. Runs on BOTH arms of
// the park_return race — a consumed unpark flag means a
// wasted-but-harmless madvise on a rare window.
//
// This is the ONLY shrink site: the preempt/yield path
// deliberately never checks (§4's bounded leak under
// saturation — syscalls must not fire when scheduler
// cycles are scarcest).
maybe_shrink_stack(slot);
if slot.word.park_return(gen) {
// RFC 007 audit: an in-site park drops its sample
// tail (nothing flushes it; on_resume re-arms).
+55 -9
View File
@@ -259,6 +259,28 @@ impl Drop for JoinHandle {
// spawn / spawn_under / self_pid
// ---------------------------------------------------------------------------
/// Per-spawn stack shape overrides (RFC 019). `None` fields resolve to the
/// runtime's [`Config`](crate::runtime::Config) defaults at spawn time, so
/// struct-update syntax works anywhere without a runtime handle:
///
/// ```
/// use smarm::SpawnOpts;
/// let opts = SpawnOpts { stack_reserve: Some(8 * 1024 * 1024), ..SpawnOpts::default() };
/// ```
///
/// Both sizes are page-rounded. The reserve is *virtual* (demand-paged):
/// an 8 MiB reserve costs address space, not memory — RSS follows touched
/// pages. The guard is PROT_NONE below the stack; raise it for FFI code
/// with unusually large C frames. Custom-shaped stacks bypass the recycle
/// pool: they are mmapped fresh at spawn and munmapped at death.
#[derive(Clone, Copy, Debug, Default, PartialEq, Eq)]
pub struct SpawnOpts {
/// Usable stack reservation. `None` ⇒ [`Config::stack_reserve`](crate::runtime::Config::stack_reserve).
pub stack_reserve: Option<usize>,
/// PROT_NONE guard below the stack. `None` ⇒ [`Config::stack_guard`](crate::runtime::Config::stack_guard).
pub guard_size: Option<usize>,
}
/// Start a new actor running `f`, and return a [`JoinHandle`] for it.
///
/// The new actor runs concurrently with its caller and with every other
@@ -281,22 +303,34 @@ pub fn spawn(f: impl FnOnce() + Send + 'static) -> JoinHandle {
spawn_under(parent, f)
}
/// [`spawn`] with per-actor stack shape overrides (RFC 019).
pub fn spawn_with(opts: SpawnOpts, f: impl FnOnce() + Send + 'static) -> JoinHandle {
let parent = current_pid().unwrap_or_else(|| {
with_runtime(|_| crate::runtime::ROOT_PID)
});
spawn_under_with(parent, opts, f)
}
/// Like [`spawn`], but explicitly attaches the new actor to `supervisor`
/// instead of the calling actor. Ordinary code should reach for [`spawn`];
/// this exists for supervision trees (see [`supervisor`](crate::supervisor))
/// and other cases that need to place a child under a specific ancestor
/// rather than its true caller.
pub fn spawn_under<A>(supervisor: Pid<A>, f: impl FnOnce() + Send + 'static) -> JoinHandle {
spawn_under_with(supervisor, SpawnOpts::default(), f)
}
/// [`spawn_under`] with per-actor stack shape overrides (RFC 019).
pub fn spawn_under_with<A>(
supervisor: Pid<A>,
opts: SpawnOpts,
f: impl FnOnce() + Send + 'static,
) -> JoinHandle {
let supervisor = supervisor.erase();
// Stack + closure boxing happen before ANY runtime lock is taken: no
// syscall and no allocation ever stalls another scheduler thread.
let stack = with_runtime(|inner| inner.stack_pool.lock().pop())
.unwrap_or_else(|| {
match crate::stack::Stack::new(crate::runtime::ACTOR_STACK_SIZE) {
Ok(stack) => stack,
Err(e) => panic!("stack allocation failed: {e}"),
}
});
// Stack + closure boxing happen before the slot locks are taken; the
// pool lock inside acquire_stack is dropped before any mmap, so no
// syscall ever stalls another scheduler thread.
let stack = with_runtime(|inner| crate::runtime::acquire_stack(inner, opts));
let sp = init_actor_stack(stack.top(), crate::actor::trampoline);
let closure: crate::runtime::Closure = Box::new(f);
@@ -336,6 +370,18 @@ pub fn spawn_addr<A: crate::pid::Addressable>(
crate::pid::assert_type::<A>(pid)
}
/// [`spawn_addr`] with per-actor stack shape overrides (RFC 019).
pub fn spawn_addr_with<A: crate::pid::Addressable>(
opts: SpawnOpts,
body: impl FnOnce(crate::channel::Receiver<A::Msg>) + Send + 'static,
) -> Pid<A> {
let (tx, rx) = crate::channel::channel::<A::Msg>();
let handle = spawn_with(opts, move || body(rx));
let pid = handle.pid();
crate::registry::install_for::<A::Msg>(pid, tx);
crate::pid::assert_type::<A>(pid)
}
use crate::context::init_actor_stack;
/// The identity of the actor currently running. Use it to hand your own
+321
View File
@@ -0,0 +1,321 @@
//! RFC 019 §7 — overflow diagnostics.
//!
//! One process-global SIGSEGV handler, installed once at [`crate::runtime::init`]
//! (before any scheduler thread exists, so the PRIOR save is unracing), plus a
//! per-scheduler-thread `sigaltstack` registered at `schedule_loop` entry — a
//! guard hit means the faulting stack has no room to run anything, so the
//! altstack is not optional.
//!
//! The handler classifies `si_addr` against the *current* actor only, reached
//! through `preempt::CURRENT_SLOT` — a const-initialized `Cell<*const Slot>`
//! whose access is a plain TLS load (no lazy init, no allocation, no dtor
//! registration), and which every scheduler thread has materialized before an
//! actor can run on it. The slot's diag atomics (`diag_stack_top` & co) are
//! written in `install_actor` before the Release publish and are only consulted
//! here while the actor is on-CPU, so they cannot be stale.
//!
//! Two classification tiers:
//! - **In-guard**: definitive. Rust frames probe pages in order
//! (`__rust_probestack`), so Rust overflow always lands here; so does any C
//! built with `-fstack-clash-protection` (distro-packaged libraries), and —
//! with the 1 MiB default guard — nearly every unprobed frame too.
//! - **Overshoot**: within [`OVERSHOOT_SLOP`] *below* the guard. An unprobed
//! frame (cargo-built C via `cc` almost never enables clash protection)
//! large enough to step over the guard in one `sub rsp`. Attribution is
//! "probable": the address is in unmapped VA that nothing else owns, an
//! actor was on-CPU, and the distance fits a frame — the diagnostic says so.
//!
//! Classified faults print one line (async-signal-safe: stack buffer +
//! `write(2)`, no fmt, no alloc, no locks) and re-raise with default
//! disposition — no unwind, no resume, no fail-soft (jarred; UB-adjacent from
//! a handler). Unclassified faults reinstate the PRIOR handler and refault, so
//! std's own "thread ... has overflowed its stack" diagnostics for OS-thread
//! stacks survive our presence. Reinstating deregisters us for good, which is
//! fine: the process is dying either way.
use std::cell::Cell;
use std::mem::MaybeUninit;
use std::sync::atomic::Ordering;
use std::sync::Once;
/// Tier-2 window below the guard. Matches the guard default (and the kernel's
/// `stack_guard_gap`): a frame that out-jumps both the guard and this window
/// in one displacement is past what a diagnostic can honestly attribute.
pub(crate) const OVERSHOOT_SLOP: usize = 1024 * 1024;
/// Per-scheduler-thread signal stack. MINSIGSTKSZ is ~11 KiB on AVX-512
/// hardware; 64 KiB leaves the formatter room without mattering to anyone.
/// One per OS thread, never freed: scheduler threads live for the process in
/// practice, and repeated `run()`s on reused threads re-use the registration
/// (the TLS flag), so the leak is bounded by the OS thread count.
const ALTSTACK_SIZE: usize = 64 * 1024;
static INSTALL: Once = Once::new();
/// The handler that was installed before ours (std's, typically). Written
/// exactly once inside INSTALL — which completes in `runtime::init` before
/// any scheduler thread (and thus any classifiable fault) can exist — and
/// only read from the handler afterwards.
static mut PRIOR: MaybeUninit<libc::sigaction> = MaybeUninit::uninit();
thread_local! {
/// Whether this OS thread has registered its altstack.
static ALTSTACK_SET: Cell<bool> = const { Cell::new(false) };
}
/// Where a fault landed relative to the current actor's stack.
#[derive(Debug, PartialEq, Eq)]
pub(crate) enum FaultClass {
/// Inside `[top − reserve − guard, top − reserve)`: the guard region.
Guard,
/// Within `OVERSHOOT_SLOP` below the guard: stepped over it. Payload is
/// the distance below `guard_lo`.
Overshoot(usize),
/// Not ours to explain.
Foreign,
}
/// Pure classifier — all edges unit-tested below. `top` is the stack's usable
/// top, `reserve`/`guard` its shape; both page-rounded by `Stack::new`.
pub(crate) fn classify(addr: usize, top: usize, reserve: usize, guard: usize) -> FaultClass {
let guard_hi = top.wrapping_sub(reserve);
let guard_lo = guard_hi.wrapping_sub(guard);
if addr >= guard_lo && addr < guard_hi {
FaultClass::Guard
} else if addr < guard_lo && addr >= guard_lo.saturating_sub(OVERSHOOT_SLOP) {
FaultClass::Overshoot(guard_lo - addr)
} else {
FaultClass::Foreign
}
}
/// Install the process-global handler. Idempotent; called from
/// `runtime::init`.
pub(crate) fn install_once() {
INSTALL.call_once(|| unsafe {
let mut sa: libc::sigaction = std::mem::zeroed();
sa.sa_sigaction = handler as *const () as usize;
sa.sa_flags = libc::SA_SIGINFO | libc::SA_ONSTACK;
libc::sigemptyset(&mut sa.sa_mask);
let prior = &mut *std::ptr::addr_of_mut!(PRIOR);
libc::sigaction(libc::SIGSEGV, &sa, prior.as_mut_ptr());
});
}
/// Register this OS thread's altstack (idempotent per thread). Called at
/// `schedule_loop` entry, so every thread that can run an actor has one.
pub(crate) fn register_altstack() {
ALTSTACK_SET.with(|set| {
if set.get() {
return;
}
unsafe {
let sp = libc::mmap(
std::ptr::null_mut(),
ALTSTACK_SIZE,
libc::PROT_READ | libc::PROT_WRITE,
libc::MAP_PRIVATE | libc::MAP_ANONYMOUS,
-1,
0,
);
if sp == libc::MAP_FAILED {
// Degrade: no altstack means a guard hit dies without the
// message (handler can't run) — the pre-RFC behavior, never
// incorrectness.
return;
}
let ss = libc::stack_t {
ss_sp: sp,
ss_flags: 0,
ss_size: ALTSTACK_SIZE,
};
libc::sigaltstack(&ss, std::ptr::null_mut());
}
set.set(true);
});
}
// ---------------------------------------------------------------------------
// The handler
// ---------------------------------------------------------------------------
unsafe extern "C" fn handler(
_sig: libc::c_int,
info: *mut libc::siginfo_t,
_ctx: *mut libc::c_void,
) {
let slot_ptr = crate::preempt::current_slot_ptr();
if !slot_ptr.is_null() {
let slot = &*slot_ptr;
let top = slot.diag_stack_top.load(Ordering::Relaxed);
if top != 0 {
let reserve = slot.diag_stack_reserve.load(Ordering::Relaxed);
let guard = slot.diag_stack_guard.load(Ordering::Relaxed);
let pid = slot.diag_pid.load(Ordering::Relaxed);
let addr = (*info).si_addr() as usize;
match classify(addr, top, reserve, guard) {
FaultClass::Guard => {
let mut b = Buf::new();
b.s("smarm: actor ");
b.pid(pid);
b.s(" overflowed its stack: fault in the guard region, depth-at-fault=");
b.u(top - addr);
b.s(" bytes (reserve=");
b.u(reserve);
b.s(", guard=");
b.u(guard);
b.s("). Raise stack_reserve (SpawnOpts or Config).\n");
b.emit();
die_by_default();
return;
}
FaultClass::Overshoot(below) => {
let mut b = Buf::new();
b.s("smarm: actor ");
b.pid(pid);
b.s(" probably overflowed its stack: fault ");
b.u(below);
b.s(" bytes below the guard - an unprobed (FFI?) frame stepped over it (reserve=");
b.u(reserve);
b.s(", guard=");
b.u(guard);
b.s("). Raise stack_guard or stack_reserve.\n");
b.emit();
die_by_default();
return;
}
FaultClass::Foreign => {}
}
}
}
// Not ours: put back whoever was there before us and refault into them.
let prior = &*std::ptr::addr_of!(PRIOR);
libc::sigaction(libc::SIGSEGV, prior.as_ptr(), std::ptr::null_mut());
}
/// Reset SIGSEGV to default disposition; returning from the handler then
/// refaults at the same instruction and the process dies the normal death
/// (core-dumpable, correct wait status), exactly as if we were never here —
/// but with the message already on stderr.
unsafe fn die_by_default() {
let mut dfl: libc::sigaction = std::mem::zeroed();
dfl.sa_sigaction = libc::SIG_DFL;
libc::sigemptyset(&mut dfl.sa_mask);
libc::sigaction(libc::SIGSEGV, &dfl, std::ptr::null_mut());
}
// ---------------------------------------------------------------------------
// Async-signal-safe formatting: fixed buffer, decimal itoa, one write(2).
// ---------------------------------------------------------------------------
struct Buf {
b: [u8; 320],
len: usize,
}
impl Buf {
fn new() -> Self {
Buf { b: [0; 320], len: 0 }
}
fn s(&mut self, s: &str) {
for &c in s.as_bytes() {
if self.len < self.b.len() {
self.b[self.len] = c;
self.len += 1;
}
}
}
fn u(&mut self, mut n: usize) {
let mut tmp = [0u8; 20];
let mut i = tmp.len();
loop {
i -= 1;
tmp[i] = b'0' + (n % 10) as u8;
n /= 10;
if n == 0 {
break;
}
}
for &c in &tmp[i..] {
if self.len < self.b.len() {
self.b[self.len] = c;
self.len += 1;
}
}
}
/// `idx.gen`, unpacked from the install-time packing.
fn pid(&mut self, packed: u64) {
self.u((packed >> 32) as usize);
self.s(".");
self.u((packed & 0xffff_ffff) as usize);
}
fn emit(&self) {
unsafe {
libc::write(2, self.b.as_ptr() as *const libc::c_void, self.len);
}
}
}
// ---------------------------------------------------------------------------
// Classifier units — the arithmetic edges, before anything integrates.
// ---------------------------------------------------------------------------
#[cfg(test)]
mod tests {
use super::{classify, FaultClass, OVERSHOOT_SLOP};
const PG: usize = 4096;
// A synthetic stack far from address-space edges: top at 1 GiB.
const TOP: usize = 1 << 30;
const RESERVE: usize = 16 * PG;
const GUARD: usize = 4 * PG;
const GUARD_HI: usize = TOP - RESERVE;
const GUARD_LO: usize = GUARD_HI - GUARD;
#[test]
fn inside_guard_both_edges() {
assert_eq!(classify(GUARD_LO, TOP, RESERVE, GUARD), FaultClass::Guard);
assert_eq!(classify(GUARD_HI - 1, TOP, RESERVE, GUARD), FaultClass::Guard);
assert_eq!(classify(GUARD_LO + GUARD / 2, TOP, RESERVE, GUARD), FaultClass::Guard);
}
#[test]
fn usable_region_is_foreign() {
// A fault inside the RW stack itself isn't a guard hit and must not
// be explained as one.
assert_eq!(classify(GUARD_HI, TOP, RESERVE, GUARD), FaultClass::Foreign);
assert_eq!(classify(TOP - 1, TOP, RESERVE, GUARD), FaultClass::Foreign);
}
#[test]
fn above_top_is_foreign() {
assert_eq!(classify(TOP, TOP, RESERVE, GUARD), FaultClass::Foreign);
assert_eq!(classify(TOP + PG, TOP, RESERVE, GUARD), FaultClass::Foreign);
}
#[test]
fn overshoot_window_edges() {
assert_eq!(
classify(GUARD_LO - 1, TOP, RESERVE, GUARD),
FaultClass::Overshoot(1)
);
assert_eq!(
classify(GUARD_LO - OVERSHOOT_SLOP, TOP, RESERVE, GUARD),
FaultClass::Overshoot(OVERSHOOT_SLOP)
);
assert_eq!(
classify(GUARD_LO - OVERSHOOT_SLOP - 1, TOP, RESERVE, GUARD),
FaultClass::Foreign
);
}
#[test]
fn low_address_stack_saturates_not_wraps() {
// A stack mapped so low that the slop window would underflow: the
// window clips to 0 instead of wrapping around the address space.
let top = RESERVE + GUARD + PG; // guard_lo == PG
assert_eq!(classify(0, top, RESERVE, GUARD), FaultClass::Overshoot(PG));
// Null-page fault still classified only because it IS within slop
// here; with a normal-height stack it is Foreign (covered above by
// the window-edge test at realistic addresses).
}
}
+227 -14
View File
@@ -1,32 +1,45 @@
//! mmap-based growable stack with a guard page below.
//! mmap-based actor stack with a PROT_NONE guard region below (RFC 019).
//!
//! Layout (low → high address):
//! [ guard page (PROT_NONE) | stack region ]
//! ^ top() — initial stack pointer
//! [ guard region (PROT_NONE) | stack region ]
//! ^ top() — initial stack pointer
//!
//! Stacks grow downward. Overflow lands in the guard page → SIGSEGV.
//! Stacks grow downward. Overflow lands in the guard region → SIGSEGV.
//!
//! Both the usable reserve and the guard are caller-chosen (page-rounded).
//! The reserve is a *virtual* reservation: anonymous mmap is demand-paged,
//! so RSS is touched-pages, not reserve × actors. The guard costs address
//! space only. A wide guard (the runtime defaults to 64 KiB) exists for
//! unprobed FFI frames: Rust frames touch pages in order (probestack), so
//! one page catches Rust overflow, but a C frame with a large local can
//! step over a single page in one `sub rsp`.
use std::io;
pub struct Stack {
/// Bottom of the entire mmap'd region (start of guard page).
/// Bottom of the entire mmap'd region (start of the guard).
base: *mut u8,
/// Total mmap'd size: guard_size + stack_size.
total_size: usize,
/// Usable stack size (excluding guard page).
/// Usable stack size (excluding the guard).
stack_size: usize,
/// PROT_NONE region below the usable stack.
guard_size: usize,
}
// Stack owns its memory; safe to send across threads.
unsafe impl Send for Stack {}
impl Stack {
/// Allocate a new stack. `stack_size` is the usable region; one page is
/// added below as a guard page. Both are rounded up to the page size.
pub fn new(stack_size: usize) -> io::Result<Self> {
/// Allocate a new stack. `stack_size` is the usable region; `guard_size`
/// is mapped PROT_NONE below it. Both are rounded up to the page size
/// and must be non-zero.
pub fn new(stack_size: usize, guard_size: usize) -> io::Result<Self> {
assert!(stack_size > 0, "stack_size must be non-zero");
assert!(guard_size > 0, "guard_size must be non-zero");
let page = page_size();
let stack_size = round_up(stack_size, page);
let guard_size = page;
let guard_size = round_up(guard_size, page);
let total_size = guard_size + stack_size;
let base = unsafe {
@@ -53,7 +66,7 @@ impl Stack {
return Err(err);
}
Ok(Self { base, total_size, stack_size })
Ok(Self { base, total_size, stack_size, guard_size })
}
/// 16-byte-aligned top of the usable region.
@@ -62,14 +75,54 @@ impl Stack {
(raw_top & !15) as *mut u8
}
/// Pointer to the bottom of the usable region (just above the guard page).
/// Pointer to the bottom of the usable region (just above the guard).
pub fn usable_base(&self) -> *mut u8 {
unsafe { self.base.add(page_size()) }
unsafe { self.base.add(self.guard_size) }
}
pub fn stack_size(&self) -> usize {
self.stack_size
}
pub fn guard_size(&self) -> usize {
self.guard_size
}
/// `(stack_size, guard_size)` after page rounding. The pool rule
/// (RFC 019 §1) compares this against the runtime defaults: only
/// default-shaped stacks are pooled.
pub fn shape(&self) -> (usize, usize) {
(self.stack_size, self.guard_size)
}
/// Pool-recycle zap (RFC 019 §6): `MADV_DONTNEED` everything below the
/// retained entry end `[top − retain, top)` — the span the next actor's
/// shallow frames land in stays resident, the dead spike below it is
/// released. The stack is unowned at the call site (its actor is dead),
/// so a synchronous eager zap races nothing and the RSS drop is
/// immediate — a museum of worst-case spikes is exactly what a pool must
/// not be; DONTNEED's ~8× per-page cost vs FREE is irrelevant off the
/// hot path. Advisory like the park-path shrink: a failure degrades to
/// "the pool keeps RSS", never to incorrectness. No-op (no syscall) when
/// `retain` covers the whole usable region — i.e. always, at the 64 KiB
/// default reserve.
pub(crate) fn recycle_zap(&self, retain: usize) {
if let Some((off, len)) = retain_range(self.stack_size, retain, page_size()) {
unsafe {
libc::madvise(
self.usable_base().add(off) as *mut libc::c_void,
len,
libc::MADV_DONTNEED,
);
}
}
}
}
/// Round `n` up to whole pages — the same rounding `Stack::new` applies, so
/// runtime defaults stored pre-rounded compare exactly against [`Stack::shape`].
pub(crate) fn round_to_pages(n: usize) -> usize {
round_up(n, page_size())
}
impl Drop for Stack {
@@ -80,10 +133,170 @@ impl Drop for Stack {
}
}
fn page_size() -> usize {
pub(crate) fn page_size() -> usize {
unsafe { libc::sysconf(libc::_SC_PAGESIZE) as usize }
}
fn round_up(n: usize, align: usize) -> usize {
(n + align - 1) & !(align - 1)
}
/// The whole-page span the park-path shrink may `MADV_FREE` (RFC 019 §3):
/// `[page_up(hwm), page_down(sp − redzone))`, or `None` if no full page fits.
///
/// `hwm` is the sampled high-water (deepest observed `sp`); everything in
/// `[hwm, sp)` is below the live frame and dead by definition. One page of
/// redzone stays resident under live `sp` — it covers the SysV 128-byte red
/// zone plus spill margin with room to spare. Rounding is inward on both
/// ends so the result can never touch the redzone, cross `sp`, or dip below
/// `hwm`; all arithmetic is checked so adversarial inputs (`sp < redzone`,
/// `hwm ≥ sp`, values near the address-space edges) collapse to `None`
/// rather than a wild or negative-length range.
pub(crate) fn shrink_range(hwm: usize, sp: usize, page: usize) -> Option<(usize, usize)> {
debug_assert!(page.is_power_of_two());
if hwm >= sp {
return None;
}
let redzone = page;
let end = sp.checked_sub(redzone)? & !(page - 1); // page_down(sp − redzone)
let start = hwm.checked_add(page - 1)? & !(page - 1); // page_up(hwm)
if end > start {
Some((start, end - start))
} else {
None
}
}
/// The `(offset_from_usable_base, len)` span the pool recycle DONTNEEDs
/// (RFC 019 §6): everything below the retained entry end. "Bottom RETAIN of
/// the stack" is read stack-wise (entry frames = highest addresses of a
/// downward stack): the retained span is `[top − page_up(retain), top)`, the
/// zapped span is the rest — retaining the low-address deep end instead
/// would keep the coldest pages and release the ones the next actor faults
/// first. `retain` rounds *up* to whole pages (retain more, zap less), so
/// with `stack_size` page-rounded by `Stack::new` the result is always
/// page-aligned. Checked math: `retain ≥ stack_size` (notably the default
/// 64 KiB reserve with the 64 KiB RETAIN) and overflow collapse to `None`.
pub(crate) fn retain_range(stack_size: usize, retain: usize, page: usize) -> Option<(usize, usize)> {
debug_assert!(page.is_power_of_two());
let retain = retain.checked_add(page - 1)? & !(page - 1); // page_up(retain)
let len = stack_size.checked_sub(retain)?;
if len == 0 {
return None;
}
Some((0, len))
}
#[cfg(test)]
mod tests {
use super::{retain_range, shrink_range};
const PG: usize = 4096;
#[test]
fn retain_covers_whole_stack_is_a_noop() {
// The default config: reserve == RETAIN == 64 KiB. No zap, no syscall.
assert_eq!(retain_range(16 * PG, 16 * PG, PG), None);
assert_eq!(retain_range(PG, PG, PG), None);
}
#[test]
fn retain_larger_than_stack_is_a_noop() {
assert_eq!(retain_range(16 * PG, 17 * PG, PG), None);
assert_eq!(retain_range(PG, usize::MAX, PG), None); // page_up overflows
}
#[test]
fn retain_zero_zaps_everything() {
assert_eq!(retain_range(16 * PG, 0, PG), Some((0, 16 * PG)));
}
#[test]
fn retain_rounds_up_zapping_less() {
// 1 byte of retain keeps a whole page.
assert_eq!(retain_range(16 * PG, 1, PG), Some((0, 15 * PG)));
assert_eq!(retain_range(16 * PG, PG + 1, PG), Some((0, 14 * PG)));
}
#[test]
fn retain_one_page_short_of_stack() {
assert_eq!(retain_range(2 * PG, PG, PG), Some((0, PG)));
}
#[test]
fn retain_range_is_page_aligned() {
for size_pg in [1usize, 2, 3, 16, 1024] {
for retain in [0usize, 1, PG - 1, PG, PG + 1, 4 * PG, size_pg * PG] {
if let Some((off, len)) = retain_range(size_pg * PG, retain, PG) {
assert_eq!(off, 0);
assert_eq!(len % PG, 0);
assert!(len <= size_pg * PG);
assert!(len > 0);
}
}
}
}
#[test]
fn empty_and_inverted_spans_are_none() {
assert_eq!(shrink_range(0x8000_0000, 0x8000_0000, PG), None); // hwm == sp
assert_eq!(shrink_range(0x8000_1000, 0x8000_0000, PG), None); // hwm > sp
}
#[test]
fn span_smaller_than_redzone_plus_page_is_none() {
let sp = 0x8000_0000;
// Everything within redzone+1 page of sp: no full page clears both
// the redzone and the page_up(hwm) rounding.
assert_eq!(shrink_range(sp - PG, sp, PG), None);
assert_eq!(shrink_range(sp - 2 * PG + 1, sp, PG), None);
}
#[test]
fn exact_two_pages_frees_one() {
let sp = 0x8000_0000;
let hwm = sp - 2 * PG;
// [hwm, hwm+PG) frees; [sp−PG, sp) is redzone.
assert_eq!(shrink_range(hwm, sp, PG), Some((hwm, PG)));
}
#[test]
fn unaligned_ends_round_inward() {
let sp = 0x8000_0123; // live sp mid-page
let hwm = 0x7f00_0abc; // high-water mid-page
let (start, len) = shrink_range(hwm, sp, PG).unwrap();
assert_eq!(start % PG, 0);
assert_eq!(len % PG, 0);
assert!(start >= hwm); // never below the sampled high-water
assert!(start + len <= (sp - PG) & !(PG - 1)); // never into the redzone
}
#[test]
fn result_never_crosses_sp() {
// Sweep hwm across every offset of the page straddling the boundary.
let sp = 0x8000_0000 + 137;
for hwm in (sp - 4 * PG)..(sp) {
if let Some((start, len)) = shrink_range(hwm, sp, PG) {
assert!(start >= hwm);
assert!(start + len + PG <= sp + PG); // end ≤ page_down(sp − PG) < sp
assert!(len > 0);
}
}
}
#[test]
fn underflow_near_zero_is_none() {
assert_eq!(shrink_range(0, PG - 1, PG), None); // sp < redzone
assert_eq!(shrink_range(0, 0, PG), None);
}
#[test]
fn big_span_frees_interior() {
let sp = 0x8000_0000;
let spike = 4 * 1024 * 1024;
let hwm = sp - spike;
let (start, len) = shrink_range(hwm, sp, PG).unwrap();
assert_eq!(start, hwm); // aligned input: starts exactly at hwm
assert_eq!(len, spike - PG); // everything but the redzone page
}
}
+5 -5
View File
@@ -23,7 +23,7 @@ extern "C-unwind" fn actor_simple() {
#[test]
fn actor_runs_and_returns_to_scheduler() {
reset_log();
let stack = Stack::new(64 * 1024).unwrap();
let stack = Stack::new(64 * 1024, 4096).unwrap();
let sp = init_actor_stack(stack.top(), actor_simple);
set_actor_sp(sp);
unsafe { switch_to_actor() };
@@ -40,7 +40,7 @@ extern "C-unwind" fn actor_two_steps() {
#[test]
fn actor_yields_and_resumes() {
reset_log();
let stack = Stack::new(64 * 1024).unwrap();
let stack = Stack::new(64 * 1024, 4096).unwrap();
let sp = init_actor_stack(stack.top(), actor_two_steps);
set_actor_sp(sp);
@@ -85,7 +85,7 @@ extern "C-unwind" fn actor_reg_check() {
#[test]
fn callee_saved_registers_survive_yield() {
let stack = Stack::new(64 * 1024).unwrap();
let stack = Stack::new(64 * 1024, 4096).unwrap();
let sp = init_actor_stack(stack.top(), actor_reg_check);
set_actor_sp(sp);
unsafe { switch_to_actor(); switch_to_actor(); }
@@ -117,8 +117,8 @@ extern "C-unwind" fn actor_b() {
#[test]
fn two_actors_dont_corrupt_each_other() {
let stack_a = Stack::new(64 * 1024).unwrap();
let stack_b = Stack::new(64 * 1024).unwrap();
let stack_a = Stack::new(64 * 1024, 4096).unwrap();
let stack_b = Stack::new(64 * 1024, 4096).unwrap();
let sp_a = init_actor_stack(stack_a.top(), actor_a);
let sp_b = init_actor_stack(stack_b.top(), actor_b);
+126
View File
@@ -237,6 +237,13 @@ fn tree_from_nests_children_and_reroots_orphans() {
overruns: 0,
messages_received: 0,
budget_cycles: 0,
stack: smarm::StackInfo {
reserve: 0,
guard: 0,
depth_high_water: 0,
parks_since_shrink: 0,
shrinks: 0,
},
};
let snap = RuntimeSnapshot {
@@ -352,3 +359,122 @@ fn budget_cycles_accumulate_when_enabled() {
h.join().unwrap();
});
}
// ---------------------------------------------------------------------------
// RFC 019 §8 — the stack introspection surface.
// ---------------------------------------------------------------------------
/// Burn ~`frames` × 4 KiB of stack with a yield at max depth, so the context
/// save samples the high-water there (RFC 019 §2: hwm is SAMPLED at
/// deschedule, not tracked continuously).
#[inline(never)]
fn burn_stack_yielding(frames: usize) -> u64 {
let mut local = [0u8; 4096];
local[0] = frames as u8;
let below = if frames == 0 {
smarm::yield_now();
0
} else {
burn_stack_yielding(frames - 1)
};
std::hint::black_box(&mut local);
below.wrapping_add(local[0] as u64)
}
#[test]
fn stack_info_reports_defaults_and_sampled_depth() {
run(|| {
let (ready_tx, ready_rx) = channel::<()>();
let (gate_tx, gate_rx) = channel::<()>();
let h = spawn(move || {
// ~32 KiB deep with a yield at the bottom: the sample point.
std::hint::black_box(burn_stack_yielding(8));
ready_tx.send(()).unwrap();
gate_rx.recv().unwrap();
});
ready_rx.recv().unwrap();
let info = spin_until(h.pid(), |a| a.state == ActorState::Parked);
let s = info.stack;
assert_eq!(s.reserve, 64 * 1024, "default reserve");
assert_eq!(s.guard, 1024 * 1024, "default guard (kernel stack_guard_gap convention)");
assert!(
s.depth_high_water >= 8 * 4096,
"hwm sampled at the deep yield: expected ≥ 32 KiB, got {}",
s.depth_high_water
);
assert!(
s.depth_high_water < s.reserve,
"depth {} cannot exceed the reserve {}",
s.depth_high_water,
s.reserve
);
// Parked at the gate right now, never shrunk (64 KiB reserve cannot
// cross the shrink threshold).
assert!(s.parks_since_shrink >= 1, "the gate park must be counted");
assert_eq!(s.shrinks, 0);
gate_tx.send(()).unwrap();
h.join().unwrap();
});
}
#[test]
fn stack_info_shrink_counters_are_live() {
use smarm::runtime::{Config, SHRINK_COOLDOWN, SHRINK_THRESHOLD};
use smarm::{spawn_with, SpawnOpts};
let rt = smarm::runtime::init(Config::exact(1));
rt.run(|| {
let (park_tx, park_rx) = channel::<()>();
let spike = 768 * 4096;
assert!(spike > SHRINK_THRESHOLD);
let worker = spawn_with(
SpawnOpts { stack_reserve: Some(8 * 1024 * 1024), ..SpawnOpts::default() },
move || {
std::hint::black_box(burn_stack_yielding(768));
for _ in 0..(SHRINK_COOLDOWN + 8) {
park_rx.recv().unwrap();
}
},
);
let wpid = worker.pid();
// Before any parks complete: the spike depth is visible.
let info = spin_until(wpid, |a| a.state == ActorState::Parked);
assert!(
info.stack.depth_high_water >= spike,
"spike should be sampled: {} < {spike}",
info.stack.depth_high_water
);
// Cross the cooldown, then read the counters live while the worker
// is parked waiting for the remaining rounds (post-join the slot is
// reclaimed and the generation check correctly hides it).
for _ in 0..(SHRINK_COOLDOWN + 2) {
spin_until(wpid, |a| a.state == ActorState::Parked);
park_tx.send(()).unwrap();
}
let info = spin_until(wpid, |a| a.state == ActorState::Parked && a.stack.shrinks >= 1);
let s = info.stack;
assert!(s.shrinks >= 1, "cooldown was crossed with a spike above threshold");
assert!(
s.parks_since_shrink < SHRINK_COOLDOWN,
"counter must reset at shrink: {}",
s.parks_since_shrink
);
assert!(
s.depth_high_water < spike,
"hwm resets to the shallow park sp at shrink; got {}",
s.depth_high_water
);
for _ in 0..6 {
spin_until(wpid, |a| a.state == ActorState::Parked);
park_tx.send(()).unwrap();
}
worker.join().unwrap();
});
}
+34
View File
@@ -517,3 +517,37 @@ fn runtime_reusable_after_root_panic() {
r.run(move || ran_t.store(true, Ordering::Relaxed));
assert!(ran.load(Ordering::Relaxed), "runtime unusable after root panic");
}
// ---------------------------------------------------------------------------
// RFC 019 — Config stack knobs
// ---------------------------------------------------------------------------
/// Burn ~`frames` × 4 KiB of stack; probestack touches pages in order so
/// exceeding the reserve would hit the guard and SIGSEGV the process.
#[inline(never)]
fn burn_stack(frames: usize) -> u64 {
let mut local = [0u8; 4096];
local[0] = frames as u8;
let below = if frames == 0 { 0 } else { burn_stack(frames - 1) };
std::hint::black_box(&mut local);
below.wrapping_add(local[0] as u64)
}
#[test]
fn config_stack_reserve_permits_deep_recursion() {
// ~256 KiB of frames: four times the old fixed 64 KiB reserve. With
// Config::stack_reserve raised this must complete; before RFC 019 it
// could only segfault.
let rt = smarm::runtime::init(Config::exact(1).stack_reserve(1024 * 1024));
let done = Arc::new(AtomicBool::new(false));
let done2 = done.clone();
rt.run(move || {
spawn(move || {
std::hint::black_box(burn_stack(64));
done2.store(true, Ordering::SeqCst);
})
.join()
.unwrap();
});
assert!(done.load(Ordering::SeqCst));
}
+212
View File
@@ -0,0 +1,212 @@
//! RFC 019 commit 2 — the `SpawnOpts` surface.
//!
//! Covers: per-spawn stack shape overrides on every spawn surface, the
//! `None ⇒ Config default` resolution, the pool rule from the outside
//! (obligation 4: a custom-shaped stack never enters the pool), and that a
//! big reserve behaviorally takes effect (deep recursion completes).
use smarm::runtime::{Config, DEFAULT_STACK_GUARD, DEFAULT_STACK_RESERVE};
use smarm::{
self_pid, spawn, spawn_under_with, spawn_with, GenServerBuilder, SpawnOpts,
};
use std::sync::atomic::{AtomicBool, Ordering};
use std::sync::Arc;
fn rt1() -> smarm::runtime::Runtime {
smarm::runtime::init(Config::exact(1))
}
#[test]
fn default_spawn_has_default_shape() {
rt1().run(|| {
let h = spawn(|| {
let shape = smarm::introspect::stack_shape(self_pid()).unwrap();
assert_eq!(shape, (DEFAULT_STACK_RESERVE, DEFAULT_STACK_GUARD));
});
h.join().unwrap();
});
}
#[test]
fn spawn_with_overrides_reserve_and_guard() {
rt1().run(|| {
let opts = SpawnOpts {
stack_reserve: Some(1024 * 1024),
guard_size: Some(256 * 1024),
};
let h = spawn_with(opts, || {
let shape = smarm::introspect::stack_shape(self_pid()).unwrap();
assert_eq!(shape, (1024 * 1024, 256 * 1024));
});
h.join().unwrap();
});
}
#[test]
fn spawn_with_partial_override_keeps_config_default_for_the_rest() {
rt1().run(|| {
let opts = SpawnOpts { stack_reserve: Some(1024 * 1024), ..SpawnOpts::default() };
let h = spawn_with(opts, || {
let shape = smarm::introspect::stack_shape(self_pid()).unwrap();
assert_eq!(shape, (1024 * 1024, DEFAULT_STACK_GUARD));
});
h.join().unwrap();
});
}
#[test]
fn spawn_with_rounds_to_pages() {
rt1().run(|| {
let opts = SpawnOpts { stack_reserve: Some(64 * 1024 + 1), guard_size: Some(4097) };
let h = spawn_with(opts, || {
let (reserve, guard) = smarm::introspect::stack_shape(self_pid()).unwrap();
assert_eq!(reserve % 4096, 0);
assert_eq!(guard % 4096, 0);
assert!(reserve >= 64 * 1024 + 1);
assert!(guard >= 4097);
});
h.join().unwrap();
});
}
#[test]
fn spawn_under_with_takes_opts() {
rt1().run(|| {
let me = self_pid();
let opts = SpawnOpts { stack_reserve: Some(128 * 1024), ..SpawnOpts::default() };
let h = spawn_under_with(me, opts, || {
let (reserve, _) = smarm::introspect::stack_shape(self_pid()).unwrap();
assert_eq!(reserve, 128 * 1024);
});
h.join().unwrap();
});
}
/// Obligation 4, from the outside: a dead custom stack must not be handed to
/// the next default spawn. The pool is LIFO, so if the custom stack had been
/// (wrongly) pushed at death, the very next default-shaped spawn on this
/// single-threaded runtime would pop it and report a custom shape.
#[test]
fn custom_stack_never_enters_the_pool() {
rt1().run(|| {
spawn_with(
SpawnOpts { stack_reserve: Some(512 * 1024), guard_size: Some(128 * 1024) },
|| {},
)
.join()
.unwrap();
let h = spawn(|| {
let shape = smarm::introspect::stack_shape(self_pid()).unwrap();
assert_eq!(shape, (DEFAULT_STACK_RESERVE, DEFAULT_STACK_GUARD));
});
h.join().unwrap();
});
}
/// The reverse direction of the pool rule: a default-shaped stack IS pooled
/// and reused (cap = threads × 4 ≥ 1 here, pool empty at start).
#[test]
fn default_stack_is_recycled() {
rt1().run(|| {
spawn(|| {}).join().unwrap();
let h = spawn(|| {
let shape = smarm::introspect::stack_shape(self_pid()).unwrap();
assert_eq!(shape, (DEFAULT_STACK_RESERVE, DEFAULT_STACK_GUARD));
});
h.join().unwrap();
});
}
/// Burn ~`frames` × 4 KiB of stack (see tests/runtime.rs twin).
#[inline(never)]
fn burn_stack(frames: usize) -> u64 {
let mut local = [0u8; 4096];
local[0] = frames as u8;
let below = if frames == 0 { 0 } else { burn_stack(frames - 1) };
std::hint::black_box(&mut local);
below.wrapping_add(local[0] as u64)
}
#[test]
fn big_reserve_behaviorally_takes_effect() {
// ~1 MiB deep on an 8 MiB per-spawn reserve, runtime default untouched.
rt1().run(|| {
let done = Arc::new(AtomicBool::new(false));
let done2 = done.clone();
spawn_with(
SpawnOpts { stack_reserve: Some(8 * 1024 * 1024), ..SpawnOpts::default() },
move || {
std::hint::black_box(burn_stack(256));
done2.store(true, Ordering::SeqCst);
},
)
.join()
.unwrap();
assert!(done.load(Ordering::SeqCst));
});
}
// ---------------------------------------------------------------------------
// Builder surfaces
// ---------------------------------------------------------------------------
struct Echo;
impl smarm::GenServer for Echo {
type Call = ();
type Reply = (usize, usize);
type Cast = ();
type Info = ();
type Timer = ();
fn handle_call(&mut self, _c: ()) -> (usize, usize) {
smarm::introspect::stack_shape(self_pid()).unwrap()
}
fn handle_cast(&mut self, _c: ()) {}
}
#[test]
fn gen_server_builder_stack_opts() {
rt1().run(|| {
let server = GenServerBuilder::new(Echo)
.stack_opts(SpawnOpts { stack_reserve: Some(256 * 1024), ..SpawnOpts::default() })
.start();
let (reserve, guard) = server.call(()).unwrap();
assert_eq!(reserve, 256 * 1024);
assert_eq!(guard, DEFAULT_STACK_GUARD);
server.shutdown();
});
}
struct Probe;
impl smarm::Machine for Probe {
type Ev = smarm::channel::Sender<(usize, usize)>;
fn state_timeout_ev() -> Self::Ev {
unreachable!("no timers in this test")
}
fn timeout_ev(_name: &'static str) -> Self::Ev {
unreachable!("no timers in this test")
}
fn on_start(&mut self, _cx: &mut smarm::Cx<Self::Ev>) {}
fn handle(
&mut self,
ev: Self::Ev,
_cx: &mut smarm::Cx<Self::Ev>,
) -> smarm::gen_statem::Step<Self::Ev> {
let _ = ev.send(smarm::introspect::stack_shape(self_pid()).unwrap());
smarm::gen_statem::Step::Stayed
}
}
#[test]
fn gen_statem_spawn_with_stack_opts() {
rt1().run(|| {
let m = smarm::gen_statem::spawn_with(
SpawnOpts { stack_reserve: Some(256 * 1024), ..SpawnOpts::default() },
Probe,
);
let (tx, rx) = smarm::channel::channel();
m.send(tx).unwrap();
let (reserve, guard) = rx.recv().unwrap();
assert_eq!(reserve, 256 * 1024);
assert_eq!(guard, DEFAULT_STACK_GUARD);
});
}
+77 -9
View File
@@ -7,13 +7,13 @@ use smarm::stack::Stack;
#[test]
fn top_is_16_byte_aligned() {
let s = Stack::new(64 * 1024).unwrap();
let s = Stack::new(64 * 1024, 4096).unwrap();
assert_eq!(s.top() as usize % 16, 0);
}
#[test]
fn top_is_within_allocation() {
let s = Stack::new(64 * 1024).unwrap();
let s = Stack::new(64 * 1024, 4096).unwrap();
let top = s.top() as usize;
let base = s.usable_base() as usize;
assert!(top > base);
@@ -22,7 +22,7 @@ fn top_is_within_allocation() {
#[test]
fn write_and_read_top_of_stack() {
let s = Stack::new(64 * 1024).unwrap();
let s = Stack::new(64 * 1024, 4096).unwrap();
let sentinel: u64 = 0xDEAD_BEEF_CAFE_1234;
unsafe {
let ptr = s.top().sub(8) as *mut u64;
@@ -33,7 +33,7 @@ fn write_and_read_top_of_stack() {
#[test]
fn write_and_read_bottom_of_usable_region() {
let s = Stack::new(64 * 1024).unwrap();
let s = Stack::new(64 * 1024, 4096).unwrap();
let sentinel: u64 = 0x0102_0304_0506_0708;
unsafe {
let ptr = s.usable_base() as *mut u64;
@@ -44,17 +44,17 @@ fn write_and_read_bottom_of_usable_region() {
#[test]
fn small_stack_allocates() {
assert!(Stack::new(4096).is_ok());
assert!(Stack::new(4096, 4096).is_ok());
}
#[test]
fn large_stack_allocates() {
assert!(Stack::new(8 * 1024 * 1024).is_ok());
assert!(Stack::new(8 * 1024 * 1024, 4096).is_ok());
}
#[test]
fn stack_size_at_least_requested() {
let s = Stack::new(64 * 1024).unwrap();
let s = Stack::new(64 * 1024, 4096).unwrap();
assert!(s.stack_size() >= 64 * 1024);
}
@@ -68,15 +68,28 @@ use std::process::Command;
fn run_as_child_if_requested() {
match env::var("SMARM_SUBTEST").as_deref() {
Ok("guard_page_direct") => {
let s = Stack::new(64 * 1024).unwrap();
let s = Stack::new(64 * 1024, 4096).unwrap();
unsafe {
let guard_ptr = s.usable_base().sub(1);
guard_ptr.write_volatile(0xAB);
}
std::process::exit(0);
}
Ok("wide_guard_top") => {
// One byte below the usable region, 64 KiB guard: must fault.
let s = Stack::new(64 * 1024, 64 * 1024).unwrap();
unsafe { s.usable_base().sub(1).write_volatile(0xAB); }
std::process::exit(0);
}
Ok("wide_guard_bottom") => {
// The very bottom page of a 64 KiB guard: an unprobed C-style
// leap over a small guard lands here — must still fault.
let s = Stack::new(64 * 1024, 64 * 1024).unwrap();
unsafe { s.usable_base().sub(64 * 1024).write_volatile(0xAB); }
std::process::exit(0);
}
Ok("stack_overflow") => {
let s = Stack::new(64 * 1024).unwrap();
let s = Stack::new(64 * 1024, 4096).unwrap();
unsafe {
let mut ptr = s.top().sub(1);
let stop = s.usable_base().sub(1);
@@ -121,3 +134,58 @@ fn stack_overflow_causes_sigsegv() {
assert_eq!(status.signal(), Some(11), "expected SIGSEGV, got: {:?}", status);
}
}
// ---------------------------------------------------------------------------
// RFC 019 — explicit shape: rounding, guard accessor, wide-guard coverage.
// ---------------------------------------------------------------------------
#[test]
fn sizes_round_up_to_page() {
let s = Stack::new(64 * 1024 + 1, 4096 + 1).unwrap();
assert_eq!(s.stack_size() % 4096, 0);
assert_eq!(s.guard_size() % 4096, 0);
assert!(s.stack_size() >= 64 * 1024 + 1);
assert!(s.guard_size() >= 4096 + 1);
}
#[test]
fn shape_reports_rounded_sizes() {
let s = Stack::new(64 * 1024, 64 * 1024).unwrap();
assert_eq!(s.shape(), (64 * 1024, 64 * 1024));
}
#[test]
fn usable_base_sits_above_guard() {
let s = Stack::new(64 * 1024, 64 * 1024).unwrap();
// The usable region must start exactly guard_size above the mapping
// base: a write at usable_base is legal, one byte below is not (the
// subprocess tests below prove the "not").
let sentinel: u64 = 0x1111_2222_3333_4444;
unsafe {
let ptr = s.usable_base() as *mut u64;
ptr.write_volatile(sentinel);
assert_eq!(ptr.read_volatile(), sentinel);
}
}
#[test]
fn wide_guard_faults_at_top() {
run_as_child_if_requested();
let status = spawn_subtest("wide_guard_top");
#[cfg(unix)]
{
use std::os::unix::process::ExitStatusExt;
assert_eq!(status.signal(), Some(11), "expected SIGSEGV, got: {:?}", status);
}
}
#[test]
fn wide_guard_faults_at_bottom() {
run_as_child_if_requested();
let status = spawn_subtest("wide_guard_bottom");
#[cfg(unix)]
{
use std::os::unix::process::ExitStatusExt;
assert_eq!(status.signal(), Some(11), "expected SIGSEGV, got: {:?}", status);
}
}
+141
View File
@@ -0,0 +1,141 @@
//! RFC 019 §7 — overflow diagnostics, observed from outside via subprocess
//! (mirrors tests/stack.rs's harness, plus stderr capture).
//!
//! Four cases:
//! - Rust recursion at defaults: probed frames walk into the guard →
//! tier-1 definitive message, death by SIGSEGV.
//! - FFI canary (96 KiB unprobed C local) at defaults: first touch lands
//! inside the 1 MiB guard → tier-1 message.
//! - FFI canary with the guard shrunk to 4 KiB: the frame steps over it
//! into unmapped VA below → tier-2 "stepped over" message. This is the
//! RFC's motivating incident (cargo-vendored gz build) reproduced.
//! - FFI canary with reserve raised to 256 KiB: fits, runs clean, exits 0 —
//! the §1 knob is the fix, proven by the same frame.
use std::env;
use std::process::Command;
unsafe extern "C" {
fn smarm_canary_burn();
}
/// Unbounded probed recursion; each frame dirties 4 KiB. black_box defeats
/// tail-call elision so the walk is real.
#[inline(never)]
#[allow(unconditional_recursion)]
fn recurse_forever(depth: u64) -> u64 {
let mut local = [0u8; 4096];
local[0] = depth as u8;
std::hint::black_box(&mut local);
recurse_forever(depth + 1).wrapping_add(local[0] as u64)
}
fn run_as_child_if_requested() {
let mode = match env::var("SMARM_DIAG_SUBTEST") {
Ok(m) => m,
Err(_) => return,
};
use smarm::runtime::Config;
use smarm::{spawn_with, SpawnOpts};
let rt = smarm::runtime::init(Config::exact(1));
rt.run(move || {
let opts = match mode.as_str() {
"rust_overflow" | "ffi_tier1" => SpawnOpts::default(),
// Small guard: the canary's 96 KiB displacement clears it.
"ffi_tier2" => SpawnOpts { guard_size: Some(4096), ..SpawnOpts::default() },
// Enough reserve: the same frame simply fits.
"ffi_clean" => SpawnOpts { stack_reserve: Some(256 * 1024), ..SpawnOpts::default() },
other => panic!("unknown subtest {other}"),
};
let is_rust = mode == "rust_overflow";
spawn_with(opts, move || {
if is_rust {
std::hint::black_box(recurse_forever(0));
} else {
unsafe { smarm_canary_burn() };
}
})
.join()
.unwrap();
});
std::process::exit(0);
}
fn spawn_subtest(name: &str) -> std::process::Output {
let exe = env::current_exe().unwrap();
Command::new(exe)
.env("SMARM_DIAG_SUBTEST", name)
.args(["--test-threads=1", "--quiet"])
.output()
.expect("failed to spawn subprocess")
}
#[cfg(unix)]
fn assert_died_sigsegv(out: &std::process::Output) {
use std::os::unix::process::ExitStatusExt;
assert_eq!(
out.status.signal(),
Some(11),
"expected death by SIGSEGV, got {:?}; stderr:\n{}",
out.status,
String::from_utf8_lossy(&out.stderr)
);
}
#[test]
fn rust_overflow_dies_with_tier1_message() {
run_as_child_if_requested();
let out = spawn_subtest("rust_overflow");
assert_died_sigsegv(&out);
let err = String::from_utf8_lossy(&out.stderr);
assert!(
err.contains("overflowed its stack") && err.contains("in the guard region"),
"missing tier-1 diagnostic; stderr:\n{err}"
);
assert!(err.contains("reserve=65536"), "wrong reserve in message:\n{err}");
assert!(err.contains("guard=1048576"), "wrong guard in message:\n{err}");
}
#[test]
fn ffi_canary_at_defaults_dies_with_tier1_message() {
run_as_child_if_requested();
let out = spawn_subtest("ffi_tier1");
assert_died_sigsegv(&out);
let err = String::from_utf8_lossy(&out.stderr);
// 96 KiB displacement from a 64 KiB reserve lands ~32 KiB into the
// 1 MiB guard: definitively classified.
assert!(
err.contains("in the guard region"),
"wide guard should catch the unprobed frame in tier 1; stderr:\n{err}"
);
}
#[test]
fn ffi_canary_over_small_guard_dies_with_tier2_message() {
run_as_child_if_requested();
let out = spawn_subtest("ffi_tier2");
assert_died_sigsegv(&out);
let err = String::from_utf8_lossy(&out.stderr);
assert!(
err.contains("stepped over it") && err.contains("below the guard"),
"expected tier-2 overshoot attribution; stderr:\n{err}"
);
assert!(err.contains("guard=4096"), "wrong guard in message:\n{err}");
}
#[test]
fn ffi_canary_with_enough_reserve_runs_clean() {
run_as_child_if_requested();
let out = spawn_subtest("ffi_clean");
assert!(
out.status.success(),
"canary should fit in 256 KiB reserve, got {:?}; stderr:\n{}",
out.status,
String::from_utf8_lossy(&out.stderr)
);
let err = String::from_utf8_lossy(&out.stderr);
assert!(
!err.contains("smarm: actor"),
"no diagnostic expected on the clean path; stderr:\n{err}"
);
}
+123
View File
@@ -0,0 +1,123 @@
//! RFC 019 commit 5 — pool recycle zaps a dead stack down to its retained
//! entry end, observed from the outside.
//!
//! A default-shaped stack that spiked deep and then died must not carry its
//! spike into the pool as resident RSS: `recycle_stack` DONTNEEDs everything
//! below the top `RECYCLE_RETAIN` bytes before pushing. The zap is
//! synchronous on the death path, so the drop is immediate — but the death
//! path itself races the observer's `join` return, hence the brief poll.
//!
//! Residency is measured with `mincore`, not smaps: a neighboring rw anon
//! mapping can land flush against the stack top and the kernel merges the
//! VMAs (observed under the full test run), so per-mapping smaps fields
//! over-count. The PROT_NONE guard below can never merge, so the usable
//! base is exactly the anchor VMA's start, and `mincore` counts pages
//! within [usable_base, usable_base + reserve) regardless of merging.
use smarm::runtime::{Config, RECYCLE_RETAIN};
use smarm::{channel, spawn, yield_now};
const RESERVE: usize = 4 * 1024 * 1024;
/// Burn ~`frames` × 4 KiB of stack, dirtying every frame.
#[inline(never)]
fn burn_stack(frames: usize) -> u64 {
let mut local = [0u8; 4096];
local[0] = frames as u8;
let below = if frames == 0 { 0 } else { burn_stack(frames - 1) };
std::hint::black_box(&mut local);
below.wrapping_add(local[0] as u64)
}
/// Resident-page count over [lo, lo + len) via mincore (len page-aligned).
fn resident_pages(lo: usize, len: usize) -> usize {
let page = 4096;
let mut vec = vec![0u8; len / page];
let ret = unsafe {
libc::mincore(lo as *mut libc::c_void, len, vec.as_mut_ptr())
};
assert_eq!(ret, 0, "mincore failed: {}", std::io::Error::last_os_error());
vec.iter().filter(|&&b| b & 1 != 0).count()
}
/// The [start, end) of the VMA containing `addr`.
fn vma_containing(addr: usize) -> (usize, usize) {
let maps = std::fs::read_to_string("/proc/self/maps").unwrap();
for line in maps.lines() {
if let Some((range, _)) = line.split_once(' ') {
if let Some((a, b)) = range.split_once('-') {
if let (Ok(start), Ok(end)) =
(usize::from_str_radix(a, 16), usize::from_str_radix(b, 16))
{
if start <= addr && addr < end {
return (start, end);
}
}
}
}
}
panic!("no VMA contains {addr:#x}");
}
fn vma_exists(addr: usize) -> bool {
let maps = std::fs::read_to_string("/proc/self/maps").unwrap();
for line in maps.lines() {
if let Some((range, _)) = line.split_once(' ') {
if let Some((a, b)) = range.split_once('-') {
if let (Ok(start), Ok(end)) =
(usize::from_str_radix(a, 16), usize::from_str_radix(b, 16))
{
if start <= addr && addr < end {
return true;
}
}
}
}
}
false
}
#[test]
fn recycle_zaps_dead_stack_down_to_retain() {
// Default reserve raised so the pool holds big stacks (default-shaped ⇒
// pooled) and the zap has something to bite; single scheduler.
let rt = smarm::runtime::init(Config::exact(1).stack_reserve(RESERVE));
rt.run(|| {
let (tx, rx) = channel::<usize>();
let h = spawn(move || {
let probe = 0u8;
let anchor = &probe as *const u8 as usize;
// The guard below is PROT_NONE and can never merge with the
// usable region, so the anchor VMA's start IS the usable base.
let (vlo, _) = vma_containing(anchor);
// Dirty ~3 MiB of the 4 MiB reserve, then die.
std::hint::black_box(burn_stack(768));
tx.send(vlo).unwrap();
});
let usable_base = rx.recv().unwrap();
h.join().unwrap();
// The zap span is everything below the retained entry end. DONTNEED
// on private anon discards synchronously and unconditionally, so
// this must go to exactly zero resident pages; the poll only covers
// the death path racing join's return.
let zap_len = RESERVE - RECYCLE_RETAIN;
let mut resident = usize::MAX;
for _ in 0..10_000 {
resident = resident_pages(usable_base, zap_len);
if resident == 0 {
break;
}
yield_now();
}
assert_eq!(
resident, 0,
"recycled stack's zap span still resident: {resident} pages in \
[{usable_base:#x}, +{zap_len:#x})"
);
// Pooled, not munmapped: the mapping must still be there.
assert!(vma_exists(usable_base), "default-shaped stack was unmapped instead of pooled");
});
}
+152
View File
@@ -0,0 +1,152 @@
//! RFC 019 commit 3 — park-path stack shrink, observed from the outside.
//!
//! The one integration-level claim of the shrink machinery: an actor that
//! spikes deep, returns shallow, and then parks past the cooldown gets its
//! dead span MADV_FREE'd — visible as `LazyFree` in `/proc/self/smaps`
//! within the stack's address range — while everything live survives.
//!
//! The high-water mark is *sampled* at context-save, so the spike yields
//! once at max depth to guarantee a sample there (in production, preemption
//! provides the quasi-random samples; a test must not rely on luck).
use smarm::runtime::{Config, SHRINK_COOLDOWN, SHRINK_THRESHOLD};
use smarm::{actor_info, channel, spawn, spawn_with, yield_now, ActorState, SpawnOpts};
/// Burn ~`frames` × 4 KiB of stack, yielding once at the bottom so the
/// context-save samples `sp` at max depth.
#[inline(never)]
fn burn_stack_yielding(frames: usize) -> u64 {
let mut local = [0u8; 4096];
local[0] = frames as u8;
let below = if frames == 0 {
yield_now();
0
} else {
burn_stack_yielding(frames - 1)
};
std::hint::black_box(&mut local);
below.wrapping_add(local[0] as u64)
}
/// Sum the `LazyFree:` kB of every smaps mapping intersecting [lo, hi).
fn lazy_free_bytes_in(lo: usize, hi: usize) -> usize {
let smaps = std::fs::read_to_string("/proc/self/smaps").unwrap();
let mut total_kb = 0usize;
let mut in_range = false;
for line in smaps.lines() {
if let Some((range, _)) = line.split_once(' ') {
if let Some((a, b)) = range.split_once('-') {
if let (Ok(start), Ok(end)) =
(usize::from_str_radix(a, 16), usize::from_str_radix(b, 16))
{
in_range = start < hi && end > lo;
continue;
}
}
}
if in_range {
if let Some(rest) = line.strip_prefix("LazyFree:") {
let kb: usize = rest.trim().trim_end_matches(" kB").trim().parse().unwrap();
total_kb += kb;
}
}
}
total_kb * 1024
}
#[test]
fn spike_then_parks_marks_lazyfree_and_keeps_live_data() {
// Single scheduler: the controller can gate on the worker being Parked.
let rt = smarm::runtime::init(Config::exact(1));
rt.run(|| {
let (park_tx, park_rx) = channel::<()>();
let (done_tx, done_rx) = channel::<(usize, u64)>();
let spike = 768 * 4096; // ~3 MiB, well past SHRINK_THRESHOLD
assert!(spike > SHRINK_THRESHOLD);
let worker = spawn_with(
SpawnOpts { stack_reserve: Some(8 * 1024 * 1024), ..SpawnOpts::default() },
move || {
// Live data that must survive the shrink, and an anchor
// address inside the stack for the smaps scan.
let live = [0xA5u8; 64];
let anchor = live.as_ptr() as usize;
// Spike: ~3 MiB deep, sampled at the bottom, unwound.
std::hint::black_box(burn_stack_yielding(768));
// Park past the cooldown. Each recv on the drained inbox is
// one park; the controller sends only when it sees us Parked.
for _ in 0..(SHRINK_COOLDOWN + 8) {
park_rx.recv().unwrap();
}
// Measure from inside: the stack spans ≤ 8 MiB below anchor.
let lazy = lazy_free_bytes_in(anchor - 8 * 1024 * 1024, anchor + 4096);
let checksum = live.iter().map(|&b| b as u64).sum();
done_tx.send((lazy, checksum)).unwrap();
},
);
let wpid = worker.pid();
for _ in 0..(SHRINK_COOLDOWN + 8) {
// Gate: send only once the worker is genuinely parked so every
// round is a real park-on-empty-mailbox.
loop {
match actor_info(wpid) {
Some(info) if info.state == ActorState::Parked => break,
Some(_) => yield_now(),
None => panic!("worker died early"),
}
}
park_tx.send(()).unwrap();
}
let (lazy, checksum) = done_rx.recv().unwrap();
// The spike was ~3 MiB; demand at least 2 MiB marked to leave slack
// for the redzone, rounding, and pages the unwind re-dirtied.
assert!(
lazy >= 2 * 1024 * 1024,
"expected ≥ 2 MiB LazyFree in the stack range, got {} bytes",
lazy
);
assert_eq!(checksum, 64 * 0xA5u64, "live stack data corrupted by shrink");
worker.join().unwrap();
});
}
/// Steady-state actors must never pay the syscall: an actor that parks a lot
/// but never spikes past the threshold ends with zero LazyFree in its stack.
#[test]
fn shallow_actor_never_shrinks() {
let rt = smarm::runtime::init(Config::exact(1));
rt.run(|| {
let (park_tx, park_rx) = channel::<()>();
let (done_tx, done_rx) = channel::<usize>();
let worker = spawn(move || {
let probe = 0u8;
let anchor = &probe as *const u8 as usize;
for _ in 0..(SHRINK_COOLDOWN + 8) {
park_rx.recv().unwrap();
}
done_tx.send(lazy_free_bytes_in(anchor - 64 * 1024, anchor + 4096)).unwrap();
});
let wpid = worker.pid();
for _ in 0..(SHRINK_COOLDOWN + 8) {
loop {
match actor_info(wpid) {
Some(info) if info.state == ActorState::Parked => break,
Some(_) => yield_now(),
None => panic!("worker died early"),
}
}
park_tx.send(()).unwrap();
}
assert_eq!(done_rx.recv().unwrap(), 0, "steady-state actor was shrunk");
worker.join().unwrap();
});
}