feat(scheduler,runtime): non-panicking try_spawn for at-capacity load shedding
allocate_slot() panics on a full slab; for a load-shedding caller (an accept loop spawning one actor per connection) that panic lands in the spawning actor, which then crash-loops under Restart::Transient into the still-full slab until its restart budget is spent — and the service stops accepting entirely. Observed live (urus slowloris scaling, 2026-08-10). A full slab is a routine overload condition for such callers, not an invariant violation. - RuntimeInner::try_allocate_slot() -> Option<u32>: the non-panicking core; a single pop under the free-list lock, so the claim is atomic (claim-or-report — no check-then-spawn TOCTOU, no headroom margin). allocate_slot() is now a thin panicking wrapper over it. - scheduler::try_spawn / try_spawn_under_with -> Result<JoinHandle, SpawnError>: parity with spawn/spawn_under_with except a full slab returns Err(SpawnError::AtCapacity) instead of panicking. Minimal surface per the agreed strategy; the remaining _with/_addr mirrors are trivial wrappers if ever needed. - Slot-first ordering on the try path (reverse of spawn's stack-first): under overload Err is the hot path, and a rejection costs one mutex pop — no mmap/pool-pop + init + recycle per shed unit of work. A drop-guard returns the claimed slot if stack allocation panics in the claim-to-install window (would otherwise leak and trip run()'s teardown slot-leak debug_assert). - SpawnError: non_exhaustive, Display + std::error::Error. - spawn and every existing call site untouched: the panic remains the correct loud invariant check at internal/bounded spawn sites. tests/try_spawn.rs: parity when slots free; exact slab accounting at capacity (Err, no panic, repeatable); custom-shape try refuses before stack allocation; self-heal after slots free; plain spawn still panics (surfaced via JoinError payload); 4-thread race for the last slots claims exactly the free count; SpawnError impl checks. Design doc: smarm-suggestion-try-spawn.md. Downstream consumer change (canned 503 on AtCapacity in urus's accept loop) is urus scope, not smarm. (cherry picked from commit 36de4b36aeaa72b2a5f9f3797b9854652656dcf6)
This commit is contained in:
committed by
Claude (sandbox)
parent
95306c7f60
commit
ca1c98336e
+20
-1
@@ -1188,10 +1188,29 @@ impl RuntimeInner {
|
||||
MonitorId(self.next_monitor_id.fetch_add(1, Ordering::Relaxed) + 1)
|
||||
}
|
||||
|
||||
/// Pop a vacant slot index, or `None` when the slab is full. The claim
|
||||
/// is atomic — a single pop under the free-list lock — so callers get
|
||||
/// claim-or-report semantics with no check-then-spawn TOCTOU: whoever
|
||||
/// gets `Some` owns that slot, full stop.
|
||||
pub(crate) fn try_allocate_slot(&self) -> Option<u32> {
|
||||
self.free.lock().pop()
|
||||
}
|
||||
|
||||
/// Return a slot claimed by [`try_allocate_slot`](Self::try_allocate_slot)
|
||||
/// that never had an actor installed into it (e.g. stack allocation
|
||||
/// panicked between claim and install). NOT for dead actors — their
|
||||
/// slots go back through `reclaim_slot`, which handles generation bump,
|
||||
/// waiter/monitor/link teardown, and stack recycling.
|
||||
pub(crate) fn return_vacant_slot(&self, idx: u32) {
|
||||
self.free.lock().push(idx);
|
||||
}
|
||||
|
||||
/// Pop a vacant slot index, or die loudly. The fixed slab is a deliberate
|
||||
/// v0.5 simplification (ROADMAP: "Deferred"); the panic names the fix.
|
||||
/// Callers that can shed load instead use [`try_allocate_slot`]
|
||||
/// (Self::try_allocate_slot) via `scheduler::try_spawn`.
|
||||
pub(crate) fn allocate_slot(&self) -> u32 {
|
||||
match self.free.lock().pop() {
|
||||
match self.try_allocate_slot() {
|
||||
Some(idx) => idx,
|
||||
None => panic!(
|
||||
"smarm: actor slot table exhausted — {} actors are live \
|
||||
|
||||
Reference in New Issue
Block a user