feat(scheduler,runtime): non-panicking try_spawn for at-capacity load shedding
allocate_slot() panics on a full slab; for a load-shedding caller (an accept loop spawning one actor per connection) that panic lands in the spawning actor, which then crash-loops under Restart::Transient into the still-full slab until its restart budget is spent — and the service stops accepting entirely. Observed live (urus slowloris scaling, 2026-08-10). A full slab is a routine overload condition for such callers, not an invariant violation. - RuntimeInner::try_allocate_slot() -> Option<u32>: the non-panicking core; a single pop under the free-list lock, so the claim is atomic (claim-or-report — no check-then-spawn TOCTOU, no headroom margin). allocate_slot() is now a thin panicking wrapper over it. - scheduler::try_spawn / try_spawn_under_with -> Result<JoinHandle, SpawnError>: parity with spawn/spawn_under_with except a full slab returns Err(SpawnError::AtCapacity) instead of panicking. Minimal surface per the agreed strategy; the remaining _with/_addr mirrors are trivial wrappers if ever needed. - Slot-first ordering on the try path (reverse of spawn's stack-first): under overload Err is the hot path, and a rejection costs one mutex pop — no mmap/pool-pop + init + recycle per shed unit of work. A drop-guard returns the claimed slot if stack allocation panics in the claim-to-install window (would otherwise leak and trip run()'s teardown slot-leak debug_assert). - SpawnError: non_exhaustive, Display + std::error::Error. - spawn and every existing call site untouched: the panic remains the correct loud invariant check at internal/bounded spawn sites. tests/try_spawn.rs: parity when slots free; exact slab accounting at capacity (Err, no panic, repeatable); custom-shape try refuses before stack allocation; self-heal after slots free; plain spawn still panics (surfaced via JoinError payload); 4-thread race for the last slots claims exactly the free count; SpawnError impl checks. Design doc: smarm-suggestion-try-spawn.md. Downstream consumer change (canned 503 on AtCapacity in urus's accept loop) is urus scope, not smarm. (cherry picked from commit 36de4b36aeaa72b2a5f9f3797b9854652656dcf6)
This commit is contained in:
committed by
Claude (sandbox)
parent
95306c7f60
commit
ca1c98336e
+3
-2
@@ -89,8 +89,9 @@ pub use runtime::{init, Config, Runtime};
|
||||
pub use scheduler::{
|
||||
block_on_io, cancel_timer, request_stop, run, self_pid, send_after, send_after_named,
|
||||
send_after_named_wall, send_after_wall, sleep, sleep_wall, spawn, spawn_addr, spawn_addr_with,
|
||||
spawn_under, spawn_under_with, spawn_with, wait_readable, wait_readable_timeout, wait_writable,
|
||||
wait_writable_timeout, yield_now, FdArm, JoinError, JoinHandle, SpawnOpts,
|
||||
spawn_under, spawn_under_with, spawn_with, try_spawn, try_spawn_under_with, wait_readable,
|
||||
wait_readable_timeout, wait_writable, wait_writable_timeout, yield_now, FdArm, JoinError,
|
||||
JoinHandle, SpawnError, SpawnOpts,
|
||||
};
|
||||
pub use supervisor::{ChildSpec, OneForOne, Restart, Signal, Strategy};
|
||||
pub use timer::TimerId;
|
||||
|
||||
Reference in New Issue
Block a user