feat(scheduler,runtime): non-panicking try_spawn for at-capacity load shedding

allocate_slot() panics on a full slab; for a load-shedding caller (an
accept loop spawning one actor per connection) that panic lands in the
spawning actor, which then crash-loops under Restart::Transient into the
still-full slab until its restart budget is spent — and the service stops
accepting entirely. Observed live (urus slowloris scaling, 2026-08-10).
A full slab is a routine overload condition for such callers, not an
invariant violation.

- RuntimeInner::try_allocate_slot() -> Option<u32>: the non-panicking
  core; a single pop under the free-list lock, so the claim is atomic
  (claim-or-report — no check-then-spawn TOCTOU, no headroom margin).
  allocate_slot() is now a thin panicking wrapper over it.
- scheduler::try_spawn / try_spawn_under_with -> Result<JoinHandle,
  SpawnError>: parity with spawn/spawn_under_with except a full slab
  returns Err(SpawnError::AtCapacity) instead of panicking. Minimal
  surface per the agreed strategy; the remaining _with/_addr mirrors are
  trivial wrappers if ever needed.
- Slot-first ordering on the try path (reverse of spawn's stack-first):
  under overload Err is the hot path, and a rejection costs one mutex
  pop — no mmap/pool-pop + init + recycle per shed unit of work. A
  drop-guard returns the claimed slot if stack allocation panics in the
  claim-to-install window (would otherwise leak and trip run()'s
  teardown slot-leak debug_assert).
- SpawnError: non_exhaustive, Display + std::error::Error.
- spawn and every existing call site untouched: the panic remains the
  correct loud invariant check at internal/bounded spawn sites.

tests/try_spawn.rs: parity when slots free; exact slab accounting at
capacity (Err, no panic, repeatable); custom-shape try refuses before
stack allocation; self-heal after slots free; plain spawn still panics
(surfaced via JoinError payload); 4-thread race for the last slots
claims exactly the free count; SpawnError impl checks.

Design doc: smarm-suggestion-try-spawn.md. Downstream consumer change
(canned 503 on AtCapacity in urus's accept loop) is urus scope, not
smarm.

(cherry picked from commit 36de4b36aeaa72b2a5f9f3797b9854652656dcf6)
This commit is contained in:
Claude (sandbox)
2026-08-13 15:03:16 +02:00
committed by Claude (sandbox)
parent 95306c7f60
commit ca1c98336e
5 changed files with 307 additions and 4 deletions
+20 -1
View File
@@ -1188,10 +1188,29 @@ impl RuntimeInner {
MonitorId(self.next_monitor_id.fetch_add(1, Ordering::Relaxed) + 1)
}
/// Pop a vacant slot index, or `None` when the slab is full. The claim
/// is atomic — a single pop under the free-list lock — so callers get
/// claim-or-report semantics with no check-then-spawn TOCTOU: whoever
/// gets `Some` owns that slot, full stop.
pub(crate) fn try_allocate_slot(&self) -> Option<u32> {
self.free.lock().pop()
}
/// Return a slot claimed by [`try_allocate_slot`](Self::try_allocate_slot)
/// that never had an actor installed into it (e.g. stack allocation
/// panicked between claim and install). NOT for dead actors — their
/// slots go back through `reclaim_slot`, which handles generation bump,
/// waiter/monitor/link teardown, and stack recycling.
pub(crate) fn return_vacant_slot(&self, idx: u32) {
self.free.lock().push(idx);
}
/// Pop a vacant slot index, or die loudly. The fixed slab is a deliberate
/// v0.5 simplification (ROADMAP: "Deferred"); the panic names the fix.
/// Callers that can shed load instead use [`try_allocate_slot`]
/// (Self::try_allocate_slot) via `scheduler::try_spawn`.
pub(crate) fn allocate_slot(&self) -> u32 {
match self.free.lock().pop() {
match self.try_allocate_slot() {
Some(idx) => idx,
None => panic!(
"smarm: actor slot table exhausted — {} actors are live \