feat(scheduler,runtime): non-panicking try_spawn for at-capacity load shedding

allocate_slot() panics on a full slab; for a load-shedding caller (an
accept loop spawning one actor per connection) that panic lands in the
spawning actor, which then crash-loops under Restart::Transient into the
still-full slab until its restart budget is spent — and the service stops
accepting entirely. Observed live (urus slowloris scaling, 2026-08-10).
A full slab is a routine overload condition for such callers, not an
invariant violation.

- RuntimeInner::try_allocate_slot() -> Option<u32>: the non-panicking
  core; a single pop under the free-list lock, so the claim is atomic
  (claim-or-report — no check-then-spawn TOCTOU, no headroom margin).
  allocate_slot() is now a thin panicking wrapper over it.
- scheduler::try_spawn / try_spawn_under_with -> Result<JoinHandle,
  SpawnError>: parity with spawn/spawn_under_with except a full slab
  returns Err(SpawnError::AtCapacity) instead of panicking. Minimal
  surface per the agreed strategy; the remaining _with/_addr mirrors are
  trivial wrappers if ever needed.
- Slot-first ordering on the try path (reverse of spawn's stack-first):
  under overload Err is the hot path, and a rejection costs one mutex
  pop — no mmap/pool-pop + init + recycle per shed unit of work. A
  drop-guard returns the claimed slot if stack allocation panics in the
  claim-to-install window (would otherwise leak and trip run()'s
  teardown slot-leak debug_assert).
- SpawnError: non_exhaustive, Display + std::error::Error.
- spawn and every existing call site untouched: the panic remains the
  correct loud invariant check at internal/bounded spawn sites.

tests/try_spawn.rs: parity when slots free; exact slab accounting at
capacity (Err, no panic, repeatable); custom-shape try refuses before
stack allocation; self-heal after slots free; plain spawn still panics
(surfaced via JoinError payload); 4-thread race for the last slots
claims exactly the free count; SpawnError impl checks.

Design doc: smarm-suggestion-try-spawn.md. Downstream consumer change
(canned 503 on AtCapacity in urus's accept loop) is urus scope, not
smarm.

(cherry picked from commit 36de4b36aeaa72b2a5f9f3797b9854652656dcf6)
This commit is contained in:
Claude (sandbox)
2026-08-13 15:03:16 +02:00
committed by Claude (sandbox)
parent 95306c7f60
commit ca1c98336e
5 changed files with 307 additions and 4 deletions
+98
View File
@@ -278,6 +278,37 @@ pub struct SpawnOpts {
pub guard_size: Option<usize>,
}
/// Why [`try_spawn`] could not start an actor.
///
/// Marked `non_exhaustive`: today the only refusal is a full slab, but a
/// future variant (say, a shutdown-in-progress refusal) must not be a
/// breaking change for shed-path `match`es.
#[non_exhaustive]
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub enum SpawnError {
/// The fixed actor slab ([`Config::max_actors`]
/// (crate::runtime::Config::max_actors)) is full: every slot is claimed
/// by a live actor. This is a routine overload condition, not an
/// invariant violation — shed the unit of work (close the socket,
/// return a 503) and try again once actors have died.
AtCapacity,
}
impl core::fmt::Display for SpawnError {
fn fmt(&self, f: &mut core::fmt::Formatter<'_>) -> core::fmt::Result {
match self {
SpawnError::AtCapacity => {
write!(
f,
"actor slab at capacity (`Config::max_actors` live actors)"
)
}
}
}
}
impl std::error::Error for SpawnError {}
/// Start a new actor running `f`, and return a [`JoinHandle`] for it.
///
/// The new actor runs concurrently with its caller and with every other
@@ -340,6 +371,73 @@ pub fn spawn_under_with<A>(
}
}
/// [`spawn`] that reports a full actor slab instead of panicking.
///
/// Behaviour parity with [`spawn`] in every case except one: when the fixed
/// slab ([`Config::max_actors`](crate::runtime::Config::max_actors)) is
/// full, this returns [`Err(SpawnError::AtCapacity)`](SpawnError::AtCapacity)
/// where `spawn` panics the calling actor. Use it at load-shedding call
/// sites — an accept loop spawning one actor per connection, a request
/// admission point — where "at capacity" is a routine overload condition to
/// handle (reject the unit of work), not an invariant violation. Internal
/// and bounded spawn sites should keep [`spawn`]: there, the panic is a
/// correct loud invariant check.
///
/// The claim is atomic (claim-or-report): there is no
/// check-then-spawn race against other spawners for the last slot, so no
/// headroom margin is needed.
pub fn try_spawn(f: impl FnOnce() + Send + 'static) -> Result<JoinHandle, SpawnError> {
let parent = current_pid().unwrap_or_else(|| with_runtime(|_| crate::runtime::ROOT_PID));
try_spawn_under_with(parent, SpawnOpts::default(), f)
}
/// [`try_spawn`] with an explicit supervisor and per-actor stack shape
/// overrides — the full-control core the other `try_` surface is built on
/// (mirrors [`spawn_under_with`]).
pub fn try_spawn_under_with<A>(
supervisor: Pid<A>,
opts: SpawnOpts,
f: impl FnOnce() + Send + 'static,
) -> Result<JoinHandle, SpawnError> {
let supervisor = supervisor.erase();
// Slot FIRST — deliberately the reverse of `spawn`'s stack-first order:
// under overload the Err arm is the HOT path, and a rejection must cost
// one mutex pop, not an mmap/pool-pop + init + recycle per shed unit of
// work. The claim is a single atomic pop (no TOCTOU; see
// `try_allocate_slot`).
let idx = match with_runtime(|inner| inner.try_allocate_slot()) {
Some(idx) => idx,
None => return Err(SpawnError::AtCapacity),
};
// Between claim and install the slot is owned by this frame alone; if
// stack allocation panics in that window the slot must go back or it
// leaks for the life of the runtime (and would trip the run()-teardown
// slot-leak debug_assert).
struct ReturnOnUnwind(Option<u32>);
impl Drop for ReturnOnUnwind {
fn drop(&mut self) {
if let Some(idx) = self.0 {
with_runtime(|inner| inner.return_vacant_slot(idx));
}
}
}
let mut claimed = ReturnOnUnwind(Some(idx));
let stack = with_runtime(|inner| crate::runtime::acquire_stack(inner, opts));
let sp = init_actor_stack(stack.top(), crate::actor::trampoline);
let closure: crate::runtime::Closure = Box::new(f);
claimed.0 = None; // install_actor takes ownership of the slot from here
let pid = with_runtime(|inner| {
crate::runtime::install_actor(inner, idx, sp, stack, supervisor, closure)
});
Ok(JoinHandle {
pid,
consumed: false,
})
}
/// Spawn an actor that other actors can message directly by its [`Pid<A>`],
/// rather than only by holding on to a channel `Sender` you passed it
/// yourself.