feat(scheduler,runtime): non-panicking try_spawn for at-capacity load shedding
allocate_slot() panics on a full slab; for a load-shedding caller (an accept loop spawning one actor per connection) that panic lands in the spawning actor, which then crash-loops under Restart::Transient into the still-full slab until its restart budget is spent — and the service stops accepting entirely. Observed live (urus slowloris scaling, 2026-08-10). A full slab is a routine overload condition for such callers, not an invariant violation. - RuntimeInner::try_allocate_slot() -> Option<u32>: the non-panicking core; a single pop under the free-list lock, so the claim is atomic (claim-or-report — no check-then-spawn TOCTOU, no headroom margin). allocate_slot() is now a thin panicking wrapper over it. - scheduler::try_spawn / try_spawn_under_with -> Result<JoinHandle, SpawnError>: parity with spawn/spawn_under_with except a full slab returns Err(SpawnError::AtCapacity) instead of panicking. Minimal surface per the agreed strategy; the remaining _with/_addr mirrors are trivial wrappers if ever needed. - Slot-first ordering on the try path (reverse of spawn's stack-first): under overload Err is the hot path, and a rejection costs one mutex pop — no mmap/pool-pop + init + recycle per shed unit of work. A drop-guard returns the claimed slot if stack allocation panics in the claim-to-install window (would otherwise leak and trip run()'s teardown slot-leak debug_assert). - SpawnError: non_exhaustive, Display + std::error::Error. - spawn and every existing call site untouched: the panic remains the correct loud invariant check at internal/bounded spawn sites. tests/try_spawn.rs: parity when slots free; exact slab accounting at capacity (Err, no panic, repeatable); custom-shape try refuses before stack allocation; self-heal after slots free; plain spawn still panics (surfaced via JoinError payload); 4-thread race for the last slots claims exactly the free count; SpawnError impl checks. Design doc: smarm-suggestion-try-spawn.md. Downstream consumer change (canned 503 on AtCapacity in urus's accept loop) is urus scope, not smarm. (cherry picked from commit 36de4b36aeaa72b2a5f9f3797b9854652656dcf6)
This commit is contained in:
committed by
Claude (sandbox)
parent
95306c7f60
commit
ca1c98336e
@@ -278,6 +278,37 @@ pub struct SpawnOpts {
|
||||
pub guard_size: Option<usize>,
|
||||
}
|
||||
|
||||
/// Why [`try_spawn`] could not start an actor.
|
||||
///
|
||||
/// Marked `non_exhaustive`: today the only refusal is a full slab, but a
|
||||
/// future variant (say, a shutdown-in-progress refusal) must not be a
|
||||
/// breaking change for shed-path `match`es.
|
||||
#[non_exhaustive]
|
||||
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
|
||||
pub enum SpawnError {
|
||||
/// The fixed actor slab ([`Config::max_actors`]
|
||||
/// (crate::runtime::Config::max_actors)) is full: every slot is claimed
|
||||
/// by a live actor. This is a routine overload condition, not an
|
||||
/// invariant violation — shed the unit of work (close the socket,
|
||||
/// return a 503) and try again once actors have died.
|
||||
AtCapacity,
|
||||
}
|
||||
|
||||
impl core::fmt::Display for SpawnError {
|
||||
fn fmt(&self, f: &mut core::fmt::Formatter<'_>) -> core::fmt::Result {
|
||||
match self {
|
||||
SpawnError::AtCapacity => {
|
||||
write!(
|
||||
f,
|
||||
"actor slab at capacity (`Config::max_actors` live actors)"
|
||||
)
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
impl std::error::Error for SpawnError {}
|
||||
|
||||
/// Start a new actor running `f`, and return a [`JoinHandle`] for it.
|
||||
///
|
||||
/// The new actor runs concurrently with its caller and with every other
|
||||
@@ -340,6 +371,73 @@ pub fn spawn_under_with<A>(
|
||||
}
|
||||
}
|
||||
|
||||
/// [`spawn`] that reports a full actor slab instead of panicking.
|
||||
///
|
||||
/// Behaviour parity with [`spawn`] in every case except one: when the fixed
|
||||
/// slab ([`Config::max_actors`](crate::runtime::Config::max_actors)) is
|
||||
/// full, this returns [`Err(SpawnError::AtCapacity)`](SpawnError::AtCapacity)
|
||||
/// where `spawn` panics the calling actor. Use it at load-shedding call
|
||||
/// sites — an accept loop spawning one actor per connection, a request
|
||||
/// admission point — where "at capacity" is a routine overload condition to
|
||||
/// handle (reject the unit of work), not an invariant violation. Internal
|
||||
/// and bounded spawn sites should keep [`spawn`]: there, the panic is a
|
||||
/// correct loud invariant check.
|
||||
///
|
||||
/// The claim is atomic (claim-or-report): there is no
|
||||
/// check-then-spawn race against other spawners for the last slot, so no
|
||||
/// headroom margin is needed.
|
||||
pub fn try_spawn(f: impl FnOnce() + Send + 'static) -> Result<JoinHandle, SpawnError> {
|
||||
let parent = current_pid().unwrap_or_else(|| with_runtime(|_| crate::runtime::ROOT_PID));
|
||||
try_spawn_under_with(parent, SpawnOpts::default(), f)
|
||||
}
|
||||
|
||||
/// [`try_spawn`] with an explicit supervisor and per-actor stack shape
|
||||
/// overrides — the full-control core the other `try_` surface is built on
|
||||
/// (mirrors [`spawn_under_with`]).
|
||||
pub fn try_spawn_under_with<A>(
|
||||
supervisor: Pid<A>,
|
||||
opts: SpawnOpts,
|
||||
f: impl FnOnce() + Send + 'static,
|
||||
) -> Result<JoinHandle, SpawnError> {
|
||||
let supervisor = supervisor.erase();
|
||||
// Slot FIRST — deliberately the reverse of `spawn`'s stack-first order:
|
||||
// under overload the Err arm is the HOT path, and a rejection must cost
|
||||
// one mutex pop, not an mmap/pool-pop + init + recycle per shed unit of
|
||||
// work. The claim is a single atomic pop (no TOCTOU; see
|
||||
// `try_allocate_slot`).
|
||||
let idx = match with_runtime(|inner| inner.try_allocate_slot()) {
|
||||
Some(idx) => idx,
|
||||
None => return Err(SpawnError::AtCapacity),
|
||||
};
|
||||
// Between claim and install the slot is owned by this frame alone; if
|
||||
// stack allocation panics in that window the slot must go back or it
|
||||
// leaks for the life of the runtime (and would trip the run()-teardown
|
||||
// slot-leak debug_assert).
|
||||
struct ReturnOnUnwind(Option<u32>);
|
||||
impl Drop for ReturnOnUnwind {
|
||||
fn drop(&mut self) {
|
||||
if let Some(idx) = self.0 {
|
||||
with_runtime(|inner| inner.return_vacant_slot(idx));
|
||||
}
|
||||
}
|
||||
}
|
||||
let mut claimed = ReturnOnUnwind(Some(idx));
|
||||
|
||||
let stack = with_runtime(|inner| crate::runtime::acquire_stack(inner, opts));
|
||||
let sp = init_actor_stack(stack.top(), crate::actor::trampoline);
|
||||
let closure: crate::runtime::Closure = Box::new(f);
|
||||
|
||||
claimed.0 = None; // install_actor takes ownership of the slot from here
|
||||
let pid = with_runtime(|inner| {
|
||||
crate::runtime::install_actor(inner, idx, sp, stack, supervisor, closure)
|
||||
});
|
||||
|
||||
Ok(JoinHandle {
|
||||
pid,
|
||||
consumed: false,
|
||||
})
|
||||
}
|
||||
|
||||
/// Spawn an actor that other actors can message directly by its [`Pid<A>`],
|
||||
/// rather than only by holding on to a channel `Sender` you passed it
|
||||
/// yourself.
|
||||
|
||||
Reference in New Issue
Block a user