• v0.6.1 ca1c98336e

    feat(scheduler,runtime): non-panicking try_spawn for at-capacity load shedding

    Markk116 released this 2026-08-13 13:03:16 +00:00 | 42 commits to master since this release

    allocate_slot() panics on a full slab; for a load-shedding caller (an
    accept loop spawning one actor per connection) that panic lands in the
    spawning actor, which then crash-loops under Restart::Transient into the
    still-full slab until its restart budget is spent — and the service stops
    accepting entirely. Observed live (urus slowloris scaling, 2026-08-10).
    A full slab is a routine overload condition for such callers, not an
    invariant violation.

    • RuntimeInner::try_allocate_slot() -> Option: the non-panicking
      core; a single pop under the free-list lock, so the claim is atomic
      (claim-or-report — no check-then-spawn TOCTOU, no headroom margin).
      allocate_slot() is now a thin panicking wrapper over it.
    • scheduler::try_spawn / try_spawn_under_with -> Result<JoinHandle,
      SpawnError>: parity with spawn/spawn_under_with except a full slab
      returns Err(SpawnError::AtCapacity) instead of panicking. Minimal
      surface per the agreed strategy; the remaining _with/_addr mirrors are
      trivial wrappers if ever needed.
    • Slot-first ordering on the try path (reverse of spawn's stack-first):
      under overload Err is the hot path, and a rejection costs one mutex
      pop — no mmap/pool-pop + init + recycle per shed unit of work. A
      drop-guard returns the claimed slot if stack allocation panics in the
      claim-to-install window (would otherwise leak and trip run()'s
      teardown slot-leak debug_assert).
    • SpawnError: non_exhaustive, Display + std::error::Error.
    • spawn and every existing call site untouched: the panic remains the
      correct loud invariant check at internal/bounded spawn sites.

    tests/try_spawn.rs: parity when slots free; exact slab accounting at
    capacity (Err, no panic, repeatable); custom-shape try refuses before
    stack allocation; self-heal after slots free; plain spawn still panics
    (surfaced via JoinError payload); 4-thread race for the last slots
    claims exactly the free count; SpawnError impl checks.

    Design doc: smarm-suggestion-try-spawn.md. Downstream consumer change
    (canned 503 on AtCapacity in urus's accept loop) is urus scope, not
    smarm.

    (cherry picked from commit 36de4b36aeaa72b2a5f9f3797b9854652656dcf6)

    Downloads