feat(runtime,io): driver-enqueues + park/wake idle path — retire the wake pipe

The swap (RFC 018). Schedulers no longer sleep on a shared level-triggered
wake pipe — the herd source that made the default 8-thread config 7x
slower than 2 threads (E1). They park on per-thread futex parkers via the
coordination layer; IO backends become producers behind a two-call
contract (make runnable, then the enqueue tail wakes exactly one parked
scheduler).

Deleted: the drain lock and the one-winner phase-1 drain; the shared
completions VecDeque; the wake pipe fds, poll_wake, drain_wake_pipe,
wake_scheduler, the FdReady/Blocking Completion enum; the 100us idle nap;
the per-pop io.lock liveness read; io.rs's as_millis timeout truncation.

Added:
- enqueue wake tail (fixes the silent enqueue): wake_one_if_idle, a fence
  + one Relaxed mask load when everyone is busy — the pure-compute hot
  path pays almost nothing.
- driver-enqueues: the pool thread stashes its result in the slot,
  decrements io_outstanding, unparks; the epoll thread removes+DELs the
  waiter under the waiters lock and unparks. Both reach the runtime via a
  Weak (no Arc cycle). The waiters map moves behind its own Arc<Mutex> so
  the epoll thread never takes the runtime io lock (teardown holds it
  while joining that thread).
- io_outstanding / io_fd_waiters atomics: the termination verdict reads
  two atomics instead of taking io.lock on every pop.
- timekeeper idle path: at most one parked scheduler holds the timer
  deadline (an expiry wakes one, not a herd); everyone else parks
  indefinitely and is woken by the enqueue tail.
- busy-path timer due-check (ratified design point (a)): under saturation
  nobody parks and no timekeeper exists, yet due timers must still fire —
  one Relaxed load of the earliest-deadline snapshot per loop, clock read
  only when a timer is armed. Maintained under the timers mutex.
- chain rule: a scheduler that pops with more work queued and a sibling
  parked wakes one, so surplus runs in parallel rather than behind it.

tests/park_wake.rs pins the two new observable properties: timers fire
under full scheduler saturation, and sub-ms sleeps are prompt (the
as_millis truncation regression). Full suite + all loom models green;
clippy --lib clean.
This commit is contained in:
smarm
2026-07-24 09:12:12 +02:00
parent 2854b560d6
commit d4839f1d81
7 changed files with 522 additions and 455 deletions
+197 -195
View File
@@ -86,21 +86,28 @@
//! # Termination (counter-based)
//!
//! The old all-clear scanned the slot table under the big lock. Now:
//! exit when `io_out == 0` (read *before* the queue lock, phase-1 ordering)
//! and, under the queue lock, the queue is empty and `live_actors == 0`.
//! `live_actors` is incremented in `spawn` before the enqueue and decremented
//! at the very END of `finalize_actor`, strictly after every wakeup that
//! finalize produces has been enqueued. The soundness crux: any enqueue
//! targets a live (not-yet-finalized) actor, so `live == 0` implies no wakeup
//! can still be in flight; combined with "spawner is itself live", observing
//! `(queue empty, live == 0)` under the queue lock means no work can ever
//! appear again.
//! exit when `io_outstanding + io_fd_waiters == 0` (two Relaxed/Acquire
//! atomic loads, read *before* the queue pop) and, under the queue lock,
//! the queue is empty and `live_actors == 0`. `live_actors` is incremented
//! in `spawn` before the enqueue and decremented at the very END of
//! `finalize_actor`, strictly after every wakeup that finalize produces has
//! been enqueued. The soundness crux: any enqueue targets a live
//! (not-yet-finalized) actor, so `live == 0` implies no wakeup can still be
//! in flight; combined with "spawner is itself live", observing
//! `(queue empty, live == 0)` means no work can ever appear again.
//!
//! # Timer / IO drain (try-lock, one-winner)
//! # Scheduler park/wake (RFC 018)
//!
//! Unchanged from phase 1: one winner per round drains due timers and IO
//! completions from their own mutexes; wakeups go through the unpark
//! protocol like everyone else's.
//! Schedulers sleep on per-thread futex parkers via the coordination layer
//! (`park.rs`), NOT on a shared wake pipe. IO backends are producers behind
//! a two-call contract — make the actor runnable (`unpark_at`), whose
//! `enqueue` tail wakes exactly one parked scheduler. The blocking pool and
//! epoll thread each route their own completions (driver-enqueues); there
//! is no shared completion queue, no drain lock, no one-winner drain phase.
//! Timers fire two ways: a busy-path due-check every loop iteration (one
//! Relaxed load of the earliest-deadline snapshot when no timer is armed),
//! and the timekeeper — at most one parked scheduler holds the timer
//! deadline, so an expiry wakes one scheduler, not a herd.
use crate::actor::{
clear_current_pid, is_actor_done, reset_actor_done, set_current_actor_box,
@@ -750,8 +757,21 @@ pub(crate) struct RuntimeInner {
pub(crate) io: Mutex<Option<IoThread>>,
/// Monotonic `MonitorId` source. Never reused.
pub(crate) next_monitor_id: AtomicU64,
/// Try-lock: exactly one scheduler thread drains timers/IO per iteration.
drain_lock: Mutex<()>,
/// RFC 018: the scheduler coordination layer — per-scheduler parkers,
/// idle mask, wake protocol, timekeeper role, earliest-deadline
/// snapshot. Arc'd because `Timers` shares it (insert-side deadline
/// notes run under the timers mutex).
pub(crate) coord: Arc<crate::park::Coordinator>,
/// `block_on_io` requests in flight. Incremented by the submitter
/// BEFORE submit (underflow-proof), decremented by the pool thread on
/// completion. Read lock-free by the idle path's termination verdict —
/// the per-pop `io.lock` of the drain era is gone.
pub(crate) io_outstanding: AtomicU32,
/// Parked fd waiters. Incremented by the registrar BEFORE
/// `epoll_register` (rolled back on error), decremented by whoever
/// consumes the registration (epoll thread on readiness, canceller on
/// an unwound wait). Same lock-free verdict read as `io_outstanding`.
pub(crate) io_fd_waiters: AtomicU32,
/// Per-thread stats, indexed by scheduler thread slot (0..N).
pub(crate) stats: Vec<SchedulerStats>,
/// Global counters for RFC 000 primitives.
@@ -802,6 +822,12 @@ impl RuntimeInner {
let slots: Box<[Slot]> = (0..max_actors).map(|_| Slot::vacant()).collect();
// Low indices on top of the stack so early spawns get low pids.
let free: Vec<u32> = (0..max_actors as u32).rev().collect();
// RFC 018: the coordination layer (asserts thread_count <= 64), and
// the timers' hook into it — every insert under the timers mutex
// notes its deadline (busy-path snapshot + timekeeper re-arm).
let coord = Arc::new(crate::park::Coordinator::new(thread_count));
let mut timers = Timers::new();
timers.attach_coordinator(coord.clone());
Arc::new(Self {
run_queue: crate::run_queue::RunQueue::new(thread_count, max_actors),
slots,
@@ -810,10 +836,12 @@ impl RuntimeInner {
root_bits: AtomicU64::new(u64::MAX),
root_exited: AtomicBool::new(false),
root_swept: AtomicBool::new(false),
timers: Mutex::new(Timers::new()),
timers: Mutex::new(timers),
io: Mutex::new(None),
next_monitor_id: AtomicU64::new(0),
drain_lock: Mutex::new(()),
coord,
io_outstanding: AtomicU32::new(0),
io_fd_waiters: AtomicU32::new(0),
stats,
io_parked: AtomicU32::new(0),
sleeping: AtomicU32::new(0),
@@ -874,6 +902,13 @@ impl RuntimeInner {
);
self.run_queue.push(pid);
crate::te!(crate::trace::Event::Enqueue(pid));
// RFC 018 enqueue wake (fixes the silent enqueue): if a scheduler
// is parked, wake exactly one. The fast path when everyone is busy
// is a fence + one Relaxed load of an unmodified line — the
// pure-compute hot path pays (almost) nothing. Bias is over-wake:
// a spurious wake costs one futex round-trip and a failed pop; a
// missed wake would cost a stranded actor.
self.coord.wake_one_if_idle();
}
/// Make `pid` runnable if it is parked; coalesce or defer otherwise.
@@ -1076,7 +1111,16 @@ impl Runtime {
self.inner.live_actors.load(Ordering::Acquire), 0,
"run() called while previous run still active"
);
let io_thread = match IoThread::start() {
// RFC 018: the IO producers reach the runtime (slot table + unpark)
// through a Weak, so no RuntimeInner → IoThread → RuntimeInner cycle
// forms. Reset the in-flight counters BEFORE the threads can touch
// them (a prior run left them at 0 on a clean exit; the asserts pin
// that).
debug_assert_eq!(self.inner.io_outstanding.load(Ordering::Acquire), 0);
debug_assert_eq!(self.inner.io_fd_waiters.load(Ordering::Acquire), 0);
self.inner.io_outstanding.store(0, Ordering::Release);
self.inner.io_fd_waiters.store(0, Ordering::Release);
let io_thread = match IoThread::start(Arc::downgrade(&self.inner)) {
Ok(io) => io,
Err(e) => panic!("failed to start IO thread: {e}"),
};
@@ -1179,6 +1223,8 @@ impl Runtime {
}
self.inner.io_parked.store(0, Ordering::Relaxed);
self.inner.sleeping.store(0, Ordering::Relaxed);
self.inner.io_outstanding.store(0, Ordering::Relaxed);
self.inner.io_fd_waiters.store(0, Ordering::Relaxed);
RUNTIME.with(|r| *r.borrow_mut() = None);
@@ -1493,6 +1539,56 @@ fn stop_live_actors(inner: &Arc<RuntimeInner>) {
}
}
// ---------------------------------------------------------------------------
// Timer firing — shared by the busy-path due-check and the timekeeper
// ---------------------------------------------------------------------------
/// Pop and dispatch every due timer. `pop_due` re-anchors the
/// earliest-deadline snapshot under the timers mutex before returning, so
/// a caller that raced a concurrent insert simply comes back on the next
/// due-check. Dispatch runs with the timers lock released.
fn fire_due_timers(inner: &Arc<RuntimeInner>, try_only: bool) {
let due = if try_only {
// Busy path: if another scheduler is already in the timers mutex
// (firing, inserting, or peeking) skip — the snapshot stays due
// until someone actually pops, so the check re-fires next loop.
match inner.timers.try_lock() {
Ok(mut t) => t.pop_due(std::time::Instant::now()),
Err(std::sync::TryLockError::WouldBlock) => return,
Err(std::sync::TryLockError::Poisoned(e)) => {
panic!("smarm: timers lock poisoned (core corrupt): {e}")
}
}
} else {
match inner.timers.lock() {
Ok(mut t) => t.pop_due(std::time::Instant::now()),
Err(e) => panic!("smarm: timers lock poisoned (core corrupt): {e}"),
}
};
for entry in due {
match entry.reason {
// A sleep expiry is just an unpark: the protocol handles
// every interleaving — Parked (re-queue), Running (the
// actor is between `timers.insert_sleep` and
// `park_current`; RunningNotified makes the upcoming park
// re-queue), or gone (no-op).
crate::timer::Reason::Sleep { epoch } => inner.unpark_at(entry.pid, epoch),
crate::timer::Reason::WaitTimeout { target, epoch } => {
// The callback may call unpark_at itself.
target.on_timeout(entry.pid, epoch);
}
// A `send_after` deadline: run the captured delivery thunk.
// It resolves the destination through the registry and
// sends now (a send can unpark a receiver) — same as any
// other in-loop unpark. The timers lock is already
// released; lock order Leaf -> Channel is preserved by the
// send itself. `pop_due` only returns still-armed Sends, so
// a cancelled one never reaches here.
crate::timer::Reason::Send { fire } => fire(),
}
}
}
// ---------------------------------------------------------------------------
// schedule_loop — runs on each scheduler OS thread
// ---------------------------------------------------------------------------
@@ -1503,120 +1599,14 @@ fn schedule_loop(inner: &Arc<RuntimeInner>, slot_idx: usize) {
loop {
// ----------------------------------------------------------------
// 1. Try to win the drain lock (timers + IO). One winner per round;
// losers skip immediately and proceed to step 2.
// 1. Busy-path timer due-check (RFC 018 design point (a)): under
// saturation nobody parks, so no timekeeper exists — due timers
// must still fire. One Relaxed load + branch when no timer is
// armed; the clock is read only when one is.
// ----------------------------------------------------------------
if let Ok(_drain_guard) = inner.drain_lock.try_lock() {
// Timers and IO live behind their own mutexes (phase 1), so the
// pure-yield / pure-compute hot path never contends a global lock
// just to discover there is nothing to drain. The clock is read
// only when the timer heap is non-empty.
let due = {
let mut t = match inner.timers.lock() {
Ok(t) => t,
Err(e) => panic!("smarm: timers lock poisoned (core corrupt): {e}"),
};
if t.is_empty() {
Vec::new()
} else {
t.pop_due(std::time::Instant::now())
}
};
let completions = match inner.io.lock() {
Ok(mut io) => io
.as_mut()
.map(|io| {
// Consume wake-pipe bytes ONLY here, under the drain
// lock and strictly before draining completions.
// Producers push their completion before writing the
// byte, so every byte consumed here has its completion
// visible to the drain below. Consuming bytes anywhere
// else — in particular after an idle poll, outside the
// lock — loses wakeups: a try_lock loser can eat the
// byte for a completion the winner never saw, leaving
// it stranded (and its EPOLLONESHOT fd disarmed) until
// an unrelated timer forces another drain pass.
crate::io::drain_wake_pipe(io.wake_fd());
io.drain_completions()
})
.unwrap_or_default(),
Err(e) => panic!("smarm: io lock poisoned (core corrupt): {e}"),
};
for entry in due {
match entry.reason {
// A sleep expiry is just an unpark: the protocol handles
// every interleaving — Parked (re-queue), Running (the
// actor is between `timers.insert_sleep` and
// `park_current`; RunningNotified makes the upcoming park
// re-queue), or gone (no-op).
crate::timer::Reason::Sleep { epoch } => {
inner.unpark_at(entry.pid, epoch)
}
crate::timer::Reason::WaitTimeout { target, epoch } => {
// The callback may call unpark_at itself.
target.on_timeout(entry.pid, epoch);
}
// A `send_after` deadline: run the captured delivery thunk.
// It resolves the destination through the registry and
// sends now (a send can unpark a receiver) — same as any
// other in-loop unpark. The timers lock is already
// released; lock order Leaf -> Channel is preserved by the
// send itself. `pop_due` only returns still-armed Sends, so
// a cancelled one never reaches here.
crate::timer::Reason::Send { fire } => fire(),
}
}
for completion in completions {
match completion {
crate::io::Completion::Blocking { pid, epoch, result } => {
match inner.io.lock() {
Ok(mut io) => {
if let Some(io) = io.as_mut() {
io.outstanding = io.outstanding.saturating_sub(1);
}
}
Err(e) => {
panic!("smarm: io lock poisoned (core corrupt): {e}")
}
}
// Stash the result under the cold lock, then unpark.
// The protocol also covers the submit→park window
// (RunningNotified), which the old code missed for
// Blocking completions — a latent lost wakeup.
if let Some(slot) = inner.slot_at(pid) {
{
let mut cold = slot.cold.lock();
if slot.generation() == pid.generation() {
cold.pending_io_result = Some(result);
} else {
// Actor died (stopped) with the op in
// flight; discard the result.
}
}
inner.unpark_at(pid, epoch);
}
}
crate::io::Completion::FdReady { fd, events: _ } => {
// Resolve the parked pid under the io lock, then wake
// through the protocol. Lock order: io before all.
let parked = match inner.io.lock() {
Ok(mut io) => io.as_mut().and_then(|io| {
let entry = io.waiters.remove(&fd);
io.epoll_deregister(fd);
entry
}),
Err(e) => {
panic!("smarm: io lock poisoned (core corrupt): {e}")
}
};
if let Some((pid, epoch)) = parked {
inner.unpark_at(pid, epoch);
}
}
}
}
} // drain_guard drops here
if inner.coord.deadline_due() {
fire_due_timers(inner, true);
}
// ----------------------------------------------------------------
// 2. Pop a runnable pid. Pop order (RFC 005): wake slot first, then
@@ -1625,7 +1615,7 @@ fn schedule_loop(inner: &Arc<RuntimeInner>, slot_idx: usize) {
// ----------------------------------------------------------------
enum Pop {
Got(Pid),
Idle { io_outstanding: u32, wake_fd: Option<std::os::fd::RawFd> },
Idle,
AllDone,
/// Root has exited and nothing is runnable: stop the parked-forever
/// remainder, then re-pop. Fires at most once per run.
@@ -1650,19 +1640,12 @@ fn schedule_loop(inner: &Arc<RuntimeInner>, slot_idx: usize) {
crate::te!(crate::trace::Event::SlotPop(pid));
pid
} else {
// Read IO liveness BEFORE the queue lock (phase-1 ordering: a
// completion resurrects an actor only via the drain path, whose
// enqueue would be visible under the queue lock we take next).
let (io_out, io_fd) = {
let io = match inner.io.lock() {
Ok(io) => io,
Err(e) => panic!("smarm: io lock poisoned (core corrupt): {e}"),
};
match io.as_ref() {
Some(io) => (io.outstanding + io.waiters.len() as u32, Some(io.wake_fd())),
None => (0, None),
}
};
// Read IO liveness BEFORE the queue pop — two atomic loads now
// (RFC 018), not a per-pop `io.lock`: a completion resurrects
// an actor via the producer's own unpark→enqueue, whose entry
// would be visible to the pop below.
let io_out = inner.io_outstanding.load(Ordering::Acquire)
+ inner.io_fd_waiters.load(Ordering::Acquire);
stats.run_queue_len.store(inner.run_queue.len(), Ordering::Relaxed);
let pop = match inner.run_queue.pop() {
@@ -1694,7 +1677,7 @@ fn schedule_loop(inner: &Arc<RuntimeInner>, slot_idx: usize) {
// the idle wait below on the next pass.
Pop::RootDrain
} else {
Pop::Idle { io_outstanding: io_out, wake_fd: io_fd }
Pop::Idle
}
}
};
@@ -1709,22 +1692,15 @@ fn schedule_loop(inner: &Arc<RuntimeInner>, slot_idx: usize) {
Ok(mut timers) => timers.clear(),
Err(e) => panic!("smarm: timers lock poisoned (core corrupt): {e}"),
}
// Terminal wake: a sibling scheduler may be blocked in its
// idle wait on a snapshot that is now terminally stale — an
// orphaned long deadline (it would sleep it out in full) or
// a stale `io_outstanding > 0` from a stop-cancelled waiter
// (it would block in poll(-1) forever; cancellation produces
// no completion, so nothing else writes the wake pipe).
// One byte wakes every poller; each re-runs the verdict,
// reaches AllDone itself, and re-wakes — idempotent.
match inner.io.lock() {
Ok(io) => {
if let Some(io) = io.as_ref() {
io.wake();
}
}
Err(e) => panic!("smarm: io lock poisoned (core corrupt): {e}"),
}
// Terminal wake (replaces the wake-pipe byte): a sibling
// may be parked on a snapshot that is now terminally
// stale — an orphaned long deadline, or a stale
// `io_fd_waiters > 0` from a stop-cancelled waiter
// (cancellation produces no completion, so nothing else
// will ever wake it). `wake_all` permits every parker;
// each sibling re-runs the verdict, reaches AllDone
// itself, and re-wakes — idempotent.
inner.coord.wake_all();
return;
}
Pop::RootDrain => {
@@ -1734,40 +1710,51 @@ fn schedule_loop(inner: &Arc<RuntimeInner>, slot_idx: usize) {
stop_live_actors(inner);
continue;
}
Pop::Idle { io_outstanding, wake_fd } => {
// Something is still in flight. Sleep on the appropriate
// source to avoid hammering the queue mutex; retry on wake.
let next_deadline = match inner.timers.lock() {
Ok(timers) => timers.peek_deadline(),
Err(e) => panic!("smarm: timers lock poisoned (core corrupt): {e}"),
};
match (next_deadline, wake_fd) {
(Some(deadline), fd_opt) => {
let now = std::time::Instant::now();
if deadline > now {
let timeout = deadline - now;
match fd_opt {
Some(fd) => {
// Wake only; the byte (if any) is
// consumed by the next drain-lock
// winner in phase 1. Level-triggered
// poll means an unconsumed byte makes
// this return immediately, so a loser
// spins briefly until the winner
// releases — never sleeps through it.
crate::io::poll_wake(fd, Some(timeout));
}
None => thread::sleep(timeout),
}
Pop::Idle => {
// Something is still in flight. Park on our own futex
// until a producer wakes us (enqueue tail), a deadline
// passes, or the re-check finds the world changed.
//
// Timekeeper (RFC 018): at most one parked scheduler
// holds the timer deadline — the first idler to arm it
// parks with a timeout, the rest park indefinitely, so a
// timer expiry wakes one scheduler, not a herd. Peek and
// arm under the timers mutex (the serialization that
// makes the insert-side re-arm race-free).
let tk_deadline = {
let timers = match inner.timers.lock() {
Ok(t) => t,
Err(e) => {
panic!("smarm: timers lock poisoned (core corrupt): {e}")
}
}
(None, Some(fd)) if io_outstanding > 0 => {
// See above: no byte consumption outside phase 1.
crate::io::poll_wake(fd, None);
}
_ => {
thread::sleep(std::time::Duration::from_micros(100));
}
};
timers
.peek_deadline()
.filter(|d| inner.coord.try_arm_timer(slot_idx, *d))
};
// The mandatory post-publish re-check: a producer that
// enqueued (or a verdict input that flipped) before it
// could see our idle bit has left us the evidence.
let _ = inner.coord.park(slot_idx, tk_deadline, || {
!inner.run_queue.is_empty()
|| (inner.live_actors.load(Ordering::Acquire) == 0
&& inner.io_outstanding.load(Ordering::Acquire) == 0
&& inner.io_fd_waiters.load(Ordering::Acquire) == 0)
|| (inner.root_exited.load(Ordering::Acquire)
&& !inner.root_swept.load(Ordering::Acquire))
|| inner.coord.deadline_due()
});
if tk_deadline.is_some() {
// Hand the role back BEFORE firing: pop_due can run
// `Send` thunks that insert new timers, and the
// insert-side re-arm check must see either no
// timekeeper (skip) or a real parked one — never us,
// awake and about to re-peek anyway.
inner.coord.disarm_timer(slot_idx);
// Woken for the deadline, for work, or to re-peek
// after an earlier insert — fire whatever is due;
// the next idle pass re-arms with the new minimum.
fire_due_timers(inner, false);
}
continue;
}
@@ -1780,6 +1767,21 @@ fn schedule_loop(inner: &Arc<RuntimeInner>, slot_idx: usize) {
// by the at-most-once-enqueued invariant nothing else can have
// changed the state of a queued actor.
// ----------------------------------------------------------------
// RFC 018 chain rule: we just took one runnable; if more remain and
// a sibling is parked, wake exactly one so the surplus runs in
// PARALLEL rather than serially behind us (without this the surplus
// is not stranded — we re-pop it after resuming — but it waits out
// our whole timeslice while an idle core sits available). Cheap: the
// queue-length check is queue-local, and `wake_one_if_idle` is a
// fence + one Relaxed mask load when nobody is parked. A Relaxed
// miss here is safe — the enqueue that created the surplus already
// issued its own wake (RFC 018 no-lost-wake); this only sharpens
// parallelism latency.
if !inner.run_queue.is_empty() {
inner.coord.wake_one_if_idle();
}
let slot = match inner.slot_at(pid) {
Some(s) => s,
None => continue, // can't happen for real pids; defensive