feat(endpoint): urus is a supervisable child — endpoint gen_server owns listeners + conns

The v0.3 shape from the spec: an app owns its runtime and root supervisor
and places urus in it as one ordered child among its own.

  your root sup
  └── ChildSpec(Permanent, urus::endpoint(cfg, pipeline)?)  <- Endpoint
      └── listener_sup  OneForOne over N listeners
          └── plain connection actors

- src/conn_registry.rs -> src/endpoint.rs. The registry gains the listener
  pool it registers for and becomes the Endpoint gen_server; ConnRegistry
  -> Endpoint. It runs inline as the ChildSpec's actor
  (NamedGenServerBuilder::run), so supervisor shutdown arrives as
  handle_shutdown and a restart re-runs init on the same still-open fds.
- Endpoint spawns its OWN listener sup in init rather than being its
  sibling: smarm's supervisor start order is not start *readiness* (spawn
  is fire-and-forget), so a sibling listener could whereis the name before
  the registry actor ran. Registrar-spawns-consumers makes it program
  order inside one init. Gap filed in smarm ROADMAP (readiness ack);
  making spawn blocking would only shrink the window, not close it —
  'has begun executing' is not 'has bound its name'.
- Listener sup is monitored: death outside shutdown = panic (loud, the
  app's supervisor decides) instead of a zombie on a dead port. Death
  during shutdown is the 'no new connections' barrier.
- DELETED: the shutdown AtomicBool, LISTENER_TICK (250ms wake per listener
  per tick, now an untimed wait_readable park), SHUTDOWN_POLL (100ms root
  poll — the root parks on the signal channel now), Restart::Transient
  (listeners are Permanent: they only exit by supervisor action, so a
  self-exit always means breakage). Verified against current smarm:
  request_stop unwinds an untimed wait_readable park, is no longer lossy
  against a QUEUED actor, and supervisor shutdown joins in ~200us.
- Config: scheduler_threads/max_actors removed (runtime knobs an
  endpoint-as-child cannot honour) -> serve_with(cfg, smarm::Config, pipe)
  and serve_with_shutdown(cfg, smarm::Config, pipe, signal). Added
  Config.name (default 'urus'): the endpoint's registry name, unique per
  endpoint, and the introspection handle via endpoint::whereis(name).
- serve* keep their meaning as the batteries-included path: they build a
  one-child tree around endpoint() with Shutdown::Infinity. Handle stays
  (a serve* caller has no RuntimeHandle to reach for) and now backs a real
  park instead of a poll.

Tests: 4 drain tests ported onto a real endpoint (bound socket, supervised
child, request_shutdown driven); new integration test boots an app tree
with an ordered sibling and asserts serve-then-drain, reverse-order
teardown and a closed port. 106 lib + 50 integration green.
This commit is contained in:
Claude
2026-08-20 12:59:57 +00:00
parent 5014b870d6
commit 099b6bc320
6 changed files with 840 additions and 689 deletions
+70 -295
View File
@@ -1,29 +1,19 @@
//! Listener pool and the `serve` entry point.
//! [`Config`] and the `serve*` entry points.
//!
//! A small fixed pool of listener actors share the same TCP listen fd (via
//! `dup`) — each blocks in non-blocking `accept4` + `wait_readable` on its
//! own copy. When a connection arrives the listener spawns a connection
//! actor with the `OwnedFd` and immediately returns to `accept`. No
//! coordination needed; the kernel serialises `accept` calls across the fds.
//! The listener pool itself lives in [`crate::endpoint`], which is the
//! real API: an endpoint is a supervisable child you place in your own
//! tree. `serve*` is the batteries-included path for a process whose only
//! job is serving HTTP — it owns the runtime and wraps one endpoint in a
//! one-child supervisor.
//!
//! Sharing via `dup` rather than the same fd is deliberate — Linux's
//! `accept4` is thread-safe on a single fd, but dup'ing per-listener keeps
//! each actor's epoll registration local to its own RawFd value (so smarm's
//! `waiters: HashMap<RawFd, Pid>` doesn't see collisions between listeners
//! waiting on "the same fd").
use crate::conn_actor::{run_connection, ConnLimits};
use crate::conn_registry::{self, ConnRegistry};
use crate::net::{accept_nonblocking, bind_and_listen, OwnedFd};
use crate::conn_actor::ConnLimits;
use crate::plug::Pipeline;
use smarm::{ChildSpec, OneForOne, Restart, GenServerRef, Strategy};
use smarm::supervisor::Shutdown;
use smarm::{ChildSpec, OneForOne, Restart};
use std::io::{self, ErrorKind};
use std::net::{SocketAddr, ToSocketAddrs};
use std::os::fd::RawFd;
use std::sync::atomic::{AtomicBool, AtomicU32, Ordering};
use std::sync::Arc;
use std::time::Duration;
// ---------------------------------------------------------------------------
@@ -67,10 +57,11 @@ pub struct Config {
/// WebSocket: cap on a complete reassembled message (spans
/// fragments; violation closes 1009).
pub max_message_bytes: usize,
/// Number of smarm scheduler OS threads. `None` means smarm's default
/// (one per CPU). Set this to a small fixed number in tests so multiple
/// concurrent test servers don't oversubscribe the host.
pub scheduler_threads: Option<usize>,
/// Registry name for this endpoint's gen_server — how it is addressed
/// from elsewhere in the app ([`crate::endpoint::whereis`]), and what
/// must be unique between two endpoints in one process (a public and
/// an admin port, say). Default `"urus"`.
pub name: &'static str,
/// Stack reserve (RFC 019 `smarm::SpawnOpts::stack_reserve`) given to
/// each per-connection actor. Request handlers routinely pull in
/// application code — DB drivers, (de)compression, templating — whose
@@ -79,12 +70,6 @@ pub struct Config {
/// Default: 256 KiB. The reserve is virtual/demand-paged, so raising it
/// costs address space, not RSS, until a handler actually uses it.
pub conn_stack_reserve: usize,
/// Maximum concurrently-live actors — smarm's fixed slot slab, allocated
/// once at init. Each connection is one actor, so this is also the hard
/// cap on concurrent connections. `None` uses smarm's default (16_384).
/// Slots are ~256 B, so raising this is cheap relative to per-connection
/// stacks; size it to peak concurrent connections.
pub max_actors: Option<usize>,
}
/// Default per-connection actor stack reserve (see [`Config::conn_stack_reserve`]).
@@ -111,13 +96,12 @@ impl Config {
drain_timeout: Duration::from_secs(30),
max_frame_payload: 1024 * 1024,
max_message_bytes: 4 * 1024 * 1024,
scheduler_threads: None,
name: "urus",
conn_stack_reserve: DEFAULT_CONN_STACK_RESERVE,
max_actors: None,
}
}
fn to_conn_limits(&self) -> ConnLimits {
pub(crate) fn to_conn_limits(&self) -> ConnLimits {
ConnLimits {
max_headers: self.max_header_count,
initial_read_buf: self.read_buf_size,
@@ -209,129 +193,14 @@ impl Config {
}
// ---------------------------------------------------------------------------
// dup helper
// Handle / ShutdownSignal — graceful shutdown plumbing for the serve* entries.
// ---------------------------------------------------------------------------
fn dup_fd(fd: RawFd) -> io::Result<OwnedFd> {
let new_fd = unsafe { libc::fcntl(fd, libc::F_DUPFD_CLOEXEC, 0) };
if new_fd < 0 {
return Err(io::Error::last_os_error());
}
Ok(OwnedFd::from_raw(new_fd))
}
// ---------------------------------------------------------------------------
// listener actor body
// ---------------------------------------------------------------------------
/// Test-only fault injection. When nonzero, the next accept-loop iteration
/// of whichever listener gets there first decrements this and panics —
/// *before* calling `accept`, so a pending connection stays in the kernel
/// backlog and must be picked up by the restarted listener. Cost when idle
/// is one relaxed load per accept-loop iteration (each of which already
/// pays a syscall). Not public API.
#[doc(hidden)]
pub static INJECT_LISTENER_PANICS: AtomicU32 = AtomicU32::new(0);
fn listener_loop(
listener: Arc<OwnedFd>,
pipeline: Pipeline,
limits: ConnLimits,
conn_stack_reserve: usize,
registry: GenServerRef<ConnRegistry>,
shutdown: Arc<AtomicBool>,
) {
let fd = listener.as_raw();
loop {
// Self-termination on shutdown — the ONLY way a listener exits at
// shutdown, and deliberately a normal return: under
// `Restart::Transient` a normal exit is terminal, so the
// supervisor's active-count drains and `sup.run()` returns on its
// own. No pid is ever `request_stop`ped, which sidesteps both the
// spawn-to-registration race of self-announced pids and smarm's
// lossy stop-while-QUEUED window (a flag is wake-free and
// race-free; a freshly spawned listener observes it on its very
// first iteration, a parked one within LISTENER_TICK).
if shutdown.load(Ordering::Relaxed) {
return;
}
if INJECT_LISTENER_PANICS
.fetch_update(Ordering::Relaxed, Ordering::Relaxed, |n| n.checked_sub(1))
.is_ok()
{
panic!("urus: injected listener panic (test hook)");
}
match accept_nonblocking(fd) {
Ok(client) => {
// Hand the fd off to a new connection actor. spawn() is
// cheap on smarm — it's a single Vec push under the
// shared lock.
let p = pipeline.clone();
let l = limits;
let r = registry.clone();
let opts = smarm::SpawnOpts {
stack_reserve: Some(conn_stack_reserve),
..smarm::SpawnOpts::default()
};
smarm::spawn_with(opts, move || run_connection(client, p, l, r));
}
Err(e) if e.kind() == ErrorKind::WouldBlock => {
// No pending connection. Park until the listener is
// readable or the tick elapses; either way we come back
// around through the shutdown-flag check above.
match smarm::wait_readable_timeout(fd, LISTENER_TICK) {
Ok(_ready) => {} // ready or tick — loop re-checks, retries accept
Err(we) => {
// epoll registration failed. Under Transient
// supervision a normal return is terminal but a
// panic restarts us — and a failed wait IS
// abnormal, so panic: a transient failure (e.g.
// EMFILE on the epoll set) heals by restart
// instead of silently shrinking the pool. (smarm
// catches actor panics in the trampoline; this is
// a Signal::Panic to the supervisor, not process
// noise.)
panic!("urus: listener wait_readable failed: {we}");
}
}
}
Err(e) if e.kind() == ErrorKind::Interrupted => {
continue;
}
Err(e) => {
// EMFILE / ENFILE / ECONNABORTED etc. Log and continue;
// the system may recover.
eprintln!("urus: accept error: {e}");
// Small backoff via smarm's sleep to avoid spinning if
// the error is sticky.
smarm::sleep(Duration::from_millis(10));
}
}
}
// The Arc clone we were started with drops here (normal exit or
// unwind), but the ChildSpec factory holds another — the fd outlives
// any one incarnation of this listener.
}
// ---------------------------------------------------------------------------
// Handle / ShutdownSignal — graceful shutdown plumbing.
// ---------------------------------------------------------------------------
/// How often the root actor polls for a shutdown signal (see
/// `serve_with_shutdown` for why this is a poll); bounds shutdown latency.
const SHUTDOWN_POLL: Duration = Duration::from_millis(100);
/// Listener accept-waits are timed at this tick; the shutdown flag is
/// observed at the top of every accept-loop iteration, so this bounds how
/// long a fully idle listener takes to notice shutdown. It is the SOLE
/// listener-shutdown mechanism (no `request_stop` — see the shutdown
/// sequence notes). Idle cost: one timer wake per listener per tick.
const LISTENER_TICK: Duration = Duration::from_millis(250);
/// A clonable trigger for graceful shutdown. Safe to use from any OS
/// thread (the send only enqueues; the serving side polls), e.g. from a
/// signal-handling thread.
/// A clonable trigger for graceful shutdown, usable from any OS thread
/// (e.g. a signal-handling thread) — the shutdown path for callers who let
/// `serve*` own the runtime and so have no [`smarm::RuntimeHandle`] of
/// their own. If you build the tree yourself with [`crate::endpoint`], use
/// `rt.handle().request_shutdown(root_sup)` instead and ignore this.
#[derive(Clone)]
pub struct Handle {
tx: smarm::Sender<()>,
@@ -340,8 +209,8 @@ pub struct Handle {
impl Handle {
/// Begin graceful shutdown: stop accepting, close idle keep-alive
/// connections, drain in-flight requests up to `Config.drain_timeout`,
/// then force-stop stragglers. `serve_with_shutdown` returns once the
/// runtime has wound down. Idempotent; extra calls are no-ops.
/// then force-stop stragglers. `serve*` returns once the runtime has
/// wound down. Idempotent; extra calls are no-ops.
pub fn shutdown(&self) {
let _ = self.tx.send(());
}
@@ -359,174 +228,80 @@ pub fn shutdown_handle() -> (Handle, ShutdownSignal) {
}
// ---------------------------------------------------------------------------
// serve_with_shutdown — main entry. Boots smarm, supervises listeners,
// blocks until shutdown.
// serve* — batteries-included entries for apps whose only job is serving.
// ---------------------------------------------------------------------------
//
// Boots an smarm runtime (one OS thread per CPU by default — see smarm's
// `Config::default()`). The root actor starts the connection registry and
// a one-for-one supervisor over the listener pool, then blocks on the
// shutdown signal. A panicking listener is restarted on the same
// (still-open) fd instead of silently shrinking the accept pool.
// Connection actors stay unsupervised bare spawns — per spec, a connection
// is cheap and its failure is local: a 500 path, not a restart path.
// (`spawn` from a listener parents the conn actor under that listener,
// which never registers a supervisor channel, so conn deaths are invisible
// to the pool supervisor by construction.)
//
// Shutdown sequence (drain-then-stop, the v0.2 chunk 2 decision):
// 1. Set the shared shutdown flag. Every listener observes it at the
// top of its accept loop (parks are timed at LISTENER_TICK) and
// returns normally. Under `Restart::Transient` a normal exit is
// terminal, so no restart happens and the supervisor's active count
// drains to zero.
// 2. Join the supervisor: `sup.run()` returns on its own once every
// listener has exited. The join is therefore the barrier "no new
// connections can ever be accepted" — listener fds are closed (the
// last Arc clones drop with the supervisor's ChildSpecs), and the
// kernel refuses new connects.
//
// Why a flag and not `request_stop`: stopping pids that announce
// themselves races the spawn-to-registration gap (a fast shutdown
// CAN beat a fresh listener to the registry), and smarm's
// `request_stop` is lossy against a QUEUED actor that then parks
// without passing an observation point — found the hard way; see
// the chunk 2 commit message. The flag is wake-free and race-free.
// 3. BeginDrain: the registry stops idle conns now and each remaining
// conn the moment it finishes its in-flight request.
// 4. Poll until no conns remain or the drain deadline passes; past the
// deadline, ForceStopConns every tick until the set empties (a conn
// accepted just before listener death may register late).
// 5. Root returns. `rt.run` itself returns only when every actor has
// exited (force-stopped conns unwind through their fd waits safely —
// smarm's 06-10 io fix — and close their sockets via OwnedFd::drop).
// These own the smarm runtime and build a one-child tree around
// `endpoint()`. An app with its own actors should call `endpoint()`
// directly and put it in its own supervision tree — that is the real API;
// everything here is a convenience wrapper over it.
/// Boot a runtime, serve until `signal` fires, then drain and return.
///
/// The tree is `root sup -> endpoint`, with the endpoint on
/// [`Shutdown::Infinity`] so its `drain_timeout` — not a supervisor
/// deadline — bounds the drain. The root actor parks on the signal
/// channel; a `Handle::shutdown` from a foreign OS thread wakes it, it
/// shuts the supervisor down (ordered, so the endpoint drains) and
/// `rt.run` returns when the last actor is gone.
pub fn serve_with_shutdown(
config: Config,
rt_config: smarm::Config,
pipeline: Pipeline,
signal: ShutdownSignal,
) -> io::Result<()> {
let listener = bind_and_listen(config.addr)?;
println!("urus: listening on {}", config.addr);
// One connection-actor-spawning loop per listener pool slot. Each gets
// its own dup'd fd so epoll registrations don't collide. Each fd is
// owned by its ChildSpec's factory closure (via `Arc`): a restarted
// listener re-enters `accept`/`wait_readable` on the same fd — no
// re-dup, no window where the slot has no fd.
let mut listener_fds = Vec::with_capacity(config.listener_pool);
listener_fds.push(Arc::new(listener)); // primary keeps the original
for _ in 1..config.listener_pool {
let dup = dup_fd(listener_fds[0].as_raw())?;
listener_fds.push(Arc::new(dup));
}
let limits = config.to_conn_limits();
let conn_stack_reserve = config.conn_stack_reserve;
let drain_timeout = config.drain_timeout;
let mut smarm_cfg = match config.scheduler_threads {
Some(n) => smarm::Config::exact(n),
None => smarm::Config::default(),
};
if let Some(m) = config.max_actors {
smarm_cfg = smarm_cfg.max_actors(m);
}
let rt = smarm::init(smarm_cfg);
// Listener self-termination flag — see the shutdown sequence below.
let shutdown_flag = Arc::new(AtomicBool::new(false));
let addr = config.addr;
let endpoint = crate::endpoint::endpoint(config, pipeline)?;
println!("urus: listening on {addr}");
let rt = smarm::init(rt_config);
rt.run(move || {
// Registry first: listeners and conns cast into it from birth.
let registry = conn_registry::start(drain_timeout);
let sup = smarm::spawn(move || {
OneForOne::new()
.child(
ChildSpec::new(Restart::Permanent, endpoint).shutdown(Shutdown::Infinity),
)
.run()
});
let mut sup = OneForOne::new().strategy(Strategy::OneForOne);
for (i, lfd) in listener_fds.into_iter().enumerate() {
let p = pipeline.clone();
let r = registry.clone();
let sf = shutdown_flag.clone();
sup = sup.child(ChildSpec::new(Restart::Transient, move || {
println!("urus: listener {} starting", i);
listener_loop(lfd.clone(), p.clone(), limits, conn_stack_reserve, r.clone(), sf.clone());
}));
}
// Default intensity (3 per 5s) applies; a listener crash-looping
// faster than that trips the cap and tears the pool down — loud
// failure over a zombie server.
let sup_h = smarm::spawn(move || sup.run());
// The old `urus.server` / `urus.listener.{i}` name registrations
// are gone with smarm's RFC 014 registry rework: `register` is now
// `(Name<M>, Sender<M>)`, self-only — a name is a typed messaging
// endpoint, not a pid tag. urus's bindings were introspection-only
// with no channel behind them, so they were dropped rather than
// faked with a unit channel. A real messageable `urus.server`
// name is in the icebox (ROADMAP.md).
// Block until told to shut down. We poll `try_recv` + `sleep`
// rather than parking in `recv`: a smarm `Sender::send` from a
// foreign OS thread (no runtime in its TLS) enqueues fine but its
// unpark is a `try_with_runtime` no-op — a parked receiver would
// never wake. The timer wake comes from inside the runtime, so the
// poll sees the message within one interval. (A cross-thread-safe
// unpark is a smarm roadmap candidate; this poll dies with it.)
// If every Handle was dropped the channel closes and no shutdown
// can ever arrive: serve forever, exactly v1's semantics.
loop {
match signal.rx.try_recv() {
Ok(Some(())) => break,
Ok(None) => smarm::sleep(SHUTDOWN_POLL),
Err(_) => smarm::sleep(Duration::from_secs(3600)),
// Park until told to shut down. If every Handle was dropped the
// channel closes and no shutdown can ever arrive: serve forever,
// exactly v1's semantics.
match signal.rx.recv() {
Ok(()) => smarm::request_shutdown(sup.pid()),
Err(_) => {
// Sender side gone. Park indefinitely; the process is
// expected to be killed externally.
loop {
smarm::sleep(Duration::from_secs(3600));
}
}
}
// ----- Shutdown. -----
// 1 + 2. Flag the listeners down and join the supervisor; the
// join returns once every listener has exited normally
// (Transient: normal exit is terminal). After this point
// no connection can ever be accepted again.
shutdown_flag.store(true, Ordering::Relaxed);
let _ = sup_h.join();
// 3 + 4. Drain: the registry owns the whole protocol now (idle
// stopped immediately, busy until its internal drain_timeout
// timer, late registrants stopped on arrival — see
// conn_registry docs). `GenServerRef::shutdown()` delivers the
// request and blocks on a monitor until the registry has
// stopped itself, which it does only once the conn set is
// empty: this line IS the "every connection is gone" barrier.
registry.shutdown();
// 5. Root returns; the runtime winds down when the last actor
// exits.
let _ = sup.join();
});
Ok(())
}
// ---------------------------------------------------------------------------
// serve_with / serve — convenience entries without a shutdown handle.
// ---------------------------------------------------------------------------
pub fn serve_with(config: Config, pipeline: Pipeline) -> io::Result<()> {
// The Handle is dropped immediately: shutdown can never be signalled
// and the server runs until externally killed (v1 semantics).
/// [`serve_with_shutdown`] without a shutdown handle: serves until the
/// process is killed.
pub fn serve_with(config: Config, rt_config: smarm::Config, pipeline: Pipeline) -> io::Result<()> {
// The Handle is dropped immediately: shutdown can never be signalled.
let (_handle, signal) = shutdown_handle();
serve_with_shutdown(config, pipeline, signal)
serve_with_shutdown(config, rt_config, pipeline, signal)
}
// ---------------------------------------------------------------------------
// serve — convenience over serve_with.
// ---------------------------------------------------------------------------
/// Defaults all round: default [`Config`], default smarm runtime (one
/// scheduler thread per CPU), serve until killed.
pub fn serve(addr: impl ToSocketAddrs, pipeline: Pipeline) -> io::Result<()> {
let addr = addr
.to_socket_addrs()?
.next()
.ok_or_else(|| io::Error::new(ErrorKind::InvalidInput, "no addresses resolved"))?;
serve_with(Config::new(addr), pipeline)
serve_with(Config::new(addr), smarm::Config::default(), pipeline)
}
#[cfg(all(test, feature = "config-file"))]
mod config_file_tests {
use super::*;