Files
urus/src/serve.rs
T
Claude 8f0da2a806 feat(pubsub,channels)!: handles are addresses — bus, hub and session registries are supervised children
smarm 0.7 (415effb, "lifetime is the actor's — refs are addresses") removed the
rule these three actors were built on: a GenServerRef no longer owns the server,
the loop holds its own inbox sender, and the inbox never closes when the last ref
drops. urus's pubsub table, channel hub bus and session registries were still
governed by that deleted rule — PubSub::new() spawned the table and the handle
owned its life — so nothing commanded them to stop. What still terminated a run
was the root-exit sweep, racing the drain: shutdown_with_open_chat_terminates and
channels_wire::shutdown_with_open_channel_terminates failed 4 times in 25
--all-features runs with "serve did not return: ... outlived the drain: Timeout".
Zero in 10 full-suite runs after this change.

The fix is not a supervisor wrapped around the old shape. Every gotcha in this
area descended from constructors that spawn: PubSub::new(), ChannelHub::new() and
PrefixRouter::channel_session() all started actors, which forced in-runtime-only
construction, which forced the Arc<OnceLock<..>> lazy-init from the first handler,
which forced the "cell must not be static" and "a relay must never hold a PubSub
clone" rules. Five documented rules propping up one inverted dependency. So:
description is separated from instantiation.

- PubSub<M> is a name, not a GenServerRef: const-constructible, Copy, spawns
  nothing, valid outside the runtime and in a static. Operations resolve through
  the registry per call, so a table restarted by its supervisor is reached
  transparently (one lookup per broadcast — bench before caching a ref, which
  would go stale across exactly the restart the supervisor exists to perform).
  PubSub::new() is gone; PubSub::new(name) + PubSub::child() replace it.
- ChannelHub::new(bus, router) returns (hub, Vec<ChildSpec>) — the bus table plus
  one registry per session route. Returning both is the point: a hub whose
  children were never started compiles and fails on the first join, so the vec is
  not left behind a method you can forget to call. #[must_use].
- channel_session gains a registry name; each session registry is separately
  named and separately supervised.
- serve_with/serve_with_shutdown take a Vec<ChildSpec> of app children and build
  the root as RestForOne[..app children, endpoint]. They start before the
  endpoint and, shutdown being ordered in reverse, stop after it has drained, so
  a request still in flight can reach the bus. RestForOne because a bus crash
  leaves live sockets addressing a table that no longer knows them.
- Deleted: the Arc<OnceLock> idiom from both examples and both test pipelines,
  and the module rules that existed only to hand-manage a refcount.

Known cost, not fixed here: channel_session("session:*", "chat-sessions", f) puts
two unrelated string literals side by side and nothing catches a transposition —
a RegistryName newtype is the obvious follow-up.

Tests: 111 lib + 50 integration + 2 doc green, clippy clean, 10/10 full-suite
runs. Unit tests poll for name binding before use — smarm's start-order-is-not-
start-readiness gap; real apps don't hit it, since a handler only runs once a
connection has been accepted.
2026-08-20 14:49:23 +00:00

380 lines
15 KiB
Rust

//! [`Config`] and the `serve*` entry points.
//!
//! The listener pool itself lives in [`crate::endpoint`], which is the
//! real API: an endpoint is a supervisable child you place in your own
//! tree. `serve*` is the batteries-included path for a process whose only
//! job is serving HTTP — it owns the runtime and wraps one endpoint in a
//! one-child supervisor.
//!
use crate::conn_actor::ConnLimits;
use crate::plug::Pipeline;
use smarm::supervisor::Shutdown;
use smarm::{ChildSpec, OneForOne, Restart, Strategy};
use std::io::{self, ErrorKind};
use std::net::{SocketAddr, ToSocketAddrs};
use std::time::Duration;
// ---------------------------------------------------------------------------
// Config
// ---------------------------------------------------------------------------
#[derive(Clone, Debug)]
pub struct Config {
pub addr: SocketAddr,
pub listener_pool: usize,
pub keep_alive_timeout: Duration,
pub max_header_count: usize,
pub read_buf_size: usize,
/// Wall-clock budget for reading the request HEAD (from first byte to
/// full head parse). Kept short — an incomplete head is the classic
/// slowloris. See `ConnLimits::head_timeout`.
pub head_timeout: Duration,
/// Absolute wall-clock cap on reading the request BODY (from head-parse
/// to full body). Sized for slow links, so much larger than
/// `head_timeout`. See `ConnLimits::body_timeout`.
pub body_timeout: Duration,
/// Burst size that resets the body stall clock. A body dribbling fewer
/// than this per `body_stall_timeout` window is evicted — the slowloris
/// / slow-legit discriminator. See `ConnLimits::body_burst_bytes`.
pub body_burst_bytes: usize,
/// Max time since the last qualifying body burst before eviction;
/// backstopped by `body_timeout`. See `ConnLimits::body_stall_timeout`.
pub body_stall_timeout: Duration,
/// Per-write budget for response bytes (the fixed head+body write, and
/// each streamed chunk). See `ConnLimits::write_timeout`.
pub write_timeout: Duration,
pub max_body_bytes: usize,
/// How long a graceful shutdown waits for in-flight requests before
/// force-stopping the remaining connections. Idle keep-alive
/// connections are closed immediately on shutdown and do not run the
/// clock out.
pub drain_timeout: Duration,
/// WebSocket: cap on a single frame's payload (header-checked
/// before buffering; violation closes 1009).
pub max_frame_payload: usize,
/// WebSocket: cap on a complete reassembled message (spans
/// fragments; violation closes 1009).
pub max_message_bytes: usize,
/// Registry name for this endpoint's gen_server — how it is addressed
/// from elsewhere in the app ([`crate::endpoint::whereis`]), and what
/// must be unique between two endpoints in one process (a public and
/// an admin port, say). Default `"urus"`.
pub name: &'static str,
/// Stack reserve (RFC 019 `smarm::SpawnOpts::stack_reserve`) given to
/// each per-connection actor. Request handlers routinely pull in
/// application code — DB drivers, (de)compression, templating — whose
/// stack needs comfortably exceed smarm's bare-actor default of 64 KiB
/// (the exact shape of bug this exists to head off; see smarm RFC 019).
/// Default: 256 KiB. The reserve is virtual/demand-paged, so raising it
/// costs address space, not RSS, until a handler actually uses it.
pub conn_stack_reserve: usize,
}
/// Default per-connection actor stack reserve (see [`Config::conn_stack_reserve`]).
pub const DEFAULT_CONN_STACK_RESERVE: usize = 256 * 1024;
impl Config {
pub fn new(addr: SocketAddr) -> Self {
let pool = std::thread::available_parallelism()
.map(|n| n.get())
.unwrap_or(2)
.max(2);
Self {
addr,
listener_pool: pool,
keep_alive_timeout: Duration::from_secs(60),
max_header_count: 64,
read_buf_size: 8 * 1024,
head_timeout: Duration::from_secs(30),
body_timeout: Duration::from_secs(300),
body_burst_bytes: 4 * 1024,
body_stall_timeout: Duration::from_secs(20),
write_timeout: Duration::from_secs(30),
max_body_bytes: 16 * 1024 * 1024,
drain_timeout: Duration::from_secs(30),
max_frame_payload: 1024 * 1024,
max_message_bytes: 4 * 1024 * 1024,
name: "urus",
conn_stack_reserve: DEFAULT_CONN_STACK_RESERVE,
}
}
pub(crate) fn to_conn_limits(&self) -> ConnLimits {
ConnLimits {
max_headers: self.max_header_count,
initial_read_buf: self.read_buf_size,
max_head_bytes: 64 * 1024,
max_body_bytes: self.max_body_bytes,
keep_alive_timeout: self.keep_alive_timeout,
head_timeout: self.head_timeout,
body_timeout: self.body_timeout,
body_burst_bytes: self.body_burst_bytes,
body_stall_timeout: self.body_stall_timeout,
write_timeout: self.write_timeout,
max_frame_payload: self.max_frame_payload,
max_message_bytes: self.max_message_bytes,
}
}
}
// ---------------------------------------------------------------------------
// config-file: TOML overlay for tuning knobs
// ---------------------------------------------------------------------------
//
// urus is a library, so it never presumes a config-file path or reads the
// environment — the embedding binary decides where a file lives and hands
// the text here. This overlays a sparse TOML document onto an existing
// `Config` (built with an addr the binary chose): only the keys present are
// applied, everything else keeps the compiled default. Durations are
// integer seconds. Unknown keys are a hard error so a typo is loud, not a
// silent no-op.
//
// Scope for now: the slowloris-tuning knobs only. Migrating the rest of the
// Config surface into the file is a separate, additive job (the loader
// mechanism is general — it just extends `TomlOverrides`).
/// Error from [`Config::with_toml_str`]: the TOML failed to parse or carried
/// an unknown/mistyped key.
#[cfg(feature = "config-file")]
#[derive(Debug)]
pub enum ConfigError {
Toml(String),
}
#[cfg(feature = "config-file")]
impl std::fmt::Display for ConfigError {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
match self {
ConfigError::Toml(m) => write!(f, "config TOML error: {m}"),
}
}
}
#[cfg(feature = "config-file")]
impl std::error::Error for ConfigError {}
#[cfg(feature = "config-file")]
#[derive(serde::Deserialize)]
#[serde(deny_unknown_fields)]
struct TomlOverrides {
head_timeout_secs: Option<u64>,
body_timeout_secs: Option<u64>,
body_burst_bytes: Option<usize>,
body_stall_timeout_secs: Option<u64>,
}
#[cfg(feature = "config-file")]
impl Config {
/// Overlay a TOML document of tuning knobs onto this config (sparse:
/// only the keys present are applied). Durations are integer seconds.
///
/// Recognized keys: `head_timeout_secs`, `body_timeout_secs`,
/// `body_burst_bytes`, `body_stall_timeout_secs`. Unknown keys error.
/// Other `Config` knobs are not yet file-configurable.
pub fn with_toml_str(mut self, s: &str) -> Result<Self, ConfigError> {
let o: TomlOverrides =
toml::from_str(s).map_err(|e| ConfigError::Toml(e.to_string()))?;
if let Some(v) = o.head_timeout_secs {
self.head_timeout = Duration::from_secs(v);
}
if let Some(v) = o.body_timeout_secs {
self.body_timeout = Duration::from_secs(v);
}
if let Some(v) = o.body_burst_bytes {
self.body_burst_bytes = v;
}
if let Some(v) = o.body_stall_timeout_secs {
self.body_stall_timeout = Duration::from_secs(v);
}
Ok(self)
}
}
// ---------------------------------------------------------------------------
// Handle / ShutdownSignal — graceful shutdown plumbing for the serve* entries.
// ---------------------------------------------------------------------------
/// A clonable trigger for graceful shutdown, usable from any OS thread
/// (e.g. a signal-handling thread) — the shutdown path for callers who let
/// `serve*` own the runtime and so have no [`smarm::RuntimeHandle`] of
/// their own. If you build the tree yourself with [`crate::endpoint`], use
/// `rt.handle().request_shutdown(root_sup)` instead and ignore this.
#[derive(Clone)]
pub struct Handle {
tx: smarm::Sender<()>,
}
impl Handle {
/// Begin graceful shutdown: stop accepting, close idle keep-alive
/// connections, drain in-flight requests up to `Config.drain_timeout`,
/// then force-stop stragglers. `serve*` returns once the runtime has
/// wound down. Idempotent; extra calls are no-ops.
pub fn shutdown(&self) {
let _ = self.tx.send(());
}
}
/// The receiving half consumed by [`serve_with_shutdown`].
pub struct ShutdownSignal {
rx: smarm::Receiver<()>,
}
/// Create a connected [`Handle`]/[`ShutdownSignal`] pair.
pub fn shutdown_handle() -> (Handle, ShutdownSignal) {
let (tx, rx) = smarm::channel();
(Handle { tx }, ShutdownSignal { rx })
}
// ---------------------------------------------------------------------------
// serve* — batteries-included entries for apps whose only job is serving.
// ---------------------------------------------------------------------------
//
// These own the smarm runtime and build a one-child tree around
// `endpoint()`. An app with its own actors should call `endpoint()`
// directly and put it in its own supervision tree — that is the real API;
// everything here is a convenience wrapper over it.
/// Boot a runtime, serve until `signal` fires, then drain and return.
///
/// The tree is `root sup -> [..app_children, endpoint]` on
/// [`Strategy::RestForOne`], with the endpoint on [`Shutdown::Infinity`]
/// so its `drain_timeout` — not a supervisor deadline — bounds the drain.
///
/// `app_children` is your application's actors: a
/// [`PubSub::child`](crate::PubSub::child), a
/// [`ChannelHub::children`](crate::channels::ChannelHub::children), your
/// own state servers. They start **before** the endpoint and, because
/// shutdown is ordered in reverse, stop **after** it has drained — so a
/// request still in flight can still reach the bus. `RestForOne` in that
/// order also means one of them crashing restarts the endpoint behind it,
/// dropping connections whose subscriptions died with it, rather than
/// leaving live sockets addressing a table that no longer knows them.
///
/// The root actor parks on the signal channel; a `Handle::shutdown` from a
/// foreign OS thread wakes it, it shuts the supervisor down and `rt.run`
/// returns when the last actor is gone.
pub fn serve_with_shutdown(
config: Config,
rt_config: smarm::Config,
pipeline: Pipeline,
app_children: Vec<ChildSpec>,
signal: ShutdownSignal,
) -> io::Result<()> {
let addr = config.addr;
let endpoint = crate::endpoint::endpoint(config, pipeline)?;
println!("urus: listening on {addr}");
let rt = smarm::init(rt_config);
rt.run(move || {
let sup = smarm::spawn(move || {
let mut sup = OneForOne::new().strategy(Strategy::RestForOne);
for child in app_children {
sup = sup.child(child);
}
sup.child(ChildSpec::new(Restart::Permanent, endpoint).shutdown(Shutdown::Infinity))
.run()
});
// Park until told to shut down. If every Handle was dropped the
// channel closes and no shutdown can ever arrive: serve forever,
// exactly v1's semantics.
match signal.rx.recv() {
Ok(()) => smarm::request_shutdown(sup.pid()),
Err(_) => {
// Sender side gone. Park indefinitely; the process is
// expected to be killed externally.
loop {
smarm::sleep(Duration::from_secs(3600));
}
}
}
let _ = sup.join();
});
Ok(())
}
/// [`serve_with_shutdown`] without a shutdown handle: serves until the
/// process is killed.
pub fn serve_with(
config: Config,
rt_config: smarm::Config,
pipeline: Pipeline,
app_children: Vec<ChildSpec>,
) -> io::Result<()> {
// The Handle is dropped immediately: shutdown can never be signalled.
let (_handle, signal) = shutdown_handle();
serve_with_shutdown(config, rt_config, pipeline, app_children, signal)
}
/// Defaults all round: default [`Config`], default smarm runtime (one
/// scheduler thread per CPU), no app children, serve until killed.
pub fn serve(addr: impl ToSocketAddrs, pipeline: Pipeline) -> io::Result<()> {
let addr = addr
.to_socket_addrs()?
.next()
.ok_or_else(|| io::Error::new(ErrorKind::InvalidInput, "no addresses resolved"))?;
serve_with(Config::new(addr), smarm::Config::default(), pipeline, Vec::new())
}
#[cfg(all(test, feature = "config-file"))]
mod config_file_tests {
use super::*;
fn base() -> Config {
Config::new("127.0.0.1:0".parse().unwrap())
}
#[test]
fn toml_empty_keeps_defaults() {
let d = base();
let c = base().with_toml_str("").unwrap();
assert_eq!(c.head_timeout, d.head_timeout);
assert_eq!(c.body_timeout, d.body_timeout);
assert_eq!(c.body_burst_bytes, d.body_burst_bytes);
assert_eq!(c.body_stall_timeout, d.body_stall_timeout);
}
#[test]
fn toml_partial_overrides_only_named() {
let d = base();
let c = base().with_toml_str("head_timeout_secs = 5").unwrap();
assert_eq!(c.head_timeout, Duration::from_secs(5)); // overridden
assert_eq!(c.body_timeout, d.body_timeout); // default kept
assert_eq!(c.body_burst_bytes, d.body_burst_bytes); // default kept
assert_eq!(c.body_stall_timeout, d.body_stall_timeout);
}
#[test]
fn toml_full_overrides_all() {
let c = base()
.with_toml_str(
"head_timeout_secs = 10\n\
body_timeout_secs = 120\n\
body_burst_bytes = 8192\n\
body_stall_timeout_secs = 15\n",
)
.unwrap();
assert_eq!(c.head_timeout, Duration::from_secs(10));
assert_eq!(c.body_timeout, Duration::from_secs(120));
assert_eq!(c.body_burst_bytes, 8192);
assert_eq!(c.body_stall_timeout, Duration::from_secs(15));
}
#[test]
fn toml_unknown_key_errors() {
// A mistyped/unknown key is a hard error, not a silent no-op.
let e = base().with_toml_str("body_timeout_sec = 120"); // typo: missing 's'
assert!(e.is_err(), "unknown key should error");
}
#[test]
fn toml_malformed_errors() {
let e = base().with_toml_str("this is not = valid = toml");
assert!(e.is_err(), "malformed TOML should error");
}
}