17 Commits
Author SHA1 Message Date
Claude ab93802d0b release: v0.3.0
Bump crate version to 0.3.0; pin smarm dependency to the v0.7.0 tag (dev path override reverted).
2026-08-21 13:10:56 +02:00
Claude 5442271229 build(examples): gate load_profile behind smarm-causal
smarm::causal is feature-gated, so examples/load_profile.rs no longer
compiles under the default feature set and broke `cargo build --examples`.
Declare it with required-features = ["smarm-causal"] so the default build
skips it and `--features smarm-causal` still builds it.

Splits the committable half out of the working-tree Cargo.toml; the
`smarm = { path = "../smarm_full" }` dev pin stays uncommitted and is
restored to a tag in the release commit.
2026-08-20 19:38:24 +00:00
Claude e02527a606 docs: record the supervised-bus cycle (v0.8) and retire the OnceLock idiom
ROADMAP gains a v0.8 entry: the smarm 415effb lifetime change as root cause,
the description/instantiation split, the tree shape, the accepted per-call
resolution cost, and what stays open (RegistryName newtype, PubSub's
unprotectable const, dynamic session actors, hammer.sh's missing feature
matrix). The v0.7 "unreproduced test failure" open item is closed out and
pointed at it — it was this hang, hidden because hammer.sh builds with default
features while the failure needs --all-features load.

README and the two module headers still taught the v0.5 rules: in-runtime-only
construction, the non-static Arc<OnceLock<..>> cell, and "a relay must never
hold a PubSub clone". None of those are true any more; they are replaced by
what actually holds now, with a note on what changed for anyone who learnt the
old shape.

session.rs also gains an honest note that its session actors are the last
lifetime in urus implied by a drop rather than stated — they are dynamic, so a
fixed ChildSpec list cannot hold them, and the trigger is at least a command
now rather than a refcount.
2026-08-20 15:01:44 +00:00
Claude 9eaa85b9df style(tests): byte-string literal in the stalled-writer probe
clippy::byte_char_slices, pre-existing and unrelated to the refactor in the
previous commit — split out so that diff stays scoped. Restores a clean
`cargo clippy --all-features --all-targets`.
2026-08-20 14:49:29 +00:00
Claude 8f0da2a806 feat(pubsub,channels)!: handles are addresses — bus, hub and session registries are supervised children
smarm 0.7 (415effb, "lifetime is the actor's — refs are addresses") removed the
rule these three actors were built on: a GenServerRef no longer owns the server,
the loop holds its own inbox sender, and the inbox never closes when the last ref
drops. urus's pubsub table, channel hub bus and session registries were still
governed by that deleted rule — PubSub::new() spawned the table and the handle
owned its life — so nothing commanded them to stop. What still terminated a run
was the root-exit sweep, racing the drain: shutdown_with_open_chat_terminates and
channels_wire::shutdown_with_open_channel_terminates failed 4 times in 25
--all-features runs with "serve did not return: ... outlived the drain: Timeout".
Zero in 10 full-suite runs after this change.

The fix is not a supervisor wrapped around the old shape. Every gotcha in this
area descended from constructors that spawn: PubSub::new(), ChannelHub::new() and
PrefixRouter::channel_session() all started actors, which forced in-runtime-only
construction, which forced the Arc<OnceLock<..>> lazy-init from the first handler,
which forced the "cell must not be static" and "a relay must never hold a PubSub
clone" rules. Five documented rules propping up one inverted dependency. So:
description is separated from instantiation.

- PubSub<M> is a name, not a GenServerRef: const-constructible, Copy, spawns
  nothing, valid outside the runtime and in a static. Operations resolve through
  the registry per call, so a table restarted by its supervisor is reached
  transparently (one lookup per broadcast — bench before caching a ref, which
  would go stale across exactly the restart the supervisor exists to perform).
  PubSub::new() is gone; PubSub::new(name) + PubSub::child() replace it.
- ChannelHub::new(bus, router) returns (hub, Vec<ChildSpec>) — the bus table plus
  one registry per session route. Returning both is the point: a hub whose
  children were never started compiles and fails on the first join, so the vec is
  not left behind a method you can forget to call. #[must_use].
- channel_session gains a registry name; each session registry is separately
  named and separately supervised.
- serve_with/serve_with_shutdown take a Vec<ChildSpec> of app children and build
  the root as RestForOne[..app children, endpoint]. They start before the
  endpoint and, shutdown being ordered in reverse, stop after it has drained, so
  a request still in flight can reach the bus. RestForOne because a bus crash
  leaves live sockets addressing a table that no longer knows them.
- Deleted: the Arc<OnceLock> idiom from both examples and both test pipelines,
  and the module rules that existed only to hand-manage a refcount.

Known cost, not fixed here: channel_session("session:*", "chat-sessions", f) puts
two unrelated string literals side by side and nothing catches a transposition —
a RegistryName newtype is the obvious follow-up.

Tests: 111 lib + 50 integration + 2 doc green, clippy clean, 10/10 full-suite
runs. Unit tests poll for name binding before use — smarm's start-order-is-not-
start-readiness gap; real apps don't hit it, since a handler only runs once a
connection has been accepted.
2026-08-20 14:49:23 +00:00
Claude 8a568c600c docs(roadmap): record the endpoint milestone (crate v0.3.0)
Full account of the refactor: the tree shape, why the endpoint spawns its
own listener sup (supervisor start order != start readiness), the config
split, what was deleted and the smarm properties re-probed to justify
deleting it. Also:
- PubSub discovery note: the non-static Arc<OnceLock> rule is now
  serve*-only; an app that owns its tree starts the table as a supervised
  sibling and addresses it by name (what crud does).
- Icebox 'operator introspection via typed names': partly delivered — the
  endpoint is a named gen_server answering Call::ConnCount; per-listener
  visibility and richer stats remain open.
- Open after this cycle: the one unreproduced hammer failure, and that
  endpoint() returns impl Fn() so a second invocation clashes on the name.
2026-08-20 13:27:27 +00:00
Claude b37888ec2c docs+examples: v0.3 endpoint API; crud converted to the app-owns-the-tree shape
- crud is now the demonstrator: app owns smarm runtime + root supervisor,
  store actor and urus::endpoint as ordered siblings (store first, so
  reverse-order shutdown drains HTTP before stopping the store),
  Shutdown::Infinity on the endpoint, shutdown via
  rt.handle().request_shutdown(root_sup) from the stdin thread.
  Deletes two kludges the old shape forced:
    * static OnceLock<Sender> spawn-on-first-use -> a supervised child
      that self-registers a typed Name; handlers use smarm::send per
      request and turn 'between incarnations' into a 503 instead of
      panicking on a dropped store.
    * static SHUTTING_DOWN AtomicBool + 250ms recv_timeout poll in the
      store loop and in the SSE ticker -> a plain park; the tree stops
      both. Smoke-tested live: CRUD round-trips, SSE stream, clean drain
      with the stream open, port closed after.
- Other examples stay short and on serve*, updated for the split config
  (serve_with(cfg, smarm::Config, pipe) / serve_with_shutdown(..., signal)).
  plain_serve's URUS_SCHED_THREADS now builds a smarm::Config.
  ws_chat's doc block explains the OnceLock is a serve*-only workaround
  and points at crud for the clean shape.
- README: new 'Your Own Supervision Tree' section (endpoint as the real
  API), graceful shutdown reframed as the serve*-only path, Config table
  loses scheduler_threads and gains name, PubSub rule 1 notes the
  supervised-sibling alternative.

111 lib + 50 integration + 2 doc tests green; clippy clean.
2026-08-20 13:20:28 +00:00
Claude 099b6bc320 feat(endpoint): urus is a supervisable child — endpoint gen_server owns listeners + conns
The v0.3 shape from the spec: an app owns its runtime and root supervisor
and places urus in it as one ordered child among its own.

  your root sup
  └── ChildSpec(Permanent, urus::endpoint(cfg, pipeline)?)  <- Endpoint
      └── listener_sup  OneForOne over N listeners
          └── plain connection actors

- src/conn_registry.rs -> src/endpoint.rs. The registry gains the listener
  pool it registers for and becomes the Endpoint gen_server; ConnRegistry
  -> Endpoint. It runs inline as the ChildSpec's actor
  (NamedGenServerBuilder::run), so supervisor shutdown arrives as
  handle_shutdown and a restart re-runs init on the same still-open fds.
- Endpoint spawns its OWN listener sup in init rather than being its
  sibling: smarm's supervisor start order is not start *readiness* (spawn
  is fire-and-forget), so a sibling listener could whereis the name before
  the registry actor ran. Registrar-spawns-consumers makes it program
  order inside one init. Gap filed in smarm ROADMAP (readiness ack);
  making spawn blocking would only shrink the window, not close it —
  'has begun executing' is not 'has bound its name'.
- Listener sup is monitored: death outside shutdown = panic (loud, the
  app's supervisor decides) instead of a zombie on a dead port. Death
  during shutdown is the 'no new connections' barrier.
- DELETED: the shutdown AtomicBool, LISTENER_TICK (250ms wake per listener
  per tick, now an untimed wait_readable park), SHUTDOWN_POLL (100ms root
  poll — the root parks on the signal channel now), Restart::Transient
  (listeners are Permanent: they only exit by supervisor action, so a
  self-exit always means breakage). Verified against current smarm:
  request_stop unwinds an untimed wait_readable park, is no longer lossy
  against a QUEUED actor, and supervisor shutdown joins in ~200us.
- Config: scheduler_threads/max_actors removed (runtime knobs an
  endpoint-as-child cannot honour) -> serve_with(cfg, smarm::Config, pipe)
  and serve_with_shutdown(cfg, smarm::Config, pipe, signal). Added
  Config.name (default 'urus'): the endpoint's registry name, unique per
  endpoint, and the introspection handle via endpoint::whereis(name).
- serve* keep their meaning as the batteries-included path: they build a
  one-child tree around endpoint() with Shutdown::Infinity. Handle stays
  (a serve* caller has no RuntimeHandle to reach for) and now backs a real
  park instead of a poll.

Tests: 4 drain tests ported onto a real endpoint (bound socket, supervised
child, request_shutdown driven); new integration test boots an app tree
with an ordered sibling and asserts serve-then-drain, reverse-order
teardown and a closed port. 106 lib + 50 integration green.
2026-08-20 12:59:57 +00:00
Claude 5014b870d6 feat(conn_registry): registry owns the drain — trapping gen_server, request_shutdown driven
The drain protocol moves wholesale into the registry (smarm >=0.7:
trap_exit + handle_shutdown + gen_server timers + StopHandle):

- handle_shutdown: flip draining, stop idle conns, arm a drain_timeout
  timer, Continue; with no conns, Exit immediately.
- ConnIdle while draining stops the conn (unchanged); ConnStarted while
  draining stops it on arrival (unchanged); ConnEnded that empties the
  set while draining = StopHandle::stop() — the registry's own normal
  exit is now the 'every connection is gone' barrier.
- handle_timer (deadline): one force-stop sweep. The old re-sweep-every-
  10ms loop existed to catch late registrants; stop-on-arrival already
  covers every post-sweep entry path, so one sweep suffices.
- Cast::{BeginDrain,ForceStopConns} deleted (internal now); ConnCount
  stays as the introspection call. start() takes drain_timeout.
- serve.rs: the root's whole drain/poll block collapses to
  registry.shutdown() (graceful, monitors until the server has stopped
  itself). SHUTDOWN_POLL + the listener flag are untouched here; they go
  with the endpoint refactor.
- Known residual window documented in the module docs: a conn spawned
  by a dying listener that has not yet registered can outlive an
  already-empty registry; it is collected by smarm's root-exit sweep.

Tests (in-lib, request_shutdown driven): empty-set immediate exit;
idle-now/busy-at-deadline ordering with exit-after; stop-on-idle
mid-drain; stop-on-arrival mid-drain. 35x hammer subset green.
2026-08-20 08:57:19 +00:00
Claude (sandbox) 8bdec97842 feat(config): optional TOML config loading behind config-file feature
A file-based way to set tuning knobs without recompiling. urus is a
library, so it never presumes a config path or reads the environment — the
binary hands the text in:

- Config::with_toml_str(&str) -> Result<Config, ConfigError>: sparse overlay
  onto an existing Config (built with the addr the binary chose). Only keys
  present are applied; durations are integer seconds; unknown keys are a
  hard error (deny_unknown_fields) so a typo is loud, not a silent no-op.
- Scope: the four slowloris knobs (head_timeout_secs, body_timeout_secs,
  body_burst_bytes, body_stall_timeout_secs). Migrating the rest of Config
  into the file is a separate, additive job — TomlOverrides just grows.
- Feature `config-file = ["dep:serde", "dep:toml"]`; the optional serde dep
  gains the derive feature. The default build is unchanged (deps + code are
  all gated).
- examples/serve_toml.rs (required-features = ["config-file"]): a `--config
  PATH` demo with no presumed default location. plain_serve and its env
  vars are left untouched.

Tests (feature-gated): empty keeps defaults, partial overrides only named,
full overrides all, unknown key errors, malformed errors. 84 lib with the
feature / 79 without; clippy --lib clean both ways; e2e smoke serves 200
from a file and rejects an unknown key loudly.
2026-08-12 13:41:57 +00:00
Claude (sandbox) f3ccb6e468 feat(serve): burst-gated body stall eviction
The body budget from the prior commit is a generous absolute cap; alone it
just hands a body-phase slowloris a bigger window. Add a stall gate under
that cap that distinguishes a slowloris trickle from a slow-but-legit
client by requiring BURSTS, not a mere average rate:

- BodyStallGate: each body read is bounded by min(body cap, mark + stall).
  The stall mark advances only when body_burst_bytes accumulate since the
  last advance, so a steady sub-burst trickle never moves it and is evicted
  at ~body_stall_timeout, while a bursty slow client keeps resetting it.
- Two words of state; one add + one compare per read. Raw socket bytes are
  counted, so chunked framing counts and an MSS-fragmented burst still
  accumulates. Reuses the existing read_some deadline plumbing.
- Wired into read_body (fixed CL) and read_chunked_body (via fill_to, the
  single choke point all chunked reads pass through).
- New knobs body_burst_bytes (4 KiB) + body_stall_timeout (20s); effective
  floor ~205 B/s enforced in bursts.

Tests: body_smooth_trickle_evicted_at_stall_timeout (chunked; active
sub-burst trickle evicted at ~stall while the cap is far away) and
bursty_slow_body_survives_stall_gate (fixed CL; real bursts with sub-stall
gaps complete intact). 79 lib + 45 integration green; clippy --lib clean.
2026-08-12 13:38:16 +00:00
Claude (sandbox) 4f06265338 feat(serve): split request read budget into head_timeout + body_timeout
The single request_timeout covered head + body under one wall clock, so a
slow-but-legit body upload (e.g. a trickling cellular IoT client) was
judged by the short head deadline and killed mid-body. Split into:

- head_timeout (default 30s): first byte -> full head parse; the classic
  slowloris surface, kept short.
- body_timeout (default 300s): head parse -> full body; an absolute cap
  sized for slow links, anchored independently once the head has parsed.

read_head no longer returns a shared deadline; run_connection anchors the
body deadline itself. ReadHeadErr::RequestTimeout -> HeadTimeout. Config
and ConnLimits gain body_timeout; request_timeout renamed to head_timeout
(breaking, but this axis is unreleased).

Tests: slow_body_outlives_head_timeout (positive: body survives past the
head clock), fixed_/chunked_body_stall_killed_at_body_timeout (body cap
still bites), slowloris_partial_head_killed_at_head_timeout (head clock
unchanged). 79 lib + 43 integration green; clippy --lib clean.
2026-08-12 13:00:22 +00:00
Claude (sandbox) 1b1ea124c8 feat(serve): plumb Config.max_actors through to smarm init
Config gained a max_actors: Option<usize> (None = smarm's DEFAULT_MAX_ACTORS
of 16_384). serve_with now applies it to the smarm runtime config, so the
per-connection actor slab can be sized to the deployment's peak concurrent
connections. Since each connection is one actor, the slab was the hard cap
on concurrent connections (previously an un-raisable 16_384) regardless of
RAM/fds; slots are ~256 B so raising it is cheap next to per-conn stacks.

Verified on the GPU box: default caps at 16_383 held connections; with
max_actors raised, a paced ramp holds 100_000 concurrent slow-header
connections at 12.4 KB RSS / 2 VMAs each (1.29 GB total) on one pinned core.
2026-08-10 05:47:22 +00:00
Claude (sandbox) b86c64d490 feat(parser): strict Transfer-Encoding framing; unknown coding -> 501
The TE arm set chunked whenever the token appeared anywhere in the value,
so 'chunked, gzip' (chunked not final) was accepted and an unknown coding
like 'bogus' was treated as no-body (h1spec #18/#19 -> 404). Collect the
ordered coding list across all TE headers and decide post-loop: TE on
HTTP/1.0 or TE+Content-Length -> 400 (the CL check now covers ANY TE, not
just chunked, closing the old TE:unknown + CL smuggling gap); chunked
present but not final -> 400; any coding other than chunked -> 501 via a
new UnknownTransferCoding variant (emit_error_response gains the 501 arm);
only a sole final chunked sets the flag. Tests cover each branch.

Note: the dead Unsupported/411 variant is left as-is (separate cleanup).
2026-08-09 08:02:04 +00:00
Claude (sandbox) 394e9b962a feat(parser): reject duplicate Content-Length (RFC 9112 §6.3)
The content-length arm ran content_length = Some(parse) per header, so a
second Content-Length silently overwrote the first with no conflict check
(CL.CL request smuggling; h1spec #21 -> 404 instead of 400). Count
occurrences and reject any duplicate post-loop, strictly (even equal
values), reusing BadContentLength (400). A single value is still required
to be one decimal integer, so a comma-list or non-numeric keeps failing at
parse as before. Tests: differing dup, equal dup, single-CL regression.
2026-08-09 07:59:25 +00:00
Claude (sandbox) 6f02cec261 feat(parser): reject missing/duplicate/invalid Host (RFC 9112 §3.2)
parse_head never inspected Host, so a missing (HTTP/1.1), duplicate, or
syntactically invalid Host all passed through to the router (h1spec #8/#9/
#10 -> 404 instead of 400). Add per-header validity (RFC 3986 host[:port]
charset via valid_host) plus a post-loop presence/uniqueness check: 1.1
MUST carry exactly one valid Host; 1.0 may omit it but a duplicate/invalid
one is still 400. Unit matrix mirrors the three h1spec cases with reg-name/
port/IPv6-literal positive controls.
2026-08-09 07:58:16 +00:00
Markk116 535f7bcc68 feat(serve): give connection actors a 256 KiB stack via smarm SpawnOpts
Connection actors were still spawned with a bare smarm::spawn(), which
gets the runtime's fixed 64 KiB default stack regardless of smarm
v0.6.0's RFC 019 SpawnOpts/stack_reserve work landing one crate down.
Any handler that leans on app code with real stack needs (DB drivers,
(de)compression, ...) blows the guard page and the connection just
dies with no response - reproduced with a CCC handler that decompresses
gzip on the identity-encoding path.

Add Config::conn_stack_reserve (default DEFAULT_CONN_STACK_RESERVE =
256 KiB) and thread it through listener_loop into a
smarm::spawn_with(SpawnOpts { stack_reserve: Some(_), .. }, ...) call
for every accepted connection. Existing Config { ..Config::new(addr) }
call sites (tests/integration.rs) pick up the new field automatically
via struct-update syntax; no call-site churn beyond that.

Bump to 0.2.2.
2026-08-08 22:39:37 +02:00
23 changed files with 2425 additions and 896 deletions
+14 -3
View File
@@ -1,27 +1,30 @@
[package]
name = "urus"
version = "0.2.1"
version = "0.3.0"
edition = "2021"
rust-version = "1.95"
description = "Cowboy/bandit-style HTTP library for the smarm actor runtime"
license = "MIT"
[dependencies]
smarm = { git = "https://git.kalsbeek.dev/Markk116/smarm", tag = "v0.6.0" }
smarm = { git = "https://git.kalsbeek.dev/Markk116/smarm", tag = "v0.7.0" }
httparse = "1.9"
libc = "0.2"
sha1_smol = "1"
# dep #4, ratified 2026-06-12: serde/serde_json behind the opt-in
# "phoenix" feature only — the "channels" core stays dependency-free.
serde = { version = "1", optional = true }
serde = { version = "1", optional = true, features = ["derive"] }
serde_json = { version = "1", optional = true }
# config-file feature: TOML loader for tuning knobs (dep #5, 2026-08-12)
toml = { version = "0.8", optional = true }
[features]
smarm-trace = ["smarm/smarm-trace"]
smarm-causal = ["smarm/smarm-causal"]
channels = []
phoenix = ["channels", "dep:serde", "dep:serde_json"]
config-file = ["dep:serde", "dep:toml"]
[dev-dependencies]
serde = { version = "1", features = ["derive"] }
@@ -47,3 +50,11 @@ path = "examples/crud.rs"
name = "channels_chat"
path = "examples/channels_chat.rs"
required-features = ["phoenix"]
[[example]]
name = "serve_toml"
required-features = ["config-file"]
[[example]]
name = "load_profile"
required-features = ["smarm-causal"]
+79 -22
View File
@@ -122,7 +122,11 @@ Handlers are closures taking `(Conn, Next) -> Conn`. Call `Next::call(c)` to con
### Server Configuration
`serve(addr, pipeline)` binds and listens on the given address. For more control, use `serve_with(config, pipeline)`:
`serve(addr, pipeline)` binds and listens on the given address. For more
control, use `serve_with(config, runtime_config, pipeline)` — the urus
`Config` holds endpoint knobs, the `smarm::Config` holds runtime knobs
(they are separate because an endpoint placed in someone else's tree
cannot dictate the runtime):
```rust
use urus::{serve_with, Config};
@@ -130,7 +134,7 @@ use std::time::Duration;
let cfg = Config {
listener_pool: 2, // Supervised accept-loop actors
scheduler_threads: Some(2), // smarm worker threads (None = one per CPU)
name: "urus", // Endpoint registry name (unique per endpoint)
keep_alive_timeout: Duration::from_secs(60), // Idle budget between requests
request_timeout: Duration::from_secs(30), // Whole-request READ deadline
write_timeout: Duration::from_secs(30), // Per-write response budget
@@ -141,7 +145,7 @@ let cfg = Config {
..Config::new("127.0.0.1:8080".parse().unwrap())
};
serve_with(cfg, pipeline).unwrap();
serve_with(cfg, smarm::Config::exact(2), pipeline).unwrap();
```
**Timeout semantics:**
@@ -162,10 +166,51 @@ serve_with(cfg, pipeline).unwrap();
write stalls past the budget (the write-side twin of slowloris). Streamed
chunks each get a fresh budget; a stream as a whole has no deadline.
### Graceful Shutdown
### Your Own Supervision Tree
`serve_with_shutdown` takes a `ShutdownSignal`; the paired `Handle` can be
triggered from anywhere (another thread, a signal handler):
The real API is `urus::endpoint(config, pipeline)`, which binds the socket
and hands back a supervisable child body. Your application owns the
runtime and the root supervisor; urus is one child among your own actors:
```rust
use smarm::{ChildSpec, OneForOne, Restart, supervisor::Shutdown};
// Binds here: an address-in-use error is a startup failure, not an actor
// crash. The fds outlive any restart of the endpoint child.
let endpoint = urus::endpoint(cfg, pipeline)?;
let rt = smarm::init(smarm::Config::default());
rt.run(move || {
let sup = smarm::spawn(move || {
OneForOne::new()
// Your state actor FIRST: reverse-order shutdown therefore
// stops it LAST, after HTTP has finished draining.
.child(ChildSpec::new(Restart::Permanent, my_store))
.child(ChildSpec::new(Restart::Permanent, endpoint)
// The endpoint bounds its own drain with drain_timeout;
// a shorter supervisor deadline would cut it in half.
.shutdown(Shutdown::Infinity))
.run()
});
let _ = sup.join();
});
```
Shutdown is then whatever your app already does — `request_shutdown` on
the root supervisor (from a SIGTERM handler via `rt.handle()`, say). The
endpoint's `handle_shutdown` stops the listener pool, closes idle
keep-alive connections, drains in-flight requests up to `drain_timeout`,
force-stops stragglers past it, and exits only when the last connection is
gone — so "the endpoint child has stopped" *is* "every connection is
gone". See `examples/crud.rs` for a complete app in this shape.
Address a running endpoint by name for introspection:
`urus::endpoint::whereis("urus")`.
### Graceful Shutdown Without a Tree
If serving is all your process does, let `serve*` own the runtime and use
`Handle`, which triggers the same sequence from any thread:
```rust
use urus::{serve_with_shutdown, shutdown_handle, Config};
@@ -177,13 +222,13 @@ std::thread::spawn(move || {
handle.shutdown();
});
serve_with_shutdown(cfg, pipeline, signal).unwrap();
serve_with_shutdown(cfg, smarm::Config::default(), pipeline, signal).unwrap();
// Returns once the runtime has fully wound down.
```
`Handle::shutdown()` is idempotent and performs, in order:
1. Stop accepting — every listener exits; no new connections.
1. Stop accepting — the listener pool is stopped; no new connections.
2. Close idle keep-alive connections immediately.
3. Drain in-flight requests for up to `Config.drain_timeout`.
4. Force-stop any stragglers past the deadline (sockets close cleanly on
@@ -313,10 +358,11 @@ independent of HTTP (it imports only smarm) and built for the WebSocket
relay pattern:
```rust
let bus: PubSub<String> = PubSub::new(); // in-runtime only!
let rx = bus.subscribe("room:lobby")?; // Receiver<Arc<String>>
bus.broadcast("room:lobby", "hi".to_string())?;
bus.broadcast_from(smarm::self_pid(), "room:lobby", "no echo".into())?;
const BUS: PubSub<String> = PubSub::new("chat"); // a name; spawns nothing
// BUS.child() goes in your supervision tree (or serve_with*'s child vec)
let rx = BUS.subscribe("room:lobby")?; // Receiver<Arc<String>>
BUS.broadcast("room:lobby", "hi".to_string())?;
BUS.broadcast_from(smarm::self_pid(), "room:lobby", "no echo".into())?;
```
One generic instance per message domain; payloads broadcast as `Arc<M>`
@@ -328,16 +374,25 @@ explicit pid for relay patterns. Mailboxes are unbounded: `broadcast`
never blocks the table, and a slow subscriber's memory bill is bounded
by the two cleanup paths above.
Two composition rules that matter (both enforced by
`shutdown_with_open_chat_terminates` in the integration suite):
Composition (enforced by `shutdown_with_open_chat_terminates` in the
integration suite):
1. `PubSub::new()` spawns an actor, so it must run in-runtime — build it
lazily from a handler via a **non-static** `Arc<OnceLock<PubSub<M>>>`
captured by the route closure. A `static` cell pins the table actor
forever and graceful shutdown never returns.
1. **The handle is an address, not the table.** `PubSub<M>` is a name:
`const`, `Copy`, spawns nothing, fine in a `static` or outside the
runtime. The actor is `PubSub::child()`, a `ChildSpec` for your
supervision tree — or for `serve_with*`'s app-children vec, which puts
it ahead of the endpoint so it stops only after the endpoint has
drained. Operations resolve the name per call, so a restarted table is
reached transparently.
2. Relay/producer actors hold the `Receiver` (plus e.g. a `WsSender`
clone) — **never a `PubSub` clone**, or relay and table keep each
other alive past shutdown.
clone). They may hold the handle too — it pins nothing — but usually
have no use for one.
Prior to v0.8 both of these read the other way round: `PubSub::new()`
spawned the table, the handle owned its life, and a non-static
`Arc<OnceLock<PubSub<M>>>` lazily built from the first handler was the
required idiom. smarm 0.7 made a server's lifetime its own, and that
whole apparatus went away with it.
See [`examples/ws_chat.rs`](examples/ws_chat.rs): rooms as topics,
`on_open` subscribes + spawns the relay, `on_message` uses
@@ -377,8 +432,10 @@ impl Channel<P> for Room {
}
}
// in-runtime, non-static — the ws_chat OnceLock pattern applies
let hub = ChannelHub::new(PrefixRouter::new().channel_default::<Room>("room:*"));
// A description: spawns nothing. `children` are the ChildSpecs it needs
// (bus table + one registry per session route) — hand them to serve_with*.
let (hub, children) =
ChannelHub::new("chat-bus", PrefixRouter::new().channel_default::<Room>("room:*"));
// route handler: hub.upgrade(conn)
// from anywhere with a hub handle: hub.broadcast("room:lobby", "news", payload)
```
+157 -6
View File
@@ -252,6 +252,10 @@ smarm has none). All six design questions ratified by the user pre-code
(crud's store shape) hangs `serve_with_shutdown`. Corollary: relays
hold `Receiver` only, never a `PubSub` clone (mutual-keepalive cycle).
Proven by `shutdown_with_open_chat_terminates`.
**Superseded by v0.7 where the app owns its tree:** start the table as a
supervised sibling of the endpoint and address it by name (what crud now
does with its store). The `OnceLock` rule still holds under `serve*`,
which owns the runtime and offers no in-runtime moment beforehand.
- **`WsHandler::on_open(&mut self, sender)`** added (defaulted,
non-breaking): without it a listen-only ws client can never be
subscribed (`on_message` never fires). Same panic contract as
@@ -349,6 +353,153 @@ reconnect cycles.
---
## v0.7 — Endpoint as a supervised child — DONE (2026-08-20)
Crate goes 0.2.x -> **0.3.0** (breaking). Unwinds the last deviation from
the spec (§2.1/§6): `serve` owning `rt.run`. The app owns the runtime and
the root supervisor; urus is one ordered child in it.
```
your root sup
└── ChildSpec(Permanent, urus::endpoint(cfg, pipeline)?) <- Endpoint gen_server
└── listener_sup OneForOne over N listeners
└── plain connection actors
```
- `src/conn_registry.rs` -> `src/endpoint.rs`; `ConnRegistry` -> `Endpoint`.
The registry absorbed the listener pool it registers for and runs inline
as the `ChildSpec`'s actor (`NamedGenServerBuilder::run`), so supervisor
shutdown arrives as `handle_shutdown` and a restart re-runs `init` on the
same still-open fds. `endpoint()` binds eagerly on the caller's thread.
- **The endpoint spawns its own listener sup rather than being its
sibling.** smarm's supervisor start *order* is not start *readiness*
(`start_child` spawns and moves on), so a sibling listener could
`whereis` the endpoint name before its actor ran. Registrar-spawns-
consumers makes that program order inside one `init`. Filed in smarm's
ROADMAP as a readiness-ack item; a blocking `spawn` would only shrink the
window ("has begun executing" != "has bound its name") at the cost of a
round-trip per accept.
- Listener sup is monitored: death outside shutdown = panic (the app's
supervisor decides) instead of a zombie on a dead port; death during
shutdown is the "no new connections" barrier.
- Drain is entirely internal and event-driven: `handle_shutdown` flips
draining, shuts the listener sup, stops idle conns, arms one
`drain_timeout` timer; conns that register or go idle mid-drain are
stopped on the spot; one force sweep at the deadline; the endpoint exits
when the sup is down and the set is empty. "Endpoint child stopped" ==
"every connection gone".
- **Deleted:** the `AtomicBool` listener flag, `LISTENER_TICK` (250ms wake
per listener per tick -> untimed `wait_readable` park), `SHUTDOWN_POLL`
(100ms root poll -> real park on the signal channel), `Restart::Transient`
for listeners (now `Permanent`: they only exit by supervisor action),
the root-side drain loop, `Cast::{BeginDrain, ForceStopConns}`.
Re-probed against current smarm: `request_stop` unwinds an untimed
`wait_readable` park, is no longer lossy against a QUEUED actor, and
supervisor shutdown joins in ~200us.
- **Config split**: `scheduler_threads` / `max_actors` are runtime knobs an
endpoint-as-child cannot honour -> `serve_with(cfg, smarm::Config, pipe)`
and `serve_with_shutdown(cfg, smarm::Config, pipe, signal)`. New
`Config.name` (default `"urus"`) is the endpoint's registry name;
`urus::endpoint::whereis(name)` for introspection. Two endpoints in one
process need distinct names.
- `serve*` stay as thin wrappers building a one-child tree with
`Shutdown::Infinity`. `Handle`/`ShutdownSignal` stay: a `serve*` caller
never sees the runtime, so it has no `RuntimeHandle` to reach for — but
the poll behind it is gone.
- `examples/crud.rs` is the demonstrator (app-owned tree, store as an
ordered sibling registered under a typed `Name`, shutdown via
`rt.handle().request_shutdown(root_sup)`); it loses its `OnceLock` store
cell and its `SHUTTING_DOWN` flag + 250ms poll. Other examples stay short
on `serve*`.
**Open after this cycle:**
- ~~One unreproduced test failure seen once in ~120 full-suite runs~~
**DIAGNOSED AND FIXED** in v0.8 below — it was the pubsub/channels drain
hang, not the `free_port()` race. It hid because `hammer.sh` builds with
default features while the failure needs `--all-features` load.
- `endpoint()` returns `impl Fn()`, so calling it twice = two endpoints
contending for one `Config.name` (second panics on the clash). Honest
failure, but the type doesn't prevent the mistake; a consume-on-first-use
newtype would. `ChannelHub::new` returning `(hub, children)` in v0.8 is
the same lesson applied: make the type refuse the mistake.
## v0.8 — Bus, hub and session registries are supervised — DONE (2026-08-20)
Same crate version (**0.3.0**); this cycle and v0.7 ship together and are
both breaking.
v0.7 put the endpoint under a supervisor. This puts everything else there
too, because smarm 0.7 left it no choice: commit `415effb` ("lifetime is
the actor's — refs are addresses") removed the rule the pubsub table,
channel hub bus and session registries were built on. A `GenServerRef` no
longer owns its server; the loop holds its own inbox sender and the inbox
never closes when the last ref drops. Nothing commanded these three to
stop, and what still terminated a run was the root-exit sweep racing the
drain.
That is the flake above: `shutdown_with_open_chat_terminates` and
`channels_wire::shutdown_with_open_channel_terminates`, 4 failures in 25
`--all-features` integration runs, all "serve did not return: ... outlived
the drain: Timeout". 0 in 10 full-suite runs after the fix.
**Description separated from instantiation.** The gotchas here all
descended from constructors that spawn — `PubSub::new()`,
`ChannelHub::new()` and `PrefixRouter::channel_session()` all started
actors, which forced in-runtime-only construction, which forced the
`Arc<OnceLock<..>>` lazy-init from the first handler, which forced the
"cell must not be `static`" and "a relay must never hold a `PubSub` clone"
rules. Five documented rules propping up one inverted dependency.
```
your root sup (RestForOne)
├── ChildSpec(Permanent, PubSub::child) <- bus table, named
├── ChildSpec(Permanent, session registry) <- one per channel_session
└── ChildSpec(Permanent, urus::endpoint(..)) <- starts last, drains first
```
- **`PubSub<M>` is a name, not a ref.** `const`-constructible, `Copy`,
spawns nothing, legal outside the runtime and in a `static`. Operations
resolve through the registry per call, so a supervisor-restarted table is
reached transparently. `PubSub::new()` is gone: `PubSub::new(name)` +
`PubSub::child()`.
- **`ChannelHub::new(bus, router) -> (hub, Vec<ChildSpec>)`**, `#[must_use]`.
Both halves come back together because a hub whose children were never
started compiles and fails on the first join. There is no `children()`
method to forget to call.
- **`channel_session` takes a registry name**; each registry is separately
named and supervised.
- **`serve_with` / `serve_with_shutdown` take a `Vec<ChildSpec>`** of app
children and build `RestForOne[..app children, endpoint]`. App children
start before the endpoint and — shutdown being ordered in reverse — stop
after it has drained, so a request still in flight can still reach the
bus. `RestForOne` because a bus crash leaves live sockets addressing a
table that no longer knows them.
- Deleted: the `Arc<OnceLock>` idiom from both examples and both test
pipelines, and every module rule that existed only to hand-manage a
refcount. `ws_chat` lost 21 lines and a struct field; examples net -25.
**Cost accepted:** one registry resolution per operation, including per
broadcast. Caching a `GenServerRef` in the handle would save it and go
stale across exactly the restart the supervisor exists to perform. Bench
before optimising.
**Open after this cycle:**
- `channel_session("session:*", "chat-sessions", f)` puts two unrelated
string literals side by side and nothing catches a transposition. A
`RegistryName` newtype makes it a type error. Not done.
- `PubSub` cannot be protected the way `ChannelHub` was: a `const` is
copied at each use, so it can't be consume-on-first-use. Forgetting
`BUS.child()` compiles and fails at runtime with `PubSubDown`. Judged
worth it — the `const` is what makes the handle pleasant.
- Session actors are still plain-spawned by their registry, not supervised
(they're dynamic, one per key — OTP would want `simple_one_for_one`).
Their exit chain is now commanded rather than refcounted: the registry is
shut down, its state drops, control senders drop, parked sessions wake on
the closed arm. Sound, but it is the last place a lifetime is implied by
a drop rather than stated.
- `hammer.sh` still passes no feature flags, which is why this hid for a
whole cycle. A feature matrix is the obvious fix.
## Known bugs
- ~~**HTTP/1.0 keep-alive: server honors but never advertises**~~ FIXED
@@ -378,9 +529,9 @@ reconnect cycles.
- **Bench suite** — `urus-bench-spec.md` exists in the artefact store;
wire it up once v0.2 lands (supervision changes the hot path not at all,
but prove it).
- **Operator introspection via typed names** — the v0.2-era
`urus.server` / `urus.listener.{i}` pid tags were dropped in the
RFC 014 port (names are messageable endpoints now, self-registered
with a real `Sender<M>`). If wanted back, do it properly: register the
serve loop's shutdown/control channel under a typed `urus.server`
name instead of faking pid tags with unit channels.
- **Operator introspection via typed names** — partly delivered in v0.7:
the endpoint is a named gen_server (`Config.name`, default `"urus"`),
reachable via `urus::endpoint::whereis(name)`, and answers
`Call::ConnCount`. Still open: per-listener visibility (the pool is
internal and anonymous), and a richer stats call (ws/channel counts,
request rates) — the endpoint is the place to hang them.
+1 -1
View File
@@ -240,7 +240,7 @@ fn main() {
let (handle, signal) = shutdown_handle();
let server = std::thread::spawn(move || {
let pipe = Pipeline::new().plug(Router::new().get("/order/:id", order));
serve_with_shutdown(Config::new(addr), pipe, signal).expect("serve");
serve_with_shutdown(Config::new(addr), smarm::Config::default(), pipe, Vec::new(), signal).expect("serve");
});
// Wait until it's accepting.
+20 -20
View File
@@ -10,10 +10,13 @@
//!
//! The shapes this demonstrates:
//!
//! - **One `ChannelHub` for the whole app**, built lazily on the first
//! connection via the NON-static `Arc<OnceLock<...>>` pattern (the
//! hub spawns the pubsub table; same in-runtime + shutdown laws as
//! `examples/ws_chat.rs`, see the docs there).
//! - **One `ChannelHub` for the whole app**, built up front. The hub is
//! a description — a bus name plus the routing table — and spawns
//! nothing; `hub.children()` is the vec of actors it needs (the bus
//! table, plus one registry per session route), handed to
//! `serve_with_shutdown` so the supervisor starts them ahead of the
//! endpoint and stops them after it has drained. Same shape as
//! `examples/ws_chat.rs`, see the docs there.
//!
//! - **A channel per joined topic, not per socket.** `Room` never sees
//! frames, refs, or the transport heartbeat — the conn-side handler
@@ -31,8 +34,6 @@
//! the buffer holds the relay path's own `Arc`s, nothing is copied.
//! The session is keyed by the `"name"` in the join payload.
use std::sync::{Arc, OnceLock};
use serde_json::{json, Value};
use urus::channels::phoenix::Json;
use urus::{
@@ -85,23 +86,16 @@ impl ChannelSession<P> for BySessionName {
}
fn main() {
let hub: Arc<OnceLock<ChannelHub<P>>> = Arc::new(OnceLock::new());
let router = Router::new().get("/socket", move |c: Conn, _n: Next| {
// Hub construction is in-runtime only (it spawns the pubsub
// table and, for session patterns, their registries); the
// non-static cell is what lets shutdown drain them all.
let hub = hub.get_or_init(|| {
ChannelHub::new(
let (hub, children) = ChannelHub::new(
"chat-bus",
PrefixRouter::new()
.channel_default::<Room>("room:*")
.channel_session::<BySessionName>("session:*", |_: &str| {
.channel_session::<BySessionName>("session:*", "chat-sessions", |_: &str| {
Box::new(Room::default()) as Box<dyn Channel<P>>
}),
)
});
hub.upgrade(c)
});
);
let router = Router::new().get("/socket", move |c: Conn, _n: Next| hub.upgrade(c));
let (handle, signal) = shutdown_handle();
std::thread::spawn(move || {
@@ -111,7 +105,13 @@ fn main() {
handle.shutdown();
});
serve_with_shutdown(Config::new("0.0.0.0:8080".parse().unwrap()), Pipeline::new().plug(router), signal)
serve_with_shutdown(
Config::new("0.0.0.0:8080".parse().unwrap()),
smarm::Config::default(),
Pipeline::new().plug(router),
children,
signal,
)
.unwrap();
println!("channels_chat: drained, bye");
}
+88 -68
View File
@@ -1,11 +1,18 @@
//! CRUD example: a tiny user database with JSON persistence.
//!
//! Demonstrates urus and the actor model together:
//! - The pipeline is shared (Arc) across all connection actors.
//! - Handlers do NOT take a lock or share mutable state directly.
//! - A single "store" actor owns the data; handlers send it a request
//! via a channel and block on the reply. Serialization is structural —
//! the store processes one request at a time, no Mutex needed.
//! Demonstrates urus and the actor model together, in the shape a real
//! application should use (v0.3):
//! - The APP owns the smarm runtime and the root supervisor. urus is one
//! child in that tree — `urus::endpoint(...)` — and the store actor is
//! an ordered sibling started BEFORE it, so the supervisor's
//! reverse-order shutdown drains HTTP first and only then stops the
//! store. No handler can be mid-request against a store that is gone.
//! - Handlers do NOT take a lock or share mutable state directly. A
//! single "store" actor owns the data; handlers address it by
//! registered name and block on a reply channel. Serialization is
//! structural — one request at a time, no Mutex.
//! - Both actors are supervised: kill the store (or let it panic) and it
//! restarts from the JSON file, with the endpoint left alone.
//! - On every mutating request the store writes the JSON file. Read
//! requests don't touch disk.
//!
@@ -23,9 +30,8 @@
//! curl -s http://localhost:8080/users/1
use serde::{Deserialize, Serialize};
use smarm::{channel, Sender};
use std::sync::OnceLock;
use urus::{serve_with_shutdown, shutdown_handle, Config, Conn, Next, Pipeline, Router};
use smarm::{channel, ChildSpec, Name, OneForOne, Restart, Sender};
use urus::{Config, Conn, Next, Pipeline, Router};
// ---------------------------------------------------------------------------
// Domain
@@ -66,7 +72,12 @@ const DB_PATH: &str = "/tmp/urus-crud.json";
// Store actor body
// ---------------------------------------------------------------------------
fn store_loop(rx: smarm::Receiver<Request>) {
fn store_loop() {
let (tx, rx) = channel::<Request>();
// Self-registration: the name is bound before the first recv, and it
// is re-bound automatically on every restart.
smarm::register(STORE, tx).expect("crud.store name already taken");
// Load on start. Missing file = empty store. Corrupt file = panic; we
// don't auto-rebuild because silently losing data is worse than failing
// loud.
@@ -78,19 +89,10 @@ fn store_loop(rx: smarm::Receiver<Request>) {
let mut next_id: u64 = users.iter().map(|u| u.id).max().unwrap_or(0) + 1;
loop {
// recv with a timeout rather than a bare recv: the store must be
// stoppable at shutdown, but a cross-thread Sender::send (from the
// stdin thread) can't wake a parked actor — same smarm limitation
// that motivates urus's SHUTDOWN_POLL. So we wake on our own timer
// and poll the flag. This poll dies with that limitation too.
let req = match rx.recv_timeout(std::time::Duration::from_millis(250)) {
// A plain park. The supervisor stops this actor at shutdown (after
// the endpoint has drained), so there is nothing to poll for.
let req = match rx.recv() {
Ok(r) => r,
Err(smarm::channel::RecvTimeoutError::Timeout) => {
if SHUTTING_DOWN.load(std::sync::atomic::Ordering::Relaxed) {
return;
}
continue;
}
Err(_) => return, // all senders dropped
};
match req {
@@ -169,25 +171,29 @@ fn persist(users: &[User]) {
// Handler helpers
// ---------------------------------------------------------------------------
//
// Once-cell trick: the store actor is spawned the first time a handler
// runs (smarm requires `spawn` to be called from inside an actor — which
// connection actors are). After that all handlers share the same Sender.
// Simpler than threading the Sender through the pipeline at startup.
// The store is a supervised child that registers its own inbox under a
// typed name; handlers resolve it per send. That replaces the old
// `OnceLock<Sender>` spawn-on-first-use trick — which had no supervisor,
// no restart, and no defined shutdown point — with a plain actor whose
// lifecycle the tree owns. A restart re-registers the same name, so
// in-flight handlers heal on their next send.
static STORE_TX: OnceLock<Sender<Request>> = OnceLock::new();
const STORE: Name<Request> = Name::new("crud.store");
// Set by the stdin thread at shutdown; the store actor polls it (see
// store_loop). An always-on app actor that never returns would otherwise
// block smarm's AllDone and keep serve_with_shutdown from returning.
static SHUTTING_DOWN: std::sync::atomic::AtomicBool =
std::sync::atomic::AtomicBool::new(false);
/// Send to the store and wait for its reply. `Err` only if the store is
/// between incarnations (restarting); the handler turns that into a 503
/// rather than pretending.
fn ask(make: impl FnOnce(Sender<(u16, Vec<u8>)>) -> Request) -> Option<(u16, Vec<u8>)> {
let (tx, rx) = channel::<(u16, Vec<u8>)>();
smarm::send(STORE, make(tx)).ok()?;
rx.recv().ok()
}
fn store() -> &'static Sender<Request> {
STORE_TX.get_or_init(|| {
let (tx, rx) = channel::<Request>();
smarm::spawn(move || store_loop(rx));
tx
})
fn reply(conn: Conn, r: Option<(u16, Vec<u8>)>) -> Conn {
match r {
Some((status, body)) => json(conn, status, body),
None => json(conn, 503, b"{\"error\":\"store unavailable\"}".to_vec()),
}
}
fn json(conn: Conn, status: u16, body: Vec<u8>) -> Conn {
@@ -206,16 +212,13 @@ fn parse_id(s: &str) -> Option<u64> {
/// SSE demo: `curl -N localhost:8080/ticker` streams a tick every second
/// (with `: keep-alive` comments if it ever goes quiet). The producer
/// exits on SseClosed (client gone / write timeout / shutdown drain) or
/// when the example is shutting down.
/// exits on SseClosed — client gone, write timeout, or the drain stopping
/// its connection actor. Nothing to flag: the tree's shutdown reaches it.
fn ticker(conn: Conn, _next: Next) -> Conn {
let (conn, events) = conn.sse();
smarm::spawn(move || {
let mut n: u64 = 0;
loop {
if SHUTTING_DOWN.load(std::sync::atomic::Ordering::Relaxed) {
return; // dropping `events` ends the stream cleanly
}
if events.send("tick", &n.to_string()).is_err() {
return; // SseClosed
}
@@ -227,18 +230,14 @@ fn ticker(conn: Conn, _next: Next) -> Conn {
}
fn list(conn: Conn, _next: Next) -> Conn {
let (tx, rx) = channel::<(u16, Vec<u8>)>();
store().send(Request::List { reply: tx }).ok();
let (status, body) = rx.recv().expect("store dropped");
json(conn, status, body)
let r = ask(|reply| Request::List { reply });
reply(conn, r)
}
fn create(conn: Conn, _next: Next) -> Conn {
let body = conn.body.as_bytes().to_vec();
let (tx, rx) = channel::<(u16, Vec<u8>)>();
store().send(Request::Create { body, reply: tx }).ok();
let (status, body) = rx.recv().expect("store dropped");
json(conn, status, body)
let r = ask(|reply| Request::Create { body, reply });
reply(conn, r)
}
fn get_one(conn: Conn, _next: Next) -> Conn {
@@ -246,10 +245,8 @@ fn get_one(conn: Conn, _next: Next) -> Conn {
Some(id) => id,
None => return json(conn, 400, b"{\"error\":\"bad id\"}".to_vec()),
};
let (tx, rx) = channel::<(u16, Vec<u8>)>();
store().send(Request::Get { id, reply: tx }).ok();
let (status, body) = rx.recv().expect("store dropped");
json(conn, status, body)
let r = ask(|reply| Request::Get { id, reply });
reply(conn, r)
}
fn update(conn: Conn, _next: Next) -> Conn {
@@ -258,10 +255,8 @@ fn update(conn: Conn, _next: Next) -> Conn {
None => return json(conn, 400, b"{\"error\":\"bad id\"}".to_vec()),
};
let body = conn.body.as_bytes().to_vec();
let (tx, rx) = channel::<(u16, Vec<u8>)>();
store().send(Request::Update { id, body, reply: tx }).ok();
let (status, body) = rx.recv().expect("store dropped");
json(conn, status, body)
let r = ask(|reply| Request::Update { id, body, reply });
reply(conn, r)
}
fn delete(conn: Conn, _next: Next) -> Conn {
@@ -269,10 +264,8 @@ fn delete(conn: Conn, _next: Next) -> Conn {
Some(id) => id,
None => return json(conn, 400, b"{\"error\":\"bad id\"}".to_vec()),
};
let (tx, rx) = channel::<(u16, Vec<u8>)>();
store().send(Request::Delete { id, reply: tx }).ok();
let (status, body) = rx.recv().expect("store dropped");
json(conn, status, body)
let r = ask(|reply| Request::Delete { id, reply });
reply(conn, r)
}
// ---------------------------------------------------------------------------
@@ -305,20 +298,47 @@ fn main() {
);
let cfg = Config::new("127.0.0.1:8080".parse().unwrap());
// Bind here, on this thread: an address-in-use error is a startup
// failure, not an actor crash. The fds outlive any restart of the
// endpoint child.
let endpoint = urus::endpoint(cfg, pipeline).expect("bind 127.0.0.1:8080");
println!("urus-crud: DB at {DB_PATH}");
println!("urus-crud: listening on 127.0.0.1:8080 — press Enter to shut down");
let rt = smarm::init(smarm::Config::default());
let handle = rt.handle();
let (sup_tx, sup_rx) = std::sync::mpsc::channel();
// Graceful shutdown on stdin-Enter: a plain OS thread blocks on
// read_line and fires the handle. No signal handling crate needed.
let (handle, signal) = shutdown_handle();
// read_line and shuts the ROOT SUPERVISOR down. No signal-handling
// crate needed, and no urus-specific shutdown plumbing — this is
// exactly what a SIGTERM handler would do.
std::thread::spawn(move || {
let sup: smarm::Pid = sup_rx.recv().expect("supervisor pid");
let mut line = String::new();
let _ = std::io::stdin().read_line(&mut line);
println!("urus-crud: shutting down (draining in-flight requests)…");
SHUTTING_DOWN.store(true, std::sync::atomic::Ordering::Relaxed);
handle.shutdown();
handle.request_shutdown(sup);
});
serve_with_shutdown(cfg, pipeline, signal).unwrap();
rt.run(move || {
let sup = smarm::spawn(move || {
OneForOne::new()
// Store FIRST: reverse-order shutdown therefore stops it
// LAST, after the endpoint has finished draining.
.child(ChildSpec::new(Restart::Permanent, store_loop))
.child(
ChildSpec::new(Restart::Permanent, endpoint)
// The endpoint bounds its own drain with
// Config::drain_timeout; a shorter supervisor
// deadline would cut that drain in half.
.shutdown(smarm::supervisor::Shutdown::Infinity),
)
.run()
});
let _ = sup_tx.send(sup.pid());
let _ = sup.join();
});
println!("urus-crud: bye");
}
+1 -1
View File
@@ -52,7 +52,7 @@ fn main() {
let (handle, signal) = shutdown_handle();
let server = std::thread::spawn(move || {
let pipe = Pipeline::new().plug(Router::new().get("/json/:id", json_id));
serve_with_shutdown(Config::new(addr), pipe, signal).expect("serve");
serve_with_shutdown(Config::new(addr), smarm::Config::default(), pipe, Vec::new(), signal).expect("serve");
});
// Readiness is the orchestrator's job (TCP probe); ours is to not
+8 -3
View File
@@ -41,8 +41,13 @@ fn main() {
// Audit line: lands in each bench cell's server.log so the effective
// scheduler count is recorded per cell, same discipline as mode-verify.
eprintln!("plain_serve: scheduler_threads={sched_threads:?}");
let mut cfg = Config::new(addr);
cfg.scheduler_threads = sched_threads;
// Scheduler threads are a RUNTIME knob, so they live in smarm::Config,
// not urus's — an endpoint placed in someone else's tree could not
// honour them anyway.
let rt_cfg = match sched_threads {
Some(n) => smarm::Config::exact(n),
None => smarm::Config::default(),
};
let pipe = Pipeline::new().plug(Router::new().get("/json/:id", json_id));
serve_with(cfg, pipe).expect("serve");
serve_with(Config::new(addr), rt_cfg, pipe, Vec::new()).expect("serve");
}
+48
View File
@@ -0,0 +1,48 @@
//! Serve with an optional TOML config overlay.
//!
//! urus is a library and never presumes a config path or reads the
//! environment for one — the binary decides where the file lives and hands
//! the text to `Config::with_toml_str`. Here that's a `--config PATH` flag;
//! with no flag, the compiled defaults are used unchanged.
//!
//! Requires the `config-file` feature:
//! cargo run --example serve_toml --features config-file -- --config urus.toml
//!
//! Example urus.toml (all keys optional, sparse override; seconds):
//! head_timeout_secs = 15
//! body_timeout_secs = 300
//! body_burst_bytes = 4096
//! body_stall_timeout_secs = 20
use std::net::SocketAddr;
use urus::{serve_with, Config, Conn, Next, Pipeline, Router};
fn json_id(conn: Conn, _next: Next) -> Conn {
let id: u64 = conn.params.get("id").and_then(|s| s.parse().ok()).unwrap_or(0);
conn.put_status(200)
.put_header("content-type", "application/json")
.put_body(format!("{{\"id\":{id}}}"))
}
fn main() {
let addr: SocketAddr = "0.0.0.0:8080".parse().expect("addr");
let mut cfg = Config::new(addr);
// Minimal flag scan: `--config PATH`. No presumed default location.
let mut args = std::env::args().skip(1);
while let Some(arg) = args.next() {
if arg == "--config" {
let path = args.next().expect("--config needs a PATH");
let toml = std::fs::read_to_string(&path)
.unwrap_or_else(|e| panic!("reading config {path}: {e}"));
cfg = cfg
.with_toml_str(&toml)
.unwrap_or_else(|e| panic!("invalid config {path}: {e}"));
eprintln!("serve_toml: loaded config from {path}");
}
}
let pipe = Pipeline::new().plug(Router::new().get("/json/:id", json_id));
serve_with(cfg, smarm::Config::default(), pipe, Vec::new()).expect("serve");
}
+31 -44
View File
@@ -1,4 +1,4 @@
//! WebSocket chat rooms (v0.5): `urus::pubsub` wired to the v0.4 duplex.
//! WebSocket chat rooms (v0.7): `urus::pubsub` wired to the v0.4 duplex.
//!
//! cargo run --example ws_chat
//! websocat ws://127.0.0.1:8080/chat/lobby (run two of these)
@@ -6,39 +6,35 @@
//! The shapes this demonstrates:
//!
//! - **One `PubSub<String>` for the whole app**, topics are rooms
//! (`room:{name}`). Built lazily on the first connection via a
//! NON-static `Arc<OnceLock<PubSub>>` captured by the route closure:
//! `PubSub::new()` spawns the table actor, which smarm only allows
//! in-runtime — and keeping the cell non-static means the table's
//! last handle drops in-runtime when the drained pipeline drops, so
//! `serve_with_shutdown` actually returns. A `static` cell would pin
//! the table forever and block smarm's all-done. (Same constraint
//! crud's store solves with its OnceLock; the non-static refinement
//! is what makes graceful shutdown compose.)
//! (`room:{name}`). `BUS` is a `const`: the handle is an address, not
//! the table. The table actor is `BUS.child()`, handed to
//! `serve_with_shutdown` as an app child, so it starts before the
//! endpoint and — shutdown being ordered in reverse — stops after the
//! endpoint has drained.
//!
//! The v0.5 idiom this replaces was a non-static `Arc<OnceLock<PubSub>>`
//! lazily initialised from the first connection, because `PubSub::new()`
//! used to spawn the table and the handle used to own its life. Neither
//! is true any more.
//!
//! - **`on_open` subscribes and spawns the relay** — a listen-only
//! client receives the room without ever sending. The subscription is
//! pinned to the CONNECTION actor (`on_open` runs inside it), so the
//! monitor cleans up exactly when the connection dies.
//!
//! - **The relay holds the `Receiver` and a `WsSender` clone — and
//! deliberately NOT a `PubSub` handle.** A relay holding the handle
//! keeps the table's inbox open while the table keeps the relay's
//! receiver open: neither ever exits, and shutdown hangs. Receiver
//! only: conn dies → monitor prunes → sender drops → relay's recv
//! errs → relay exits. Every link in that chain is in-runtime.
use std::sync::{Arc, OnceLock};
//! - **The relay holds the `Receiver` and a `WsSender` clone.** It may
//! also hold `BUS` — a handle pins nothing now — but it has no use for
//! one. Conn dies → monitor prunes → sender drops → relay's recv errs
//! → relay exits.
use urus::{
serve_with_shutdown, shutdown_handle, Config, Conn, Message, Next, Pipeline, PubSub, Router,
WsHandler, WsSender,
};
type Bus = PubSub<String>;
const BUS: PubSub<String> = PubSub::new("chat");
struct ChatHandler {
bus: Arc<OnceLock<Bus>>,
room: String,
}
@@ -46,32 +42,23 @@ impl ChatHandler {
fn topic(&self) -> String {
format!("room:{}", self.room)
}
fn bus(&self) -> &Bus {
// First connection anywhere spawns the table; we are inside the
// connection actor here, so the spawn is legal.
self.bus.get_or_init(PubSub::new)
}
}
impl WsHandler for ChatHandler {
fn on_open(&mut self, sender: &WsSender) {
let topic = self.topic();
let rx = match self.bus().subscribe(&topic) {
let rx = match BUS.subscribe(&topic) {
Ok(rx) => rx,
Err(_) => {
let _ = sender.close(1011, "chat bus down");
return;
}
};
let _ = self
.bus()
.broadcast_from(smarm::self_pid(), &topic, format!("* someone joined {topic}"));
let _ = BUS.broadcast_from(smarm::self_pid(), &topic, format!("* someone joined {topic}"));
// The relay: room messages -> this socket. Exits when the
// subscription is pruned (conn death / unsubscribe) or the
// socket is gone (WsClosed). See the module docs for why it
// must not capture a Bus handle.
// socket is gone (WsClosed).
let out = sender.clone();
smarm::spawn(move || {
while let Ok(msg) = rx.recv() {
@@ -90,16 +77,12 @@ impl WsHandler for ChatHandler {
// broadcast_from: the sender's own relay is skipped — no echo.
// self_pid() here is the connection actor, the pid on_open
// subscribed as.
let _ = self
.bus()
.broadcast_from(smarm::self_pid(), self.topic(), text);
let _ = BUS.broadcast_from(smarm::self_pid(), self.topic(), text);
}
fn on_close(&mut self, _code: Option<u16>, _reason: &str) {
let topic = self.topic();
let _ = self
.bus()
.broadcast_from(smarm::self_pid(), &topic, format!("* someone left {topic}"));
let _ = BUS.broadcast_from(smarm::self_pid(), &topic, format!("* someone left {topic}"));
// No explicit unsubscribe: the connection actor is about to
// exit and the monitor prunes the subscription (which is also
// what stops the relay).
@@ -107,12 +90,9 @@ impl WsHandler for ChatHandler {
}
fn main() {
let bus: Arc<OnceLock<Bus>> = Arc::new(OnceLock::new());
let pipeline = Pipeline::new().plug(Router::new().get("/chat/:room", move |c: Conn, _n: Next| {
let pipeline = Pipeline::new().plug(Router::new().get("/chat/:room", |c: Conn, _n: Next| {
let room = c.params.get("room").unwrap_or("lobby").to_string();
let bus = bus.clone();
c.upgrade(ChatHandler { bus, room })
c.upgrade(ChatHandler { room })
}));
let (handle, signal) = shutdown_handle();
@@ -123,6 +103,13 @@ fn main() {
handle.shutdown();
});
serve_with_shutdown(Config::new("0.0.0.0:8080".parse().unwrap()), pipeline, signal).unwrap();
serve_with_shutdown(
Config::new("0.0.0.0:8080".parse().unwrap()),
smarm::Config::default(),
pipeline,
vec![BUS.child()],
signal,
)
.unwrap();
println!("ws_chat: drained, bye");
}
+1 -1
View File
@@ -98,6 +98,6 @@ fn main() {
handle.shutdown();
});
serve_with_shutdown(cfg, pipeline, signal).unwrap();
serve_with_shutdown(cfg, smarm::Config::default(), pipeline, Vec::new(), signal).unwrap();
println!("ws_echo: bye");
}
+128 -38
View File
@@ -9,12 +9,12 @@
//!
//! # Topology
//!
//! - The app builds one [`ChannelHub`] (a [`TopicRouter`] + a
//! `PubSub<Broadcast<P>>`). **In-runtime only** — `PubSub::new`
//! spawns the table actor — and the hub must be NON-static (the
//! `Arc<OnceLock<..>>`-in-the-route-closure pattern from v0.5;
//! a `static` hub pins the pubsub table and hangs
//! `serve_with_shutdown`).
//! - The app builds one [`ChannelHub`] (a [`TopicRouter`] + the name of
//! a `PubSub<Broadcast<P>>`). It is a description and spawns nothing,
//! so it may be built anywhere; [`ChannelHub::new`] hands back the
//! `ChildSpec`s for the actors it needs — the bus table, plus one
//! registry per session route — which go in the supervision tree
//! ahead of the endpoint.
//! - [`ChannelHub::upgrade`] turns an HTTP `Conn` into a channel
//! socket: a [`WsHandler`] running in the connection actor that
//! decodes frames and routes them by topic.
@@ -63,7 +63,7 @@
use crate::pubsub::PubSub;
use crate::ws::{Message, WsHandler, WsSender};
use smarm::{Pid, Receiver, Sender};
use smarm::{ChildSpec, Pid, Receiver, Sender};
use std::cell::RefCell;
use std::collections::HashMap;
@@ -199,6 +199,16 @@ pub trait Channel<P: Send + Sync + 'static>: Send + 'static {
pub trait ChannelFactory<P: Send + Sync + 'static>: Send + Sync {
fn create(&self, topic: &str) -> Box<dyn Channel<P>>;
/// Supervised actors this factory needs running before it can
/// deploy. Empty for the default (ephemeral) factory, which spawns
/// its channel actor per join under the connection; the session
/// factory returns its registry's spec. Internal seam, collected by
/// [`ChannelHub::children`].
#[doc(hidden)]
fn children(&self) -> Vec<ChildSpec> {
Vec::new()
}
/// How an accepted-routing join becomes a running channel actor.
/// Internal seam — the default (ephemeral actor, linked to the
/// connection, cold start per join) is the contract; only the
@@ -258,6 +268,13 @@ impl<P: Send + Sync + 'static, C: Channel<P> + Default> ChannelFactory<P> for De
/// rejects the join (status error).
pub trait TopicRouter<P: Send + Sync + 'static>: Send + Sync + 'static {
fn route(&self, topic: &str) -> Option<Arc<dyn ChannelFactory<P>>>;
/// Every supervised actor the routes need, gathered for
/// [`ChannelHub::children`]. Override only if your router holds
/// factories it does not surface through [`route`](Self::route).
fn children(&self) -> Vec<ChildSpec> {
Vec::new()
}
}
/// The shipped router: exact topics and `head:*` prefix patterns.
@@ -301,31 +318,46 @@ impl<P: Send + Sync + 'static> PrefixRouter<P> {
/// [`channel`](Self::channel), with opt-in session persistence
/// keyed and configured by `S`'s [`ChannelSession`] impl (usually
/// `S` is the channel type itself). **In-runtime only**: this
/// spawns the pattern's session-registry actor, same law as
/// [`ChannelHub::new`].
pub fn channel_session<S>(self, pattern: &str, factory: impl ChannelFactory<P> + 'static) -> Self
/// `S` is the channel type itself).
///
/// `registry` names the pattern's session-registry actor, which is
/// supervised: it appears in [`ChannelHub::children`] and must be
/// unique across the app. Spawns nothing here.
pub fn channel_session<S>(
self,
pattern: &str,
registry: &'static str,
factory: impl ChannelFactory<P> + 'static,
) -> Self
where
S: ChannelSession<P>,
S::Key: Clone,
P: Encode + Decode,
{
self.channel(pattern, session::SessionFactory::new::<S>(Arc::new(factory)))
self.channel(pattern, session::SessionFactory::new::<S>(registry, Arc::new(factory)))
}
/// [`channel_session`](Self::channel_session) with a
/// [`Default`]-built impl that is its own session config.
pub fn channel_session_default<C>(self, pattern: &str) -> Self
pub fn channel_session_default<C>(self, pattern: &str, registry: &'static str) -> Self
where
C: Channel<P> + ChannelSession<P> + Default,
C::Key: Clone,
P: Encode + Decode,
{
self.channel_session::<C>(pattern, DefaultFactory::<C>(PhantomData))
self.channel_session::<C>(pattern, registry, DefaultFactory::<C>(PhantomData))
}
}
impl<P: Send + Sync + 'static> TopicRouter<P> for PrefixRouter<P> {
fn children(&self) -> Vec<ChildSpec> {
self.exact
.values()
.chain(self.prefix.values())
.flat_map(|f| f.children())
.collect()
}
fn route(&self, topic: &str) -> Option<Arc<dyn ChannelFactory<P>>> {
if let Some(f) = self.exact.get(topic) {
return Some(f.clone());
@@ -427,13 +459,20 @@ impl<P: Encode + Decode + Send + Sync + 'static> ChannelSocket<P> {
// ChannelHub — the app-facing entry point
// ---------------------------------------------------------------------------
/// One per app (or per channel namespace): the router plus the pubsub
/// bus every channel broadcasts on.
/// One per app (or per channel namespace): the router plus the address
/// of the pubsub bus every channel broadcasts on.
///
/// **In-runtime only** (spawns the pubsub table) and **must not live in
/// a `static`** — use the non-static `Arc<OnceLock<ChannelHub<P>>>`
/// captured by the route closure, exactly the v0.5 pubsub pattern, or
/// graceful shutdown will hang waiting for the table to exit.
/// Spawns nothing — build it wherever you like, including out of the
/// runtime and alongside the pipeline. The actors it needs (the bus
/// table, plus one registry per session route) come from
/// [`children`](Self::children); hand that vec to `serve_with*` or splice
/// it into your own supervision tree ahead of the endpoint.
///
/// ```ignore
/// let (hub, children) =
/// ChannelHub::new("chat-bus", PrefixRouter::new().channel_default::<Room>("room:*"));
/// serve_with(cfg, rt_cfg, pipeline, children)?;
/// ```
pub struct ChannelHub<P: Send + Sync + 'static> {
bus: PubSub<Broadcast<P>>,
router: Arc<dyn TopicRouter<P>>,
@@ -441,13 +480,27 @@ pub struct ChannelHub<P: Send + Sync + 'static> {
impl<P: Send + Sync + 'static> Clone for ChannelHub<P> {
fn clone(&self) -> Self {
Self { bus: self.bus.clone(), router: self.router.clone() }
Self { bus: self.bus, router: self.router.clone() }
}
}
impl<P: Encode + Decode + Send + Sync + 'static> ChannelHub<P> {
pub fn new(router: impl TopicRouter<P>) -> Self {
Self { bus: PubSub::new(), router: Arc::new(router) }
/// Address a hub whose bus table is registered under `bus`, together
/// with every actor it needs running: the bus table first, then one
/// registry per session route. Spawns nothing.
///
/// The two come back together on purpose. A hub whose children were
/// never started compiles fine and fails on the first join, so the
/// constructor hands you the vec rather than leaving it behind a
/// method you can forget to call. Start them ahead of the endpoint —
/// `RestForOne` in that order means a bus crash also restarts the
/// endpoint, dropping connections whose subscriptions died with it.
#[must_use]
pub fn new(bus: &'static str, router: impl TopicRouter<P>) -> (Self, Vec<ChildSpec>) {
let hub = Self { bus: PubSub::new(bus), router: Arc::new(router) };
let mut children = vec![hub.bus.child()];
children.extend(hub.router.children());
(hub, children)
}
/// Accept the WebSocket upgrade on `conn` and speak channels over
@@ -455,7 +508,7 @@ impl<P: Encode + Decode + Send + Sync + 'static> ChannelHub<P> {
/// [`Conn::upgrade`](crate::Conn::upgrade) apply unchanged.
pub fn upgrade(&self, conn: crate::Conn) -> crate::Conn {
conn.upgrade(SocketHandler {
bus: self.bus.clone(),
bus: self.bus,
router: self.router.clone(),
joined: HashMap::new(),
})
@@ -567,7 +620,7 @@ impl<P: Encode + Decode + Send + Sync + 'static> WsHandler for SocketHandler<P>
reference: frame.reference,
payload,
ws: sender.clone(),
bus: self.bus.clone(),
bus: self.bus,
});
self.joined
.insert(frame.topic, Joined { join_ref: frame.join_ref, tx: inbox.0 });
@@ -649,7 +702,7 @@ fn run_channel<P: Encode + Decode + Send + Sync + 'static>(
let socket = ChannelSocket {
ws,
bus: bus.clone(),
bus,
topic: topic.clone(),
join_ref,
cur_ref: RefCell::new(None),
@@ -799,12 +852,39 @@ mod tests {
}
fn hub(terminated: &Arc<AtomicBool>) -> ChannelHub<TP> {
const TEST_BUS: &str = "chan-unit-bus";
fn hub(terminated: &Arc<AtomicBool>) -> (ChannelHub<TP>, Vec<ChildSpec>) {
ChannelHub::new(
TEST_BUS,
PrefixRouter::new().channel("room:*", RoomFactory { terminated: terminated.clone() }),
)
}
/// Start a hub's children under a supervisor and wait for the bus to
/// bind its name (smarm's start-order-is-not-start-readiness gap:
/// `start_child` spawns and moves on). Returns the sup's pid.
pub(crate) fn start_hub<P: Encode + Decode + Send + Sync + 'static>(
hub: &ChannelHub<P>,
children: Vec<ChildSpec>,
) -> Pid {
let sup = smarm::spawn(move || {
let mut sup = smarm::OneForOne::new();
for c in children {
sup = sup.child(c);
}
sup.run()
});
let bus = hub.bus;
for _ in 0..200 {
if bus.pid().is_some() {
return sup.pid();
}
smarm::sleep(std::time::Duration::from_millis(10));
}
panic!("hub children never came up");
}
#[test]
fn prefix_router_exact_and_prefix_and_miss() {
@@ -825,7 +905,8 @@ mod tests {
let out = Arc::new(Mutex::new(Vec::<String>::new()));
let (out2, t) = (out.clone(), Arc::new(AtomicBool::new(false)));
smarm::run(move || {
let hub = hub(&t);
let (hub, children) = hub(&t);
let _sup = start_hub(&hub, children);
let mut h = Harness::new(&hub);
h.send("j1|r1|room:a|event|phx_join|hi");
out2.lock().unwrap().push(h.recv());
@@ -843,7 +924,8 @@ mod tests {
let (out2, t) = (out.clone(), Arc::new(AtomicBool::new(false)));
let t2 = t.clone();
smarm::run(move || {
let hub = hub(&t2);
let (hub, children) = hub(&t2);
let _sup = start_hub(&hub, children);
let mut h = Harness::new(&hub);
h.send("j1|r1|room:locked|event|phx_join|hi");
*out2.lock().unwrap() = h.recv();
@@ -859,7 +941,8 @@ mod tests {
let out = Arc::new(Mutex::new(String::new()));
let (out2, t) = (out.clone(), Arc::new(AtomicBool::new(false)));
smarm::run(move || {
let hub = hub(&t);
let (hub, children) = hub(&t);
let _sup = start_hub(&hub, children);
let mut h = Harness::new(&hub);
h.send("j1|r1|hall:a|event|phx_join|hi");
*out2.lock().unwrap() = h.recv();
@@ -872,7 +955,8 @@ mod tests {
let out = Arc::new(Mutex::new(String::new()));
let (out2, t) = (out.clone(), Arc::new(AtomicBool::new(false)));
smarm::run(move || {
let hub = hub(&t);
let (hub, children) = hub(&t);
let _sup = start_hub(&hub, children);
let mut h = Harness::new(&hub);
h.send("-|hb1|phoenix|event|heartbeat|");
*out2.lock().unwrap() = h.recv();
@@ -885,7 +969,8 @@ mod tests {
let out = Arc::new(Mutex::new(String::new()));
let (out2, t) = (out.clone(), Arc::new(AtomicBool::new(false)));
smarm::run(move || {
let hub = hub(&t);
let (hub, children) = hub(&t);
let _sup = start_hub(&hub, children);
let mut h = Harness::new(&hub);
h.send("j1|r1|room:a|event|ping|x");
*out2.lock().unwrap() = h.recv();
@@ -898,7 +983,8 @@ mod tests {
let got = Arc::new(Mutex::new(Vec::<(u8, String)>::new()));
let (got2, t) = (got.clone(), Arc::new(AtomicBool::new(false)));
smarm::run(move || {
let hub = hub(&t);
let (hub, children) = hub(&t);
let _sup = start_hub(&hub, children);
let mut a = Harness::new(&hub);
let mut b = Harness::new(&hub);
a.send("j1|r1|room:a|event|phx_join|A");
@@ -930,7 +1016,8 @@ mod tests {
let (out2, t) = (out.clone(), Arc::new(AtomicBool::new(false)));
let t2 = t.clone();
smarm::run(move || {
let hub = hub(&t2);
let (hub, children) = hub(&t2);
let _sup = start_hub(&hub, children);
let mut h = Harness::new(&hub);
h.send("j1|r1|room:a|event|phx_join|hi");
out2.lock().unwrap().push(h.recv());
@@ -954,8 +1041,9 @@ mod tests {
let t = Arc::new(AtomicBool::new(false));
let t2 = t.clone();
smarm::run(move || {
let hub = hub(&t2);
let bus = hub.bus.clone();
let (hub, children) = hub(&t2);
let _sup = start_hub(&hub, children);
let bus = hub.bus;
{
let mut h = Harness::new(&hub);
h.send("j1|r1|room:a|event|phx_join|hi");
@@ -974,7 +1062,8 @@ mod tests {
let (out2, t) = (out.clone(), Arc::new(AtomicBool::new(false)));
let t2 = t.clone();
smarm::run(move || {
let hub = hub(&t2);
let (hub, children) = hub(&t2);
let _sup = start_hub(&hub, children);
let mut h = Harness::new(&hub);
h.send("j1|r1|room:a|event|phx_join|one");
out2.lock().unwrap().push(h.recv());
@@ -998,7 +1087,8 @@ mod tests {
let out = Arc::new(Mutex::new(String::new()));
let (out2, t) = (out.clone(), Arc::new(AtomicBool::new(false)));
smarm::run(move || {
let hub = hub(&t);
let (hub, children) = hub(&t);
let _sup = start_hub(&hub, children);
let mut h = Harness::new(&hub);
h.send("j1|r1|room:a|event|phx_join|hi");
h.recv();
+69 -26
View File
@@ -22,10 +22,17 @@
//! the transport is the point. Its exits: explicit leave, rejected
//! (re)join, TTL expiry, buffer cap, pubsub relay death, or the
//! registry's control sender dropping — which is exactly the shutdown
//! chain (conns die -> router `Arc`s drop -> registry's inbox closes
//! -> registry exits -> control senders drop -> every parked session
//! wakes on the closed control arm, terminates, and exits; `AllDone`
//! composes without links).
//! chain (the supervisor shuts the registry down after the endpoint
//! has drained -> the registry's state drops -> control senders drop
//! -> every parked session wakes on the closed control arm,
//! terminates, and exits; `AllDone` composes without links).
//!
//! Note this is the last lifetime in urus implied by a drop rather
//! than stated: sessions are dynamic (one per key), so a fixed
//! `ChildSpec` list cannot hold them. The trigger is now a command
//! rather than a refcount, which is what the v0.8 cycle was about, but
//! a dynamic supervisor (OTP's `simple_one_for_one`) is the honest
//! shape if smarm grows one.
//!
//! As-landed decisions (veto by diff):
//! - **Every attach calls `ch.join()` again** on the same instance —
@@ -48,8 +55,8 @@
use super::*;
use smarm::gen_server::{self, GenServer, GenServerCtx};
use smarm::{Down, GenServerRef, Watcher};
use smarm::gen_server::{self, GenServer, GenServerBuilder, GenServerCtx, GenServerName};
use smarm::{Down, Restart, Watcher};
use std::collections::VecDeque;
use std::hash::Hash;
@@ -87,7 +94,14 @@ pub trait ChannelSession<P>: Send + 'static {
pub(super) struct SessionFactory<P: Encode + Decode + Send + Sync + 'static, K: SessionKey> {
inner: Arc<dyn ChannelFactory<P>>,
keyfn: fn(&str, &P) -> K,
registry: GenServerRef<Registry<P, K>>,
/// The registry's registered name. Resolved per deploy, so a
/// registry restarted by the supervisor is reached transparently —
/// with its session map empty, which is the honest outcome: the
/// session actors it tracked died with their control senders.
registry: GenServerName<Registry<P, K>>,
/// Everything the registry's `ChildSpec` needs to build it again.
cap: usize,
ttl: Duration,
}
/// The registry's working bounds for a session key.
@@ -95,17 +109,19 @@ pub(super) trait SessionKey: Eq + Hash + Clone + Send + 'static {}
impl<K: Eq + Hash + Clone + Send + 'static> SessionKey for K {}
impl<P: Encode + Decode + Send + Sync + 'static, K: SessionKey> SessionFactory<P, K> {
/// **In-runtime only** — spawns the registry gen_server.
pub(super) fn new<S: ChannelSession<P, Key = K>>(inner: Arc<dyn ChannelFactory<P>>) -> Self {
let registry = gen_server::start(Registry {
factory: inner.clone(),
/// Describe the session route. Spawns nothing; the registry actor is
/// started from [`children`](ChannelFactory::children).
pub(super) fn new<S: ChannelSession<P, Key = K>>(
registry: &'static str,
inner: Arc<dyn ChannelFactory<P>>,
) -> Self {
SessionFactory {
inner,
keyfn: S::session_key,
registry: GenServerName::new(registry),
cap: S::buffer_cap(),
ttl: S::ttl(),
sessions: HashMap::new(),
pid_key: HashMap::new(),
watcher: None,
});
SessionFactory { inner, keyfn: S::session_key, registry }
}
}
}
@@ -116,6 +132,23 @@ impl<P: Encode + Decode + Send + Sync + 'static, K: SessionKey> ChannelFactory<P
self.inner.create(topic)
}
fn children(&self) -> Vec<ChildSpec> {
let (name, factory, cap, ttl) = (self.registry, self.inner.clone(), self.cap, self.ttl);
vec![ChildSpec::new(Restart::Permanent, move || {
let state = Registry::<P, K> {
factory: factory.clone(),
cap,
ttl,
sessions: HashMap::new(),
pid_key: HashMap::new(),
watcher: None,
};
if let Err(e) = GenServerBuilder::new(state).named(name).run() {
panic!("urus session registry '{}': name already taken: {e:?}", name.as_str());
}
})]
}
fn deploy(&self, hs: JoinHandshake<P>) -> ChannelInbox<P>
where
P: Encode + Decode,
@@ -125,7 +158,7 @@ impl<P: Encode + Decode + Send + Sync + 'static, K: SessionKey> ChannelFactory<P
// Cast: the actor acks the join straight to the socket; the
// conn actor has nothing to wait for. A dead registry can only
// mean shutdown — the failed join is moot.
let _ = self.registry.cast(Join { key, hs, inbound: rx });
let _ = gen_server::cast(self.registry, Join { key, hs, inbound: rx });
ChannelInbox(tx)
}
}
@@ -568,10 +601,14 @@ mod tests {
fn session_hub<S: ChannelSession<TP, Key = String>>(
term: &Arc<AtomicBool>,
reject_rejoin: bool,
) -> ChannelHub<TP> {
) -> (ChannelHub<TP>, Vec<smarm::ChildSpec>) {
ChannelHub::new(
PrefixRouter::new()
.channel_session::<S>("room:*", counter_factory(term, reject_rejoin)),
"sess-unit-bus",
PrefixRouter::new().channel_session::<S>(
"room:*",
"sess-unit-registry",
counter_factory(term, reject_rejoin),
),
)
}
@@ -580,7 +617,8 @@ mod tests {
let out = Arc::new(Mutex::new(Vec::<String>::new()));
let (out2, term) = (out.clone(), Arc::new(AtomicBool::new(false)));
smarm::run(move || {
let hub = session_hub::<Cfg>(&term, false);
let (hub, children) = session_hub::<Cfg>(&term, false);
let _sup = crate::channels::tests::start_hub(&hub, children);
let mut h1 = Harness::new(&hub);
h1.send("j1|r1|room:a|event|phx_join|u");
out2.lock().unwrap().push(h1.recv());
@@ -611,7 +649,8 @@ mod tests {
let (out2, term) = (out.clone(), Arc::new(AtomicBool::new(false)));
let term2 = term.clone();
smarm::run(move || {
let hub = session_hub::<CfgTtl>(&term2, false);
let (hub, children) = session_hub::<CfgTtl>(&term2, false);
let _sup = crate::channels::tests::start_hub(&hub, children);
let mut h1 = Harness::new(&hub);
h1.send("j1|r1|room:a|event|phx_join|u");
out2.lock().unwrap().push(h1.recv());
@@ -632,7 +671,8 @@ mod tests {
let out = Arc::new(Mutex::new(Vec::<String>::new()));
let (out2, term) = (out.clone(), Arc::new(AtomicBool::new(false)));
smarm::run(move || {
let hub = session_hub::<CfgCap>(&term, false);
let (hub, children) = session_hub::<CfgCap>(&term, false);
let _sup = crate::channels::tests::start_hub(&hub, children);
let mut h1 = Harness::new(&hub);
h1.send("j1|r1|room:a|event|phx_join|u");
out2.lock().unwrap().push(h1.recv());
@@ -657,7 +697,8 @@ mod tests {
let out = Arc::new(Mutex::new(Vec::<String>::new()));
let (out2, term) = (out.clone(), Arc::new(AtomicBool::new(false)));
smarm::run(move || {
let hub = session_hub::<Cfg>(&term, false);
let (hub, children) = session_hub::<Cfg>(&term, false);
let _sup = crate::channels::tests::start_hub(&hub, children);
let mut h = Harness::new(&hub);
h.send("j1|r1|room:a|event|phx_join|u");
out2.lock().unwrap().push(h.recv());
@@ -681,7 +722,8 @@ mod tests {
let out = Arc::new(Mutex::new(Vec::<String>::new()));
let (out2, term) = (out.clone(), Arc::new(AtomicBool::new(false)));
smarm::run(move || {
let hub = session_hub::<Cfg>(&term, false);
let (hub, children) = session_hub::<Cfg>(&term, false);
let _sup = crate::channels::tests::start_hub(&hub, children);
let mut h1 = Harness::new(&hub);
h1.send("j1|r1|room:a|event|phx_join|u");
out2.lock().unwrap().push(h1.recv());
@@ -714,7 +756,8 @@ mod tests {
let out = Arc::new(Mutex::new(Vec::<String>::new()));
let (out2, term) = (out.clone(), Arc::new(AtomicBool::new(false)));
smarm::run(move || {
let hub = session_hub::<Cfg>(&term, true);
let (hub, children) = session_hub::<Cfg>(&term, true);
let _sup = crate::channels::tests::start_hub(&hub, children);
let mut h1 = Harness::new(&hub);
h1.send("j1|r1|room:a|event|phx_join|u");
out2.lock().unwrap().push(h1.recv());
+1 -1
View File
@@ -75,7 +75,7 @@ impl Harness {
let (tx, out_rx) = smarm::channel();
Self {
handler: SocketHandler {
bus: hub.bus.clone(),
bus: hub.bus,
router: hub.router.clone(),
joined: HashMap::new(),
},
+144 -55
View File
@@ -14,7 +14,7 @@
//! during those parks, other connection actors progress freely.
use crate::conn::{Body, Conn, HttpVersion, RespBody, StreamBody};
use crate::conn_registry::{Cast, ConnRegistry, DeregisterGuard};
use crate::endpoint::{Cast, DeregisterGuard, Endpoint};
use crate::net::OwnedFd;
use crate::parser::{self, ParseError};
use crate::plug::Pipeline;
@@ -46,14 +46,34 @@ pub struct ConnLimits {
/// connection). Expiry closes the connection silently — nothing is
/// owed to a client that isn't talking.
pub keep_alive_timeout: Duration,
/// Per-request wall-clock budget, measured from the first byte of a
/// request until the request (head + body) is fully read. Expiry
/// mid-head gets a best-effort 408; expiry mid-body just closes.
/// Pipeline run time is NOT covered — that's the handler's business.
/// Covers the READ phase only; the write phase has its own
/// per-write budget (`write_timeout`) so a streaming response can
/// legitimately outlive any whole-request clock.
pub request_timeout: Duration,
/// Wall-clock budget for reading the request HEAD, measured from the
/// first byte of a request until the head is fully parsed. Expiry
/// mid-head gets a best-effort 408. Kept short: an incomplete head is
/// the classic slowloris, and a legitimate client sends its head in a
/// single burst. The BODY has its own, larger budget (`body_timeout`)
/// so a slow-but-legit upload is not judged by the head clock.
pub head_timeout: Duration,
/// Absolute wall-clock cap on reading the request BODY, measured from
/// the moment the head finished parsing until the body is fully read.
/// Sized for slow links (e.g. a trickling cellular IoT client), so it
/// is much larger than `head_timeout`. Expiry mid-body just closes —
/// nothing is owed to a client this far gone. Pipeline run time is NOT
/// covered (that's the handler's business); the write phase has its own
/// per-write budget (`write_timeout`).
pub body_timeout: Duration,
/// Burst-gated body stall eviction: the bytes that must accumulate
/// since the last advance to count as a "burst" and reset the stall
/// clock. A body that dribbles fewer than this per `body_stall_timeout`
/// window is evicted — the discriminator between a slowloris trickle
/// (near-zero, smooth) and a slow-but-legit client (delivers real
/// bursts). The pair implies an effective floor of
/// body_burst_bytes / body_stall_timeout, enforced in bursts.
pub body_burst_bytes: usize,
/// Max time since the last qualifying burst (`body_burst_bytes`)
/// before a stalled body read is evicted. Must comfortably exceed a
/// legit client's worst quiet gap (e.g. cellular RRC/handover/DRX
/// stalls). The absolute `body_timeout` always backstops it.
pub body_stall_timeout: Duration,
/// Per-write budget for response bytes: every `write_all` (the fixed
/// head+body, and each streamed chunk) must complete within this.
/// A client that stops reading mid-response is dropped when its
@@ -76,7 +96,10 @@ impl Default for ConnLimits {
max_head_bytes: 64 * 1024,
max_body_bytes: 16 * 1024 * 1024,
keep_alive_timeout: Duration::from_secs(60),
request_timeout: Duration::from_secs(30),
head_timeout: Duration::from_secs(30),
body_timeout: Duration::from_secs(300),
body_burst_bytes: 4 * 1024,
body_stall_timeout: Duration::from_secs(20),
write_timeout: Duration::from_secs(30),
max_frame_payload: 1024 * 1024,
max_message_bytes: 4 * 1024 * 1024,
@@ -92,7 +115,7 @@ pub fn run_connection(
fd: OwnedFd,
pipeline: Pipeline,
limits: ConnLimits,
registry: GenServerRef<ConnRegistry>,
registry: GenServerRef<Endpoint>,
) {
// The OwnedFd cleans up via Drop on any exit path (panic, error, or
// normal close). No explicit close calls below.
@@ -102,7 +125,7 @@ pub fn run_connection(
// Self-register (initially idle: no request head parsed yet) and arm
// the deregistration guard. Both casts come from this actor, so
// Started always precedes Ended in the registry's inbox — see
// conn_registry module docs for why the listener must not do this.
// endpoint module docs for why the listener must not do this.
let me = smarm::self_pid();
let _ = registry.cast(Cast::ConnStarted(me));
let _guard = DeregisterGuard::new(registry.clone(), me, Cast::ConnEnded);
@@ -111,7 +134,7 @@ pub fn run_connection(
// ----- 1. Read until we have a full request head. -----
// We are idle until a head parses: stoppable by a draining
// registry while parked here.
let (parsed, request_deadline) = match read_head(raw, &mut buf, &limits) {
let parsed = match read_head(raw, &mut buf, &limits) {
Ok(p) => p,
Err(ReadHeadErr::ClientClosed) => {
// Clean EOF between requests (or before any request). Normal.
@@ -122,8 +145,8 @@ pub fn run_connection(
// a request. Nothing is owed; close silently.
return;
}
Err(ReadHeadErr::RequestTimeout) => {
// request_timeout expired mid-head (slowloris and friends).
Err(ReadHeadErr::HeadTimeout) => {
// head_timeout expired mid-head (slowloris and friends).
// Best-effort 408 WITHOUT parking on writability — a client
// that stalls reads must not defeat the timeout by making
// the 408 write park forever.
@@ -142,6 +165,11 @@ pub fn run_connection(
let _ = registry.cast(Cast::ConnBusy(me));
// ----- 2. Read body. -----
// The body has its OWN absolute budget, anchored here (head just
// parsed) and independent of the head clock — a slow-but-legit
// upload must not be judged by the short head deadline. Expiry
// closes the connection (nothing owed mid-body).
let body_deadline = Instant::now() + limits.body_timeout;
// Content-Length pre-check only applies to fixed bodies; a chunked
// body is bounded incrementally by the decoder.
let body_len = parsed.content_length.unwrap_or(0);
@@ -169,7 +197,7 @@ pub fn run_connection(
// bottom of the loop must drop exactly this much to land on the
// next pipelined request.
let (body, consumed_past_head) = if parsed.chunked {
match read_chunked_body(raw, &mut buf, parsed.head_len, &limits, request_deadline) {
match read_chunked_body(raw, &mut buf, parsed.head_len, &limits, body_deadline) {
Ok(ok) => ok,
Err(ChunkedBodyErr::TooLarge) => {
let _ = write_all(
@@ -191,7 +219,7 @@ pub fn run_connection(
Err(ChunkedBodyErr::Io(_)) => return,
}
} else {
match read_body(raw, &mut buf, parsed.head_len, body_len, request_deadline) {
match read_body(raw, &mut buf, parsed.head_len, body_len, &limits, body_deadline) {
Ok(b) => (b, body_len),
// Timeout mid-body (and any other body io error) -> just
// close; there's no point talking HTTP to a client this far
@@ -330,9 +358,9 @@ enum ReadHeadErr {
/// keep_alive_timeout expired while waiting for the first byte of a
/// request. Close silently.
IdleTimeout,
/// request_timeout expired after the request had started arriving.
/// Best-effort 408.
RequestTimeout,
/// head_timeout expired after the request had started arriving but
/// before the head finished parsing. Best-effort 408.
HeadTimeout,
Io(io::Error),
Parse(ParseError),
}
@@ -347,22 +375,22 @@ enum ReadHeadErr {
/// - while `buf` is empty and nothing has arrived, we are *idle* and the
/// wait is bounded by `keep_alive_timeout`;
/// - the instant the request has started (first byte read, or pipelined
/// bytes already in `buf` at entry), the *request* clock starts: an
/// `Instant` deadline of `request_timeout` from that moment, which also
/// covers body reads — it is returned alongside the parsed head so the
/// caller can thread it into `read_body`.
/// bytes already in `buf` at entry), the *head* clock starts: an
/// `Instant` deadline of `head_timeout` from that moment. This budget
/// covers the HEAD only; the body has its own budget (`body_timeout`),
/// which the caller anchors once the head has parsed.
fn read_head(
fd: RawFd,
buf: &mut Vec<u8>,
limits: &ConnLimits,
) -> Result<(parser::ParsedHead, Instant), ReadHeadErr> {
) -> Result<parser::ParsedHead, ReadHeadErr> {
let entry = Instant::now();
let idle_deadline = entry + limits.keep_alive_timeout;
// Pipelined leftovers count as a started request.
let mut request_deadline: Option<Instant> = if buf.is_empty() {
let mut head_deadline: Option<Instant> = if buf.is_empty() {
None
} else {
Some(entry + limits.request_timeout)
Some(entry + limits.head_timeout)
};
loop {
@@ -375,11 +403,7 @@ fn read_head(
parser::parse_head(buf, limits.max_headers)
};
match head {
Ok(h) => {
let deadline = request_deadline
.unwrap_or_else(|| Instant::now() + limits.request_timeout);
return Ok((h, deadline));
}
Ok(h) => return Ok(h),
Err(ParseError::Incomplete) => {} // need more bytes
Err(e) => return Err(ReadHeadErr::Parse(e)),
}
@@ -390,19 +414,19 @@ fn read_head(
}
// Read more, bounded by whichever budget is active.
let deadline = request_deadline.unwrap_or(idle_deadline);
let deadline = head_deadline.unwrap_or(idle_deadline);
match read_some(fd, buf, limits.initial_read_buf, deadline) {
Ok(0) => return Err(ReadHeadErr::ClientClosed),
Ok(_) => {
if request_deadline.is_none() {
// First byte(s) of this request: the request clock
if head_deadline.is_none() {
// First byte(s) of this request: the head clock
// starts now.
request_deadline = Some(Instant::now() + limits.request_timeout);
head_deadline = Some(Instant::now() + limits.head_timeout);
}
}
Err(e) if e.kind() == ErrorKind::TimedOut => {
return Err(if request_deadline.is_some() {
ReadHeadErr::RequestTimeout
return Err(if head_deadline.is_some() {
ReadHeadErr::HeadTimeout
} else {
ReadHeadErr::IdleTimeout
});
@@ -412,6 +436,58 @@ fn read_head(
}
}
// ---------------------------------------------------------------------------
// BodyStallGate — burst-gated stall eviction for body reads
// ---------------------------------------------------------------------------
//
// Each body read is bounded by the SOONER of two deadlines: the absolute
// body cap (`body_timeout`, passed in as `cap`) and a sliding stall window
// (`mark + body_stall_timeout`). The stall mark only advances when the
// client delivers a full burst (`body_burst_bytes` accumulated since the
// last advance) — so a steady sub-burst trickle never moves the mark and
// is evicted at ~body_stall_timeout, while a bursty slow-but-legit client
// keeps resetting it and survives up to the absolute cap.
//
// State is two words (`mark`, `since_mark`); the per-read cost is one add
// and one compare. Bytes counted are RAW socket bytes (progress = the
// client is sending *something*), so chunked framing counts too, and a
// burst that the kernel fragments into several reads still accumulates.
struct BodyStallGate {
cap: Instant,
stall_timeout: Duration,
burst_bytes: usize,
mark: Instant,
since_mark: usize,
}
impl BodyStallGate {
fn new(cap: Instant, limits: &ConnLimits, now: Instant) -> Self {
Self {
cap,
stall_timeout: limits.body_stall_timeout,
burst_bytes: limits.body_burst_bytes,
mark: now,
since_mark: 0,
}
}
/// Deadline for the next read: the sooner of the absolute cap and the
/// current stall window.
fn deadline(&self) -> Instant {
(self.mark + self.stall_timeout).min(self.cap)
}
/// Record `n` freshly-read raw body bytes; advance the stall mark if a
/// full burst has accumulated since the last advance.
fn record(&mut self, n: usize, now: Instant) {
self.since_mark += n;
if self.since_mark >= self.burst_bytes {
self.mark = now;
self.since_mark = 0;
}
}
}
// ---------------------------------------------------------------------------
// read_body
// ---------------------------------------------------------------------------
@@ -421,7 +497,8 @@ fn read_body(
buf: &mut Vec<u8>,
head_len: usize,
body_len: usize,
deadline: Instant,
limits: &ConnLimits,
cap: Instant,
) -> io::Result<Vec<u8>> {
// Bytes already in `buf` past the head belong to the body.
let already = buf.len().saturating_sub(head_len);
@@ -433,13 +510,17 @@ fn read_body(
return Ok(buf[head_len..head_len + body_len].to_vec());
}
// Read until we have the rest, on the same request budget that the
// head was read under.
// Read until we have the rest, bounded by the body cap AND the
// burst-gated stall window (whichever is sooner).
let mut gate = BodyStallGate::new(cap, limits, Instant::now());
let mut total_read = already;
while total_read < body_len {
match read_some(fd, buf, 8 * 1024, deadline) {
match read_some(fd, buf, 8 * 1024, gate.deadline()) {
Ok(0) => return Err(io::Error::new(ErrorKind::UnexpectedEof, "client closed during body")),
Ok(n) => total_read += n,
Ok(n) => {
total_read += n;
gate.record(n, Instant::now());
}
Err(e) => return Err(e),
}
}
@@ -451,8 +532,9 @@ fn read_body(
// ---------------------------------------------------------------------------
//
// Decodes `Transfer-Encoding: chunked` from `buf[head_len..]`, reading more
// from the socket as needed on the SAME request deadline the head was read
// under. Returns (decoded_body, raw_bytes_consumed_past_head) — the raw
// from the socket as needed on the body deadline (anchored by the caller
// when the head finished parsing, independent of the head clock).
// Returns (decoded_body, raw_bytes_consumed_past_head) — the raw
// count includes all framing and the trailer section, so the caller's
// keep-alive drain lands exactly on the next pipelined request.
//
@@ -479,23 +561,25 @@ fn read_chunked_body(
limits: &ConnLimits,
deadline: Instant,
) -> Result<(Vec<u8>, usize), ChunkedBodyErr> {
// Ensure `buf` holds at least `until` bytes, reading on the request
// deadline. Io(TimedOut) on expiry, UnexpectedEof on early close.
// Ensure `buf` holds at least `until` bytes, reading under the body
// stall gate (absolute cap AND burst-gated stall window). Io(TimedOut)
// on expiry, UnexpectedEof on early close. All chunked socket reads
// funnel through here, so recording bytes here covers the whole path.
fn fill_to(
fd: RawFd,
buf: &mut Vec<u8>,
until: usize,
deadline: Instant,
gate: &mut BodyStallGate,
) -> Result<(), ChunkedBodyErr> {
while buf.len() < until {
match read_some(fd, buf, 8 * 1024, deadline) {
match read_some(fd, buf, 8 * 1024, gate.deadline()) {
Ok(0) => {
return Err(ChunkedBodyErr::Io(io::Error::new(
ErrorKind::UnexpectedEof,
"client closed during chunked body",
)))
}
Ok(_) => {}
Ok(n) => gate.record(n, Instant::now()),
Err(e) => return Err(ChunkedBodyErr::Io(e)),
}
}
@@ -510,7 +594,7 @@ fn read_chunked_body(
buf: &mut Vec<u8>,
from: usize,
max_line: usize,
deadline: Instant,
gate: &mut BodyStallGate,
) -> Result<usize, ChunkedBodyErr> {
let mut scan = from;
loop {
@@ -523,16 +607,19 @@ fn read_chunked_body(
return Err(ChunkedBodyErr::Malformed);
}
}
fill_to(fd, buf, buf.len() + 1, deadline)?;
fill_to(fd, buf, buf.len() + 1, gate)?;
}
}
// `deadline` is the absolute body cap; the gate layers the burst-gated
// stall window under it. All reads below go through find_crlf/fill_to.
let mut gate = BodyStallGate::new(deadline, limits, Instant::now());
let mut pos = head_len;
let mut decoded: Vec<u8> = Vec::new();
loop {
// ----- size line: HEX[;extensions]\r\n -----
let line_end = find_crlf(fd, buf, pos, MAX_SIZE_LINE, deadline)?;
let line_end = find_crlf(fd, buf, pos, MAX_SIZE_LINE, &mut gate)?;
let line = &buf[pos..line_end];
let size_str = match line.iter().position(|&b| b == b';') {
Some(i) => &line[..i], // chunk extensions: ignored
@@ -549,7 +636,7 @@ fn read_chunked_body(
// ----- trailer section: zero or more header lines, then CRLF -----
let trailer_start = pos;
loop {
let t_end = find_crlf(fd, buf, pos, MAX_SIZE_LINE.max(1024), deadline)?;
let t_end = find_crlf(fd, buf, pos, MAX_SIZE_LINE.max(1024), &mut gate)?;
let empty = t_end == pos;
pos = t_end + 2;
if empty {
@@ -566,7 +653,7 @@ fn read_chunked_body(
}
// ----- chunk payload + trailing CRLF -----
fill_to(fd, buf, pos + size + 2, deadline)?;
fill_to(fd, buf, pos + size + 2, &mut gate)?;
decoded.extend_from_slice(&buf[pos..pos + size]);
if &buf[pos + size..pos + size + 2] != b"\r\n" {
return Err(ChunkedBodyErr::Malformed);
@@ -773,6 +860,8 @@ fn emit_error_response(fd: RawFd, err: &ParseError, deadline: Instant) {
b"HTTP/1.1 400 Bad Request\r\ncontent-length: 0\r\nconnection: close\r\n\r\n",
ParseError::Unsupported =>
b"HTTP/1.1 411 Length Required\r\ncontent-length: 0\r\nconnection: close\r\n\r\n",
ParseError::UnknownTransferCoding =>
b"HTTP/1.1 501 Not Implemented\r\ncontent-length: 0\r\nconnection: close\r\n\r\n",
// Incomplete and Malformed both lead here; Incomplete shouldn't
// appear (read_head loops on it).
_ =>
-168
View File
@@ -1,168 +0,0 @@
//! Connection registry — the shutdown coordinator (v0.2 chunk 2).
//!
//! A `gen_server` tracking live connection pids and their busy/idle
//! state. Connection actors self-register as their first action and
//! self-deregister via a drop guard (so panic unwinds deregister too);
//! both casts come from the same sender, so Started always precedes
//! Ended in the inbox. (The roadmap sketched the *listener* casting
//! `{Started, pid}`, but then a short-lived conn's Ended could overtake
//! its Started and leak a dead pid into the set forever. Self-
//! registration makes the order a per-sender FIFO guarantee instead of
//! a race.)
//!
//! Listeners are deliberately NOT tracked here: they shut down via a
//! shared flag + timed accept-waits (see `serve`), never via
//! `request_stop` — stopping pids that announce themselves is racy (the
//! spawn-to-registration gap), and smarm's `request_stop` is lossy
//! against an actor that is QUEUED and then parks without passing an
//! observation point (see the v0.2 shutdown notes in the commit
//! message). Connections don't suffer this in practice: every stop the
//! registry issues targets a pid that just sent us a cast (so it is
//! running or parked, both covered), and the force-stop path re-sweeps
//! until the set empties.
//!
//! Drain protocol: `BeginDrain` stops every idle connection immediately
//! and flips the registry into draining mode, in which any connection
//! that *becomes* idle (finishes its in-flight request) is stopped on the
//! spot. Busy connections are left to finish; `ForceStopConns` (sent by
//! `serve` at the drain deadline) stops whatever remains. `request_stop`
//! unwinds a conn actor parked in `wait_readable` safely (smarm's 06-10
//! io fix) and `OwnedFd::drop` closes its socket on the way out.
//!
//! This server is also the planned introspection point for ws/channels
//! (roadmap v0.4+), which is why it exists as its own module rather than
//! being inlined into `serve`.
use smarm::{GenServer, Pid, GenServerBuilder, GenServerRef};
use std::collections::HashMap;
// ---------------------------------------------------------------------------
// Messages
// ---------------------------------------------------------------------------
pub enum Cast {
/// A connection actor started (self-registered, initially idle: it has
/// not parsed a request head yet).
ConnStarted(Pid),
/// Parsed a request head; a response is now owed.
ConnBusy(Pid),
/// Response written; parked (or about to park) waiting for the next
/// keep-alive request.
ConnIdle(Pid),
ConnEnded(Pid),
/// Stop idle conns now and stop each remaining conn as it goes idle.
BeginDrain,
/// Drain deadline passed: stop every remaining conn.
ForceStopConns,
}
pub enum Call {
ConnCount,
}
pub enum Reply {
ConnCount(usize),
}
// ---------------------------------------------------------------------------
// Server
// ---------------------------------------------------------------------------
#[derive(Clone, Copy, PartialEq, Eq)]
enum ConnState {
Busy,
Idle,
}
#[derive(Default)]
pub struct ConnRegistry {
conns: HashMap<Pid, ConnState>,
draining: bool,
}
impl GenServer for ConnRegistry {
type Call = Call;
type Reply = Reply;
type Cast = Cast;
type Info = ();
type Timer = ();
fn handle_call(&mut self, request: Call) -> Reply {
match request {
Call::ConnCount => Reply::ConnCount(self.conns.len()),
}
}
fn handle_cast(&mut self, request: Cast) {
match request {
Cast::ConnStarted(pid) => {
self.conns.insert(pid, ConnState::Idle);
if self.draining {
// Listener-stop race: this conn was accepted just
// before its listener died. Drain means no new work.
smarm::request_stop(pid);
}
}
Cast::ConnBusy(pid) => {
if let Some(s) = self.conns.get_mut(&pid) {
*s = ConnState::Busy;
}
}
Cast::ConnIdle(pid) => {
if let Some(s) = self.conns.get_mut(&pid) {
*s = ConnState::Idle;
if self.draining {
// Finished its in-flight request; nothing more is
// owed. The pid leaves the map via its drop
// guard's ConnEnded once the unwind completes.
smarm::request_stop(pid);
}
}
}
Cast::ConnEnded(pid) => { self.conns.remove(&pid); }
Cast::BeginDrain => {
self.draining = true;
for (pid, state) in &self.conns {
if *state == ConnState::Idle {
smarm::request_stop(*pid);
}
}
}
Cast::ForceStopConns => {
for pid in self.conns.keys() {
smarm::request_stop(*pid);
}
}
}
}
}
pub fn start() -> GenServerRef<ConnRegistry> {
GenServerBuilder::new(ConnRegistry::default()).start()
}
// ---------------------------------------------------------------------------
// Drop guard — self-deregistration on any exit path.
// ---------------------------------------------------------------------------
/// Casts `make(pid)` on drop. Runs on normal return, `request_stop`
/// unwind, and panic unwind alike; the cast is infallible from the
/// guard's perspective (a dead registry just returns an ignored Err).
pub struct DeregisterGuard {
registry: GenServerRef<ConnRegistry>,
pid: Pid,
make: fn(Pid) -> Cast,
}
impl DeregisterGuard {
pub fn new(registry: GenServerRef<ConnRegistry>, pid: Pid, make: fn(Pid) -> Cast) -> Self {
Self { registry, pid, make }
}
}
impl Drop for DeregisterGuard {
fn drop(&mut self) {
let _ = self.registry.cast((self.make)(self.pid));
}
}
+656
View File
@@ -0,0 +1,656 @@
//! The endpoint: one supervised gen_server that owns a listening socket,
//! the pool of listener actors accepting on it, and the set of live
//! connections.
//!
//! # Shape (v0.3)
//!
//! ```text
//! your root supervisor
//! └── ChildSpec(Permanent, urus::endpoint(config, pipeline)?) <- Endpoint gen_server
//! └── listener_sup OneForOne over N listener actors (spawned in init)
//! └── (listeners spawn plain connection actors)
//! ```
//!
//! The endpoint runs *inline as the ChildSpec's actor*
//! (`NamedGenServerBuilder::run`), so your supervisor's ordered shutdown
//! reaches it as [`GenServer::handle_shutdown`] and a restart re-runs the
//! factory on the same still-open listen fds.
//!
//! ## Why the endpoint spawns its own listener supervisor
//!
//! The obvious tree is registry and listener-sup as *siblings* under a
//! `RestForOne`, with listeners resolving the registry by name. It has a
//! boot race: smarm's supervisor starts children with a fire-and-forget
//! spawn, so start *order* is not start *readiness* — a listener can look
//! the name up before the registry actor has run. Rather than paper over
//! that with a retry loop, the registrar spawns its consumers: everything
//! here is program order inside one actor's `init`, not a cross-actor
//! guarantee. (The general gap is filed in smarm's ROADMAP as a
//! supervisor readiness-ack item; when it lands, the sibling shape becomes
//! available, but this one costs nothing and is not waiting on it.)
//!
//! ## Connection accounting
//!
//! Connection actors self-register as their first action and self-
//! deregister via a drop guard (so panic unwinds deregister too); both
//! casts come from the same sender, so `ConnStarted` always precedes
//! `ConnEnded` in the inbox. Were the *listener* to announce the pid
//! instead, a short-lived conn's `Ended` could overtake its `Started` and
//! leak a dead pid into the set forever.
//!
//! ## Shutdown
//!
//! `handle_shutdown` (i.e. a `request_shutdown` from your supervisor, or
//! from anywhere):
//!
//! 1. flip `draining` — from here on, a connection that registers or goes
//! idle is stopped on the spot;
//! 2. `request_shutdown` the listener supervisor, which stops its
//! listeners in reverse order and exits; its monitored death is the
//! "no new connections can ever be accepted" barrier;
//! 3. stop every currently-idle connection;
//! 4. arm one `drain_timeout` timer;
//! 5. return `Continue` — the endpoint exits normally once the listener
//! sup is down *and* the connection set is empty, so "the endpoint
//! actor has exited" is exactly "every connection is gone". With no
//! connections and (once the Down lands) nothing to wait for, that is
//! immediate.
//!
//! At the drain deadline one force sweep `request_stop`s whatever remains
//! (plus, defensively, the listener sup). A single sweep suffices: every
//! path a connection can enter the set on afterwards stops it at the point
//! of entry. Force-stopped conns unwind safely out of their fd waits and
//! close their sockets via `OwnedFd::drop`.
//!
//! Give the endpoint [`Shutdown::Infinity`](smarm::supervisor::Shutdown) in
//! your child spec: it bounds itself with `drain_timeout`, and a
//! supervisor-imposed deadline shorter than that would kill the drain
//! halfway and orphan connections.
//!
//! ## Known residual window
//!
//! A connection *spawned* by a listener in the instant before that
//! listener is stopped, which has not yet run, has not registered. If the
//! set was already empty the endpoint can exit before its `ConnStarted`
//! arrives, and that connection serves on as a forest root until smarm's
//! root-exit sweep collects it. Inherent to self-registration (the
//! alternative loses `Started`/`Ended` ordering, which is worse);
//! documented rather than defended.
use crate::conn_actor::{run_connection, ConnLimits};
use crate::net::{accept_nonblocking, bind_and_listen, OwnedFd};
use crate::plug::Pipeline;
use crate::serve::Config;
use smarm::gen_server::{GenServerName, ShutdownAction, StopHandle, TimerHandle};
use smarm::{
ChildSpec, GenServer, GenServerBuilder, GenServerCtx, GenServerRef, OneForOne, Pid, Restart,
Strategy,
};
use std::collections::HashMap;
use std::io::{self, ErrorKind};
use std::os::fd::RawFd;
use std::sync::atomic::{AtomicU32, Ordering};
use std::sync::Arc;
use std::time::Duration;
// ---------------------------------------------------------------------------
// Messages
// ---------------------------------------------------------------------------
pub enum Cast {
/// A connection actor started (self-registered, initially idle: it has
/// not parsed a request head yet).
ConnStarted(Pid),
/// Parsed a request head; a response is now owed.
ConnBusy(Pid),
/// Response written; parked (or about to park) waiting for the next
/// keep-alive request.
ConnIdle(Pid),
ConnEnded(Pid),
}
pub enum Call {
/// Live connection count — introspection (and the planned hook for
/// ws/channels stats).
ConnCount,
}
pub enum Reply {
ConnCount(usize),
}
// ---------------------------------------------------------------------------
// Endpoint
// ---------------------------------------------------------------------------
#[derive(Clone, Copy, PartialEq, Eq)]
enum ConnState {
Busy,
Idle,
}
/// Everything an endpoint incarnation needs to (re)build its listener
/// pool. Shared behind an `Arc` by the factory closure so a restart reuses
/// the same already-bound fds — no re-bind, no window where the port is
/// unclaimed.
struct Boot {
listener_fds: Vec<Arc<OwnedFd>>,
pipeline: Pipeline,
limits: ConnLimits,
conn_stack_reserve: usize,
name: &'static str,
}
pub struct Endpoint {
boot: Arc<Boot>,
conns: HashMap<Pid, ConnState>,
draining: bool,
drain_timeout: Duration,
/// The listener supervisor, spawned in `init` and monitored. `None`
/// once it is down.
listener_sup: Option<Pid>,
stop: Option<StopHandle<Self>>,
timer: Option<TimerHandle<Self>>,
}
impl Endpoint {
fn name(&self) -> GenServerName<Self> {
GenServerName::new(self.boot.name)
}
/// Exit normally once nothing can arrive and nothing is left: the
/// listener sup is down and the conn set is empty. This exit *is* the
/// barrier a supervisor's ordered shutdown waits on.
fn stop_if_drained(&self) {
if self.draining && self.listener_sup.is_none() && self.conns.is_empty() {
self.stop.as_ref().expect("init ran first").stop();
}
}
}
impl GenServer for Endpoint {
type Call = Call;
type Reply = Reply;
type Cast = Cast;
type Info = ();
/// One meaning: the drain deadline elapsed — force-stop the stragglers.
type Timer = ();
fn init(&mut self, ctx: &GenServerCtx<Self>) {
ctx.trap_exit(); // shutdown arrives as handle_shutdown, not a kill
self.stop = Some(ctx.stop_handle());
self.timer = Some(ctx.timer());
// Our own name is bound (from inside `run`) before `init` runs, so
// this resolves to us. Listeners are handed the resolved ref, not
// the name: no per-connection registry lookup on the accept path.
let me: GenServerRef<Self> =
smarm::gen_server::whereis_server(self.name()).expect("endpoint name bound before init");
let boot = self.boot.clone();
let sup = smarm::spawn(move || {
let mut sup = OneForOne::new().strategy(Strategy::OneForOne);
for lfd in &boot.listener_fds {
let lfd = lfd.clone();
let pipeline = boot.pipeline.clone();
let limits = boot.limits;
let reserve = boot.conn_stack_reserve;
let me = me.clone();
// Permanent: a listener only ever exits by supervisor
// action now (its normal-exit-as-shutdown flag is gone),
// so "exited on its own" always means something broke and
// always deserves a restart.
sup = sup.child(ChildSpec::new(Restart::Permanent, move || {
listener_loop(lfd.clone(), pipeline.clone(), limits, reserve, me.clone());
}));
}
// Default intensity (3 restarts / 5s) applies; a listener
// crash-looping faster than that trips the cap and the pool
// tears down — which our monitor turns into a loud endpoint
// failure rather than a zombie server on a dead port.
sup.run();
});
ctx.watch(smarm::monitor(sup.pid()));
self.listener_sup = Some(sup.pid());
}
fn handle_call(&mut self, request: Call) -> Reply {
match request {
Call::ConnCount => Reply::ConnCount(self.conns.len()),
}
}
fn handle_cast(&mut self, request: Cast) {
match request {
Cast::ConnStarted(pid) => {
self.conns.insert(pid, ConnState::Idle);
if self.draining {
// Accepted just before its listener was stopped.
// Draining means no new work; it leaves the map via
// its guard's ConnEnded.
smarm::request_stop(pid);
}
}
Cast::ConnBusy(pid) => {
if let Some(s) = self.conns.get_mut(&pid) {
*s = ConnState::Busy;
}
}
Cast::ConnIdle(pid) => {
if let Some(s) = self.conns.get_mut(&pid) {
*s = ConnState::Idle;
if self.draining {
smarm::request_stop(pid); // in-flight request finished
}
}
}
Cast::ConnEnded(pid) => {
self.conns.remove(&pid);
self.stop_if_drained();
}
}
}
fn handle_down(&mut self, down: smarm::Down) {
if self.listener_sup != Some(down.pid) {
return;
}
self.listener_sup = None;
if !self.draining {
// The pool died on its own (restart intensity exceeded, or the
// supervisor itself failed). Nothing is listening any more, so
// this endpoint is a zombie: fail loudly and let the caller's
// supervisor decide (restart re-runs init on the same fds).
panic!("urus endpoint '{}': listener pool died: {:?}", self.boot.name, down.reason);
}
// Expected during shutdown: the accept side is now provably gone.
self.stop_if_drained();
}
fn handle_shutdown(&mut self) -> ShutdownAction {
self.draining = true;
if let Some(sup) = self.listener_sup {
// Stops listeners in reverse start order and exits; the
// monitored Down is our "no new connections" barrier.
smarm::request_shutdown(sup);
}
for (pid, state) in &self.conns {
if *state == ConnState::Idle {
smarm::request_stop(*pid);
}
}
// Busy conns get until the deadline; then handle_timer sweeps.
self.timer
.as_ref()
.expect("init ran first")
.arm_after(self.drain_timeout, ());
ShutdownAction::Continue
}
fn handle_timer(&mut self, _deadline: ()) {
// Drain deadline: force-stop everything left. One sweep — late
// registrants are already stopped on arrival (see module docs).
for pid in self.conns.keys() {
smarm::request_stop(*pid);
}
if let Some(sup) = self.listener_sup {
// Defensive: a listener wedged past its supervisor's own grace
// period must not hold the whole shutdown open.
smarm::request_stop(sup);
}
// The map empties via each conn's guard ConnEnded; stop_if_drained
// fires on the last one (or on the sup's Down, whichever is last).
}
}
// ---------------------------------------------------------------------------
// endpoint() — the public constructor
// ---------------------------------------------------------------------------
/// Bind `config.addr` and return the endpoint's supervisable body: pass it
/// to [`ChildSpec::new`] under your own supervisor, alongside your
/// application's other children.
///
/// The bind happens **here**, eagerly, so an address-in-use error surfaces
/// on the caller's thread rather than inside an actor — and the fds
/// outlive any restart of the child.
///
/// ```ignore
/// let endpoint = urus::endpoint(Config::new(addr), pipeline)?;
/// let rt = smarm::init(smarm::Config::default());
/// rt.run(move || {
/// smarm::OneForOne::new()
/// .child(ChildSpec::new(Restart::Permanent, my_app_state))
/// .child(ChildSpec::new(Restart::Permanent, endpoint)
/// .shutdown(Shutdown::Infinity))
/// .run()
/// });
/// ```
///
/// The returned closure is `Clone`, so one endpoint definition can be
/// handed to more than one place; each *invocation* is one running
/// endpoint, and two live at once under the same [`Config::name`] is a
/// name clash (panic on the second).
pub fn endpoint(
config: Config,
pipeline: Pipeline,
) -> io::Result<impl Fn() + Clone + Send + Sync + 'static> {
// One dup'd fd per listener: `accept4` is thread-safe on a single fd,
// but a per-listener RawFd keeps each actor's epoll registration
// distinct in smarm's `waiters: HashMap<RawFd, Pid>`.
let mut listener_fds = Vec::with_capacity(config.listener_pool);
listener_fds.push(Arc::new(bind_and_listen(config.addr)?));
for _ in 1..config.listener_pool {
let dup = dup_fd(listener_fds[0].as_raw())?;
listener_fds.push(Arc::new(dup));
}
let boot = Arc::new(Boot {
listener_fds,
pipeline,
limits: config.to_conn_limits(),
conn_stack_reserve: config.conn_stack_reserve,
name: config.name,
});
let drain_timeout = config.drain_timeout;
Ok(move || {
let state = Endpoint {
boot: boot.clone(),
conns: HashMap::new(),
draining: false,
drain_timeout,
listener_sup: None,
stop: None,
timer: None,
};
let name = GenServerName::<Endpoint>::new(state.boot.name);
// Inline: this actor *is* the server, so a supervisor's shutdown
// reaches handle_shutdown and a restart re-binds the name.
if let Err(e) = GenServerBuilder::new(state).named(name).run() {
panic!("urus endpoint '{}': name already taken: {e:?}", name.as_str());
}
})
}
/// Resolve a running endpoint by [`Config::name`] — for `ConnCount` and
/// other introspection from elsewhere in the app.
pub fn whereis(name: &'static str) -> Option<GenServerRef<Endpoint>> {
smarm::gen_server::whereis_server(GenServerName::<Endpoint>::new(name))
}
// ---------------------------------------------------------------------------
// Listener actors
// ---------------------------------------------------------------------------
fn dup_fd(fd: RawFd) -> io::Result<OwnedFd> {
let new_fd = unsafe { libc::fcntl(fd, libc::F_DUPFD_CLOEXEC, 0) };
if new_fd < 0 {
return Err(io::Error::last_os_error());
}
Ok(OwnedFd::from_raw(new_fd))
}
/// Test-only fault injection. When nonzero, the next accept-loop iteration
/// of whichever listener gets there first decrements this and panics —
/// *before* calling `accept`, so a pending connection stays in the kernel
/// backlog and must be picked up by the restarted listener. Cost when idle
/// is one relaxed load per accept-loop iteration (each of which already
/// pays a syscall). Not public API.
#[doc(hidden)]
pub static INJECT_LISTENER_PANICS: AtomicU32 = AtomicU32::new(0);
/// Accept forever, spawning one connection actor per connection. Exits
/// only by supervisor action: `request_stop` unwinds the untimed
/// `wait_readable` park below (smarm's 06-10 io fix), which is why there
/// is no shutdown flag and no tick — the park is genuinely open-ended and
/// costs nothing while idle.
fn listener_loop(
listener: Arc<OwnedFd>,
pipeline: Pipeline,
limits: ConnLimits,
conn_stack_reserve: usize,
endpoint: GenServerRef<Endpoint>,
) {
let fd = listener.as_raw();
loop {
if INJECT_LISTENER_PANICS
.fetch_update(Ordering::Relaxed, Ordering::Relaxed, |n| n.checked_sub(1))
.is_ok()
{
panic!("urus: injected listener panic (test hook)");
}
match accept_nonblocking(fd) {
Ok(client) => {
// Hand the fd off to a new connection actor. spawn() is
// cheap on smarm — it's a single Vec push under the
// shared lock.
let p = pipeline.clone();
let l = limits;
let e = endpoint.clone();
let opts = smarm::SpawnOpts {
stack_reserve: Some(conn_stack_reserve),
..smarm::SpawnOpts::default()
};
smarm::spawn_with(opts, move || run_connection(client, p, l, e));
}
Err(e) if e.kind() == ErrorKind::WouldBlock => {
if let Err(we) = smarm::wait_readable(fd) {
// epoll registration failed — abnormal, so panic: a
// transient failure (e.g. EMFILE on the epoll set)
// heals by restart instead of silently shrinking the
// pool. smarm catches actor panics in the trampoline;
// this is a Signal::Panic to the supervisor, not
// process noise.
panic!("urus: listener wait_readable failed: {we}");
}
}
Err(e) if e.kind() == ErrorKind::Interrupted => continue,
Err(e) => {
// EMFILE / ENFILE / ECONNABORTED etc. Log and back off
// briefly in case the error is sticky; the system may
// recover.
eprintln!("urus: accept error: {e}");
smarm::sleep(Duration::from_millis(10));
}
}
}
// The Arc clone we were started with drops on unwind, but the
// ChildSpec factory holds another — the fd outlives any one
// incarnation of this listener.
}
// ---------------------------------------------------------------------------
// Deregistration guard
// ---------------------------------------------------------------------------
/// Casts `make(pid)` on drop. Runs on normal return, `request_stop`
/// unwind, and panic unwind alike; the cast is infallible from the
/// guard's perspective (a dead endpoint just returns an ignored Err).
pub struct DeregisterGuard {
endpoint: GenServerRef<Endpoint>,
pid: Pid,
make: fn(Pid) -> Cast,
}
impl DeregisterGuard {
pub fn new(endpoint: GenServerRef<Endpoint>, pid: Pid, make: fn(Pid) -> Cast) -> Self {
Self { endpoint, pid, make }
}
}
impl Drop for DeregisterGuard {
fn drop(&mut self) {
let _ = self.endpoint.cast((self.make)(self.pid));
}
}
// ---------------------------------------------------------------------------
// Tests — the drain protocol, driven purely by request_shutdown.
// ---------------------------------------------------------------------------
#[cfg(test)]
mod tests {
use super::*;
use crate::plug::Pipeline;
use std::time::Instant;
const NAME: &str = "urus.test.endpoint";
/// Spawn a real endpoint (bound to an ephemeral port) as a supervised
/// child, exactly as an application would, and hand back its
/// supervisor's pid.
fn spawn_endpoint(drain: Duration) -> Pid {
let cfg = Config {
listener_pool: 1,
drain_timeout: drain,
name: NAME,
..Config::new("127.0.0.1:0".parse().unwrap())
};
let body = endpoint(cfg, Pipeline::new()).expect("bind");
let sup = smarm::spawn(move || {
OneForOne::new()
.child(
ChildSpec::new(Restart::Permanent, body)
.shutdown(smarm::supervisor::Shutdown::Infinity),
)
.run()
});
await_pred("endpoint up", Duration::from_secs(2), || {
whereis(NAME).is_some()
});
sup.pid()
}
/// A stand-in connection actor: registers with the endpoint, optionally
/// reports busy, then parks forever. Only a `request_stop` ends it; the
/// guard's ConnEnded runs on the unwind.
fn fake_conn(ep: GenServerRef<Endpoint>, busy: bool) {
let me = smarm::self_pid();
let _ = ep.cast(Cast::ConnStarted(me));
let _guard = DeregisterGuard::new(ep.clone(), me, Cast::ConnEnded);
if busy {
let _ = ep.cast(Cast::ConnBusy(me));
}
loop {
smarm::sleep(Duration::from_secs(3600));
}
}
fn spawn_fake_conn(busy: bool) -> Pid {
let ep = whereis(NAME).expect("endpoint live");
smarm::spawn(move || fake_conn(ep, busy)).pid()
}
/// Live conn count, or `Err` once the endpoint is gone.
fn conn_count() -> Result<usize, ()> {
match whereis(NAME) {
Some(ep) => match ep.call(Call::ConnCount) {
Ok(Reply::ConnCount(n)) => Ok(n),
Err(_) => Err(()),
},
None => Err(()),
}
}
fn await_pred(what: &str, deadline: Duration, mut pred: impl FnMut() -> bool) {
let end = Instant::now() + deadline;
while !pred() {
assert!(Instant::now() < end, "timed out waiting for: {what}");
smarm::sleep(Duration::from_millis(5));
}
}
#[test]
fn shutdown_with_no_conns_exits_promptly() {
smarm::run(|| {
let sup = spawn_endpoint(Duration::from_secs(30));
assert_eq!(conn_count(), Ok(0));
let t0 = Instant::now();
smarm::request_shutdown(sup);
// Nothing to drain: gone well inside the (huge) drain window,
// i.e. the idle path does not run the clock out.
await_pred("endpoint exit", Duration::from_secs(3), || {
conn_count().is_err()
});
assert!(t0.elapsed() < Duration::from_secs(3));
});
}
#[test]
fn drain_stops_idle_now_busy_at_deadline_then_exits() {
smarm::run(|| {
let drain = Duration::from_millis(300);
let sup = spawn_endpoint(drain);
spawn_fake_conn(false); // idle
spawn_fake_conn(true); // busy
await_pred("both registered", Duration::from_secs(2), || {
conn_count() == Ok(2)
});
let t0 = Instant::now();
smarm::request_shutdown(sup);
// Idle conn goes promptly, well before the deadline.
await_pred("idle stopped", drain / 2, || conn_count() == Ok(1));
// Busy conn holds until the force sweep, then the endpoint winds up.
await_pred("busy swept + endpoint exit", drain * 6, || {
conn_count().is_err()
});
assert!(
t0.elapsed() >= drain,
"busy conn must not be stopped before the drain deadline"
);
});
}
#[test]
fn conn_finishing_mid_drain_is_stopped_on_idle() {
smarm::run(|| {
// Long deadline: this passes only via the ConnIdle path, not
// the force sweep.
let sup = spawn_endpoint(Duration::from_secs(30));
let conn = spawn_fake_conn(true);
await_pred("registered busy", Duration::from_secs(2), || {
conn_count() == Ok(1)
});
smarm::request_shutdown(sup);
smarm::sleep(Duration::from_millis(50));
assert_eq!(conn_count(), Ok(1), "busy conn survives the idle sweep");
// "Request finishes": the conn reports idle.
let ep = whereis(NAME).expect("endpoint still draining");
let _ = ep.cast(Cast::ConnIdle(conn));
await_pred("stopped on idle + endpoint exit", Duration::from_secs(3), || {
conn_count().is_err()
});
});
}
#[test]
fn conn_registering_mid_drain_is_stopped_on_arrival() {
smarm::run(|| {
let sup = spawn_endpoint(Duration::from_secs(30));
let holder = spawn_fake_conn(true); // keeps the drain open
await_pred("holder registered", Duration::from_secs(2), || {
conn_count() == Ok(1)
});
smarm::request_shutdown(sup);
smarm::sleep(Duration::from_millis(20));
// The listener-death race: a fresh conn registers mid-drain.
spawn_fake_conn(false);
// It is stopped on arrival — the count returns to just the holder.
await_pred("late arrival stopped", Duration::from_secs(2), || {
conn_count() == Ok(1)
});
// Cleanup: release the holder so the run can end.
let ep = whereis(NAME).expect("endpoint still draining");
let _ = ep.cast(Cast::ConnIdle(holder));
});
}
}
+4 -1
View File
@@ -24,7 +24,7 @@ pub mod router;
pub mod parser;
pub mod net;
pub mod conn_actor;
pub mod conn_registry;
pub mod endpoint;
pub mod serve;
pub mod sse;
pub mod ws;
@@ -41,6 +41,9 @@ pub use pubsub::{PubSub, PubSubDown};
#[cfg(feature = "channels")]
pub use channels::{Channel, ChannelHub, ChannelSession, ChannelSocket, PrefixRouter, Status, TopicRouter};
pub use ws::{Message, WsClosed, WsHandler, WsSender};
pub use endpoint::endpoint;
pub use serve::{
serve, serve_with, serve_with_shutdown, shutdown_handle, Config, Handle, ShutdownSignal,
};
#[cfg(feature = "config-file")]
pub use serve::ConfigError;
+231 -12
View File
@@ -9,8 +9,10 @@
//! - No body header — empty body.
//! - `Transfer-Encoding: chunked` (HTTP/1.1) — flagged in `ParsedHead`;
//! the connection actor decodes incrementally (`read_chunked_body`).
//! Chunked + Content-Length together, or chunked on HTTP/1.0, is
//! Malformed (request-smuggling ambiguity; RFC 7230 §3.3.3).
//! TE is 1.1-only and overrides Content-Length: TE on HTTP/1.0, or TE
//! together with a Content-Length, is Malformed (400). `chunked` must be
//! the final coding (non-final -> 400); any other coding is unimplemented
//! (-> 501). Only a sole final `chunked` sets the flag (RFC 9112 §6.1/§6.3).
use crate::conn::{Body, Conn, HeaderMap, HttpVersion, Method, RespBody};
@@ -33,6 +35,10 @@ pub enum ParseError {
/// (chunked decoding landed in v0.3); kept for future unsupported
/// framings. Connection actor responds 411 + close.
Unsupported,
/// `Transfer-Encoding` names a transfer coding we don't implement
/// (`chunked` is the only one urus decodes). Connection actor responds
/// 501 Not Implemented + close (RFC 9112 §6.1, §7).
UnknownTransferCoding,
}
// ---------------------------------------------------------------------------
@@ -100,6 +106,11 @@ pub fn parse_head(buf: &[u8], max_headers: usize) -> Result<ParsedHead, ParseErr
let mut connection_hdr = None;
let mut chunked = false;
let mut expect_100 = false;
let mut host_count = 0usize;
let mut host_ok = true;
let mut cl_count = 0usize;
let mut te_present = false;
let mut te_codings: Vec<String> = Vec::new();
for h in req.headers.iter() {
let name_lower = h.name.to_ascii_lowercase();
@@ -107,6 +118,10 @@ pub fn parse_head(buf: &[u8], max_headers: usize) -> Result<ParsedHead, ParseErr
match name_lower.as_str() {
"content-length" => {
// Count occurrences; duplicates (even equal) are rejected
// post-loop. A single value must be one decimal integer —
// a comma-list ("5, 5") or non-numeric fails parse here.
cl_count += 1;
content_length = Some(
value.trim()
.parse::<usize>()
@@ -114,10 +129,17 @@ pub fn parse_head(buf: &[u8], max_headers: usize) -> Result<ParsedHead, ParseErr
);
}
"transfer-encoding" => {
// We only care whether it includes "chunked". Multiple codings
// can appear; chunked is the only one we'd need to decode.
if value.to_ascii_lowercase().split(',').any(|t| t.trim() == "chunked") {
chunked = true;
// Collect the ordered coding list across any number of TE
// headers; finality/known-ness is decided post-loop. Empty
// list elements (legacy `#rule`, e.g. a trailing comma) are
// skipped; a wholly empty value leaves te_codings empty and
// is caught below.
te_present = true;
for coding in value.split(',') {
let c = coding.trim().to_ascii_lowercase();
if !c.is_empty() {
te_codings.push(c);
}
}
}
"connection" => {
@@ -126,19 +148,68 @@ pub fn parse_head(buf: &[u8], max_headers: usize) -> Result<ParsedHead, ParseErr
"expect" if value.eq_ignore_ascii_case("100-continue") => {
expect_100 = true;
}
"host" => {
// Presence/uniqueness enforced post-loop; validity here.
host_count += 1;
if !valid_host(value) {
host_ok = false;
}
}
_ => {}
}
headers.append(&name_lower, value.to_string());
}
if chunked {
// Transfer-Encoding is an HTTP/1.1 mechanism; a 1.0 request
// carrying it is malformed. And a request carrying BOTH a
// Content-Length and TE: chunked is the classic request-smuggling
// ambiguity — RFC 7230 §3.3.3 lets a server reject it, and we do.
if version == HttpVersion::Http10 || content_length.is_some() {
// Host (RFC 9112 §3.2): an HTTP/1.1 request MUST carry exactly one valid
// Host; a missing, duplicate, or malformed Host is a 400. HTTP/1.0 may
// omit Host, but a duplicate or invalid one is still rejected on any
// version (ambiguous / malformed authority).
if host_count > 1 || !host_ok {
return Err(ParseError::Malformed);
}
if version == HttpVersion::Http11 && host_count == 0 {
return Err(ParseError::Malformed);
}
// Content-Length (RFC 9112 §6.3): more than one Content-Length is an
// unrecoverable framing ambiguity (CL.CL request smuggling). We are
// strict — reject any duplicate, not only differing values.
if cl_count > 1 {
return Err(ParseError::BadContentLength);
}
// Transfer-Encoding (RFC 9112 §6.1/§6.3). TE is a 1.1 mechanism and
// overrides Content-Length; only `chunked` is implemented here.
if te_present {
// TE on HTTP/1.0 is malformed (no 1.0 chunked).
if version == HttpVersion::Http10 {
return Err(ParseError::Malformed);
}
// TE together with Content-Length is the classic smuggling
// ambiguity; TE overrides CL and we reject rather than forward.
if content_length.is_some() {
return Err(ParseError::Malformed);
}
// A Transfer-Encoding header that carries no coding frames nothing.
if te_codings.is_empty() {
return Err(ParseError::Malformed);
}
let last_is_chunked = te_codings.last().map(String::as_str) == Some("chunked");
let has_chunked = te_codings.iter().any(|c| c == "chunked");
if has_chunked && !last_is_chunked {
// chunked present but not final: body length isn't reliably
// determinable -> 400.
return Err(ParseError::Malformed);
}
if te_codings.iter().any(|c| c != "chunked") {
// Some coding we don't implement (chunked is the only decodable
// one). Whether or not chunked is final, we can't apply it -> 501.
return Err(ParseError::UnknownTransferCoding);
}
// Sole, final `chunked`: the connection actor decodes the body.
chunked = true;
}
// Keep-alive logic, RFC 7230 §6.3:
@@ -163,6 +234,32 @@ pub fn parse_head(buf: &[u8], max_headers: usize) -> Result<ParsedHead, ParseErr
})
}
/// Conservative RFC 3986 check for a `Host` field-value: non-empty and every
/// byte drawn from the `host[:port]` productions (reg-name / IP-literal
/// brackets / port colon). This is charset-level, not full structural
/// validation (no bracket matching, no pct-encoding well-formedness) — enough
/// to reject the smuggling-relevant garbage (whitespace, controls, `@`, `/`,
/// `?`, `#`) while accepting every legitimate host. Tighter structural checks
/// (bracketed IPv6, single port colon) are a possible follow-up.
fn valid_host(value: &str) -> bool {
!value.is_empty()
&& value.bytes().all(|b| {
b.is_ascii_alphanumeric()
|| matches!(
b,
// unreserved punctuation
b'-' | b'.' | b'_' | b'~'
// sub-delims
| b'!' | b'$' | b'&' | b'\'' | b'(' | b')'
| b'*' | b'+' | b',' | b';' | b'='
// pct-encoded lead
| b'%'
// IP-literal brackets + port separator
| b'[' | b']' | b':'
)
})
}
// ---------------------------------------------------------------------------
// Conn assembly
// ---------------------------------------------------------------------------
@@ -403,6 +500,128 @@ mod tests {
}
}
// --- Host (RFC 9112 §3.2) -------------------------------------------
#[test]
fn parse_missing_host_http11_is_malformed() {
let req = b"GET / HTTP/1.1\r\n\r\n";
match parse_head(req, 64) {
Err(ParseError::Malformed) => {}
_ => panic!("expected Malformed for missing Host on 1.1"),
}
}
#[test]
fn parse_missing_host_http10_is_allowed() {
// Host is optional in HTTP/1.0.
let req = b"GET / HTTP/1.0\r\n\r\n";
assert!(parse_head(req, 64).is_ok(), "1.0 may omit Host");
}
#[test]
fn parse_duplicate_host_is_malformed() {
let req = b"GET / HTTP/1.1\r\nHost: a\r\nHost: b\r\n\r\n";
match parse_head(req, 64) {
Err(ParseError::Malformed) => {}
_ => panic!("expected Malformed for duplicate Host"),
}
}
#[test]
fn parse_invalid_host_value_is_malformed() {
// Embedded whitespace — invalid in an RFC 3986 authority.
let req = b"GET / HTTP/1.1\r\nHost: bad host\r\n\r\n";
match parse_head(req, 64) {
Err(ParseError::Malformed) => {}
_ => panic!("expected Malformed for invalid Host"),
}
}
#[test]
fn parse_valid_hosts_accepted() {
// Positive controls: reg-name, reg-name:port, and IPv6-literal:port.
for req in [
b"GET / HTTP/1.1\r\nHost: example.com\r\n\r\n".as_slice(),
b"GET / HTTP/1.1\r\nHost: example.com:8080\r\n\r\n".as_slice(),
b"GET / HTTP/1.1\r\nHost: [::1]:443\r\n\r\n".as_slice(),
] {
assert!(parse_head(req, 64).is_ok(), "should accept a valid Host");
}
}
// --- Content-Length (RFC 9112 §6.3) ---------------------------------
#[test]
fn parse_conflicting_content_length_is_rejected() {
// Two differing Content-Length values — classic CL.CL smuggling.
let req = b"POST / HTTP/1.1\r\nHost: x\r\nContent-Length: 5\r\nContent-Length: 7\r\n\r\nhello!!";
match parse_head(req, 64) {
Err(ParseError::BadContentLength) => {}
_ => panic!("expected BadContentLength for conflicting CL"),
}
}
#[test]
fn parse_duplicate_equal_content_length_is_rejected() {
// Strict: even identical duplicates are rejected.
let req = b"POST / HTTP/1.1\r\nHost: x\r\nContent-Length: 5\r\nContent-Length: 5\r\n\r\nhello";
match parse_head(req, 64) {
Err(ParseError::BadContentLength) => {}
_ => panic!("expected BadContentLength for duplicate CL"),
}
}
#[test]
fn parse_single_content_length_still_ok() {
// Regression: the ordinary single-CL path is unchanged.
let req = b"POST / HTTP/1.1\r\nHost: x\r\nContent-Length: 5\r\n\r\nhello";
let head = parse_head(req, 64).unwrap();
assert_eq!(head.content_length, Some(5));
}
// --- Transfer-Encoding (RFC 9112 §6.1/§6.3) -------------------------
#[test]
fn parse_non_final_chunked_is_malformed() {
// chunked must be the FINAL coding.
let req = b"POST / HTTP/1.1\r\nHost: x\r\nTransfer-Encoding: chunked, gzip\r\n\r\n";
match parse_head(req, 64) {
Err(ParseError::Malformed) => {}
_ => panic!("expected Malformed for non-final chunked"),
}
}
#[test]
fn parse_unknown_transfer_coding_is_unimplemented() {
// A coding urus doesn't implement, no chunked at all -> 501.
let req = b"POST / HTTP/1.1\r\nHost: x\r\nTransfer-Encoding: nonsense\r\n\r\n";
match parse_head(req, 64) {
Err(ParseError::UnknownTransferCoding) => {}
_ => panic!("expected UnknownTransferCoding for unknown coding"),
}
}
#[test]
fn parse_gzip_then_chunked_is_unimplemented() {
// chunked IS final, but gzip is still a coding we can't apply -> 501.
let req = b"POST / HTTP/1.1\r\nHost: x\r\nTransfer-Encoding: gzip, chunked\r\n\r\n";
match parse_head(req, 64) {
Err(ParseError::UnknownTransferCoding) => {}
_ => panic!("expected UnknownTransferCoding for gzip,chunked"),
}
}
#[test]
fn parse_te_with_content_length_is_malformed() {
// ANY Transfer-Encoding + Content-Length -> reject (smuggling),
// not only chunked+CL. This closes the old TE:unknown + CL gap.
let req = b"POST / HTTP/1.1\r\nHost: x\r\nTransfer-Encoding: bogus\r\nContent-Length: 5\r\n\r\nhello";
match parse_head(req, 64) {
Err(ParseError::Malformed) => {}
_ => panic!("expected Malformed for TE + CL"),
}
}
#[test]
fn serialise_basic_200() {
let conn = Conn::new().put_status(200).put_body("hi");
+123 -55
View File
@@ -38,37 +38,46 @@
//! but stays alive keeps its (one-shot, inert) monitor until death;
//! that's a bounded bookkeeping entry, not a leak.
//!
//! # Construction must happen in-runtime
//! # The handle is an address, not an owner (v0.7)
//!
//! [`PubSub::new`] spawns the table actor, which smarm only permits from
//! inside `Runtime::run`. Pipelines are built *before* `serve*` boots the
//! runtime, so the working pattern (same constraint crud's store hits) is
//! lazy init from the first handler — but with a **non-static**
//! `Arc<OnceLock<PubSub<M>>>` captured by the route closure, NOT a
//! `static`: when the drained pipeline drops at graceful shutdown, the
//! cell (and thus the last handle) drops in-runtime, the table's inbox
//! closes, and the table exits — `serve_with_shutdown` returns. A
//! `static` pins the table forever and blocks smarm's all-done. The same
//! reasoning forbids long-lived consumer actors (relays, producers) from
//! holding a `PubSub` clone: relay holds table's inbox open, table holds
//! relay's receiver open, neither exits. Relays take the `Receiver` only.
//! See `examples/ws_chat.rs` and the `shutdown_with_open_chat_terminates`
//! integration test for the full chain.
//! [`PubSub::new`] takes a **name** and spawns nothing. It is `const`,
//! costs a `&'static str`, and may be built anywhere — out of the
//! runtime, in a `static`, in a route closure, held by a long-lived
//! relay. Every operation resolves the name through smarm's registry, so
//! a table restarted by its supervisor is reached transparently.
//!
//! # Why there is no `register(name)` helper (yet)
//! The table actor is started by [`PubSub::child`], a `ChildSpec` you put
//! in your supervision tree (or hand to `serve_with*`, which puts it in
//! the root ahead of the endpoint). Its lifetime is the supervisor's:
//! it stops when the supervisor shuts it down, in reverse start order,
//! after the endpoint has drained.
//!
//! smarm's registry maps `name → Pid`, but a `Pid` cannot be turned back
//! into a `GenServerRef` (the ref *is* the inbox sender). A useful named
//! lookup therefore needs either smarm support (registry-held senders)
//! or a process-global type-erased map here — both against the grain of
//! the ratified design. Deferred; pass the handle.
//! ```ignore
//! const BUS: PubSub<Event> = PubSub::new("events");
//! serve_with(cfg, rt_cfg, pipeline, vec![BUS.child()])?;
//! ```
//!
//! This replaces the v0.5–v0.6 idiom where `PubSub::new()` spawned the
//! table and the handle owned its life — a non-static
//! `Arc<OnceLock<PubSub<M>>>` lazily initialised from the first handler,
//! with matching rules that the cell must not be `static` and that
//! relays must never hold a clone. All of that existed to hand-manage a
//! refcount. smarm 0.7 made a server's lifetime its own (refs are
//! addresses; the inbox no longer closes when the last ref drops), which
//! both removed the mechanism those rules relied on and made the named
//! lookup below possible.
//!
//! Cost: one registry resolution per operation, including per broadcast.
//! Caching a `GenServerRef` in the handle would save it and go stale
//! across exactly the restart the supervisor exists to perform. Measure
//! before optimising — see the bench item in `ROADMAP.md`.
use std::collections::{HashMap, HashSet};
use std::fmt;
use std::sync::Arc;
use smarm::gen_server::{self, GenServer, GenServerCtx};
use smarm::{channel, Down, Pid, Receiver, Sender, GenServerRef, Watcher};
use smarm::gen_server::{self, GenServer, GenServerBuilder, GenServerCtx, GenServerName};
use smarm::{channel, ChildSpec, Down, GenServerRef, Pid, Receiver, Restart, Sender, Watcher};
// ---------------------------------------------------------------------------
// Public handle
@@ -87,32 +96,58 @@ impl fmt::Display for PubSubDown {
impl std::error::Error for PubSubDown {}
/// A clonable handle to one pub/sub instance (one topic table actor).
/// The address of one pub/sub instance: a name, resolved per operation.
///
/// Cheap (`Copy`, a `&'static str`), constructible anywhere including
/// `const` context, and owns nothing — the table actor it addresses is
/// started by [`child`](Self::child) under a supervisor. Every operation
/// returns [`PubSubDown`] if no live table currently holds the name.
///
/// Payloads are broadcast as `Arc<M>`: one allocation per broadcast, not
/// per subscriber.
pub struct PubSub<M: Send + Sync + 'static> {
server: GenServerRef<Table<M>>,
name: GenServerName<Table<M>>,
}
impl<M: Send + Sync + 'static> Clone for PubSub<M> {
fn clone(&self) -> Self {
PubSub { server: self.server.clone() }
*self
}
}
impl<M: Send + Sync + 'static> Default for PubSub<M> {
fn default() -> Self {
Self::new()
}
}
impl<M: Send + Sync + 'static> Copy for PubSub<M> {}
impl<M: Send + Sync + 'static> PubSub<M> {
/// Start a fresh topic table. Must run inside the smarm runtime (it
/// spawns the table's gen_server actor). The table lives until the
/// last handle is dropped.
pub fn new() -> Self {
PubSub { server: gen_server::start(Table::new()) }
/// Address the topic table registered under `name`. Spawns nothing
/// and never fails: the name is resolved at each use.
pub const fn new(name: &'static str) -> Self {
PubSub { name: GenServerName::new(name) }
}
/// The `ChildSpec` that runs this instance's table actor. Put it in
/// your supervision tree ahead of anything that broadcasts —
/// `serve_with*` takes a `Vec<ChildSpec>` for exactly this.
///
/// `Permanent`: a table that dies is a bug, and its subscribers'
/// receivers died with it, so the restart is only half a repair —
/// pair it with a `RestForOne` parent (as `serve_with*` does) so the
/// endpoint restarts behind it and connections re-subscribe.
pub fn child(&self) -> ChildSpec {
let name = self.name;
ChildSpec::new(Restart::Permanent, move || {
if let Err(e) = GenServerBuilder::new(Table::<M>::new()).named(name).run() {
panic!("urus pubsub '{}': name already taken: {e:?}", name.as_str());
}
})
}
/// The registry key this handle resolves.
pub const fn name(&self) -> &'static str {
self.name.as_str()
}
fn server(&self) -> Result<GenServerRef<Table<M>>, PubSubDown> {
gen_server::whereis_server(self.name).ok_or(PubSubDown)
}
/// Subscribe the **calling actor** to `topic`. Returns the receiving
@@ -139,7 +174,7 @@ impl<M: Send + Sync + 'static> PubSub<M> {
let (tx, rx) = channel();
// A call, not a cast: when this returns the table is updated, so
// a broadcast issued right after by the same caller is seen.
match self.server.call(Call::Subscribe { topic: topic.into(), pid, tx }) {
match self.server()?.call(Call::Subscribe { topic: topic.into(), pid, tx }) {
Ok(_) => Ok(rx),
Err(_) => Err(PubSubDown),
}
@@ -153,7 +188,7 @@ impl<M: Send + Sync + 'static> PubSub<M> {
/// [`unsubscribe`](Self::unsubscribe) for an explicit pid.
pub fn unsubscribe_as(&self, pid: Pid, topic: impl Into<String>) -> Result<(), PubSubDown> {
self.server
self.server()?
.cast(Cast::Unsubscribe { topic: topic.into(), pid })
.map_err(|_| PubSubDown)
}
@@ -181,20 +216,21 @@ impl<M: Send + Sync + 'static> PubSub<M> {
/// table — a subscriber whose receiver was dropped but hasn't been
/// pruned yet (no broadcast since, still alive) is still counted.
pub fn subscriber_count(&self, topic: impl Into<String>) -> Result<usize, PubSubDown> {
match self.server.call(Call::Count { topic: topic.into() }) {
match self.server()?.call(Call::Count { topic: topic.into() }) {
Ok(Reply::Count(n)) => Ok(n),
Ok(Reply::Subscribed) => unreachable!("Count call answered with Subscribed"),
Err(_) => Err(PubSubDown),
}
}
/// The topic table actor's pid — for introspection / registry use.
pub fn pid(&self) -> Pid {
self.server.pid()
/// The topic table actor's pid, or `None` if no live table holds the
/// name (not started yet, or between a crash and its restart).
pub fn pid(&self) -> Option<Pid> {
self.server().ok().map(|s| s.pid())
}
fn cast_broadcast(&self, topic: String, msg: M, skip: Option<Pid>) -> Result<(), PubSubDown> {
self.server
self.server()?
.cast(Cast::Broadcast { topic, msg: Arc::new(msg), skip })
.map_err(|_| PubSubDown)
}
@@ -328,15 +364,33 @@ mod tests {
false
}
/// Start `ps`'s table under a supervisor and hand back the sup's pid.
///
/// The readiness poll is smarm's "start order is not start readiness"
/// gap (see smarm ROADMAP): `start_child` spawns and moves on, so the
/// name may not be bound when this returns. Real apps don't hit it —
/// a handler only runs once a connection has been accepted, long
/// after the tree is up — but a test that broadcasts immediately does.
fn start_table<M: Send + Sync + 'static>(ps: PubSub<M>) -> Pid {
let sup = smarm::spawn(move || smarm::OneForOne::new().child(ps.child()).run());
assert!(eventually(|| ps.pid().is_some()), "table never bound its name");
sup.pid()
}
/// Ordered shutdown of the tree `start_table` built.
fn stop_table(sup: Pid) {
smarm::request_shutdown(sup);
}
#[test]
fn broadcast_reaches_all_subscribers_once() {
let got = Arc::new(Mutex::new(Vec::<(u32, String)>::new()));
let got2 = got.clone();
smarm::run(move || {
let ps = PubSub::<String>::new();
let ps = PubSub::<String>::new("test-bus");
let sup = start_table(ps);
let mut handles = Vec::new();
for i in 0..2u32 {
let ps = ps.clone();
let got = got2.clone();
handles.push(smarm::spawn(move || {
let rx = ps.subscribe("room:a").unwrap();
@@ -351,6 +405,7 @@ mod tests {
for h in handles {
h.join().unwrap();
}
stop_table(sup);
});
let mut v = got.lock().unwrap().clone();
v.sort();
@@ -362,10 +417,10 @@ mod tests {
let ptrs = Arc::new(Mutex::new(Vec::<usize>::new()));
let ptrs2 = ptrs.clone();
smarm::run(move || {
let ps = PubSub::<Vec<u8>>::new();
let ps = PubSub::<Vec<u8>>::new("test-bus");
let sup = start_table(ps);
let mut handles = Vec::new();
for _ in 0..2 {
let ps = ps.clone();
let ptrs = ptrs2.clone();
handles.push(smarm::spawn(move || {
let rx = ps.subscribe("t").unwrap();
@@ -378,6 +433,7 @@ mod tests {
for h in handles {
h.join().unwrap();
}
stop_table(sup);
});
let v = ptrs.lock().unwrap();
assert_eq!(v.len(), 2);
@@ -389,8 +445,9 @@ mod tests {
let got = Arc::new(Mutex::new(Vec::<String>::new()));
let got2 = got.clone();
smarm::run(move || {
let ps = PubSub::<String>::new();
let ps_loud = ps.clone();
let ps = PubSub::<String>::new("test-bus");
let sup = start_table(ps);
let ps_loud = ps;
let got = got2.clone();
let loud = smarm::spawn(move || {
let rx = ps_loud.subscribe("room").unwrap();
@@ -410,6 +467,7 @@ mod tests {
.unwrap()
.push(format!("root got {} then {}", *first, *second));
loud.join().unwrap();
stop_table(sup);
});
let v = got.lock().unwrap();
assert!(v.contains(&"loud got for everyone".to_string()), "{v:?}");
@@ -426,7 +484,8 @@ mod tests {
let ok = Arc::new(Mutex::new(false));
let ok2 = ok.clone();
smarm::run(move || {
let ps = PubSub::<u32>::new();
let ps = PubSub::<u32>::new("test-bus");
let sup = start_table(ps);
let rx = ps.subscribe("t").unwrap();
ps.unsubscribe("t").unwrap();
// Unsubscribe is a cast; the table holds the only sender, so
@@ -435,6 +494,7 @@ mod tests {
assert!(rx.recv().is_err());
ps.broadcast("t", 7).unwrap(); // no subscribers: no-op, no panic
*ok2.lock().unwrap() = true;
stop_table(sup);
});
assert!(*ok.lock().unwrap());
}
@@ -444,7 +504,8 @@ mod tests {
let got = Arc::new(Mutex::new((0u32, false)));
let got2 = got.clone();
smarm::run(move || {
let ps = PubSub::<u32>::new();
let ps = PubSub::<u32>::new("test-bus");
let sup = start_table(ps);
let rx_old = ps.subscribe("t").unwrap();
let rx_new = ps.subscribe("t").unwrap();
assert_eq!(ps.subscriber_count("t").unwrap(), 1, "idempotent per (pid, topic)");
@@ -452,6 +513,7 @@ mod tests {
let v = *rx_new.recv().unwrap();
let old_closed = rx_old.recv().is_err();
*got2.lock().unwrap() = (v, old_closed);
stop_table(sup);
});
assert_eq!(*got.lock().unwrap(), (42, true));
}
@@ -461,8 +523,9 @@ mod tests {
let ok = Arc::new(Mutex::new(false));
let ok2 = ok.clone();
smarm::run(move || {
let ps = PubSub::<u32>::new();
let ps2 = ps.clone();
let ps = PubSub::<u32>::new("test-bus");
let sup = start_table(ps);
let ps2 = ps;
let h = smarm::spawn(move || {
let _rx = ps2.subscribe("t").unwrap();
// Exit without ever receiving: rx drops with the stack.
@@ -471,6 +534,7 @@ mod tests {
// No broadcast issued — this MUST be the monitor path, not
// prune-on-send-failure.
*ok2.lock().unwrap() = eventually(|| ps.subscriber_count("t").unwrap() == 0);
stop_table(sup);
});
assert!(*ok.lock().unwrap(), "monitor Down never pruned the dead subscriber");
}
@@ -480,7 +544,8 @@ mod tests {
let counts = Arc::new(Mutex::new((0usize, 0usize)));
let counts2 = counts.clone();
smarm::run(move || {
let ps = PubSub::<u32>::new();
let ps = PubSub::<u32>::new("test-bus");
let sup = start_table(ps);
let rx = ps.subscribe("t").unwrap();
drop(rx);
let before = ps.subscriber_count("t").unwrap();
@@ -493,6 +558,7 @@ mod tests {
after = if eventually(|| ps.subscriber_count("t").unwrap() == 0) { 0 } else { after };
}
*counts2.lock().unwrap() = (before, after);
stop_table(sup);
});
assert_eq!(*counts.lock().unwrap(), (1, 0));
}
@@ -502,11 +568,13 @@ mod tests {
let got = Arc::new(Mutex::new(Vec::<u32>::new()));
let got2 = got.clone();
smarm::run(move || {
let ps = PubSub::<u32>::new();
let ps = PubSub::<u32>::new("test-bus");
let sup = start_table(ps);
let rx_a = ps.subscribe("a").unwrap();
ps.broadcast("b", 99).unwrap(); // nobody on b; must not reach a
ps.broadcast("a", 1).unwrap();
got2.lock().unwrap().push(*rx_a.recv().unwrap());
stop_table(sup);
});
assert_eq!(*got.lock().unwrap(), vec![1]);
}
+245 -290
View File
@@ -1,29 +1,19 @@
//! Listener pool and the `serve` entry point.
//! [`Config`] and the `serve*` entry points.
//!
//! A small fixed pool of listener actors share the same TCP listen fd (via
//! `dup`) — each blocks in non-blocking `accept4` + `wait_readable` on its
//! own copy. When a connection arrives the listener spawns a connection
//! actor with the `OwnedFd` and immediately returns to `accept`. No
//! coordination needed; the kernel serialises `accept` calls across the fds.
//! The listener pool itself lives in [`crate::endpoint`], which is the
//! real API: an endpoint is a supervisable child you place in your own
//! tree. `serve*` is the batteries-included path for a process whose only
//! job is serving HTTP — it owns the runtime and wraps one endpoint in a
//! one-child supervisor.
//!
//! Sharing via `dup` rather than the same fd is deliberate — Linux's
//! `accept4` is thread-safe on a single fd, but dup'ing per-listener keeps
//! each actor's epoll registration local to its own RawFd value (so smarm's
//! `waiters: HashMap<RawFd, Pid>` doesn't see collisions between listeners
//! waiting on "the same fd").
use crate::conn_actor::{run_connection, ConnLimits};
use crate::conn_registry::{self, Call, Cast, ConnRegistry, Reply};
use crate::net::{accept_nonblocking, bind_and_listen, OwnedFd};
use crate::conn_actor::ConnLimits;
use crate::plug::Pipeline;
use smarm::{ChildSpec, OneForOne, Restart, GenServerRef, Strategy};
use smarm::supervisor::Shutdown;
use smarm::{ChildSpec, OneForOne, Restart, Strategy};
use std::io::{self, ErrorKind};
use std::net::{SocketAddr, ToSocketAddrs};
use std::os::fd::RawFd;
use std::sync::atomic::{AtomicBool, AtomicU32, Ordering};
use std::sync::Arc;
use std::time::Duration;
// ---------------------------------------------------------------------------
@@ -37,7 +27,21 @@ pub struct Config {
pub keep_alive_timeout: Duration,
pub max_header_count: usize,
pub read_buf_size: usize,
pub request_timeout: Duration,
/// Wall-clock budget for reading the request HEAD (from first byte to
/// full head parse). Kept short — an incomplete head is the classic
/// slowloris. See `ConnLimits::head_timeout`.
pub head_timeout: Duration,
/// Absolute wall-clock cap on reading the request BODY (from head-parse
/// to full body). Sized for slow links, so much larger than
/// `head_timeout`. See `ConnLimits::body_timeout`.
pub body_timeout: Duration,
/// Burst size that resets the body stall clock. A body dribbling fewer
/// than this per `body_stall_timeout` window is evicted — the slowloris
/// / slow-legit discriminator. See `ConnLimits::body_burst_bytes`.
pub body_burst_bytes: usize,
/// Max time since the last qualifying body burst before eviction;
/// backstopped by `body_timeout`. See `ConnLimits::body_stall_timeout`.
pub body_stall_timeout: Duration,
/// Per-write budget for response bytes (the fixed head+body write, and
/// each streamed chunk). See `ConnLimits::write_timeout`.
pub write_timeout: Duration,
@@ -53,12 +57,24 @@ pub struct Config {
/// WebSocket: cap on a complete reassembled message (spans
/// fragments; violation closes 1009).
pub max_message_bytes: usize,
/// Number of smarm scheduler OS threads. `None` means smarm's default
/// (one per CPU). Set this to a small fixed number in tests so multiple
/// concurrent test servers don't oversubscribe the host.
pub scheduler_threads: Option<usize>,
/// Registry name for this endpoint's gen_server — how it is addressed
/// from elsewhere in the app ([`crate::endpoint::whereis`]), and what
/// must be unique between two endpoints in one process (a public and
/// an admin port, say). Default `"urus"`.
pub name: &'static str,
/// Stack reserve (RFC 019 `smarm::SpawnOpts::stack_reserve`) given to
/// each per-connection actor. Request handlers routinely pull in
/// application code — DB drivers, (de)compression, templating — whose
/// stack needs comfortably exceed smarm's bare-actor default of 64 KiB
/// (the exact shape of bug this exists to head off; see smarm RFC 019).
/// Default: 256 KiB. The reserve is virtual/demand-paged, so raising it
/// costs address space, not RSS, until a handler actually uses it.
pub conn_stack_reserve: usize,
}
/// Default per-connection actor stack reserve (see [`Config::conn_stack_reserve`]).
pub const DEFAULT_CONN_STACK_RESERVE: usize = 256 * 1024;
impl Config {
pub fn new(addr: SocketAddr) -> Self {
let pool = std::thread::available_parallelism()
@@ -71,24 +87,31 @@ impl Config {
keep_alive_timeout: Duration::from_secs(60),
max_header_count: 64,
read_buf_size: 8 * 1024,
request_timeout: Duration::from_secs(30),
head_timeout: Duration::from_secs(30),
body_timeout: Duration::from_secs(300),
body_burst_bytes: 4 * 1024,
body_stall_timeout: Duration::from_secs(20),
write_timeout: Duration::from_secs(30),
max_body_bytes: 16 * 1024 * 1024,
drain_timeout: Duration::from_secs(30),
max_frame_payload: 1024 * 1024,
max_message_bytes: 4 * 1024 * 1024,
scheduler_threads: None,
name: "urus",
conn_stack_reserve: DEFAULT_CONN_STACK_RESERVE,
}
}
fn to_conn_limits(&self) -> ConnLimits {
pub(crate) fn to_conn_limits(&self) -> ConnLimits {
ConnLimits {
max_headers: self.max_header_count,
initial_read_buf: self.read_buf_size,
max_head_bytes: 64 * 1024,
max_body_bytes: self.max_body_bytes,
keep_alive_timeout: self.keep_alive_timeout,
request_timeout: self.request_timeout,
head_timeout: self.head_timeout,
body_timeout: self.body_timeout,
body_burst_bytes: self.body_burst_bytes,
body_stall_timeout: self.body_stall_timeout,
write_timeout: self.write_timeout,
max_frame_payload: self.max_frame_payload,
max_message_bytes: self.max_message_bytes,
@@ -97,124 +120,87 @@ impl Config {
}
// ---------------------------------------------------------------------------
// dup helper
// config-file: TOML overlay for tuning knobs
// ---------------------------------------------------------------------------
//
// urus is a library, so it never presumes a config-file path or reads the
// environment — the embedding binary decides where a file lives and hands
// the text here. This overlays a sparse TOML document onto an existing
// `Config` (built with an addr the binary chose): only the keys present are
// applied, everything else keeps the compiled default. Durations are
// integer seconds. Unknown keys are a hard error so a typo is loud, not a
// silent no-op.
//
// Scope for now: the slowloris-tuning knobs only. Migrating the rest of the
// Config surface into the file is a separate, additive job (the loader
// mechanism is general — it just extends `TomlOverrides`).
fn dup_fd(fd: RawFd) -> io::Result<OwnedFd> {
let new_fd = unsafe { libc::fcntl(fd, libc::F_DUPFD_CLOEXEC, 0) };
if new_fd < 0 {
return Err(io::Error::last_os_error());
/// Error from [`Config::with_toml_str`]: the TOML failed to parse or carried
/// an unknown/mistyped key.
#[cfg(feature = "config-file")]
#[derive(Debug)]
pub enum ConfigError {
Toml(String),
}
#[cfg(feature = "config-file")]
impl std::fmt::Display for ConfigError {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
match self {
ConfigError::Toml(m) => write!(f, "config TOML error: {m}"),
}
}
}
#[cfg(feature = "config-file")]
impl std::error::Error for ConfigError {}
#[cfg(feature = "config-file")]
#[derive(serde::Deserialize)]
#[serde(deny_unknown_fields)]
struct TomlOverrides {
head_timeout_secs: Option<u64>,
body_timeout_secs: Option<u64>,
body_burst_bytes: Option<usize>,
body_stall_timeout_secs: Option<u64>,
}
#[cfg(feature = "config-file")]
impl Config {
/// Overlay a TOML document of tuning knobs onto this config (sparse:
/// only the keys present are applied). Durations are integer seconds.
///
/// Recognized keys: `head_timeout_secs`, `body_timeout_secs`,
/// `body_burst_bytes`, `body_stall_timeout_secs`. Unknown keys error.
/// Other `Config` knobs are not yet file-configurable.
pub fn with_toml_str(mut self, s: &str) -> Result<Self, ConfigError> {
let o: TomlOverrides =
toml::from_str(s).map_err(|e| ConfigError::Toml(e.to_string()))?;
if let Some(v) = o.head_timeout_secs {
self.head_timeout = Duration::from_secs(v);
}
if let Some(v) = o.body_timeout_secs {
self.body_timeout = Duration::from_secs(v);
}
if let Some(v) = o.body_burst_bytes {
self.body_burst_bytes = v;
}
if let Some(v) = o.body_stall_timeout_secs {
self.body_stall_timeout = Duration::from_secs(v);
}
Ok(self)
}
Ok(OwnedFd::from_raw(new_fd))
}
// ---------------------------------------------------------------------------
// listener actor body
// Handle / ShutdownSignal — graceful shutdown plumbing for the serve* entries.
// ---------------------------------------------------------------------------
/// Test-only fault injection. When nonzero, the next accept-loop iteration
/// of whichever listener gets there first decrements this and panics —
/// *before* calling `accept`, so a pending connection stays in the kernel
/// backlog and must be picked up by the restarted listener. Cost when idle
/// is one relaxed load per accept-loop iteration (each of which already
/// pays a syscall). Not public API.
#[doc(hidden)]
pub static INJECT_LISTENER_PANICS: AtomicU32 = AtomicU32::new(0);
fn listener_loop(
listener: Arc<OwnedFd>,
pipeline: Pipeline,
limits: ConnLimits,
registry: GenServerRef<ConnRegistry>,
shutdown: Arc<AtomicBool>,
) {
let fd = listener.as_raw();
loop {
// Self-termination on shutdown — the ONLY way a listener exits at
// shutdown, and deliberately a normal return: under
// `Restart::Transient` a normal exit is terminal, so the
// supervisor's active-count drains and `sup.run()` returns on its
// own. No pid is ever `request_stop`ped, which sidesteps both the
// spawn-to-registration race of self-announced pids and smarm's
// lossy stop-while-QUEUED window (a flag is wake-free and
// race-free; a freshly spawned listener observes it on its very
// first iteration, a parked one within LISTENER_TICK).
if shutdown.load(Ordering::Relaxed) {
return;
}
if INJECT_LISTENER_PANICS
.fetch_update(Ordering::Relaxed, Ordering::Relaxed, |n| n.checked_sub(1))
.is_ok()
{
panic!("urus: injected listener panic (test hook)");
}
match accept_nonblocking(fd) {
Ok(client) => {
// Hand the fd off to a new connection actor. spawn() is
// cheap on smarm — it's a single Vec push under the
// shared lock.
let p = pipeline.clone();
let l = limits;
let r = registry.clone();
smarm::spawn(move || run_connection(client, p, l, r));
}
Err(e) if e.kind() == ErrorKind::WouldBlock => {
// No pending connection. Park until the listener is
// readable or the tick elapses; either way we come back
// around through the shutdown-flag check above.
match smarm::wait_readable_timeout(fd, LISTENER_TICK) {
Ok(_ready) => {} // ready or tick — loop re-checks, retries accept
Err(we) => {
// epoll registration failed. Under Transient
// supervision a normal return is terminal but a
// panic restarts us — and a failed wait IS
// abnormal, so panic: a transient failure (e.g.
// EMFILE on the epoll set) heals by restart
// instead of silently shrinking the pool. (smarm
// catches actor panics in the trampoline; this is
// a Signal::Panic to the supervisor, not process
// noise.)
panic!("urus: listener wait_readable failed: {we}");
}
}
}
Err(e) if e.kind() == ErrorKind::Interrupted => {
continue;
}
Err(e) => {
// EMFILE / ENFILE / ECONNABORTED etc. Log and continue;
// the system may recover.
eprintln!("urus: accept error: {e}");
// Small backoff via smarm's sleep to avoid spinning if
// the error is sticky.
smarm::sleep(Duration::from_millis(10));
}
}
}
// The Arc clone we were started with drops here (normal exit or
// unwind), but the ChildSpec factory holds another — the fd outlives
// any one incarnation of this listener.
}
// ---------------------------------------------------------------------------
// Handle / ShutdownSignal — graceful shutdown plumbing.
// ---------------------------------------------------------------------------
/// How often the root actor polls for a shutdown signal (see
/// `serve_with_shutdown` for why this is a poll); bounds shutdown latency.
const SHUTDOWN_POLL: Duration = Duration::from_millis(100);
/// Listener accept-waits are timed at this tick; the shutdown flag is
/// observed at the top of every accept-loop iteration, so this bounds how
/// long a fully idle listener takes to notice shutdown. It is the SOLE
/// listener-shutdown mechanism (no `request_stop` — see the shutdown
/// sequence notes). Idle cost: one timer wake per listener per tick.
const LISTENER_TICK: Duration = Duration::from_millis(250);
/// A clonable trigger for graceful shutdown. Safe to use from any OS
/// thread (the send only enqueues; the serving side polls), e.g. from a
/// signal-handling thread.
/// A clonable trigger for graceful shutdown, usable from any OS thread
/// (e.g. a signal-handling thread) — the shutdown path for callers who let
/// `serve*` own the runtime and so have no [`smarm::RuntimeHandle`] of
/// their own. If you build the tree yourself with [`crate::endpoint`], use
/// `rt.handle().request_shutdown(root_sup)` instead and ignore this.
#[derive(Clone)]
pub struct Handle {
tx: smarm::Sender<()>,
@@ -223,8 +209,8 @@ pub struct Handle {
impl Handle {
/// Begin graceful shutdown: stop accepting, close idle keep-alive
/// connections, drain in-flight requests up to `Config.drain_timeout`,
/// then force-stop stragglers. `serve_with_shutdown` returns once the
/// runtime has wound down. Idempotent; extra calls are no-ops.
/// then force-stop stragglers. `serve*` returns once the runtime has
/// wound down. Idempotent; extra calls are no-ops.
pub fn shutdown(&self) {
let _ = self.tx.send(());
}
@@ -242,183 +228,152 @@ pub fn shutdown_handle() -> (Handle, ShutdownSignal) {
}
// ---------------------------------------------------------------------------
// serve_with_shutdown — main entry. Boots smarm, supervises listeners,
// blocks until shutdown.
// serve* — batteries-included entries for apps whose only job is serving.
// ---------------------------------------------------------------------------
//
// Boots an smarm runtime (one OS thread per CPU by default — see smarm's
// `Config::default()`). The root actor starts the connection registry and
// a one-for-one supervisor over the listener pool, then blocks on the
// shutdown signal. A panicking listener is restarted on the same
// (still-open) fd instead of silently shrinking the accept pool.
// Connection actors stay unsupervised bare spawns — per spec, a connection
// is cheap and its failure is local: a 500 path, not a restart path.
// (`spawn` from a listener parents the conn actor under that listener,
// which never registers a supervisor channel, so conn deaths are invisible
// to the pool supervisor by construction.)
//
// Shutdown sequence (drain-then-stop, the v0.2 chunk 2 decision):
// 1. Set the shared shutdown flag. Every listener observes it at the
// top of its accept loop (parks are timed at LISTENER_TICK) and
// returns normally. Under `Restart::Transient` a normal exit is
// terminal, so no restart happens and the supervisor's active count
// drains to zero.
// 2. Join the supervisor: `sup.run()` returns on its own once every
// listener has exited. The join is therefore the barrier "no new
// connections can ever be accepted" — listener fds are closed (the
// last Arc clones drop with the supervisor's ChildSpecs), and the
// kernel refuses new connects.
//
// Why a flag and not `request_stop`: stopping pids that announce
// themselves races the spawn-to-registration gap (a fast shutdown
// CAN beat a fresh listener to the registry), and smarm's
// `request_stop` is lossy against a QUEUED actor that then parks
// without passing an observation point — found the hard way; see
// the chunk 2 commit message. The flag is wake-free and race-free.
// 3. BeginDrain: the registry stops idle conns now and each remaining
// conn the moment it finishes its in-flight request.
// 4. Poll until no conns remain or the drain deadline passes; past the
// deadline, ForceStopConns every tick until the set empties (a conn
// accepted just before listener death may register late).
// 5. Root returns. `rt.run` itself returns only when every actor has
// exited (force-stopped conns unwind through their fd waits safely —
// smarm's 06-10 io fix — and close their sockets via OwnedFd::drop).
// These own the smarm runtime and build a one-child tree around
// `endpoint()`. An app with its own actors should call `endpoint()`
// directly and put it in its own supervision tree — that is the real API;
// everything here is a convenience wrapper over it.
/// Boot a runtime, serve until `signal` fires, then drain and return.
///
/// The tree is `root sup -> [..app_children, endpoint]` on
/// [`Strategy::RestForOne`], with the endpoint on [`Shutdown::Infinity`]
/// so its `drain_timeout` — not a supervisor deadline — bounds the drain.
///
/// `app_children` is your application's actors: a
/// [`PubSub::child`](crate::PubSub::child), a
/// [`ChannelHub::children`](crate::channels::ChannelHub::children), your
/// own state servers. They start **before** the endpoint and, because
/// shutdown is ordered in reverse, stop **after** it has drained — so a
/// request still in flight can still reach the bus. `RestForOne` in that
/// order also means one of them crashing restarts the endpoint behind it,
/// dropping connections whose subscriptions died with it, rather than
/// leaving live sockets addressing a table that no longer knows them.
///
/// The root actor parks on the signal channel; a `Handle::shutdown` from a
/// foreign OS thread wakes it, it shuts the supervisor down and `rt.run`
/// returns when the last actor is gone.
pub fn serve_with_shutdown(
config: Config,
rt_config: smarm::Config,
pipeline: Pipeline,
app_children: Vec<ChildSpec>,
signal: ShutdownSignal,
) -> io::Result<()> {
let listener = bind_and_listen(config.addr)?;
println!("urus: listening on {}", config.addr);
// One connection-actor-spawning loop per listener pool slot. Each gets
// its own dup'd fd so epoll registrations don't collide. Each fd is
// owned by its ChildSpec's factory closure (via `Arc`): a restarted
// listener re-enters `accept`/`wait_readable` on the same fd — no
// re-dup, no window where the slot has no fd.
let mut listener_fds = Vec::with_capacity(config.listener_pool);
listener_fds.push(Arc::new(listener)); // primary keeps the original
for _ in 1..config.listener_pool {
let dup = dup_fd(listener_fds[0].as_raw())?;
listener_fds.push(Arc::new(dup));
}
let limits = config.to_conn_limits();
let drain_timeout = config.drain_timeout;
let smarm_cfg = match config.scheduler_threads {
Some(n) => smarm::Config::exact(n),
None => smarm::Config::default(),
};
let rt = smarm::init(smarm_cfg);
// Listener self-termination flag — see the shutdown sequence below.
let shutdown_flag = Arc::new(AtomicBool::new(false));
let addr = config.addr;
let endpoint = crate::endpoint::endpoint(config, pipeline)?;
println!("urus: listening on {addr}");
let rt = smarm::init(rt_config);
rt.run(move || {
// Registry first: listeners and conns cast into it from birth.
let registry = conn_registry::start();
let mut sup = OneForOne::new().strategy(Strategy::OneForOne);
for (i, lfd) in listener_fds.into_iter().enumerate() {
let p = pipeline.clone();
let r = registry.clone();
let sf = shutdown_flag.clone();
sup = sup.child(ChildSpec::new(Restart::Transient, move || {
println!("urus: listener {} starting", i);
listener_loop(lfd.clone(), p.clone(), limits, r.clone(), sf.clone());
}));
let sup = smarm::spawn(move || {
let mut sup = OneForOne::new().strategy(Strategy::RestForOne);
for child in app_children {
sup = sup.child(child);
}
// Default intensity (3 per 5s) applies; a listener crash-looping
// faster than that trips the cap and tears the pool down — loud
// failure over a zombie server.
let sup_h = smarm::spawn(move || sup.run());
// The old `urus.server` / `urus.listener.{i}` name registrations
// are gone with smarm's RFC 014 registry rework: `register` is now
// `(Name<M>, Sender<M>)`, self-only — a name is a typed messaging
// endpoint, not a pid tag. urus's bindings were introspection-only
// with no channel behind them, so they were dropped rather than
// faked with a unit channel. A real messageable `urus.server`
// name is in the icebox (ROADMAP.md).
sup.child(ChildSpec::new(Restart::Permanent, endpoint).shutdown(Shutdown::Infinity))
.run()
});
// Block until told to shut down. We poll `try_recv` + `sleep`
// rather than parking in `recv`: a smarm `Sender::send` from a
// foreign OS thread (no runtime in its TLS) enqueues fine but its
// unpark is a `try_with_runtime` no-op — a parked receiver would
// never wake. The timer wake comes from inside the runtime, so the
// poll sees the message within one interval. (A cross-thread-safe
// unpark is a smarm roadmap candidate; this poll dies with it.)
// If every Handle was dropped the channel closes and no shutdown
// can ever arrive: serve forever, exactly v1's semantics.
// Park until told to shut down. If every Handle was dropped the
// channel closes and no shutdown can ever arrive: serve forever,
// exactly v1's semantics.
match signal.rx.recv() {
Ok(()) => smarm::request_shutdown(sup.pid()),
Err(_) => {
// Sender side gone. Park indefinitely; the process is
// expected to be killed externally.
loop {
match signal.rx.try_recv() {
Ok(Some(())) => break,
Ok(None) => smarm::sleep(SHUTDOWN_POLL),
Err(_) => smarm::sleep(Duration::from_secs(3600)),
smarm::sleep(Duration::from_secs(3600));
}
}
// ----- Shutdown. -----
// 1 + 2. Flag the listeners down and join the supervisor; the
// join returns once every listener has exited normally
// (Transient: normal exit is terminal). After this point
// no connection can ever be accepted again.
shutdown_flag.store(true, Ordering::Relaxed);
let _ = sup_h.join();
// 3 + 4. Drain. Same sweep discipline as listeners on the force-
// stop path: a conn accepted just before its listener died may
// register after the deadline, so keep force-stopping until the
// set is empty (each pass kills everything registered; new
// registrants are a strictly shrinking population once listeners
// are gone).
let _ = registry.cast(Cast::BeginDrain);
let deadline = std::time::Instant::now() + drain_timeout;
let mut force = false;
loop {
match registry.call(Call::ConnCount) {
Ok(Reply::ConnCount(0)) => break,
Ok(_) => {}
Err(_) => break, // registry gone; nothing left to track
}
let now = std::time::Instant::now();
if force || now >= deadline {
force = true;
let _ = registry.cast(Cast::ForceStopConns);
smarm::sleep(Duration::from_millis(10));
} else {
smarm::sleep(Duration::from_millis(50).min(deadline - now));
}
}
// 5. Our GenServerRef drops here. The registry's inbox closes once
// the last conn's clone drops with it, and the runtime winds
// down when the last actor exits.
let _ = sup.join();
});
Ok(())
}
// ---------------------------------------------------------------------------
// serve_with / serve — convenience entries without a shutdown handle.
// ---------------------------------------------------------------------------
pub fn serve_with(config: Config, pipeline: Pipeline) -> io::Result<()> {
// The Handle is dropped immediately: shutdown can never be signalled
// and the server runs until externally killed (v1 semantics).
/// [`serve_with_shutdown`] without a shutdown handle: serves until the
/// process is killed.
pub fn serve_with(
config: Config,
rt_config: smarm::Config,
pipeline: Pipeline,
app_children: Vec<ChildSpec>,
) -> io::Result<()> {
// The Handle is dropped immediately: shutdown can never be signalled.
let (_handle, signal) = shutdown_handle();
serve_with_shutdown(config, pipeline, signal)
serve_with_shutdown(config, rt_config, pipeline, app_children, signal)
}
// ---------------------------------------------------------------------------
// serve — convenience over serve_with.
// ---------------------------------------------------------------------------
/// Defaults all round: default [`Config`], default smarm runtime (one
/// scheduler thread per CPU), no app children, serve until killed.
pub fn serve(addr: impl ToSocketAddrs, pipeline: Pipeline) -> io::Result<()> {
let addr = addr
.to_socket_addrs()?
.next()
.ok_or_else(|| io::Error::new(ErrorKind::InvalidInput, "no addresses resolved"))?;
serve_with(Config::new(addr), pipeline)
serve_with(Config::new(addr), smarm::Config::default(), pipeline, Vec::new())
}
#[cfg(all(test, feature = "config-file"))]
mod config_file_tests {
use super::*;
fn base() -> Config {
Config::new("127.0.0.1:0".parse().unwrap())
}
#[test]
fn toml_empty_keeps_defaults() {
let d = base();
let c = base().with_toml_str("").unwrap();
assert_eq!(c.head_timeout, d.head_timeout);
assert_eq!(c.body_timeout, d.body_timeout);
assert_eq!(c.body_burst_bytes, d.body_burst_bytes);
assert_eq!(c.body_stall_timeout, d.body_stall_timeout);
}
#[test]
fn toml_partial_overrides_only_named() {
let d = base();
let c = base().with_toml_str("head_timeout_secs = 5").unwrap();
assert_eq!(c.head_timeout, Duration::from_secs(5)); // overridden
assert_eq!(c.body_timeout, d.body_timeout); // default kept
assert_eq!(c.body_burst_bytes, d.body_burst_bytes); // default kept
assert_eq!(c.body_stall_timeout, d.body_stall_timeout);
}
#[test]
fn toml_full_overrides_all() {
let c = base()
.with_toml_str(
"head_timeout_secs = 10\n\
body_timeout_secs = 120\n\
body_burst_bytes = 8192\n\
body_stall_timeout_secs = 15\n",
)
.unwrap();
assert_eq!(c.head_timeout, Duration::from_secs(10));
assert_eq!(c.body_timeout, Duration::from_secs(120));
assert_eq!(c.body_burst_bytes, 8192);
assert_eq!(c.body_stall_timeout, Duration::from_secs(15));
}
#[test]
fn toml_unknown_key_errors() {
// A mistyped/unknown key is a hard error, not a silent no-op.
let e = base().with_toml_str("body_timeout_sec = 120"); // typo: missing 's'
assert!(e.is_err(), "unknown key should error");
}
#[test]
fn toml_malformed_errors() {
let e = base().with_toml_str("this is not = valid = toml");
assert!(e.is_err(), "malformed TOML should error");
}
}
+2 -1
View File
@@ -33,7 +33,8 @@
//! write — event or heartbeat — stalls past `write_timeout`; the conn
//! actor then drops the stream and the producer's next [`EventSender`]
//! call returns `Err(SseClosed)`. There is no request clock on an SSE
//! response: `request_timeout` covers only the read phase, by design.
//! response: the head/body read budgets cover only the read phase, by
//! design.
use crate::conn::{Conn, RespBody, StreamBody};
+366 -72
View File
@@ -26,15 +26,19 @@ fn free_port() -> u16 {
}
fn spawn_server(pipeline: Pipeline) -> u16 {
spawn_server_children(pipeline, Vec::new())
}
/// [`spawn_server`] with app children started ahead of the endpoint.
fn spawn_server_children(pipeline: Pipeline, app_children: Vec<smarm::ChildSpec>) -> u16 {
let port = free_port();
let addr: SocketAddr = format!("127.0.0.1:{port}").parse().unwrap();
std::thread::spawn(move || {
let cfg = Config {
listener_pool: 2,
scheduler_threads: Some(2),
..Config::new(addr)
};
serve_with(cfg, pipeline).unwrap();
serve_with(cfg, smarm::Config::exact(2), pipeline, app_children).unwrap();
});
// Wait for the server to actually be listening.
for _ in 0..50 {
@@ -252,10 +256,9 @@ fn panicking_listener_restarts() {
std::thread::spawn(move || {
let cfg = Config {
listener_pool: 1,
scheduler_threads: Some(2),
..Config::new(addr)
};
serve_with(cfg, pipe).unwrap();
serve_with(cfg, smarm::Config::exact(2), pipe, Vec::new()).unwrap();
});
for _ in 0..50 {
if TcpStream::connect(addr).is_ok() {
@@ -271,13 +274,13 @@ fn panicking_listener_restarts() {
// Arm the fault: the listener's next accept-loop iteration panics
// *before* accepting, so our connection waits in the kernel backlog
// until the restarted listener picks it up.
urus::serve::INJECT_LISTENER_PANICS.store(1, std::sync::atomic::Ordering::Relaxed);
urus::endpoint::INJECT_LISTENER_PANICS.store(1, std::sync::atomic::Ordering::Relaxed);
let resp = send_request(port, b"GET /ping HTTP/1.1\r\nHost: x\r\nConnection: close\r\n\r\n");
assert_eq!(http_status(&resp), 200, "request pending across the panic was not served");
assert_eq!(http_body(&resp), b"pong");
assert_eq!(
urus::serve::INJECT_LISTENER_PANICS.load(std::sync::atomic::Ordering::Relaxed),
urus::endpoint::INJECT_LISTENER_PANICS.load(std::sync::atomic::Ordering::Relaxed),
0,
"fault was never consumed — listener didn't wake for the connection"
);
@@ -298,6 +301,16 @@ fn panicking_listener_restarts() {
fn spawn_server_with_handle(
pipeline: Pipeline,
drain: Duration,
) -> (u16, urus::Handle, std::sync::mpsc::Receiver<()>) {
spawn_server_with_children(pipeline, drain, Vec::new())
}
/// [`spawn_server_with_handle`] with app children started ahead of the
/// endpoint — the pubsub table, a channel hub's actors, and so on.
fn spawn_server_with_children(
pipeline: Pipeline,
drain: Duration,
app_children: Vec<smarm::ChildSpec>,
) -> (u16, urus::Handle, std::sync::mpsc::Receiver<()>) {
let port = free_port();
let addr: SocketAddr = format!("127.0.0.1:{port}").parse().unwrap();
@@ -306,11 +319,11 @@ fn spawn_server_with_handle(
std::thread::spawn(move || {
let cfg = Config {
listener_pool: 2,
scheduler_threads: Some(2),
drain_timeout: drain,
..Config::new(addr)
};
urus::serve_with_shutdown(cfg, pipeline, signal).unwrap();
urus::serve_with_shutdown(cfg, smarm::Config::exact(2), pipeline, app_children, signal)
.unwrap();
let _ = done_tx.send(());
});
for _ in 0..50 {
@@ -437,19 +450,52 @@ fn shutdown_force_stops_at_drain_deadline() {
fn spawn_server_with_timeouts(
pipeline: Pipeline,
keep_alive: Duration,
request: Duration,
head: Duration,
body: Duration,
) -> u16 {
let port = free_port();
let addr: SocketAddr = format!("127.0.0.1:{port}").parse().unwrap();
std::thread::spawn(move || {
let cfg = Config {
listener_pool: 2,
scheduler_threads: Some(2),
keep_alive_timeout: keep_alive,
request_timeout: request,
head_timeout: head,
body_timeout: body,
..Config::new(addr)
};
serve_with(cfg, pipeline).unwrap();
serve_with(cfg, smarm::Config::exact(2), pipeline, Vec::new()).unwrap();
});
for _ in 0..50 {
if TcpStream::connect(addr).is_ok() {
return port;
}
std::thread::sleep(Duration::from_millis(50));
}
panic!("server didn't come up on {addr}");
}
/// Spawn a server with the body stall-gate knobs under test; keep-alive
/// and head budgets are set out of the way so only the body path matters.
fn spawn_server_with_body_gate(
pipeline: Pipeline,
head: Duration,
body: Duration,
burst_bytes: usize,
stall: Duration,
) -> u16 {
let port = free_port();
let addr: SocketAddr = format!("127.0.0.1:{port}").parse().unwrap();
std::thread::spawn(move || {
let cfg = Config {
listener_pool: 2,
keep_alive_timeout: Duration::from_secs(30),
head_timeout: head,
body_timeout: body,
body_burst_bytes: burst_bytes,
body_stall_timeout: stall,
..Config::new(addr)
};
serve_with(cfg, smarm::Config::exact(2), pipeline, Vec::new()).unwrap();
});
for _ in 0..50 {
if TcpStream::connect(addr).is_ok() {
@@ -486,7 +532,8 @@ fn idle_keepalive_reaped_at_keep_alive_timeout() {
let port = spawn_server_with_timeouts(
pipe,
Duration::from_millis(300), // keep_alive_timeout under test
Duration::from_secs(10), // request_timeout out of the way
Duration::from_secs(10), // head_timeout out of the way
Duration::from_secs(10), // body_timeout out of the way
);
let mut s = TcpStream::connect(("127.0.0.1", port)).unwrap();
@@ -511,17 +558,18 @@ fn idle_keepalive_reaped_at_keep_alive_timeout() {
}
/// A slowloris client that sends a partial head and then stalls is killed
/// at request_timeout with a best-effort 408, even though the (large)
/// keep-alive budget hasn't expired.
/// at head_timeout with a best-effort 408, even though the (large)
/// keep-alive and body budgets haven't expired.
#[test]
fn slowloris_partial_head_killed_at_request_timeout() {
fn slowloris_partial_head_killed_at_head_timeout() {
let pipe = Pipeline::new().plug(
Router::new().get("/", |c: Conn, _n: Next| c.put_status(200))
);
let port = spawn_server_with_timeouts(
pipe,
Duration::from_secs(10), // keep_alive_timeout out of the way
Duration::from_millis(300), // request_timeout under test
Duration::from_millis(300), // head_timeout under test
Duration::from_secs(10), // body_timeout out of the way
);
let mut s = TcpStream::connect(("127.0.0.1", port)).unwrap();
@@ -776,11 +824,10 @@ fn stalled_reader_killed_at_write_timeout() {
std::thread::spawn(move || {
let cfg = Config {
listener_pool: 2,
scheduler_threads: Some(2),
write_timeout: Duration::from_millis(300),
..Config::new(addr)
};
serve_with(cfg, pipe).unwrap();
serve_with(cfg, smarm::Config::exact(2), pipe, Vec::new()).unwrap();
});
for _ in 0..50 {
if TcpStream::connect(addr).is_ok() {
@@ -866,11 +913,10 @@ fn chunked_request_over_limit_413() {
std::thread::spawn(move || {
let cfg = Config {
listener_pool: 2,
scheduler_threads: Some(2),
max_body_bytes: 8, // tiny
..Config::new(addr)
};
serve_with(cfg, pipe).unwrap();
serve_with(cfg, smarm::Config::exact(2), pipe, Vec::new()).unwrap();
});
for _ in 0..50 {
if TcpStream::connect(addr).is_ok() { break; }
@@ -913,14 +959,16 @@ fn chunked_plus_content_length_400() {
assert_eq!(http_status(&resp), 400);
}
/// A chunked body that stalls mid-stream is killed by the request
/// deadline: the connection just closes (no response owed mid-body).
/// A chunked body that stalls mid-stream is killed by the BODY deadline
/// (head budget generous): the connection just closes (no response owed
/// mid-body).
#[test]
fn chunked_request_stall_killed_at_request_timeout() {
fn chunked_body_stall_killed_at_body_timeout() {
let port = spawn_server_with_timeouts(
echo_pipeline(),
Duration::from_secs(30),
Duration::from_millis(400), // request_timeout
Duration::from_secs(30), // keep_alive_timeout out of the way
Duration::from_secs(30), // head_timeout out of the way
Duration::from_millis(400), // body_timeout under test
);
let mut s = TcpStream::connect(("127.0.0.1", port)).unwrap();
s.set_read_timeout(Some(Duration::from_secs(5))).unwrap();
@@ -936,6 +984,139 @@ fn chunked_request_stall_killed_at_request_timeout() {
assert!(start.elapsed() < Duration::from_secs(3), "close took too long");
}
/// The core of the head/body split: a client that sends a COMPLETE head
/// promptly and then trickles its (small) body over a span LONGER than
/// head_timeout still succeeds, because the body runs on its own, larger
/// budget. Under the old shared request clock this would have been killed
/// mid-body at head_timeout. This is the slow-but-legit IoT upload we must
/// not punish.
#[test]
fn slow_body_outlives_head_timeout() {
let port = spawn_server_with_timeouts(
echo_pipeline(),
Duration::from_secs(30), // keep_alive_timeout out of the way
Duration::from_millis(500), // head_timeout: SHORT
Duration::from_secs(8), // body_timeout: generous
);
let mut s = TcpStream::connect(("127.0.0.1", port)).unwrap();
s.set_read_timeout(Some(Duration::from_secs(10))).unwrap();
// Full head at once (parses well within head_timeout), Connection:
// close so the server closes after responding and read_to_end lands
// the whole response.
s.write_all(
b"POST /echo HTTP/1.1\r\nHost: x\r\nContent-Length: 4\r\nConnection: close\r\n\r\n",
)
.unwrap();
// Trickle the 4-byte body at 250ms/byte => ~1s total, well past the
// 500ms head_timeout but inside the 8s body_timeout.
for b in b"test" {
std::thread::sleep(Duration::from_millis(250));
s.write_all(&[*b]).unwrap();
}
let mut resp = Vec::new();
s.read_to_end(&mut resp).expect("expected full response");
assert_eq!(http_status(&resp), 200, "resp: {:?}", String::from_utf8_lossy(&resp));
assert!(
resp.ends_with(b"test"),
"expected echoed body 'test', got: {:?}", String::from_utf8_lossy(&resp)
);
}
/// A fixed-Content-Length body that stalls before completing is killed by
/// the BODY deadline (head budget generous): silent close, nothing owed
/// mid-body. The fixed-path twin of chunked_body_stall_killed_at_body_timeout.
#[test]
fn fixed_body_stall_killed_at_body_timeout() {
let port = spawn_server_with_timeouts(
echo_pipeline(),
Duration::from_secs(30), // keep_alive_timeout out of the way
Duration::from_secs(30), // head_timeout out of the way
Duration::from_millis(400), // body_timeout under test
);
let mut s = TcpStream::connect(("127.0.0.1", port)).unwrap();
s.set_read_timeout(Some(Duration::from_secs(5))).unwrap();
// Promises 100 bytes, sends a few, then stalls forever.
s.write_all(b"POST /echo HTTP/1.1\r\nHost: x\r\nContent-Length: 100\r\n\r\npartial")
.unwrap();
let start = std::time::Instant::now();
let mut resp = Vec::new();
s.read_to_end(&mut resp).unwrap(); // server closes; EOF
assert!(resp.is_empty(), "expected silent close, got: {:?}", String::from_utf8_lossy(&resp));
assert!(start.elapsed() < Duration::from_secs(3), "close took too long");
}
/// Burst gate, NEGATIVE (chunked path): a client that ACTIVELY but SMOOTHLY
/// trickles sub-burst bytes is evicted at ~body_stall_timeout — even though
/// the absolute body_timeout is far away and the client never goes fully
/// silent. This is the slowloris-body case the gate exists to catch, and
/// exercises the fill_to gate in read_chunked_body.
#[test]
fn body_smooth_trickle_evicted_at_stall_timeout() {
let port = spawn_server_with_body_gate(
echo_pipeline(),
Duration::from_secs(30), // head_timeout out of the way
Duration::from_secs(30), // body_timeout out of the way (prove it's the STALL gate)
4096, // body_burst_bytes
Duration::from_millis(800), // body_stall_timeout under test
);
let mut s = TcpStream::connect(("127.0.0.1", port)).unwrap();
s.set_read_timeout(Some(Duration::from_secs(5))).unwrap();
// Head + a chunk-size line announcing a 4096-byte chunk, then trickle
// its payload one byte at a time: never a full burst, so the stall mark
// never advances.
s.write_all(b"POST /echo HTTP/1.1\r\nHost: x\r\nTransfer-Encoding: chunked\r\n\r\n1000\r\n")
.unwrap();
let start = std::time::Instant::now();
let mut evicted = false;
for _ in 0..200 { // up to ~20s; eviction expected at ~800ms
if s.write_all(b"x").is_err() {
evicted = true; // server closed on us -> write failed
break;
}
std::thread::sleep(Duration::from_millis(100));
}
assert!(evicted, "server never evicted the smooth sub-burst trickle");
assert!(
start.elapsed() < Duration::from_secs(3),
"eviction took {:?}, expected ~800ms (stall gate, not the 30s cap)", start.elapsed()
);
}
/// Burst gate, POSITIVE (fixed-CL path): a slow-but-legit client that
/// delivers real bursts with gaps SHORTER than body_stall_timeout keeps
/// resetting the stall mark and completes intact. This is the slow IoT
/// upload the gate must NOT punish; exercises the read_body gate.
#[test]
fn bursty_slow_body_survives_stall_gate() {
let port = spawn_server_with_body_gate(
echo_pipeline(),
Duration::from_secs(30), // head_timeout out of the way
Duration::from_secs(30), // body_timeout out of the way
4096, // body_burst_bytes
Duration::from_secs(2), // body_stall_timeout: gaps stay under this
);
let mut s = TcpStream::connect(("127.0.0.1", port)).unwrap();
s.set_read_timeout(Some(Duration::from_secs(10))).unwrap();
// Promise 3 * 4096 bytes, Connection: close so read_to_end lands the
// full echo.
let burst = vec![b'x'; 4096];
s.write_all(b"POST /echo HTTP/1.1\r\nHost: x\r\nContent-Length: 12288\r\nConnection: close\r\n\r\n")
.unwrap();
for i in 0..3 {
s.write_all(&burst).unwrap();
if i < 2 {
std::thread::sleep(Duration::from_millis(500)); // < 2s stall window
}
}
let mut resp = Vec::new();
s.read_to_end(&mut resp).expect("expected full response");
assert_eq!(http_status(&resp), 200, "resp head: {:?}", String::from_utf8_lossy(&resp[..resp.len().min(120)]));
let body_at = resp.windows(4).position(|w| w == b"\r\n\r\n").expect("no head terminator") + 4;
let body = &resp[body_at..];
assert_eq!(body.len(), 12288, "echoed body truncated: {} bytes", body.len());
assert!(body.iter().all(|&b| b == b'x'), "echoed body corrupted");
}
// ---------------------------------------------------------------------------
// SSE (v0.3 chunk 3)
// ---------------------------------------------------------------------------
@@ -1273,15 +1454,14 @@ fn shutdown_force_stops_open_ws() {
/// Minimal chat handler mirroring the example: on_open subscribes (conn
/// actor pid) + spawns a relay holding ONLY the Receiver and a WsSender
/// clone; on_message broadcast_from's, skipping the sender's own relay.
struct Chat {
bus: std::sync::Arc<std::sync::OnceLock<urus::PubSub<String>>>,
}
const CHAT_BUS: urus::PubSub<String> = urus::PubSub::new("chat-test");
struct Chat;
impl urus::WsHandler for Chat {
fn on_open(&mut self, sender: &urus::WsSender) {
let bus = self.bus.get_or_init(urus::PubSub::new);
let rx = bus.subscribe("room").unwrap();
let _ = bus.broadcast_from(smarm::self_pid(), "room", "joined".into());
let rx = CHAT_BUS.subscribe("room").unwrap();
let _ = CHAT_BUS.broadcast_from(smarm::self_pid(), "room", "joined".into());
let out = sender.clone();
smarm::spawn(move || {
while let Ok(msg) = rx.recv() {
@@ -1303,19 +1483,14 @@ impl urus::WsHandler for Chat {
let _ = sender.send(urus::Message::Text("synced".into()));
return;
}
let bus = self.bus.get_or_init(urus::PubSub::new);
let _ = bus.broadcast_from(smarm::self_pid(), "room", t);
let _ = CHAT_BUS.broadcast_from(smarm::self_pid(), "room", t);
}
}
}
fn chat_pipeline() -> Pipeline {
let bus: std::sync::Arc<std::sync::OnceLock<urus::PubSub<String>>> =
std::sync::Arc::new(std::sync::OnceLock::new());
Pipeline::new().plug(Router::new().get("/ws", move |c: Conn, _n: Next| {
let bus = bus.clone();
c.upgrade(Chat { bus })
}))
Pipeline::new()
.plug(Router::new().get("/ws", |c: Conn, _n: Next| c.upgrade(Chat)))
}
/// Two clients in a room: a broadcast reaches the other client and (via
@@ -1324,7 +1499,7 @@ fn chat_pipeline() -> Pipeline {
/// having sent a single frame when A's first message arrives.
#[test]
fn ws_chat_broadcast_reaches_other_client_not_sender() {
let port = spawn_server(chat_pipeline());
let port = spawn_server_children(chat_pipeline(), vec![CHAT_BUS.child()]);
let mut a = ws_connect(port);
let mut a_buf = Vec::new();
@@ -1367,7 +1542,11 @@ fn ws_chat_broadcast_reaches_other_client_not_sender() {
#[test]
fn shutdown_with_open_chat_terminates() {
let (port, handle, done_rx) =
spawn_server_with_handle(chat_pipeline(), Duration::from_millis(300));
spawn_server_with_children(
chat_pipeline(),
Duration::from_millis(300),
vec![CHAT_BUS.child()],
);
let mut a = ws_connect(port);
let mut a_buf = Vec::new();
@@ -1427,21 +1606,19 @@ mod channels_wire {
}
}
fn channels_pipeline() -> Pipeline {
let hub: std::sync::Arc<std::sync::OnceLock<ChannelHub<P>>> =
std::sync::Arc::new(std::sync::OnceLock::new());
Pipeline::new().plug(Router::new().get("/socket", move |c: Conn, _n: Next| {
// Hub construction is in-runtime only (it spawns the pubsub
// table); the NON-static OnceLock is what lets shutdown
// drain the table (the v0.5 pattern).
let hub = hub.clone();
let hub = hub.get_or_init(|| {
ChannelHub::new(PrefixRouter::new().channel("room:*", |_: &str| {
Box::new(Lobby) as Box<dyn Channel<P>>
}))
});
hub.upgrade(c)
}))
/// The hub is a description: no actors, so it is built once here and
/// its `children()` handed to the supervisor separately.
fn channels_hub() -> (ChannelHub<P>, Vec<smarm::ChildSpec>) {
ChannelHub::new(
"chan-test",
PrefixRouter::new()
.channel("room:*", |_: &str| Box::new(Lobby) as Box<dyn Channel<P>>),
)
}
fn channels_pipeline(hub: ChannelHub<P>) -> Pipeline {
Pipeline::new()
.plug(Router::new().get("/socket", move |c: Conn, _n: Next| hub.upgrade(c)))
}
const CHAN_HANDSHAKE: &[u8] =
@@ -1469,7 +1646,8 @@ mod channels_wire {
#[test]
fn channels_join_heartbeat_event_broadcast_leave() {
let port = spawn_server(channels_pipeline());
let (hub, children) = channels_hub();
let port = spawn_server_children(channels_pipeline(hub), children);
let mut a = chan_connect(port);
let mut a_buf = Vec::new();
@@ -1555,7 +1733,8 @@ mod channels_wire {
#[test]
fn channels_codec_garbage_closes_1002() {
let port = spawn_server(channels_pipeline());
let (hub, children) = channels_hub();
let port = spawn_server_children(channels_pipeline(hub), children);
let mut s = chan_connect(port);
let mut buf = Vec::new();
ws_send(&mut s, &Frame::new(Opcode::Text, "not a v2 frame"));
@@ -1572,7 +1751,14 @@ mod channels_wire {
// is monitor-pruned -> the non-static hub's last handle drops
// with the drained pipeline -> pubsub table exits -> AllDone.
let (port, handle, done_rx) =
spawn_server_with_handle(channels_pipeline(), Duration::from_millis(300));
{
let (hub, children) = channels_hub();
spawn_server_with_children(
channels_pipeline(hub),
Duration::from_millis(300),
children,
)
};
let mut a = chan_connect(port);
let mut a_buf = Vec::new();
@@ -1600,19 +1786,20 @@ mod channels_wire {
}
}
fn session_pipeline() -> Pipeline {
let hub: std::sync::Arc<std::sync::OnceLock<ChannelHub<P>>> =
std::sync::Arc::new(std::sync::OnceLock::new());
Pipeline::new().plug(Router::new().get("/socket", move |c: Conn, _n: Next| {
let hub = hub.clone();
let hub = hub.get_or_init(|| {
ChannelHub::new(PrefixRouter::new().channel_session::<LobbySession>(
fn session_hub() -> (ChannelHub<P>, Vec<smarm::ChildSpec>) {
ChannelHub::new(
"sess-test-bus",
PrefixRouter::new().channel_session::<LobbySession>(
"room:*",
"sess-test-registry",
|_: &str| Box::new(Lobby) as Box<dyn Channel<P>>,
))
});
hub.upgrade(c)
}))
),
)
}
fn session_pipeline(hub: ChannelHub<P>) -> Pipeline {
Pipeline::new()
.plug(Router::new().get("/socket", move |c: Conn, _n: Next| hub.upgrade(c)))
}
#[test]
@@ -1625,7 +1812,14 @@ mod channels_wire {
// terminates, and exits -> AllDone. No links anywhere in that
// chain; this test is the proof it composes.
let (port, handle, done_rx) =
spawn_server_with_handle(session_pipeline(), Duration::from_millis(300));
{
let (hub, children) = session_hub();
spawn_server_with_children(
session_pipeline(hub),
Duration::from_millis(300),
children,
)
};
let mut a = chan_connect(port);
let mut a_buf = Vec::new();
@@ -1644,3 +1838,103 @@ mod channels_wire {
.expect("detached session outlived the drain: the registry-drop chain is broken");
}
}
// ---------------------------------------------------------------------------
// Endpoint as a supervised child (v0.3) — the spec §6 tree shape.
// ---------------------------------------------------------------------------
/// The whole point of v0.3: the APP owns the runtime and the root
/// supervisor; urus is one ordered child among the app's own. A
/// `RuntimeHandle::request_shutdown` on the root sup (the SIGTERM shape)
/// winds the tree down in reverse start order and `rt.run` returns.
#[test]
fn endpoint_as_supervised_child_serves_and_drains() {
use std::sync::atomic::{AtomicBool, Ordering};
use std::sync::Arc;
let port = free_port();
let addr: SocketAddr = format!("127.0.0.1:{port}").parse().unwrap();
let app_child_shut_down = Arc::new(AtomicBool::new(false));
let pipeline = Pipeline::new()
.plug(Router::new().get("/", |c: Conn, _n: Next| c.put_status(200).put_body("app+urus")));
// Endpoint construction is eager about the bind: errors surface here,
// on the app's thread, not inside some actor.
let endpoint = urus::endpoint(
Config {
listener_pool: 2,
drain_timeout: Duration::from_secs(5),
..Config::new(addr)
},
pipeline,
)
.expect("bind");
let (handle_tx, handle_rx) = std::sync::mpsc::channel();
let (pid_tx, pid_rx) = std::sync::mpsc::channel();
let (done_tx, done_rx) = std::sync::mpsc::channel();
let flag = app_child_shut_down.clone();
std::thread::spawn(move || {
let rt = smarm::init(smarm::Config::exact(2));
handle_tx.send(rt.handle()).unwrap();
rt.run(move || {
let sup = smarm::spawn(move || {
smarm::OneForOne::new()
.strategy(smarm::Strategy::RestForOne)
// An app child started BEFORE the endpoint: reverse-order
// shutdown must take the endpoint down first, so when
// this child's guard runs the port must already be dead.
.child(smarm::ChildSpec::new(smarm::Restart::Permanent, {
let flag = flag.clone();
move || {
struct G(Arc<AtomicBool>);
impl Drop for G {
fn drop(&mut self) {
self.0.store(true, Ordering::SeqCst);
}
}
let _g = G(flag.clone());
loop {
smarm::sleep(Duration::from_secs(3600));
}
}
}))
.child(smarm::ChildSpec::new(smarm::Restart::Permanent, endpoint.clone()))
.run()
});
// Export the sup pid for the "signal thread" below.
pid_tx.send(sup.pid()).unwrap();
sup.join().expect("root sup returns normally after shutdown");
});
let _ = done_tx.send(());
});
// Stand-in signal thread state.
let handle = handle_rx.recv().unwrap();
let sup_pid = pid_rx.recv().unwrap();
// Server is up and serving through the supervised endpoint.
for _ in 0..50 {
if TcpStream::connect(addr).is_ok() {
break;
}
std::thread::sleep(Duration::from_millis(50));
}
let resp = send_request(port, b"GET / HTTP/1.1\r\nhost: x\r\nconnection: close\r\n\r\n");
assert!(resp.windows(8).any(|w| w == b"app+urus"), "endpoint serves");
// SIGTERM shape: an outside thread shuts the root sup down.
handle.request_shutdown(sup_pid);
done_rx
.recv_timeout(Duration::from_secs(5))
.expect("rt.run returned after root-sup shutdown");
assert!(
app_child_shut_down.load(Ordering::SeqCst),
"app sibling wound down"
);
assert!(
TcpStream::connect(addr).is_err(),
"listen fds closed: no new connections after shutdown"
);
}