52 Commits
Author SHA1 Message Date
Claude ab93802d0b release: v0.3.0
Bump crate version to 0.3.0; pin smarm dependency to the v0.7.0 tag (dev path override reverted).
2026-08-21 13:10:56 +02:00
Claude 5442271229 build(examples): gate load_profile behind smarm-causal
smarm::causal is feature-gated, so examples/load_profile.rs no longer
compiles under the default feature set and broke `cargo build --examples`.
Declare it with required-features = ["smarm-causal"] so the default build
skips it and `--features smarm-causal` still builds it.

Splits the committable half out of the working-tree Cargo.toml; the
`smarm = { path = "../smarm_full" }` dev pin stays uncommitted and is
restored to a tag in the release commit.
2026-08-20 19:38:24 +00:00
Claude e02527a606 docs: record the supervised-bus cycle (v0.8) and retire the OnceLock idiom
ROADMAP gains a v0.8 entry: the smarm 415effb lifetime change as root cause,
the description/instantiation split, the tree shape, the accepted per-call
resolution cost, and what stays open (RegistryName newtype, PubSub's
unprotectable const, dynamic session actors, hammer.sh's missing feature
matrix). The v0.7 "unreproduced test failure" open item is closed out and
pointed at it — it was this hang, hidden because hammer.sh builds with default
features while the failure needs --all-features load.

README and the two module headers still taught the v0.5 rules: in-runtime-only
construction, the non-static Arc<OnceLock<..>> cell, and "a relay must never
hold a PubSub clone". None of those are true any more; they are replaced by
what actually holds now, with a note on what changed for anyone who learnt the
old shape.

session.rs also gains an honest note that its session actors are the last
lifetime in urus implied by a drop rather than stated — they are dynamic, so a
fixed ChildSpec list cannot hold them, and the trigger is at least a command
now rather than a refcount.
2026-08-20 15:01:44 +00:00
Claude 9eaa85b9df style(tests): byte-string literal in the stalled-writer probe
clippy::byte_char_slices, pre-existing and unrelated to the refactor in the
previous commit — split out so that diff stays scoped. Restores a clean
`cargo clippy --all-features --all-targets`.
2026-08-20 14:49:29 +00:00
Claude 8f0da2a806 feat(pubsub,channels)!: handles are addresses — bus, hub and session registries are supervised children
smarm 0.7 (415effb, "lifetime is the actor's — refs are addresses") removed the
rule these three actors were built on: a GenServerRef no longer owns the server,
the loop holds its own inbox sender, and the inbox never closes when the last ref
drops. urus's pubsub table, channel hub bus and session registries were still
governed by that deleted rule — PubSub::new() spawned the table and the handle
owned its life — so nothing commanded them to stop. What still terminated a run
was the root-exit sweep, racing the drain: shutdown_with_open_chat_terminates and
channels_wire::shutdown_with_open_channel_terminates failed 4 times in 25
--all-features runs with "serve did not return: ... outlived the drain: Timeout".
Zero in 10 full-suite runs after this change.

The fix is not a supervisor wrapped around the old shape. Every gotcha in this
area descended from constructors that spawn: PubSub::new(), ChannelHub::new() and
PrefixRouter::channel_session() all started actors, which forced in-runtime-only
construction, which forced the Arc<OnceLock<..>> lazy-init from the first handler,
which forced the "cell must not be static" and "a relay must never hold a PubSub
clone" rules. Five documented rules propping up one inverted dependency. So:
description is separated from instantiation.

- PubSub<M> is a name, not a GenServerRef: const-constructible, Copy, spawns
  nothing, valid outside the runtime and in a static. Operations resolve through
  the registry per call, so a table restarted by its supervisor is reached
  transparently (one lookup per broadcast — bench before caching a ref, which
  would go stale across exactly the restart the supervisor exists to perform).
  PubSub::new() is gone; PubSub::new(name) + PubSub::child() replace it.
- ChannelHub::new(bus, router) returns (hub, Vec<ChildSpec>) — the bus table plus
  one registry per session route. Returning both is the point: a hub whose
  children were never started compiles and fails on the first join, so the vec is
  not left behind a method you can forget to call. #[must_use].
- channel_session gains a registry name; each session registry is separately
  named and separately supervised.
- serve_with/serve_with_shutdown take a Vec<ChildSpec> of app children and build
  the root as RestForOne[..app children, endpoint]. They start before the
  endpoint and, shutdown being ordered in reverse, stop after it has drained, so
  a request still in flight can reach the bus. RestForOne because a bus crash
  leaves live sockets addressing a table that no longer knows them.
- Deleted: the Arc<OnceLock> idiom from both examples and both test pipelines,
  and the module rules that existed only to hand-manage a refcount.

Known cost, not fixed here: channel_session("session:*", "chat-sessions", f) puts
two unrelated string literals side by side and nothing catches a transposition —
a RegistryName newtype is the obvious follow-up.

Tests: 111 lib + 50 integration + 2 doc green, clippy clean, 10/10 full-suite
runs. Unit tests poll for name binding before use — smarm's start-order-is-not-
start-readiness gap; real apps don't hit it, since a handler only runs once a
connection has been accepted.
2026-08-20 14:49:23 +00:00
Claude 8a568c600c docs(roadmap): record the endpoint milestone (crate v0.3.0)
Full account of the refactor: the tree shape, why the endpoint spawns its
own listener sup (supervisor start order != start readiness), the config
split, what was deleted and the smarm properties re-probed to justify
deleting it. Also:
- PubSub discovery note: the non-static Arc<OnceLock> rule is now
  serve*-only; an app that owns its tree starts the table as a supervised
  sibling and addresses it by name (what crud does).
- Icebox 'operator introspection via typed names': partly delivered — the
  endpoint is a named gen_server answering Call::ConnCount; per-listener
  visibility and richer stats remain open.
- Open after this cycle: the one unreproduced hammer failure, and that
  endpoint() returns impl Fn() so a second invocation clashes on the name.
2026-08-20 13:27:27 +00:00
Claude b37888ec2c docs+examples: v0.3 endpoint API; crud converted to the app-owns-the-tree shape
- crud is now the demonstrator: app owns smarm runtime + root supervisor,
  store actor and urus::endpoint as ordered siblings (store first, so
  reverse-order shutdown drains HTTP before stopping the store),
  Shutdown::Infinity on the endpoint, shutdown via
  rt.handle().request_shutdown(root_sup) from the stdin thread.
  Deletes two kludges the old shape forced:
    * static OnceLock<Sender> spawn-on-first-use -> a supervised child
      that self-registers a typed Name; handlers use smarm::send per
      request and turn 'between incarnations' into a 503 instead of
      panicking on a dropped store.
    * static SHUTTING_DOWN AtomicBool + 250ms recv_timeout poll in the
      store loop and in the SSE ticker -> a plain park; the tree stops
      both. Smoke-tested live: CRUD round-trips, SSE stream, clean drain
      with the stream open, port closed after.
- Other examples stay short and on serve*, updated for the split config
  (serve_with(cfg, smarm::Config, pipe) / serve_with_shutdown(..., signal)).
  plain_serve's URUS_SCHED_THREADS now builds a smarm::Config.
  ws_chat's doc block explains the OnceLock is a serve*-only workaround
  and points at crud for the clean shape.
- README: new 'Your Own Supervision Tree' section (endpoint as the real
  API), graceful shutdown reframed as the serve*-only path, Config table
  loses scheduler_threads and gains name, PubSub rule 1 notes the
  supervised-sibling alternative.

111 lib + 50 integration + 2 doc tests green; clippy clean.
2026-08-20 13:20:28 +00:00
Claude 099b6bc320 feat(endpoint): urus is a supervisable child — endpoint gen_server owns listeners + conns
The v0.3 shape from the spec: an app owns its runtime and root supervisor
and places urus in it as one ordered child among its own.

  your root sup
  └── ChildSpec(Permanent, urus::endpoint(cfg, pipeline)?)  <- Endpoint
      └── listener_sup  OneForOne over N listeners
          └── plain connection actors

- src/conn_registry.rs -> src/endpoint.rs. The registry gains the listener
  pool it registers for and becomes the Endpoint gen_server; ConnRegistry
  -> Endpoint. It runs inline as the ChildSpec's actor
  (NamedGenServerBuilder::run), so supervisor shutdown arrives as
  handle_shutdown and a restart re-runs init on the same still-open fds.
- Endpoint spawns its OWN listener sup in init rather than being its
  sibling: smarm's supervisor start order is not start *readiness* (spawn
  is fire-and-forget), so a sibling listener could whereis the name before
  the registry actor ran. Registrar-spawns-consumers makes it program
  order inside one init. Gap filed in smarm ROADMAP (readiness ack);
  making spawn blocking would only shrink the window, not close it —
  'has begun executing' is not 'has bound its name'.
- Listener sup is monitored: death outside shutdown = panic (loud, the
  app's supervisor decides) instead of a zombie on a dead port. Death
  during shutdown is the 'no new connections' barrier.
- DELETED: the shutdown AtomicBool, LISTENER_TICK (250ms wake per listener
  per tick, now an untimed wait_readable park), SHUTDOWN_POLL (100ms root
  poll — the root parks on the signal channel now), Restart::Transient
  (listeners are Permanent: they only exit by supervisor action, so a
  self-exit always means breakage). Verified against current smarm:
  request_stop unwinds an untimed wait_readable park, is no longer lossy
  against a QUEUED actor, and supervisor shutdown joins in ~200us.
- Config: scheduler_threads/max_actors removed (runtime knobs an
  endpoint-as-child cannot honour) -> serve_with(cfg, smarm::Config, pipe)
  and serve_with_shutdown(cfg, smarm::Config, pipe, signal). Added
  Config.name (default 'urus'): the endpoint's registry name, unique per
  endpoint, and the introspection handle via endpoint::whereis(name).
- serve* keep their meaning as the batteries-included path: they build a
  one-child tree around endpoint() with Shutdown::Infinity. Handle stays
  (a serve* caller has no RuntimeHandle to reach for) and now backs a real
  park instead of a poll.

Tests: 4 drain tests ported onto a real endpoint (bound socket, supervised
child, request_shutdown driven); new integration test boots an app tree
with an ordered sibling and asserts serve-then-drain, reverse-order
teardown and a closed port. 106 lib + 50 integration green.
2026-08-20 12:59:57 +00:00
Claude 5014b870d6 feat(conn_registry): registry owns the drain — trapping gen_server, request_shutdown driven
The drain protocol moves wholesale into the registry (smarm >=0.7:
trap_exit + handle_shutdown + gen_server timers + StopHandle):

- handle_shutdown: flip draining, stop idle conns, arm a drain_timeout
  timer, Continue; with no conns, Exit immediately.
- ConnIdle while draining stops the conn (unchanged); ConnStarted while
  draining stops it on arrival (unchanged); ConnEnded that empties the
  set while draining = StopHandle::stop() — the registry's own normal
  exit is now the 'every connection is gone' barrier.
- handle_timer (deadline): one force-stop sweep. The old re-sweep-every-
  10ms loop existed to catch late registrants; stop-on-arrival already
  covers every post-sweep entry path, so one sweep suffices.
- Cast::{BeginDrain,ForceStopConns} deleted (internal now); ConnCount
  stays as the introspection call. start() takes drain_timeout.
- serve.rs: the root's whole drain/poll block collapses to
  registry.shutdown() (graceful, monitors until the server has stopped
  itself). SHUTDOWN_POLL + the listener flag are untouched here; they go
  with the endpoint refactor.
- Known residual window documented in the module docs: a conn spawned
  by a dying listener that has not yet registered can outlive an
  already-empty registry; it is collected by smarm's root-exit sweep.

Tests (in-lib, request_shutdown driven): empty-set immediate exit;
idle-now/busy-at-deadline ordering with exit-after; stop-on-idle
mid-drain; stop-on-arrival mid-drain. 35x hammer subset green.
2026-08-20 08:57:19 +00:00
Claude (sandbox) 8bdec97842 feat(config): optional TOML config loading behind config-file feature
A file-based way to set tuning knobs without recompiling. urus is a
library, so it never presumes a config path or reads the environment — the
binary hands the text in:

- Config::with_toml_str(&str) -> Result<Config, ConfigError>: sparse overlay
  onto an existing Config (built with the addr the binary chose). Only keys
  present are applied; durations are integer seconds; unknown keys are a
  hard error (deny_unknown_fields) so a typo is loud, not a silent no-op.
- Scope: the four slowloris knobs (head_timeout_secs, body_timeout_secs,
  body_burst_bytes, body_stall_timeout_secs). Migrating the rest of Config
  into the file is a separate, additive job — TomlOverrides just grows.
- Feature `config-file = ["dep:serde", "dep:toml"]`; the optional serde dep
  gains the derive feature. The default build is unchanged (deps + code are
  all gated).
- examples/serve_toml.rs (required-features = ["config-file"]): a `--config
  PATH` demo with no presumed default location. plain_serve and its env
  vars are left untouched.

Tests (feature-gated): empty keeps defaults, partial overrides only named,
full overrides all, unknown key errors, malformed errors. 84 lib with the
feature / 79 without; clippy --lib clean both ways; e2e smoke serves 200
from a file and rejects an unknown key loudly.
2026-08-12 13:41:57 +00:00
Claude (sandbox) f3ccb6e468 feat(serve): burst-gated body stall eviction
The body budget from the prior commit is a generous absolute cap; alone it
just hands a body-phase slowloris a bigger window. Add a stall gate under
that cap that distinguishes a slowloris trickle from a slow-but-legit
client by requiring BURSTS, not a mere average rate:

- BodyStallGate: each body read is bounded by min(body cap, mark + stall).
  The stall mark advances only when body_burst_bytes accumulate since the
  last advance, so a steady sub-burst trickle never moves it and is evicted
  at ~body_stall_timeout, while a bursty slow client keeps resetting it.
- Two words of state; one add + one compare per read. Raw socket bytes are
  counted, so chunked framing counts and an MSS-fragmented burst still
  accumulates. Reuses the existing read_some deadline plumbing.
- Wired into read_body (fixed CL) and read_chunked_body (via fill_to, the
  single choke point all chunked reads pass through).
- New knobs body_burst_bytes (4 KiB) + body_stall_timeout (20s); effective
  floor ~205 B/s enforced in bursts.

Tests: body_smooth_trickle_evicted_at_stall_timeout (chunked; active
sub-burst trickle evicted at ~stall while the cap is far away) and
bursty_slow_body_survives_stall_gate (fixed CL; real bursts with sub-stall
gaps complete intact). 79 lib + 45 integration green; clippy --lib clean.
2026-08-12 13:38:16 +00:00
Claude (sandbox) 4f06265338 feat(serve): split request read budget into head_timeout + body_timeout
The single request_timeout covered head + body under one wall clock, so a
slow-but-legit body upload (e.g. a trickling cellular IoT client) was
judged by the short head deadline and killed mid-body. Split into:

- head_timeout (default 30s): first byte -> full head parse; the classic
  slowloris surface, kept short.
- body_timeout (default 300s): head parse -> full body; an absolute cap
  sized for slow links, anchored independently once the head has parsed.

read_head no longer returns a shared deadline; run_connection anchors the
body deadline itself. ReadHeadErr::RequestTimeout -> HeadTimeout. Config
and ConnLimits gain body_timeout; request_timeout renamed to head_timeout
(breaking, but this axis is unreleased).

Tests: slow_body_outlives_head_timeout (positive: body survives past the
head clock), fixed_/chunked_body_stall_killed_at_body_timeout (body cap
still bites), slowloris_partial_head_killed_at_head_timeout (head clock
unchanged). 79 lib + 43 integration green; clippy --lib clean.
2026-08-12 13:00:22 +00:00
Claude (sandbox) 1b1ea124c8 feat(serve): plumb Config.max_actors through to smarm init
Config gained a max_actors: Option<usize> (None = smarm's DEFAULT_MAX_ACTORS
of 16_384). serve_with now applies it to the smarm runtime config, so the
per-connection actor slab can be sized to the deployment's peak concurrent
connections. Since each connection is one actor, the slab was the hard cap
on concurrent connections (previously an un-raisable 16_384) regardless of
RAM/fds; slots are ~256 B so raising it is cheap next to per-conn stacks.

Verified on the GPU box: default caps at 16_383 held connections; with
max_actors raised, a paced ramp holds 100_000 concurrent slow-header
connections at 12.4 KB RSS / 2 VMAs each (1.29 GB total) on one pinned core.
2026-08-10 05:47:22 +00:00
Claude (sandbox) b86c64d490 feat(parser): strict Transfer-Encoding framing; unknown coding -> 501
The TE arm set chunked whenever the token appeared anywhere in the value,
so 'chunked, gzip' (chunked not final) was accepted and an unknown coding
like 'bogus' was treated as no-body (h1spec #18/#19 -> 404). Collect the
ordered coding list across all TE headers and decide post-loop: TE on
HTTP/1.0 or TE+Content-Length -> 400 (the CL check now covers ANY TE, not
just chunked, closing the old TE:unknown + CL smuggling gap); chunked
present but not final -> 400; any coding other than chunked -> 501 via a
new UnknownTransferCoding variant (emit_error_response gains the 501 arm);
only a sole final chunked sets the flag. Tests cover each branch.

Note: the dead Unsupported/411 variant is left as-is (separate cleanup).
2026-08-09 08:02:04 +00:00
Claude (sandbox) 394e9b962a feat(parser): reject duplicate Content-Length (RFC 9112 §6.3)
The content-length arm ran content_length = Some(parse) per header, so a
second Content-Length silently overwrote the first with no conflict check
(CL.CL request smuggling; h1spec #21 -> 404 instead of 400). Count
occurrences and reject any duplicate post-loop, strictly (even equal
values), reusing BadContentLength (400). A single value is still required
to be one decimal integer, so a comma-list or non-numeric keeps failing at
parse as before. Tests: differing dup, equal dup, single-CL regression.
2026-08-09 07:59:25 +00:00
Claude (sandbox) 6f02cec261 feat(parser): reject missing/duplicate/invalid Host (RFC 9112 §3.2)
parse_head never inspected Host, so a missing (HTTP/1.1), duplicate, or
syntactically invalid Host all passed through to the router (h1spec #8/#9/
#10 -> 404 instead of 400). Add per-header validity (RFC 3986 host[:port]
charset via valid_host) plus a post-loop presence/uniqueness check: 1.1
MUST carry exactly one valid Host; 1.0 may omit it but a duplicate/invalid
one is still 400. Unit matrix mirrors the three h1spec cases with reg-name/
port/IPv6-literal positive controls.
2026-08-09 07:58:16 +00:00
Markk116 535f7bcc68 feat(serve): give connection actors a 256 KiB stack via smarm SpawnOpts
Connection actors were still spawned with a bare smarm::spawn(), which
gets the runtime's fixed 64 KiB default stack regardless of smarm
v0.6.0's RFC 019 SpawnOpts/stack_reserve work landing one crate down.
Any handler that leans on app code with real stack needs (DB drivers,
(de)compression, ...) blows the guard page and the connection just
dies with no response - reproduced with a CCC handler that decompresses
gzip on the identity-encoding path.

Add Config::conn_stack_reserve (default DEFAULT_CONN_STACK_RESERVE =
256 KiB) and thread it through listener_loop into a
smarm::spawn_with(SpawnOpts { stack_reserve: Some(_), .. }, ...) call
for every accepted connection. Existing Config { ..Config::new(addr) }
call sites (tests/integration.rs) pick up the new field automatically
via struct-update syntax; no call-site churn beyond that.

Bump to 0.2.2.
2026-08-08 22:39:37 +02:00
Claude (sandbox)andClaude b137c646b0 release: v0.2.1 — switch smarm to the pinned v0.6.0 git tag
RFC 019 lands upstream: per-actor stacks, park-path shrink, recycle zap,
SIGSEGV diagnostics, introspect surface. No urus code changes required —
E1 interleaved A/B on the box shows every ka cell within +0.3..+2.9% of
the v0.5.0 pin (t8-c4 close-mode control is bistable either side; see
smarm v0.6.0 release notes). Tag must exist upstream before this builds:
push smarm master + v0.6.0 first.
2026-08-08 22:17:32 +02:00
Markk116 792897d3e4 License under MIT
- LICENSE: MIT text
- Cargo.toml: license = "MIT"
2026-08-08 16:37:57 +02:00
Markk116 b77448191e release: v0.2.0 — switch smarm to the pinned v0.5.0 git tag
The gen_server API port itself already landed upstream (078072b, tracking
smarm HEAD efbc254 pre-git-dep). This just moves the dependency off the
local path checkout onto the git remote, pinned to the smarm v0.5.0
release tag (a03a7ca) rather than a floating branch HEAD.

Verified: 90 lib + 45 integration + 2 doc tests green, examples build,
under default and --all-features.
2026-08-08 11:50:51 +02:00
Claude 34730930e7 feat(bench): E1 scheduler-count sweep orchestrator
Discriminates herd-contention vs stranding for the ka low-c latency
floor: sweeps URUS_SCHED_THREADS {1,2,4,8} x CONNS {4,8} over the plain
matrix (PLAIN_ONLY; close rides along as control). Per-cell audit
asserts the knob reached the server via plain_serve's stderr line.
Summary parser (req/s + p50/75/90/99, unit-normalized) validated
against the 96b40ad5 sweep output.
2026-07-23 15:00:29 +00:00
Claude 72064c7a79 feat(bench): URUS_SCHED_THREADS knob in plain_serve
E1 (herd-vs-stranding discrimination) sweeps smarm scheduler count at
fixed low concurrency. Strict parse — a malformed value panics instead
of silently reverting to default, and the effective value is logged to
stderr so every bench cell's server.log records it (same per-cell
verification discipline as mode-verify).
2026-07-23 14:59:45 +00:00
claude 14be21db15 feat(bench): PLAIN_ONLY knob — skip causal cells for concurrency sweeps
The c=64 matrix run (job 7dbc1ede) showed plain-mode parity but a 2.2x ka
latency penalty; the decisive test is a CONNS sweep, which only needs the
plain cells. PLAIN_ONLY=1 skips the two causal cells; the summary already
tolerates their absence.
2026-07-20 20:43:32 +00:00
claude 182f4fe602 feat(bench): ka/close A/B matrix — plain_serve baseline + box orchestrator
The jar's ka-vs-close throughput inversion was observed on the box but its
run configuration died with the job workspaces, so it can't be re-derived —
this commit makes the comparison a committed, controlled experiment instead
of a lost one-off.

- examples/plain_serve: the load_profile request path with zero causal
  machinery (same route/handler/Config), serving until killed. The plain
  half of the A/B; throughput measured externally by wrk.
- scripts/ka-close-matrix.sh: {ka,close} x {plain,causal} on fresh ports,
  byte-identical wrk args per mode pair except the Connection: close
  header, disjoint-core pinning for server vs wrk, per-cell curl
  verification of the negotiated connection behavior, ss/env capture, and
  a parsed summary with a close/ka ratio verdict. The causal cells double
  as the forgiveness-fix box validation (v13 pending item, pull-forward
  agreed): key on forgiven staying ~ injected x parked-actors in
  close-mode churn, not process-lifetime phantoms.
2026-07-20 07:01:14 +00:00
Claude 17cd4a5ceb feat(causal): load_profile example — external-loadgen causal target
Server-only discovery tool: real hot path, no planted bottleneck, no
in-process loadgen. wrk (or any external generator, pinned to other
cores) drives GET /json/:id; the sweep runs the causal_bench plan over
the five lib sites. Trust gates only (traffic flowed, ledger printed) —
no known-answer verdict, this one asks the question instead.
2026-07-17 19:36:04 +00:00
Claude 288b52d89c feat(causal): webserver bench with a planted, known-answer bottleneck
examples/causal_bench.rs (required-features smarm-causal): the Phase 2
RFC 007 test case — a real urus server replacing the synthetic demo as
validation workload.

Shape: 16 plain-OS-thread clients hammer GET /order/:id over blocking
loopback keep-alive TCP (OS threads can't absorb injected virtual delay
— the established trick, so only server code is delayable). The handler
calls a single store actor (crud's once-cell pattern, serialization
structural) burning 400µs of calibrated *work* under
causal_site!("store") — work-shaped, not timed, or injection reads as a
no-op. 50µs handler render lands under the enclosing pipeline site.

Known answer: store serialized + saturated => ceiling 2500 rps; speedup
p => x1/(1-p): +33% @25, +100% @50. The five lib sites are tens of µs
and parallel across conns => ~0. Verdict mirrors the demo: store @25 >
+15%, others < +10%, SKIPPED under 4 cores; magnitudes trusted to ~±15%
per the smarm-side validation notes, ranking is the hard check.

1-core sandbox smoke (verdict SKIPPED but numbers indicative): store
+26.2% @25 / +74.0% @50, all other sites within ±2%. Unlike the demo
this workload keeps its bottleneck structure on one core — everything
but the store is IO-parked — and the theory shortfall there is plain
core contention (render/parse share the store's CPU), which the 24-core
sweep should close.
2026-07-13 08:59:49 +00:00
Claude ecaddc579a feat(causal): instrument the request hot path — five sites and a progress point
RFC 007 causal-site coverage for the Phase 2 bench (and any downstream
profiling): parse (head parsing attempts in read_head), router (dispatch
matching only — the winning handler and the next fall-through run outside
the guard, so the site measures dispatch, not what it dispatches to;
call() restructured match-then-dispatch for that, semantics preserved
incl. first-path+method-match-wins and the 405 two-pass), pipeline (the
plug chain inside catch_unwind; nested sites like router take over
attribution during their span, so this reads as chain overhead + handler
code outside inner sites), serialize (response head+bytes-body
serialisation; chunked stream framing is not covered — it interleaves
with writes in pump_stream and the bench doesn't stream), and
socket-write (all of write_all, writability parks included: a park
inside a site is exactly what the park-gated resume credit attributes).

progress!("responses") fires once per response fully on the wire, at
both completion points (the common path and the HTTP/1.0 EOF-stream
early return).

Site guards are #[cfg(feature = "smarm-causal")]-gated: the macros are
no-ops featureless, but binding the unit expansion would trip clippy's
let_unit_value. progress! is bare — its featureless expansion is an
empty block, free and lint-clean. Featureless build byte-behavior is
unchanged; both configs clippy-clean, full suite green both ways.
2026-07-13 08:30:53 +00:00
Claude 072ee126f9 feat: smarm-causal passthrough feature
Mirrors the smarm-trace precedent: urus itself gains no causal code
paths yet — this just lets downstream binaries (the Phase 2 causal
bench lives in examples/) flip smarm's instrumentation on through the
urus dep. Whole suite is green with the feature enabled: the causal
runtime is behaviorally inert while no experiment is active.
2026-07-13 08:25:49 +00:00
Claude 0f824635d1 chore(clippy): appease 1.97 lints — derive Default, collapse ifs, drop needless borrow
Pre-existing, surfaced by the toolchain bump (rust-version is 1.95; the
sandbox gates with stable 1.97). All four are cargo clippy --fix output
with the mechanical-collapse indentation hand-tidied to house style;
the parser change is semantics-preserving (a non-100-continue Expect
value now falls to the _ arm instead of an empty if — headers.append
still runs after the match either way). No fmt pass: rustfmt would
clobber the aligned-assignment style, so only the touched lines moved.
2026-07-13 08:24:31 +00:00
Claude 078072b527 port(smarm): track HEAD efbc254 — RFC 014/015 API sync
- gen_server rename (RFC 015, 3e31606): ServerRef/ServerCtx/ServerBuilder
  -> GenServerRef/GenServerCtx/GenServerBuilder across conn_actor,
  conn_registry, serve, pubsub, channels::session.
- type Timer = () on the three GenServer impls (RFC 015, 57eadb5);
  handle_timer/handle_idle/tick_every stay defaulted — opt-in later.
- Watcher and GenServerCtx are now generic over the server type: watcher
  fields typed Watcher<Table<M>> / Watcher<Registry<P, K>>; Registry's
  struct bounds strengthened to its GenServer impl bounds
  (P: Encode + Decode + Send + Sync + 'static, K: SessionKey) so the
  Watcher field's G: GenServer bound is satisfiable at the declaration.
- RFC 014 (a866e34) registry: register is (Name<M>, Sender<M>), self-only
  — a name is a typed messaging endpoint, not a pid tag. The
  introspection-only urus.server / urus.listener.{i} bindings are dropped
  rather than faked with unit channels; the whereis integration test is
  deleted; a proper messageable-name design is icebox'd in ROADMAP.md.
- Audit vs smarm 6c2b7e9 (queued messages dropped when Receiver drops):
  pubsub's prune-on-send-failure retain still holds — send Errs once
  receiver_alive is false, so a dropped rx prunes on next broadcast,
  exactly what subscriber_count's doc already promised. Freeing stranded
  Arc<M> broadcasts is strictly good. No change needed.
- The 7x E0283 in channels/mod.rs were cascade fallout of the generics
  changes; dissolved with the port, as discovery predicted.

Suite: 90 lib + 45 integration + 2 doc, green under default, smarm-trace,
phoenix, and all-features. 3 pre-existing clippy lints (conn.rs,
parser.rs, conn_actor.rs; clippy 1.97 strictness) deferred to a follow-up
chore commit to keep this diff pure.
2026-07-13 08:22:40 +00:00
Claude 86c8d31b93 docs(channels): example, README section, ROADMAP v0.6 mark-done — chunk 4
examples/channels_chat.rs (required-features phoenix): ephemeral
room:* and persistent session:* side by side over phoenix.js V2 JSON,
stdin-driven graceful shutdown (the ws_chat pattern). Smoke-tested
live: rejoin after transport drop lands on the same instance (state
counter persists across three transports), broadcasts carry the new
join generation's ref.

README: channels section (feature split, core trait walkthrough,
session persistence semantics incl. the re-entrant-join contract).
ROADMAP: v0.6 marked done with the full as-landed decision list,
chunked as committed.
2026-06-12 15:26:21 +00:00
Claude 0de2baa72d feat(channels): opt-in session persistence — v0.6 chunk 3
ChannelSession<P> (src/channels/session.rs): a channel actor that
outlives its transport, buffering outbound broadcasts as the same
Arc<Broadcast<P>>s the relay path carries (zero re-allocation) until
rejoin, buffer cap, or TTL. Registered per pattern via
PrefixRouter::channel_session::<S>(pattern, factory) /
channel_session_default::<C>(pattern); one registry gen_server per
registration, monomorphic over S::Key (no type erasure).

Machinery: deploy() computes the key on the conn actor and casts the
join handshake to the registry; the registry maps key -> (pid, control
sender), forwards reattaches, spawns-and-monitors fresh actors, and
prunes on Down (pid-guarded against replaced entries). The session
actor selects [inbound, bus, ctl] attached and select_timeout([ctl,
bus], ttl-remaining) detached; detach triggers are the closed inbound
arm and ws send failure (the failed broadcast is the buffer's first
entry — nothing lost). Drain re-encodes under the new join generation.

As-landed decisions (each in module docs):
- ChannelFactory gained a #[doc(hidden)] deploy() seam (default =
  ephemeral linked actor, the chunk-1 behavior verbatim); JoinHandshake
  and ChannelInbox are opaque pub structs. Channel construction moved
  from run_channel into deploy — same linked blast radius.
- Every attach calls ch.join() again (rejoin acks need a payload and
  channels get to re-auth); session channels must treat join as
  re-entrant.
- A rejected rejoin ENDS the session (no zombie sessions held open for
  unauthorized clients); next join is a cold start.
- A second transport EVICTS the first (best-effort phx_close); a
  session is single-transport.
- Explicit leave ends the session, not just the attachment.
- TTL per detach episode; subscribe once, on first successful join;
  cap.max(1) semantics (cap 0 = first buffered message tears down,
  but an empty buffer still waits out the TTL).
- Key: Clone beyond the ratified Eq+Hash+Send — the registry keeps a
  pid -> key reverse index for Down pruning.
- The registry spawns eagerly at registration (channel_session is
  in-runtime-only, same law as ChannelHub::new): a lazy OnceLock would
  block std-sync under green threads on first-join races.
- Shutdown composes WITHOUT links: conns die -> router Arcs drop ->
  registry inbox closes -> registry exits -> ctl senders drop ->
  detached sessions wake on the closed ctl arm, terminate, exit.
  shutdown_with_detached_session_terminates is the proof (ttl 30s vs
  5s budget — only the chain can reap it).

Tests: 6 session unit tests (buffer-and-drain ordering across
reconnect, TTL expiry, cap teardown, leave, eviction incl. stale
conn-entry self-prune, rejected rejoin) + the shutdown integration
test. Shared toy codec + conn harness extracted to
channels/testkit.rs (mod.rs tests now import it; no behavior change).

Hammer: 35x lifecycle subset (+channels +session) + 3x full (phoenix)
+ 1x full (phoenix+smarm-trace) + featureless + channels-only, all
green; dbg-grep clean. Suite: 90 unit + 46 integration + 2 doc under
--features phoenix.
2026-06-12 15:22:54 +00:00
Claude 39e4f70e9d feat(channels): phoenix V2 JSON codec — v0.6 chunk 2
Json<T> newtype codec (src/channels/phoenix.rs), feature "phoenix":
speaks the Phoenix V2 wire format [join_ref, ref, topic, event, payload]
for any T: Serialize + DeserializeOwned. Replies encode as phx_reply
with the {"status", "response"} wrapper; payload-less control frames
encode payload as {}.

Deviation from roadmap sketch (documented in module docs): the roadmap
sketched a blanket Encode/Decode impl on bare P: Serialize. That is
coherence poison — with additive features, enabling "phoenix" anywhere
in the crate graph would forbid every user-written Encode impl. The
blanket therefore lives behind the Json<T> newtype: zero-boilerplate
(Deref/DerefMut/From), but opt-in per payload type.

Encode failure panics and rides the linked-teardown contract (channel
actor is linked to the conn; a payload that cannot serialize is a
programming error, not a protocol event). Decode strictness is exactly
T's Deserialize — use serde defaults or Json<serde_json::Value> for
lenient channels.

Integration (tests/integration.rs, cfg(feature = "phoenix")):
- channels_join_heartbeat_event_broadcast_leave: join ack ordering,
  conn-side heartbeat, unrouted-topic error, join rejection, broadcast
  incl. sender, correlated echo reply, leave -> ok + phx_close, and
  post-leave isolation.
- channels_codec_garbage_closes_1002: undecodable text frame closes
  the socket with protocol-error status.
- shutdown_with_open_channel_terminates: drain force-stop tears down
  the linked channel actor and the runtime reaches AllDone in budget.

Hammer: 35x lifecycle subset (+channels filter) + 3x full (phoenix) +
1x full (phoenix+smarm-trace) + featureless + channels-only, all green.
Suite: 84 unit + 45 integration + 2 doc under --features phoenix.
2026-06-12 14:56:26 +00:00
Claude 0531390613 feat(channels): core protocol machinery — v0.6 chunk 1
Feature 'channels' (zero new deps): wire-neutral ChannelFrame/FrameRef +
Encode/Decode codec traits (one method each, whole-envelope — the codec
owns the wire format end to end); Channel trait; ChannelFactory +
TopicRouter + shipped PrefixRouter (exact map + head:* prefix map);
ChannelSocket (reply/push/broadcast/broadcast_from/broadcast_to);
ChannelHub (router + PubSub<Broadcast<P>>, in-runtime + non-static per
the v0.5 pattern, hub.upgrade(conn) entry, server-side hub.broadcast).

Topology as ratified: one channel actor per (socket, topic), linked to
the conn actor, select(&[&inbound, &bus_rx]) with inbound at index 0;
subscription pinned to the channel actor pid. Subscribe-before-ok-reply
makes the join ack first-class (the v0.5 on_open race, answered
structurally). Cleanup is the sender-drop chain (conn handler drops ->
inbound closes -> closed arm wakes -> terminate); bus arm dropped from
the select set once observed closed (closed arms stay ready forever).
Heartbeat answered conn-side on topic 'phoenix'; rejoin replaces the
old generation with a phx_close to the old join_ref; lazy prune of dead
channel entries on first failed forward (the pubsub discipline).

As-landed deviations from the ratified sketch (module docs, veto by
diff): join takes &mut self (trait-object callable; construction moved
to ChannelFactory which still gets the full raw topic); reply(status,
payload) with the ref threaded invisibly (the sketch listed both);
broadcast gained the event parameter (unencodable without it); route()
returns Arc not Box. Outbound payloads encode from borrows (FrameRef)
so bus fan-out never clones P; payload-less control frames
(phx_close, error/heartbeat replies) encode as the codec's empty
payload rather than conjuring a P.

11 runtime-backed unit tests with a toy pipe codec (core provably
JSON-free). Suites: 78 w/ feature, 67 featureless, both green.
2026-06-12 13:53:52 +00:00
Claude a676c01adc docs(roadmap): ratify v0.6 Channels design
User-ratified 2026-06-12. Resolves the v0.6 fork left by the v0.5
handoff: dep #4 is serde/serde_json, but ONLY behind an opt-in
'phoenix' feature (V2 codec + JsonCodec). The 'channels' core takes
zero new deps — payload generic over urus Encode/Decode traits.
Channel trait shape, ChannelSocket, TopicRouter/PrefixRouter, and
opt-in ChannelSession persistence (zero-copy Arc<M> buffer drain)
all specified in the v0.6 section.
2026-06-12 13:42:35 +00:00
Claude ad19848db3 feat(ws+pubsub): on_open hook + ws_chat — v0.5 chunk 2
WsHandler::on_open(&mut self, sender) — defaulted, non-breaking,
added mid-chunk for veto-by-diff (the v0.3 "push" precedent for
defaults-level decisions): without it a listen-only ws client can
never be subscribed, since on_message never fires for a client that
doesn't send. Runs exactly once, first, in the conn actor, before any
buffered pipelined frame is decoded. Panic contract identical to
on_message: check_cancelled re-raise dance, else 1011 and no
on_close.

DISCOVERY 1 (documented, not fixed — inherent): the 101 reaches the
client BEFORE on_open runs, so "is my subscription live yet" is a
race client-side. App-level ack is the answer (any reply proves
on_open completed; callbacks are sequential). The integration tests
hit this immediately (first read of a join notice flaked into a 5s
read-timeout) and use a sync/synced ack; same reason Phoenix joins
reply.

DISCOVERY 2 (the structural one): PubSub::new() is in-runtime only,
but pipelines are built pre-runtime. crud's static-OnceLock bootstrap
DOES NOT COMPOSE with serve_with_shutdown here: a static pins the
table actor's last ServerRef forever, its inbox never closes, the
gen_server never exits, AllDone is unreachable — serve hangs (the
cross-thread-unpark limitation closes the workaround routes). The
pattern that works: NON-static Arc<OnceLock<PubSub<M>>> captured by
the route closure. Drain stops conns+listeners -> last Arc<Pipeline>
drops IN-runtime -> cell+handle drop -> inbox closes -> table exits.
Corollary: relays hold the Receiver only, NEVER a PubSub clone
(relay holds table's inbox open, table holds relay's receiver open:
mutual keepalive, shutdown hangs). Both rules in module docs, README,
and enforced by the new shutdown_with_open_chat_terminates test:
two open chat sockets + live relays + live table, shutdown must
return within 5s.

examples/ws_chat.rs: rooms as topics (room:{name}), on_open
subscribes (conn-actor pid: the monitor scopes cleanup to the
session) + spawns the relay (rx -> WsSender clone; exits on prune or
WsClosed), on_message broadcast_from(self_pid()) so the speaker is
never echoed, on_close broadcasts the leave notice and relies on the
monitor for unsubscribe.

Integration: ws_chat_broadcast_reaches_other_client_not_sender
(no-echo proven orderingly, no timeout reads: B speaks after A, A's
next frame must be B's) + shutdown_with_open_chat_terminates.

hammer.sh default subset now includes ws_ (v0.4 ran it ad hoc; ws IS
conn-lifecycle). Hammered: 35x subset (incl. both new tests) + 3x
full + 1x full under smarm-trace, all green. dbg!-grep clean. smarm
PRISTINE. README pub/sub + on_open sections; ROADMAP v0.5 as-landed.

Suite: 67 unit + 42 integration + 2 doc.
2026-06-12 10:30:08 +00:00
Claude 341d20ab45 feat(pubsub): topic table gen_server — v0.5 chunk 1
urus::pubsub, phoenix_pubsub-shaped, local node only. Independent of
HTTP (imports smarm only); crate-extraction candidate.

Design as ratified (handoff Q1-Q5, user said "push" on the leans):
- Q1 generic: PubSub<M> per instance, zero-cost, no downcasts.
  Heterogeneous events = app-side enum M. v0.6 channels build on it.
- Q2 Receiver: subscribe(topic) -> Receiver<Arc<M>>, fresh channel per
  subscription, subscriber owns its loop. Fan-in via a Sender-passing
  subscribe_with is a compatible later addition, deliberately not v1.
  subscribe_as(pid, topic) pins the subscription lifetime to an
  explicit pid for relay patterns (session actor subscribes, spawned
  relay receives) — without it the monitor would watch the wrong actor.
- Q3 unbounded: broadcast never blocks the table (bounded-block =
  head-of-line across ALL topics; bounded-drop = silent loss). Memory
  risk documented; TWO cleanup paths: monitor Down on subscriber death
  (eager) + prune-on-send-failure for dropped receivers (lazy, on next
  broadcast). Bench before sharding.
- Q4 user-owned ServerRef wrapper. DISCOVERY: the ratified optional
  register-by-name helper is NOT implementable against smarm's
  registry — it maps name -> Pid, and a Pid cannot be turned back into
  a ServerRef (the ref is the inbox sender). Needs smarm support or a
  global type-erased map; deferred, noted in module docs. pid()
  exposed for introspection.
- Q5 unique per (pid, topic): HashMap<Topic, HashMap<Pid, Sender>>,
  subscribe idempotent (replaces sender; stale receiver's channel
  closes). One monitor per live pid (monitored: HashSet<Pid>, retired
  in handle_down so slot-reuse pids re-monitor). handle_down does the
  full-scan cleanup as ratified; pid->topics reverse index is the
  documented optimization if deaths ever measure hot.

Subscribe is a call (table updated before return: a broadcast issued
right after by the same caller is seen); unsubscribe/broadcast are
casts. subscriber_count(topic) added as a call — needed by the tests
to observe async cleanup, legitimate API anyway.

8 runtime-backed unit tests (smarm::run + collect-outside-assert-after
pattern from smarm's own gen_server tests), including: Arc payload
ptr-equality across subscribers, monitor-path pruning with NO
broadcast issued (isolates it from the lazy path), broadcast_from
self-skip, resubscribe replacement.

Suite: 67 unit + 40 integration + 2 doc, green. smarm PRISTINE.
2026-06-12 10:17:50 +00:00
Claude ffc579c500 feat(ws): duplex wiring + handler API + echo example (v0.4 chunk 3)
Topology (b) as agreed: ONE actor per connection via smarm RFC 008 fd
arms. After the 101, ws::duplex::run_duplex replaces the HTTP loop:
try_select(&[&outbound_rx, &FdArm::readable(fd)]) per iteration,
outbound at index 0 (priority order = owed writes drain before reads;
honest backpressure). Accepted gap documented: mid-write of a large
outbound frame the actor isn't reading.

Handler API (deferred questions, as answered this session):
- WsHandler { on_message, on_close } runs IN the conn actor's select
  loop; concurrency = spawn an actor with a WsSender clone (the SSE
  producer pattern). Conn::upgrade(self, handler) flat signature;
  WsUpgrade grows from marker to Box<dyn WsHandler> payload.
- WsSender: Clone; send/text/binary/ping/close -> Err(WsClosed).
  Unbounded channel like SSE; slow client bounded by write_timeout.
- Control frames invisible v1: auto-pong inline (5.5.2), peer close
  auto-echoed (5.5.1; server closes TCP first per 7.1.1) surfacing
  only as on_close(Some(code), reason); wordless endings (EOF, write
  failure, drain timeout) = on_close(None, ""). on_ping/on_pong can
  land later as default methods, non-breaking.
- Server-initiated close (WsSender::close, FrameError->1002/1009/1007,
  handler panic->1011): close out, bounded drain (one write_timeout
  budget) for the peer echo, data discarded (1.4). Handler panics
  re-raise smarm's stop sentinel first (the pipeline catch_unwind
  dance, replicated).
- Caps: Config.max_frame_payload (1 MiB) / max_message_bytes (4 MiB).

Conn actor: 101 branch now hands fd + leftover buf (pipelined first
frame carries over; tested) + boxed handler to run_duplex; registry
entry stays Busy for the ws lifetime, so graceful shutdown force-stops
the conn out of the select park at the drain deadline (tested — the
in-runtime request_stop wake covers select parks too).

Tests: chunk-1 EOF test rewritten into a 9-test duplex suite (echo,
pipelined first frame, ping/pong, both close directions, 1002 unmasked,
1009 header-cap, fragmentation with interleaved ping, shutdown
force-stop). Suite 59u+40i+2d. Hammer: 35x lifecycle+ws subset + 3 full
+ 1 full under smarm-trace, all green; no AlreadyExists out of
try_select (the fresh eager-cleanup path held). Validated against
python websocket-client (echo, pong payload, close 1000).
examples/ws_echo.rs: /echo in-actor + /clock producer-spawning.
2026-06-12 09:44:35 +00:00
Claude 64137cccf0 fix(http): advertise keep-alive on HTTP/1.0 responses we keep open
1.0 defaults to close: honoring 'Connection: Keep-Alive' without echoing
it back left spec-following clients waiting for an EOF that never came
(ab -k deadlocked to its poll timeout on response 1).

As agreed:
- serialise_response echoes 'connection: keep-alive' when keep_alive
  is on AND version is 1.0. 1.1 keep-alive stays implicit.
- Error statuses (>=400) on 1.0 force keep_alive off in the conn actor
  (joins the existing 1.0-stream force-close): we don't advertise reuse
  on a 4xx/5xx to a protocol generation where reuse is opt-in.
- HEAD needs nothing: the echo is framing-neutral.
- Parse-error responses already carried 'connection: close'.

Tests: unit pair (1.0 echo / 1.1 implicit) + integration pair (socket
actually reused on 1.0; 404 forces close + EOF). Validated against the
original repro: ab -k -n 5000 -c 50 GET /users -> 5000/5000 keep-alive,
0 failed, ~63k rps (~18.9k without keep-alive).
2026-06-12 09:15:11 +00:00
Claude c171be7ddc chore: scripts/hammer.sh — flake-hammer for conn-lifecycle discipline
Default run encodes the pre-commit ritual: 35x the lifecycle-sensitive
integration subset (shutdown timeout reaped slowloris streaming chunked
sse stalled), fail-fast, then 3x the full suite. -n/-f override counts,
'-- <words>' overrides the filter. Failing run's output lands in
target/hammer-fail.log.
2026-06-12 09:10:17 +00:00
Claude f5ccd640fb docs(roadmap): known bug — HTTP/1.0 keep-alive honored but not advertised
Found benching with ab -k: a 1.0 request with Connection: Keep-Alive is
kept alive (second pipelined request answered on the same socket) but
the response never echoes the header. 1.0 clients only reuse on explicit
server opt-in, so they wait for EOF that never comes; ab -k deadlocks to
apr_pollset_poll timeout on the first response. Recorded with repro,
fix direction (echo header in serialise_response on live 1.0 conns, or
force-close 1.0), and the no-keep-alive baseline: GET /users, release,
loopback, logging silenced, c=50 -> ~18.9k req/s, p50 3ms / p99 5ms,
0 failed. Fix deliberately not started: scope (error responses, HEAD)
is a user call.
2026-06-12 06:58:31 +00:00
Claude d96ba88ec8 docs(roadmap): v0.4 chunks 1-2 as-landed, dep #3 spent on sha1_smol 2026-06-12 05:41:44 +00:00
Claude 286915329a feat(ws): RFC 6455 frame codec — decode/encode, close payloads, assembler (v0.4 chunk 2)
Pure module (src/ws/frame.rs): bytes in, frames out, zero io/actor
machinery, so the whole RFC 6455 §5 edge surface lives in fast unit
tests ahead of the chunk-3 conn-lifecycle wiring.

- decode(buf, require_masked, max_payload) -> Ok(Some((frame,
  consumed))) | Ok(None: incomplete) | Err(fatal). Server mode requires
  client masking (§5.1); the flag keeps the codec reusable for a client
  mode. Enforced: RSV=0, reserved opcodes, control frames FIN+<=125,
  MINIMAL length encodings (16/64-bit), 64-bit MSB clear. The payload
  cap is checked from the header BEFORE the payload is buffered — a
  hostile 8-byte length can't make the conn actor allocate toward it.
- Frame::encode is unmask-only (server MUST NOT mask); encode_masked
  exists for the client side of tests + future client mode.
- Close payloads: parse_close_payload validates the §7.4 wire-code
  ranges (1004/1005/1006/1015 and <1000/1012-2999/>=5000 rejected,
  3000-4999 app space allowed), 1-byte payload rejected, UTF-8 reason
  required; close_payload(code, reason) truncates the reason to the
  125-byte control cap on a char boundary.
- Assembler (data frames only; control frames interleave at the caller
  per §5.4): orphan continuation and mid-fragmentation data frames are
  protocol errors, message cap spans fragments (and the unfragmented
  fast path), text UTF-8 validated once on completion (whole-message
  delivery makes incremental validation pointless), reusable after
  every message/error.
- FrameError -> close_code mapping: Protocol=1002, TooLarge=1009,
  BadUtf8=1007.

Tests pin all five §5.7 worked examples (masked Hello bytes verbatim),
round-trip the 125/126 and 65535/65536 boundaries, and prove decode
returns None at EVERY proper prefix of a frame. 57 unit tests total.
2026-06-12 05:40:54 +00:00
Claude 49538ccc83 feat(ws): RFC 6455 opening handshake — Conn::upgrade, 101 short-circuit (v0.4 chunk 1)
Design (as agreed in the v0.4 pass):
- Dep #3 spent: sha1_smol 1.0.1 (zero transitive deps) instead of a
  vendored SHA-1 — user call: no hand-rolled crypto. Consequence noted
  in ROADMAP: the v0.6 wire-format JSON question now costs dep #4.
- Base64 stays in-tree but ENCODE-ONLY (RFC 4648 vectors tested): it is
  an encoding, not crypto, and the decode direction (where parsing bugs
  live) is deliberately not implemented.
- src/ws/{mod,handshake}.rs: validate() implements §4.2.1 (GET, 1.1,
  Host, upgrade/connection token lists case-insensitively, key shape =
  base64 of exactly 16 bytes, version 13). Version mismatch -> 426 +
  sec-websocket-version: 13 (§4.2.2); everything else -> 400. Rejected
  upgrades stay plain HTTP with keep-alive intact (tested: 400 then 101
  on the same connection).
- Conn grows pub upgrade: Option<WsUpgrade> (opaque marker; chunk 3
  turns it into the duplex/handler handoff payload — shape deliberately
  uncommitted while the handler API is open). Conn::upgrade(self) is the
  commit point: 101 + accept key + marker + halt, or the rejection.
- conn_actor: upgrade marker + status 101 -> write head, leave the HTTP
  loop. Until chunk 3 that drops the fd (clients see 101 then EOF);
  status check is defensive against a post-handler plug clobbering 101.
- serialise_response: 1xx no longer get content-length injected
  (RFC 7230 §3.3.2). 204 left alone on purpose — separate conversation.

Accept-key path pinned to the §1.3 worked example end-to-end (unit +
wire). Suite: 34 unit + 33 integration + 2 doc.
2026-06-12 05:38:14 +00:00
Claude 43bd5a91a4 docs: README streaming/SSE sections, ROADMAP v0.3 done, crud SSE ticker (v0.3 chunk 4)
README: write_timeout in the Config sample, read/write timeout-split
semantics spelled out, new Streaming Responses and Server-Sent Events
sections (producer contract, 1.0 fallback, heartbeat liveness), roadmap
summary bumped. ROADMAP: v0.3 marked done with the as-landed decisions
(pull shape per the user's fork answer; Q3 timeout split as proposed)
and verified exit criteria. crud example: GET /ticker SSE route — one
event/second, producer exits on SseClosed or the example's shutdown
flag (so an open ticker doesn't block AllDone past the drain).
2026-06-12 04:59:39 +00:00
Claude d14fb7c331 feat(sse): Conn::sse() sugar — EventSender + pump heartbeats (v0.3 chunk 3)
New src/sse.rs. Conn::sse() (and sse_with_heartbeat) turns the response
into a text/event-stream: status 200 if unset, content-type +
cache-control: no-cache, and a RespBody::Stream whose heartbeat is wired
to the new StreamBody.heartbeat field. Returns (Conn, EventSender);
EventSender is Clone (stream ends when the LAST clone drops) with
send(event, data) / data(data) / comment(text), all returning
Err(SseClosed) once the conn actor has dropped the stream — the producer
exit signal. Multi-line data becomes one data: line per line (pure
format_event, unit-tested).

Heartbeat lives in the pump, as the roadmap leaned: with
StreamBody.heartbeat set, pump_stream waits recv_timeout(interval) and on
expiry writes the payload (": keep-alive" comment chunk) and keeps
waiting. Liveness inversion is deliberate: no request clock governs an
SSE response (request_timeout is read-phase only since chunk 1); a dead
client is detected when an event or heartbeat write stalls past
write_timeout, and a draining shutdown force-stops the recv park like any
stream.

Tests: 4 unit (3 framing, 1 sse() wiring) + 2 integration (event round
trip over decoded chunked stream incl. named + unnamed events and clean
0-chunk EOS on sender drop; >=2 heartbeats observed across 450ms of
producer silence at a 100ms interval).
2026-06-12 04:58:13 +00:00
Claude 4482f47265 feat(parser,conn): chunked request bodies (v0.3 chunk 2)
parser: Transfer-Encoding: chunked no longer 411s — it sets
ParsedHead.chunked and the conn actor decodes. Two hard rejections at
parse time, both Malformed/400: chunked together with Content-Length
(the request-smuggling ambiguity; RFC 7230 §3.3.3 permits rejection)
and chunked on HTTP/1.0 (TE is a 1.1 mechanism). ParseError::Unsupported
is now unconstructed; kept for future framings.

conn_actor: read_chunked_body decodes incrementally from buf[head_len..],
reading on the SAME request deadline the head came in under (a stalled
chunked body dies at request_timeout exactly like a stalled CL body).
Chunk extensions are ignored; trailers are consumed and discarded.
Bounds: decoded size capped at max_body_bytes -> 413 the moment the cap
would be crossed (the CL pre-check can't see a chunked body's size up
front); size lines capped at 128 bytes and the trailer section at 8 KiB
-> 400, so framing spam can't grow the read buffer unboundedly.

Returns (decoded, raw_consumed_past_head): the keep-alive drain at the
loop bottom now drops head_len + RAW framing length (not decoded length),
so pipelined requests behind a chunked body land exactly — covered by
the trailers+pipelining test.

Tests: 3 parser unit (flagging, CL+TE reject, 1.0 reject; the old
chunked-is-Unsupported test replaced) + 6 integration (decode with
extensions, trailers + pipelined next request, 413 over decoded cap,
400 malformed size line, 400 CL+TE on the wire, stalled chunked body
killed at request_timeout with silent close).
2026-06-12 04:56:09 +00:00
Claude 42a2464743 feat(conn): streaming response bodies — RespBody::Stream + chunked TE (v0.3 chunk 1)
Design decisions (per handoff Q1 user answer + Q3 proposal):

- Producer shape is PULL: the handler returns a smarm Receiver<Vec<u8>>
  (RespBody::Stream / StreamBody, From<Receiver<Vec<u8>>> for ergonomics);
  the producer is an actor the handler spawned. The conn actor pumps in
  pump_stream(): it keeps sole ownership of the socket and of write
  deadlines, and the recv() park between chunks is stoppable by the
  draining registry's request_stop (stop sentinel unwinds out of
  park_current; fd + registry guards clean up) — so an infinite stream is
  force-stoppable at the drain deadline like any in-flight request, with
  no polling. End of stream = every Sender dropped = terminating 0-chunk.
  Producer contract: a send() error means the conn died; exit.

- Framing: HTTP/1.1 gets transfer-encoding: chunked and the connection
  stays reusable after the terminator (keep-alive after chunked). HTTP/1.0
  has no chunked TE: bytes go raw, keep-alive is forced off, EOF delimits.
  User-set content-length/transfer-encoding headers are DROPPED for stream
  bodies — we own the framing, and CL+chunked is a smuggling vector.
  Empty producer chunks are skipped (a 0-length chunk would terminate the
  framing early).

- request_timeout stays READ-phase only (unchanged). NEW Config/ConnLimits
  field write_timeout (default 30s) gives every response write a per-write
  budget: the fixed head+body write, each streamed chunk, and the error
  paths (413/100-continue/4xx) all go through the now deadline-bounded
  write_all (wait_writable_timeout, mirroring read_some). Each chunk gets
  a FRESH budget — streams may outlive any whole-response clock; a single
  stalled write may not (write-side slowloris). Naming/default were
  flagged as a user call in the handoff: veto here if write_timeout(30s)
  isn't it.

- RespBody loses derive(Clone) (Receiver isn't Clone; Clone was unused)
  and gets a manual Debug.

Tests: 3 serialiser unit tests (chunked head shape, user-framing-header
stripping, 1.0 fallback) + 5 integration (chunked round-trip with decoder,
keep-alive after chunked, 1.0 EOF-delimited, shutdown force-stops an
infinite stream at drain deadline, stalled reader killed at write_timeout
with producer observing the closed channel).
2026-06-12 04:53:27 +00:00
Claude 8065ea561e feat(serve): registry names + docs (v0.2 chunk 4)
Registration:
- Each listener registers "urus.listener.{i}" from inside its ChildSpec
  factory. On a restart the stale binding points at a dead pid; smarm's
  registry evicts those lazily on register/whereis, so re-registering
  the same name is safe. Result ignored — a registry hiccup must not
  take the listener down.
- "urus.server" -> supervisor pid, registered via the JoinHandle's
  .pid() immediately after spawn rather than inside the closure: the
  binding exists before the supervisor runs an instruction, so an early
  whereis can't observe None.

Test: server_and_listeners_are_registered probes whereis from a route
handler — whereis must run inside the runtime, and the test thread is a
foreign OS thread with no runtime in its TLS.

Docs:
- README: Config documented in full (timeout semantics: keep_alive =
  idle-before-first-byte, request = Instant deadline first-byte ->
  head+body, slowloris-proof), Handle/shutdown_handle/serve_with_shutdown
  section with the 4-step shutdown sequence, named-actors note, test
  coverage list refreshed, Roadmap section now points at ROADMAP.md.
- ROADMAP: v0.2 chunks 1-4 marked landed; design detail stays in the
  chunk commit bodies.

Verified: cargo build --features smarm-trace clean; full suite (32
tests) green 3x; debug-grep clean.

This closes v0.2. Exit criteria all verified across chunks 1-3: crud
drains on stdin-Enter (~100ms), idle keep-alive reaped at
keep_alive_timeout, panicking listener restarts under load.
2026-06-11 22:24:02 +00:00
Claude a60687bfec feat(conn): keep-alive and per-request read timeouts (v0.2 chunk 3)
Two wall-clock budgets in the conn read path, both via
smarm::wait_readable_timeout:

- keep_alive_timeout: idle budget between requests — how long we park
  waiting for the FIRST byte of a request (including the first request on
  a fresh connection). Expiry closes silently; nothing is owed to a
  client that isn't talking.
- request_timeout: per-request budget as an Instant deadline from the
  first byte of a request until the request (head + body) is fully read.
  Pipelined leftovers in the buffer at cycle start count as a started
  request. The deadline is returned from read_head and threaded into
  read_body so head and body share one budget. Pipeline run time is NOT
  covered. Each read wait uses whichever budget is currently active;
  deadlines are Instants, so EAGAIN retries don't reset the clock.

Expiry mid-head gets a best-effort 408 via try_write_once: a single
non-parking write syscall — a client that stalls its read side must not
defeat the timeout by making the 408 write park forever. Expiry mid-body
just closes.

examples/crud.rs gains stdin-Enter graceful shutdown (plain OS thread on
read_line -> handle.shutdown(); no signal crate). Doing so surfaced a
pattern worth knowing: crud's lazily-spawned store actor parked forever
in recv(), which blocks smarm's AllDone, so serve_with_shutdown never
returned (gdb: scheduler idle in poll_wake, store the only live actor).
It can't be messaged awake from the stdin thread either — a cross-thread
send's unpark is a no-op without runtime TLS, the same limitation behind
SHUTDOWN_POLL. Fix: store_loop recv_timeout(250ms) + a SHUTTING_DOWN
AtomicBool set by the stdin thread before handle.shutdown(). This poll
dies with the cross-thread-unpark limitation; recorded in smarm docs.

Tests: idle keep-alive conn reaped at a small keep_alive_timeout;
slowloris partial-head stall killed at request_timeout with the 408
observed. Timeout+shutdown subset hammered 30x clean; full suite 3x;
smarm-trace build and suite clean.
2026-06-11 22:07:18 +00:00
Claude c658ac06b2 feat(serve): graceful shutdown + connection registry (v0.2 chunk 2)
Shutdown is drain-then-force-stop: flag.store(true) -> sup_h.join() (the
no-more-accepts barrier) -> cast BeginDrain -> poll ConnCount every 50ms
until 0 or drain_timeout; past the deadline each tick casts ForceStopConns
and re-sweeps, so conns that registered after the deadline are still
caught.

Listener shutdown is redesigned vs chunk 1: Restart::Permanent ->
Transient, wait-failure return -> panic (Transient restarts it), and a
shared Arc<AtomicBool> flag checked at loop top with a 250ms
wait_readable_timeout tick instead of request_stop — no request_stop on
the supervisor or listeners anywhere. Historical note: the redesign was
originally forced by smarm's then-lossy stop against QUEUED actors (fixed
upstream in 7bab4d2); it is kept because flag-based listener shutdown is
simpler and stop-semantics-free.

Conns self-register with the registry (ConnStarted from the conn actor,
not a listener cast): per-sender FIFO ordering is the only thing that
prevents ConnEnded overtaking ConnStarted; see conn_registry module docs.

conn_actor: in the catch_unwind(pipeline.run) Err branch,
smarm::preempt::check_cancelled() runs BEFORE composing the 500 — smarm's
stop is an undowncastable panic_any(StopSentinel), so the catch_unwind
would otherwise swallow a stop and 500-and-keep-running. The stop flag is
persistent, so check_cancelled re-raises cleanly.

Handle.shutdown() wakes the root via a 100ms try_recv poll
(SHUTDOWN_POLL): a cross-thread Sender::send's unpark is a
try_with_runtime no-op without runtime TLS on the sending thread (still
true on smarm 8e5b754; cross-thread unpark is a recorded smarm roadmap
candidate — this poll dies with it).

Two smarm bugs were found during this chunk and fixed upstream: lossy
stop against QUEUED actors (7bab4d2) and the terminal-wake shutdown
stall/hang (eddf3fe); post-mortem in artefact smarm-bug-terminal-wake.md.
2026-06-11 21:45:11 +00:00
Claude 5fe696992a feat(serve): supervised listener pool (v0.2 chunk 1)
OneForOne supervisor on the root actor, Restart::Permanent per listener.
ChildSpec factory owns its fd via Arc<OwnedFd> (OwnedFd gains Sync);
restarts reuse the fd — no re-dup, no fd-less window. wait_readable
failure now exits-and-restarts instead of being fatal. Conn actors stay
unsupervised bare spawns. Test hook INJECT_LISTENER_PANICS panics before
accept so a pending connection must survive into the restarted listener.

Roadmap: chunk 1 marked done; chunk 3 fork resolved to smarm-native
timed waits (RFC 008 landed at 393cdd0, reaper option removed); chunk 2
drain-then-stop decided; v0.4-3(b) unblocked.
2026-06-11 12:07:19 +00:00
35 changed files with 10301 additions and 362 deletions
+1
View File
@@ -1,2 +1,3 @@
target target
Cargo.lock Cargo.lock
smarm_trace.json
+33 -3
View File
@@ -1,22 +1,39 @@
[package] [package]
name = "urus" name = "urus"
version = "0.1.0" version = "0.3.0"
edition = "2021" edition = "2021"
rust-version = "1.95" rust-version = "1.95"
description = "Cowboy/bandit-style HTTP library for the smarm actor runtime" description = "Cowboy/bandit-style HTTP library for the smarm actor runtime"
license = "MIT"
[dependencies] [dependencies]
smarm = { path = "../smarm" } smarm = { git = "https://git.kalsbeek.dev/Markk116/smarm", tag = "v0.7.0" }
httparse = "1.9" httparse = "1.9"
libc = "0.2" libc = "0.2"
sha1_smol = "1"
# dep #4, ratified 2026-06-12: serde/serde_json behind the opt-in
# "phoenix" feature only — the "channels" core stays dependency-free.
serde = { version = "1", optional = true, features = ["derive"] }
serde_json = { version = "1", optional = true }
# config-file feature: TOML loader for tuning knobs (dep #5, 2026-08-12)
toml = { version = "0.8", optional = true }
[features] [features]
smarm-trace = ["smarm/smarm-trace"] smarm-trace = ["smarm/smarm-trace"]
smarm-causal = ["smarm/smarm-causal"]
channels = []
phoenix = ["channels", "dep:serde", "dep:serde_json"]
config-file = ["dep:serde", "dep:toml"]
[dev-dependencies] [dev-dependencies]
serde = { version = "1", features = ["derive"] } serde = { version = "1", features = ["derive"] }
serde_json = "1" serde_json = "1"
[[example]]
name = "causal_bench"
required-features = ["smarm-causal"]
[profile.dev] [profile.dev]
panic = "unwind" panic = "unwind"
@@ -28,3 +45,16 @@ codegen-units = 1
[[example]] [[example]]
name = "crud" name = "crud"
path = "examples/crud.rs" path = "examples/crud.rs"
[[example]]
name = "channels_chat"
path = "examples/channels_chat.rs"
required-features = ["phoenix"]
[[example]]
name = "serve_toml"
required-features = ["config-file"]
[[example]]
name = "load_profile"
required-features = ["smarm-causal"]
+21
View File
@@ -0,0 +1,21 @@
MIT License
Copyright (c) 2026 Mark Kalsbeek
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
+360 -15
View File
@@ -122,21 +122,365 @@ Handlers are closures taking `(Conn, Next) -> Conn`. Call `Next::call(c)` to con
### Server Configuration ### Server Configuration
`serve(addr, pipeline)` binds and listens on the given address. For more control, use `serve_with(config, pipeline)`: `serve(addr, pipeline)` binds and listens on the given address. For more
control, use `serve_with(config, runtime_config, pipeline)` — the urus
`Config` holds endpoint knobs, the `smarm::Config` holds runtime knobs
(they are separate because an endpoint placed in someone else's tree
cannot dictate the runtime):
```rust ```rust
use urus::{serve_with, Config}; use urus::{serve_with, Config};
use std::net::SocketAddr; use std::time::Duration;
let cfg = Config { let cfg = Config {
listener_pool: 2, // Number of OS threads handling accept() listener_pool: 2, // Supervised accept-loop actors
scheduler_threads: Some(2), // Number of smarm worker threads name: "urus", // Endpoint registry name (unique per endpoint)
keep_alive_timeout: Duration::from_secs(60), // Idle budget between requests
request_timeout: Duration::from_secs(30), // Whole-request READ deadline
write_timeout: Duration::from_secs(30), // Per-write response budget
max_header_count: 64,
read_buf_size: 8 * 1024,
max_body_bytes: 1 << 20,
drain_timeout: Duration::from_secs(30), // Graceful-shutdown drain budget
..Config::new("127.0.0.1:8080".parse().unwrap()) ..Config::new("127.0.0.1:8080".parse().unwrap())
}; };
serve_with(cfg, pipeline).unwrap(); serve_with(cfg, smarm::Config::exact(2), pipeline).unwrap();
``` ```
**Timeout semantics:**
- `keep_alive_timeout` — how long a connection may sit idle waiting for the
*first byte* of a request. Expiry closes the socket silently (nothing was
in flight). Pipelined leftover bytes count as a started request, not idle.
- `request_timeout` — a wall-clock deadline from a request's first byte
until its head and body are fully read. Expiry mid-head gets a
best-effort `408 Request Timeout`; expiry mid-body just closes. Deadlines
are absolute instants, so a trickling (slowloris-style) client can't
reset its budget by sending one byte at a time. Covers the **read phase
only** — a streaming response may legitimately outlive any whole-request
clock.
- `write_timeout` — a per-write budget for response bytes: the fixed
head+body write, and *each* streamed chunk, must complete within it. A
client that stops reading is dropped once its socket buffer fills and a
write stalls past the budget (the write-side twin of slowloris). Streamed
chunks each get a fresh budget; a stream as a whole has no deadline.
### Your Own Supervision Tree
The real API is `urus::endpoint(config, pipeline)`, which binds the socket
and hands back a supervisable child body. Your application owns the
runtime and the root supervisor; urus is one child among your own actors:
```rust
use smarm::{ChildSpec, OneForOne, Restart, supervisor::Shutdown};
// Binds here: an address-in-use error is a startup failure, not an actor
// crash. The fds outlive any restart of the endpoint child.
let endpoint = urus::endpoint(cfg, pipeline)?;
let rt = smarm::init(smarm::Config::default());
rt.run(move || {
let sup = smarm::spawn(move || {
OneForOne::new()
// Your state actor FIRST: reverse-order shutdown therefore
// stops it LAST, after HTTP has finished draining.
.child(ChildSpec::new(Restart::Permanent, my_store))
.child(ChildSpec::new(Restart::Permanent, endpoint)
// The endpoint bounds its own drain with drain_timeout;
// a shorter supervisor deadline would cut it in half.
.shutdown(Shutdown::Infinity))
.run()
});
let _ = sup.join();
});
```
Shutdown is then whatever your app already does — `request_shutdown` on
the root supervisor (from a SIGTERM handler via `rt.handle()`, say). The
endpoint's `handle_shutdown` stops the listener pool, closes idle
keep-alive connections, drains in-flight requests up to `drain_timeout`,
force-stops stragglers past it, and exits only when the last connection is
gone — so "the endpoint child has stopped" *is* "every connection is
gone". See `examples/crud.rs` for a complete app in this shape.
Address a running endpoint by name for introspection:
`urus::endpoint::whereis("urus")`.
### Graceful Shutdown Without a Tree
If serving is all your process does, let `serve*` own the runtime and use
`Handle`, which triggers the same sequence from any thread:
```rust
use urus::{serve_with_shutdown, shutdown_handle, Config};
let (handle, signal) = shutdown_handle();
std::thread::spawn(move || {
// e.g. wait for SIGTERM / stdin / an admin endpoint...
handle.shutdown();
});
serve_with_shutdown(cfg, smarm::Config::default(), pipeline, signal).unwrap();
// Returns once the runtime has fully wound down.
```
`Handle::shutdown()` is idempotent and performs, in order:
1. Stop accepting — the listener pool is stopped; no new connections.
2. Close idle keep-alive connections immediately.
3. Drain in-flight requests for up to `Config.drain_timeout`.
4. Force-stop any stragglers past the deadline (sockets close cleanly on
unwind via `OwnedFd::drop`).
If every `Handle` is dropped, shutdown can never be signalled and the
server runs forever — exactly `serve_with`'s semantics (it does this
internally).
### Streaming Responses
A handler can return a body it doesn't have yet: hand urus the read end of
a channel and feed it from a producer actor (the conn actor owns the
socket and all write deadlines; the producer never touches the fd):
```rust
use urus::{Conn, Next};
|c: Conn, _n: Next| {
let (tx, rx) = smarm::channel::<Vec<u8>>();
smarm::spawn(move || {
for part in ["chunk one", "chunk two"] {
if tx.send(part.as_bytes().to_vec()).is_err() {
return; // connection died — stop producing
}
}
// tx drops: end of stream.
});
c.put_status(200).put_body(rx)
}
```
On HTTP/1.1 the response uses `Transfer-Encoding: chunked` and the
connection is reusable afterwards; on HTTP/1.0 (no chunked framing) bytes
go out raw and the connection closes to delimit the body. End of stream is
signalled by dropping every `Sender`. The producer contract: a `send`
error means the connection is gone — exit. Chunked **request** bodies are
decoded transparently; handlers see `conn.body` either way.
### Server-Sent Events
`Conn::sse()` is the streaming body with the SSE headers and a heartbeat
wired up:
```rust
|c: Conn, _n: Next| {
let (c, events) = c.sse(); // or c.sse_with_heartbeat(interval)
smarm::spawn(move || {
loop {
if events.send("tick", "data").is_err() {
return; // SseClosed: client gone or server draining
}
smarm::sleep(std::time::Duration::from_secs(1));
}
});
c
}
```
While the producer is silent the connection actor emits `: keep-alive`
comment chunks every `HEARTBEAT_INTERVAL` (15s default). Liveness comes
from that ping: a vanished client is detected when a write stalls past
`write_timeout`. No request clock applies to an open SSE stream, and a
graceful shutdown force-stops it at the drain deadline.
### Named Actors
For introspection and debugging, the server registers itself in smarm's
process registry: `urus.server` (the listener-pool supervisor) and
`urus.listener.{i}` (each accept loop). `smarm::whereis(name)` resolves
them from any actor inside the runtime.
### WebSocket (v0.4)
A route handler accepts the upgrade by handing `Conn::upgrade` a
`WsHandler`; after the `101` the connection actor leaves HTTP and drives
the handler from a duplex select loop (smarm RFC 008 fd arms: first-of
outbound-channel / fd-readable, one actor, no fd dup):
```rust
struct Echo;
impl urus::WsHandler for Echo {
fn on_message(&mut self, msg: urus::Message, sender: &urus::WsSender) {
let _ = sender.send(msg); // or .text(..) / .binary(..) / .ping() / .close(code, reason)
}
fn on_close(&mut self, code: Option<u16>, reason: &str) {
// what the PEER said: Some(code) iff its close frame arrived;
// None for EOF / write failure / handshake timeout.
let _ = (code, reason);
}
}
Router::new().get("/echo", |c: Conn, _n: Next| c.upgrade(Echo))
```
Callbacks run **inside the connection actor** — a slow `on_message`
stops reads, which is backpressure, not a bug. A handler that wants
concurrency spawns its own actor and hands it a `WsSender` clone (the
SSE-producer pattern); every `WsSender` method returns `Err(WsClosed)`
once the connection ends. While a large outbound frame is mid-write the
actor isn't reading — accepted v1 trade-off of the one-actor topology.
Control frames are invisible to the handler in v1: pings are
auto-ponged, a peer close is auto-echoed and surfaces as `on_close`.
Protocol violations close with the mapped RFC 6455 code (1002/1009/1007;
handler panic → 1011). Caps: `Config.max_frame_payload` (1 MiB default,
enforced from the frame header before payload buffers) and
`Config.max_message_bytes` (4 MiB default, spans fragments).
Like SSE, an open WebSocket has no request clock: idle is fine, a dead
client is caught when a write stalls past `write_timeout`, and graceful
shutdown force-stops the connection at the drain deadline. See
`examples/ws_echo.rs`.
The `WsHandler` trait also has a defaulted `on_open(&mut self, sender)`,
called once before the frame loop — the place to subscribe or spawn
producers for clients that may never send. Note the client sees the 101
*before* `on_open` runs; a client that must know its subscriptions are
live uses an application-level ack (any reply proves `on_open` completed,
callbacks being sequential).
## Pub/Sub (v0.5)
`urus::pubsub` is a local-node, phoenix_pubsub-shaped topic broadcaster —
independent of HTTP (it imports only smarm) and built for the WebSocket
relay pattern:
```rust
const BUS: PubSub<String> = PubSub::new("chat"); // a name; spawns nothing
// BUS.child() goes in your supervision tree (or serve_with*'s child vec)
let rx = BUS.subscribe("room:lobby")?; // Receiver<Arc<String>>
BUS.broadcast("room:lobby", "hi".to_string())?;
BUS.broadcast_from(smarm::self_pid(), "room:lobby", "no echo".into())?;
```
One generic instance per message domain; payloads broadcast as `Arc<M>`
(one allocation per broadcast). `subscribe` is idempotent per
`(pid, topic)` and pins the subscription to the **calling actor** — its
death prunes the entry via a monitor; a dropped `Receiver` is pruned
lazily on the next broadcast. `subscribe_as(pid, topic)` pins to an
explicit pid for relay patterns. Mailboxes are unbounded: `broadcast`
never blocks the table, and a slow subscriber's memory bill is bounded
by the two cleanup paths above.
Composition (enforced by `shutdown_with_open_chat_terminates` in the
integration suite):
1. **The handle is an address, not the table.** `PubSub<M>` is a name:
`const`, `Copy`, spawns nothing, fine in a `static` or outside the
runtime. The actor is `PubSub::child()`, a `ChildSpec` for your
supervision tree — or for `serve_with*`'s app-children vec, which puts
it ahead of the endpoint so it stops only after the endpoint has
drained. Operations resolve the name per call, so a restarted table is
reached transparently.
2. Relay/producer actors hold the `Receiver` (plus e.g. a `WsSender`
clone). They may hold the handle too — it pins nothing — but usually
have no use for one.
Prior to v0.8 both of these read the other way round: `PubSub::new()`
spawned the table, the handle owned its life, and a non-static
`Arc<OnceLock<PubSub<M>>>` lazily built from the first handler was the
required idiom. smarm 0.7 made a server's lifetime its own, and that
whole apparatus went away with it.
See [`examples/ws_chat.rs`](examples/ws_chat.rs): rooms as topics,
`on_open` subscribes + spawns the relay, `on_message` uses
`broadcast_from` so the speaker isn't echoed.
## Channels (v0.6)
Phoenix-style join/leave/event multiplexing over the WebSocket duplex,
pubsub underneath. Two additive features:
- `channels` — the wire-neutral core (`Channel`, `ChannelHub`,
`PrefixRouter`, `ChannelSocket`, sessions). Zero new dependencies;
bring your own codec by implementing `Encode`/`Decode` for your
payload type.
- `phoenix` — implies `channels`, adds serde/serde_json and the
phoenix.js V2 wire format via the `Json<T>` newtype.
```rust
use urus::{Channel, ChannelHub, ChannelSocket, PrefixRouter, Status};
use urus::channels::phoenix::Json;
use serde_json::{json, Value};
type P = Json<Value>;
#[derive(Default)]
struct Room;
impl Channel<P> for Room {
fn join(&mut self, topic: &str, payload: P, _s: &ChannelSocket<P>) -> Result<P, P> {
Ok(Json(json!({ "joined": topic }))) // Err(..) rejects
}
fn handle_in(&mut self, event: &str, payload: P, s: &ChannelSocket<P>) {
match event {
"shout" => s.broadcast("shouted", payload),
_ => s.reply(Status::Ok, Json(json!({ "echo": event }))),
}
}
}
// A description: spawns nothing. `children` are the ChildSpecs it needs
// (bus table + one registry per session route) — hand them to serve_with*.
let (hub, children) =
ChannelHub::new("chat-bus", PrefixRouter::new().channel_default::<Room>("room:*"));
// route handler: hub.upgrade(conn)
// from anywhere with a hub handle: hub.broadcast("room:lobby", "news", payload)
```
One channel actor per joined topic per socket, linked to the
connection: callbacks are sequential, a channel panic is a socket
event, and the conn-side handler answers transport heartbeats and
routes joins — `Channel` impls never see frames or refs (`reply`
correlates to the in-flight event invisibly).
**Session persistence (opt-in).** By default channels cold-start per
join, Phoenix-server style. Implementing `ChannelSession` and
registering with `channel_session` makes the channel actor outlive its
transport: broadcasts that arrive while detached are buffered (the same
`Arc`s the relay path carries — zero re-serialisation) and drained, in
order, under the rejoin's new generation. Buffer cap or TTL exceeded,
explicit leave, or a rejected rejoin tears the session down; the next
join is a cold start. A second transport claiming the same key evicts
the first.
```rust
struct ByToken;
impl urus::ChannelSession<P> for ByToken {
type Key = String;
fn session_key(topic: &str, join_payload: &P) -> String {
format!("{topic}/{}", join_payload.0["token"].as_str().unwrap_or(""))
}
// fn buffer_cap() -> usize { 128 } // defaults
// fn ttl() -> Duration { Duration::from_secs(30) }
}
PrefixRouter::new().channel_session::<ByToken>("session:*", |_: &str| {
Box::new(Room::default()) as Box<dyn Channel<P>>
})
```
Session channels must treat `join` as re-entrant: every reattach calls
it again on the same instance (the rejoin ack needs a payload, and the
channel gets to re-auth).
See [`examples/channels_chat.rs`](examples/channels_chat.rs)
(`cargo run --example channels_chat --features phoenix`): ephemeral
`room:*` and persistent `session:*` side by side over phoenix.js V2
JSON.
## Examples ## Examples
### CRUD with Actor Ownership ### CRUD with Actor Ownership
@@ -165,11 +509,12 @@ cargo test
``` ```
Tests in `tests/integration.rs` cover: Tests in `tests/integration.rs` cover:
- Basic routing (GET, POST, PUT, DELETE) - Basic routing, request bodies, headers, path parameters, status codes
- Request body echoing - Keep-alive and pipelining
- HTTP header parsing - Listener supervision (a panicking listener restarts without dropping a pending accept)
- Path parameter extraction - Graceful shutdown (idle close, in-flight drain, force-stop at the drain deadline)
- Status code responses - Timeouts (idle keep-alive reaping, slowloris-style slow headers/body)
- Registry names (`urus.server`, `urus.listener.{i}` resolve via `whereis`)
## Architecture ## Architecture
@@ -205,11 +550,11 @@ This avoids the complexity of async ecosystems while maintaining full concurrenc
## Roadmap ## Roadmap
v1 covers HTTP/1.1, the plug pipeline, and a built-in router. Future versions may include: See [`ROADMAP.md`](ROADMAP.md). Summary: v0.2 (supervised listener pool,
- Middleware ecosystem (auth, logging, compression, etc.). graceful shutdown, enforced timeouts, registry names), v0.3 (streaming
- WebSocket support. bodies, chunked requests, SSE), v0.4 (WebSocket: handshake, frame
- Benchmarking suite (see `urus-bench-spec.md`). codec, duplex handler API), v0.5 (PubSub) and v0.6 (Phoenix-style
- Additional utilities and examples. channels with opt-in session persistence) are done.
Refer to `urus-spec.md` and `urus-v1-build-notes.md` in the artifact persistence for the original design and implementation notes. Refer to `urus-spec.md` and `urus-v1-build-notes.md` in the artifact persistence for the original design and implementation notes.
+443 -96
View File
@@ -7,7 +7,8 @@
## State assessment (2026-06-11) ## State assessment (2026-06-11)
Verified against smarm master (`v0.8`, slot slab + park-epoch + select): Verified against smarm master (`393cdd0`: v0.8 slot slab + park-epoch +
select, plus RFC 008 phase 1 fd arms / timed fd waits):
- **Builds clean, 24/24 tests pass.** The scheduler surface urus consumes - **Builds clean, 24/24 tests pass.** The scheduler surface urus consumes
(`spawn`, `wait_readable`, `wait_writable`, `sleep`, `init`, `Config`, (`spawn`, `wait_readable`, `wait_writable`, `sleep`, `init`, `Config`,
@@ -27,65 +28,85 @@ Verified against smarm master (`v0.8`, slot slab + park-epoch + select):
- **No streaming response path.** `RespBody` is `Empty | Bytes`. Blocks SSE - **No streaming response path.** `RespBody` is `Empty | Bytes`. Blocks SSE
(a spec v1 goal that slipped) and chunked responses. (a spec v1 goal that slipped) and chunked responses.
- **Full-duplex constraint (matters for WebSocket).** smarm's epoll model is - **Full-duplex constraint (matters for WebSocket).** smarm's epoll model is
one waiter per RawFd (`AlreadyExists` on the second registration) and one one waiter per RawFd (`AlreadyExists`, now surfaced as `Err` via the
direction per parked actor. Concurrent read+write on one connection needs `try_select` twins) and one direction per parked actor. RFC 008 phase 1
either the dup-fd trick (serve.rs already uses it for the listener pool) (landed `393cdd0`) gives first-of(fd, channel) via fd arms in select —
or a new smarm primitive. Decide in v0.2 design, before the frame codec. the duplex options in v0.4-3 are both open now. Decide in v0.4 design,
before the frame codec.
--- ---
## v0.2 — OTP refresh ⏳ ## v0.2 — OTP refresh ✅ (landed 2026-06, chunks 1-4 / commits 5fe6969..HEAD)
Adopt the primitives urus predates. No protocol surface changes; `Plug`/`Conn` Adopt the primitives urus predates. No protocol surface changes; `Plug`/`Conn`
untouched. Chunks are independently landable, in order: untouched. Chunks are independently landable, in order:
### 1. Supervised listener pool ### 1. Supervised listener pool ✅
Replace bare spawns in `serve_with` with a supervisor: Landed: `serve_with` runs a `Strategy::OneForOne` supervisor on the root
`Strategy::OneForOne`, one `ChildSpec` per dup'd listener fd, actor, one `ChildSpec` per dup'd listener fd, `Restart::Permanent`. A
`Restart::Permanent`. A listener panic restarts that listener; the pool listener panic (or a transient `wait_readable` failure, which now exits
and restarts instead of being fatal) restarts that listener; the pool
never silently shrinks. Connection actors stay unsupervised bare spawns never silently shrinks. Connection actors stay unsupervised bare spawns
(per spec: a connection is cheap and its failure is local — a 500 path, (per spec: a connection is cheap and its failure is local — a 500 path,
not a restart path). not a restart path); `spawn` parents them under the listener, which has no
supervisor channel, so their deaths are invisible to the pool supervisor
by construction.
Open question: should the restarted child re-dup from listener 0's fd, or Open question resolved as leaned: the ChildSpec factory owns the fd via
own its fd across restarts via the ChildSpec closure? Leaning: closure `Arc<OwnedFd>` (`OwnedFd` gained `Sync` for this) and every (re)start
captures the `OwnedFd`, re-registers on restart. No re-dup needed. reuses it. No re-dup, no window where a pool slot has no fd. Covered by
`tests/integration.rs::panicking_listener_restarts` (pool of 1, panic
injected before `accept` so the pending connection must survive into the
restarted listener).
### 2. Graceful shutdown ### 2. Graceful shutdown ✅
Landed (c658ac0): `serve_with_shutdown(config, pipeline, signal)` +
`shutdown_handle()`. Drain-then-stop as decided below; design details in
the commit body. `serve_with` keeps v1 run-forever semantics by dropping
its own Handle.
`serve_with` gains a shutdown signal (an `urus::Handle` with `.shutdown()`, `serve_with` gains a shutdown signal (an `urus::Handle` with `.shutdown()`,
backed by a channel + `request_stop` fan-out): backed by a channel + `request_stop` fan-out):
1. Stop listeners (no new accepts; listener fds close). 1. Stop listeners (no new accepts; listener fds close).
2. Stop connection actors — either drain-then-stop with a deadline, or 2. **Decided: drain-then-stop with a deadline** — in-flight requests get
immediate `request_stop` (unwind closes fds cleanly via `OwnedFd::drop`). to finish; `request_stop` after the deadline (unwind closes fds cleanly
via `OwnedFd::drop`). Drain also keeps the door open for an
apache-reload-style restart later.
3. Supervisor winds down; `rt.run` returns; `serve_with` returns. 3. Supervisor winds down; `rt.run` returns; `serve_with` returns.
Requires tracking live connection pids. Cheapest honest version: a Requires tracking live connection pids. Cheapest honest version: a
`gen_server` connection registry — listener casts `{Started, pid}`, conn `gen_server` connection registry — listener casts `{Started, pid}`, conn
actor monitors-free self-deregisters via a drop guard cast. This registry actor monitors-free self-deregisters via a drop guard cast. Pid set only
is also chunk 3's reaper and the future ws/channels introspection point, (no activity stamps — chunk 3 went per-conn timed waits, no reaper); still
so build it once here. the future ws/channels introspection point, so build it once here.
### 3. Enforce the timeouts we advertise ### 3. Enforce the timeouts we advertise ✅
**Decision fork — pick one before implementing:** Landed (a60687b): timed waits as decided below; the leaning held — a
per-request `Instant` deadline keeps `request_timeout` an honest
whole-request wall clock (mid-head expiry: best-effort non-parking 408;
mid-body: close). Details in the commit body.
**Decided: smarm-native timed waits** (the fork's option (a), a reaper in
urus, died when RFC 008 landed at `393cdd0` — `wait_readable_timeout` /
`wait_writable_timeout` exist upstream now, so there is nothing for the
reaper to be the stopgap for):
- **(a) Reaper in urus.** The chunk-2 registry tracks per-conn `read_some` parks in `wait_readable_timeout` instead of `wait_readable`;
`last_activity`; a ticker (`smarm::sleep` loop) sweeps and `Ok(false)` means the deadline passed and the conn actor closes. Per-conn
`request_stop`s expired conns. Works today, zero smarm changes. semantics, no activity stamping, no sweep ticker. Semantics note: a timed
Granularity = sweep interval; conns must stamp activity (one cast per read wait enforces *idle-between-reads* — `keep_alive_timeout` between
request — measurable but small). requests, and slowloris dies because each stalled read has a deadline.
- **(b) `wait_readable_timeout` in smarm.** A first-of(fd, timer) wait. `request_timeout` as a whole-request wall clock is NOT what this gives;
The v0.7 epoch-stamped consuming-unpark machinery exists precisely for decide in implementation whether to track a per-request deadline across
multi-arm waits (`select_timeout` is the channel-flavoured proof), but reads (cheap: one `Instant` in the conn loop, pass the min of both
fd arms aren't `Selectable` today — this is real smarm work in budgets to each wait) or redefine `request_timeout` as per-read. Leaning
`wait_fd`/io thread. Cleaner per-conn semantics, benefits every future the `Instant`: it keeps the advertised config honest.
smarm io consumer, no reaper bookkeeping.
Leaning **(a) now, (b) later**: (a) ships the security fix this cycle and urus is the first real consumer of the timed wrappers — note the RFC 008
the registry exists anyway; (b) becomes a smarm roadmap item whose landing reviewer flag (`wait_fd` wraps in `NoPreempt`, select arms don't); the
lets the reaper's stamps go away. But if you'd rather not build throwaway: crud example under load doubles as that confirmation.
(b) first is defensible — say the word.
### 4. Hygiene ### 4. Hygiene ✅
Landed: registry names + whereis test (in-runtime probe via a handler —
tests are foreign threads), README rewrite, smarm-trace compile verified.
- Registry names: `urus.server`, `urus.listener.{n}` via `smarm::register` - Registry names: `urus.server`, `urus.listener.{n}` via `smarm::register`
(debuggability; `whereis` in tests). (debuggability; `whereis` in tests).
- README: rewrite Roadmap section to point here; document `Handle`. - README: rewrite Roadmap section to point here; document `Handle`.
@@ -98,81 +119,403 @@ panicking listener restarts under load without dropped accepts (test each).
--- ---
## v0.3 — Streaming bodies + SSE ## v0.3 — Streaming bodies + SSE ✅ (landed 2026-06, chunks 1-3)
The slipped v1 goal, and the WebSocket prerequisite (an upgraded connection The slipped v1 goal, and the WebSocket prerequisite (an upgraded connection
is "a response that never ends"). is "a response that never ends").
1. `RespBody::Stream(...)` — handler hands back a writer callback or a 1. `RespBody::Stream(StreamBody)` ✅ — **decided: pull** (user call at the
`Receiver<Vec<u8>>`; conn actor switches to chunked transfer-encoding v0.3 design fork): the handler hands back a `smarm::Receiver<Vec<u8>>`
(HTTP/1.1) and pumps until the stream closes. Keep-alive after a and the producer is an actor the handler spawned; the conn actor pumps
completed chunked response. (`pump_stream`) and keeps sole ownership of the socket and write
2. Chunked **request** bodies (`parser.rs` currently returns 411) — same deadlines. The recv() park between chunks is stoppable by the draining
cycle, the codec knowledge is fresh. registry, so infinite streams die at the drain deadline with no
3. SSE sugar: `Conn::sse()` → an `EventSender` (`.send(event, data)`), polling. HTTP/1.1: chunked TE, keep-alive after the terminator;
sets `content-type: text/event-stream`, disables buffering, heartbeats HTTP/1.0: raw bytes, forced close, EOF-delimited. End of stream =
via comment lines on a timer. An SSE conn is just a Stream body that every Sender dropped = 0-chunk. Timeout split (handoff Q3, as
the reaper exempts from `request_timeout` (keep-alive ping instead). proposed): `request_timeout` covers the READ phase only; new
`Config.write_timeout` (default 30s) bounds every response write —
fresh budget per streamed chunk, and the fixed-body + error-response
writes went deadline-bounded too (write-side slowloris closed).
2. Chunked **request** bodies ✅ — `read_chunked_body` decodes
incrementally on the request deadline; extensions ignored, trailers
consumed + discarded, decoded size capped at max_body_bytes (413),
size-line/trailer spam capped (400). CL + TE: chunked, or chunked on
HTTP/1.0, rejected 400 at parse (smuggling ambiguity). Keep-alive
drain accounts RAW framing length, so pipelined requests behind a
chunked body land exactly.
3. SSE sugar ✅ — `Conn::sse()` / `sse_with_heartbeat` → `EventSender`
(`send(event, data)` / `data` / `comment`, Clone, `Err(SseClosed)`
once the stream is gone). Heartbeat lives in the pump as leaned: with
`StreamBody.heartbeat` set, the chunk wait is `recv_timeout(interval)`
and expiry writes a `: keep-alive` comment chunk. Liveness from the
ping (dead client = heartbeat write stalls past write_timeout), not
from any request clock.
**Exit criteria (verified):** chunked round-trips both directions with
keep-alive and pipelining intact; an infinite SSE stream heartbeats while
silent, is killed by write_timeout when the client stops reading, and is
force-stopped at the drain deadline on shutdown; full suite green.
--- ---
## v0.4 — WebSocket upgrade path ## v0.4 — WebSocket upgrade path ✅ (landed 2026-06, chunks 1-3)
1. **Handshake**: detect `Connection: Upgrade` + `Upgrade: websocket` in 1. **Handshake** ✅ (49538cc) — `Conn::upgrade()` (handler arg comes
the pipeline; `Conn::upgrade(handler)` short-circuits — conn actor with chunk 3): §4.2.1 validation, 101 + accept key, marker field on
validates `Sec-WebSocket-Key`, emits 101, exits the HTTP loop. `Conn`, actor leaves the HTTP loop (fd drops until chunk 3 takes
(SHA-1 + base64 needed; vendor a tiny SHA-1 rather than pulling a over). Rejections stay plain HTTP (400 / 426 + version header) with
crate — urus has two deps and it should stay that way.) keep-alive intact. **Dep decision changed at the design pass**: SHA-1
2. **Frame codec**: RFC 6455 framing — fragmentation, masking (client is `sha1_smol` (zero transitive deps) — dep #3 SPENT, no hand-rolled
frames are always masked), control frames (ping/pong/close), 125-byte crypto (user call). The v0.6 JSON question is therefore a dep #4
control-payload rule, close-code handshake. question now. Base64 stays in-tree, encode-only.
3. **Duplex topology** (the design decision from the assessment): 2. **Frame codec** ✅ (2869153) — pure `ws::frame`: incremental
- **(a) Two actors, dup'd fd**: reader actor on the dup, writer actor `decode` (mask-required server mode), minimal-length + RSV + control
on the original, writer owns a `Receiver<Frame>`; the user handler rules enforced, header-derived cap check before payload buffering,
talks to both via channels. Proven trick (listener pool), works today. close-code wire validation, `Assembler` for fragmentation with
- **(b) One actor, smarm first-of(fd-readable, channel)**: needs fd whole-message UTF-8 check, `FrameError`→close-code map
arms in select — the same smarm work as v0.2-3(b). Simpler topology, (1002/1009/1007). §5.7 examples pinned verbatim.
no dup, but blocked on smarm. 3. **Duplex wiring** ✅ — topology decision: **(b) one actor, RFC 008
Default (a); revisit if (b) lands for timeouts anyway. fd arms** (urus exists to put smarm through its paces; first fd-arm
4. Handler API sketch (settle in design pass): trait with consumer is the point). As landed:
`on_message/on_close` + a `WsSender` handle — gen_server-shaped, so a - `ws::duplex::run_duplex` replaces the HTTP loop after the 101:
ws connection *is* an actor you can monitor, link, and register. `try_select(&[&outbound_rx, &FdArm::readable(fd)])` per iteration —
outbound at index 0 (select is priority-ordered: owed writes drain
before reading more). Accepted gap documented: mid-write of a large
outbound frame the actor isn't reading.
- Handler API (the deferred questions, as answered): in-actor
callbacks — `WsHandler { on_message, on_close }` runs inside the
select loop; concurrency = spawn your own actor with a `WsSender`
clone (the SSE-producer pattern). `Conn::upgrade(self, handler)`
flat signature (builder deferred until per-upgrade options exist,
e.g. subprotocols). `WsSender`: Clone, `send/text/binary/ping/
close → Err(WsClosed)`, unbounded channel like SSE (slow client
bounded by write_timeout). Control frames invisible v1: auto-pong
inline (§5.5.2), peer close auto-echoed (§5.5.1, server closes TCP
first per §7.1.1) surfacing only as `on_close(Some(code), reason)`;
wordless endings (EOF, write failure, handshake timeout) =
`on_close(None, "")`. `on_ping`/`on_pong` can land later as default
trait methods, non-breaking.
- Server-initiated close (WsSender::close or FrameError→1002/1009/
1007, handler panic→1011): close frame out, then a bounded drain
(one write_timeout budget) for the peer's echo, discarding data
(§1.4). Handler panics re-raise smarm's stop sentinel before being
treated as panics — the pipeline's catch_unwind dance, replicated.
- Caps: `Config.max_frame_payload` (1 MiB) / `max_message_bytes`
(4 MiB) — Config fields per the max_body_bytes precedent; header
cap checked before payload buffers.
- Buf handover: bytes pipelined past the upgrade head carry into the
duplex loop (tested).
- Registry: conn stays **Busy** for the ws lifetime (an open ws is
in-flight work); graceful shutdown force-stops it out of the select
park at the drain deadline (tested — the in-runtime request_stop
wake covers select parks, not just channel recv).
- Validated against a real client (python websocket-client): echo,
ping/pong payload, close handshake. `examples/ws_echo.rs` shows
both handler shapes (/echo in-actor, /clock producer-spawning).
4. **Handler API** ✅ — folded into chunk 3 above (was: settle in design
pass; settled: gen_server-shaped trait + WsSender handle).
## v0.5 — PubSub ## v0.5 — PubSub ✅ (landed 2026-06, chunks 1-2)
`urus::pubsub`, phoenix_pubsub-shaped, local node only (no distribution — `urus::pubsub`, phoenix_pubsub-shaped, local node only (no distribution —
smarm has none): smarm has none). All six design questions ratified by the user pre-code
("push" on the stated leans); as-landed:
- `subscribe(topic)`, `unsubscribe(topic)`, `broadcast(topic, msg)`, - **Generic `PubSub<M>`** per instance (Q1); heterogeneous buses are an
`broadcast_from(self, topic, msg)`. app-side enum. `Arc<M>` payloads: one allocation per broadcast.
- Topic table: gen_server owning `HashMap<Topic, Vec<(Pid, Sender<Arc<M>>)>>`; - **`subscribe(topic) -> Receiver<Arc<M>>`** (Q2), fresh channel per
shard N ways by topic hash if the single server measures hot (don't subscription; `subscribe_as(pid, ..)` pins lifetime to an explicit pid
pre-shard; bench first — smarm channel sends are cheap). for relay shapes. Sender-passing fan-in (`subscribe_with`) deliberately
- Dead subscriber cleanup via `monitor` + `handle_down` (this is exactly deferred — compatible addition.
what 06-10's gen_server `handle_down` was built for). - **Unbounded mailboxes** (Q3): broadcast never blocks the table. Cleanup
- `Arc<M>` payloads: one allocation per broadcast, not per subscriber — is monitor `Down` (eager, on subscriber death) + prune-on-send-failure
the shared-heap advantage over BEAM, take it. (lazy, dropped receivers). Bench before sharding.
- **User-owned handle** (Q4). The ratified register-by-name helper turned
out NOT implementable: smarm's registry maps name → Pid and a Pid can't
be turned back into a `ServerRef`. Deferred pending smarm support (or a
type-erased global, rejected). `pid()` exposed.
- **Unique per `(pid, topic)`** (Q5): subscribe idempotent (replaces the
sender; the stale receiver closes). One monitor per live pid.
`handle_down` is a full table scan as ratified; pid→topics reverse
index is the documented optimization if deaths measure hot.
- Table: gen_server `HashMap<Topic, HashMap<Pid, Sender<Arc<M>>>>` —
single server, not sharded (bench first). `subscriber_count(topic)`
added (tests need to observe async cleanup; legitimate API).
Deliberately independent of HTTP — usable from any smarm app. Separate **Discoveries this cycle:**
module now, candidate for crate extraction later. - `PubSub::new()` is in-runtime only (spawns the table actor) but
pipelines are built pre-runtime → the composition pattern is a
NON-static `Arc<OnceLock<PubSub<M>>>` in the route closure. Non-static
is load-bearing: the drained pipeline drops the cell in-runtime, the
table's inbox closes, the table exits, AllDone is reachable. A `static`
(crud's store shape) hangs `serve_with_shutdown`. Corollary: relays
hold `Receiver` only, never a `PubSub` clone (mutual-keepalive cycle).
Proven by `shutdown_with_open_chat_terminates`.
**Superseded by v0.7 where the app owns its tree:** start the table as a
supervised sibling of the endpoint and address it by name (what crud now
does with its store). The `OnceLock` rule still holds under `serve*`,
which owns the runtime and offers no in-runtime moment beforehand.
- **`WsHandler::on_open(&mut self, sender)`** added (defaulted,
non-breaking): without it a listen-only ws client can never be
subscribed (`on_message` never fires). Same panic contract as
`on_message` (1011, no `on_close`). The 101 reaches the client before
`on_open` runs — subscription liveness needs an app-level ack
(documented; the integration tests use a `sync`/`synced` ack).
- hammer.sh default subset now includes `ws_`.
## v0.6 — Channels Chunk 2's ws_chat: rooms as topics, `on_open` subscribes (conn-actor
pid) + spawns the relay, `broadcast_from(self_pid(), ..)` for no-echo.
Phoenix channels: the join/leave/event protocol over WebSocket transport, ## v0.6 — Channels — DONE (2026-06-12)
PubSub underneath.
- `Channel` trait: `join(topic, payload, socket)`, `handle_in(event, Join/leave/event protocol over WebSocket transport, PubSub underneath.
payload, socket)`, `handle_out`, `terminate`. Design decisions ratified pre-code (2026-06-12):
- Topic router (`"room:*"` patterns) → channel actor per (conn, topic),
spawned under the ws connection, linked so conn death reaps channels. **Feature flags.** Two additive features:
- Wire format: Phoenix V2 JSON serializer - `"channels"` — the `Channel` trait, `TopicRouter` trait, `PrefixRouter`,
(`[join_ref, ref, topic, event, payload]`) — free interop with `ChannelSocket`, session machinery. Zero new deps; pubsub is already
phoenix.js clients is the killer feature; needs a JSON dep or a always-on (it only exists because channels needs it).
hand-rolled mini-codec (decision point — this would be dep #3). - `"phoenix"` — implies `"channels"` + `serde`/`serde_json` + the V2 frame
- Presence: explicitly out of scope until distribution exists somewhere. codec. Opt-in phoenix.js interop; users who want protobuf or any other
wire format take `"channels"` only and supply their own codec.
**Payload generics.** The `Channel` trait is generic over payload:
`P: Encode + Decode` where `Encode`/`Decode` are urus codec traits (one
method each). The `"phoenix"` feature provides a `JsonCodec` blanket impl
via serde. No JSON anywhere in the `"channels"` surface.
**`Channel` trait shape.**
```
trait Channel<P>: Send + 'static {
fn join(topic: &str, payload: P, socket: &ChannelSocket<P>) -> Result<P, P>;
fn handle_in(&mut self, event: &str, payload: P, socket: &ChannelSocket<P>);
fn terminate(&mut self) {}
}
```
`handle_out` deferred — can land as a defaulted method, non-breaking.
**`ChannelSocket`.** Newtype over `WsSender` + a `PubSub` handle.
Exposes `reply(ref, status, payload)`, `push(event, payload)`,
`broadcast(topic, payload)`. The `ref` from the incoming frame is threaded
in by the channel actor loop, invisible to the impl. Coupling to `PubSub`
is intentional: channels owns pubsub, and broadcast is the core use case.
**Topic router.** `TopicRouter` trait: `fn route(&self, topic: &str) ->
Option<Box<dyn ChannelFactory<P>>>`. Factory receives the full raw topic
string so wildcard-segment extraction is always possible. urus ships
`PrefixRouter` as the default impl: scan to `'*'`, discard the rest,
single `HashMap` lookup on the prefix. No regex dep, no glob machinery —
users who need that implement `TopicRouter` themselves.
**Channel actor lifetime + session persistence.**
Default: channel actor linked to the ws connection actor, cold-start on
every reconnect (phoenix-server-compatible behavior). Opt-in persistence
via a `ChannelSession` trait:
```
trait ChannelSession: Send + 'static {
type Key: Eq + Hash + Send + 'static;
fn session_key(topic: &str, join_payload: &P) -> Self::Key;
fn buffer_cap() -> usize { 128 }
fn ttl() -> Duration { Duration::from_secs(30) }
}
```
When implemented, the router consults a session registry (gen_server,
same pattern as the connection registry) before spawning: an existing
actor is reattached, and buffered outbound messages (`Vec<Arc<M>>`, no
re-serialisation) are drained to the reconnecting transport. Buffer full
or TTL expired → actor tears down; next join is a cold start. This
diverges from Phoenix server conventions (client-re-syncs) but is
invisible to phoenix.js at the wire level. Zero-copy drain is the smarm
motivation: messages are `Arc<M>` in the buffer and in the PubSub relay
path, one allocation per broadcast regardless of subscriber count or
reconnect cycles.
**Presence:** explicitly out of scope until distribution exists.
**As-landed (2026-06-12), beyond the ratified text — veto by diff:**
- chunk 1: `Channel::join` takes `&mut self` (state from join, and
session rejoins need the same instance); the in-flight ref is
threaded invisibly through the socket; `broadcast` gained an `event`
parameter; `TopicRouter::route` returns `Option<Arc<dyn
ChannelFactory>>`.
- chunk 2: the serde blanket lives behind the `Json<T>` newtype
(a blanket on bare `P: Serialize` is coherence poison under additive
features); encode failure panics into the linked-teardown contract.
- chunk 3: `ChannelFactory` gained a `#[doc(hidden)] deploy()` seam
(default = the ephemeral linked actor); every attach re-calls
`join()` (re-entrant by contract); rejected rejoin / explicit leave /
cap / TTL all end the session; second transport evicts the first;
`Key: Clone` (registry's pid→key reverse index); registry spawns
eagerly at registration (in-runtime law, same as `ChannelHub::new`);
shutdown composes linklessly via the registry-drop chain — proven by
`shutdown_with_detached_session_terminates`.
--- ---
## v0.7 — Endpoint as a supervised child — DONE (2026-08-20)
Crate goes 0.2.x -> **0.3.0** (breaking). Unwinds the last deviation from
the spec (§2.1/§6): `serve` owning `rt.run`. The app owns the runtime and
the root supervisor; urus is one ordered child in it.
```
your root sup
└── ChildSpec(Permanent, urus::endpoint(cfg, pipeline)?) <- Endpoint gen_server
└── listener_sup OneForOne over N listeners
└── plain connection actors
```
- `src/conn_registry.rs` -> `src/endpoint.rs`; `ConnRegistry` -> `Endpoint`.
The registry absorbed the listener pool it registers for and runs inline
as the `ChildSpec`'s actor (`NamedGenServerBuilder::run`), so supervisor
shutdown arrives as `handle_shutdown` and a restart re-runs `init` on the
same still-open fds. `endpoint()` binds eagerly on the caller's thread.
- **The endpoint spawns its own listener sup rather than being its
sibling.** smarm's supervisor start *order* is not start *readiness*
(`start_child` spawns and moves on), so a sibling listener could
`whereis` the endpoint name before its actor ran. Registrar-spawns-
consumers makes that program order inside one `init`. Filed in smarm's
ROADMAP as a readiness-ack item; a blocking `spawn` would only shrink the
window ("has begun executing" != "has bound its name") at the cost of a
round-trip per accept.
- Listener sup is monitored: death outside shutdown = panic (the app's
supervisor decides) instead of a zombie on a dead port; death during
shutdown is the "no new connections" barrier.
- Drain is entirely internal and event-driven: `handle_shutdown` flips
draining, shuts the listener sup, stops idle conns, arms one
`drain_timeout` timer; conns that register or go idle mid-drain are
stopped on the spot; one force sweep at the deadline; the endpoint exits
when the sup is down and the set is empty. "Endpoint child stopped" ==
"every connection gone".
- **Deleted:** the `AtomicBool` listener flag, `LISTENER_TICK` (250ms wake
per listener per tick -> untimed `wait_readable` park), `SHUTDOWN_POLL`
(100ms root poll -> real park on the signal channel), `Restart::Transient`
for listeners (now `Permanent`: they only exit by supervisor action),
the root-side drain loop, `Cast::{BeginDrain, ForceStopConns}`.
Re-probed against current smarm: `request_stop` unwinds an untimed
`wait_readable` park, is no longer lossy against a QUEUED actor, and
supervisor shutdown joins in ~200us.
- **Config split**: `scheduler_threads` / `max_actors` are runtime knobs an
endpoint-as-child cannot honour -> `serve_with(cfg, smarm::Config, pipe)`
and `serve_with_shutdown(cfg, smarm::Config, pipe, signal)`. New
`Config.name` (default `"urus"`) is the endpoint's registry name;
`urus::endpoint::whereis(name)` for introspection. Two endpoints in one
process need distinct names.
- `serve*` stay as thin wrappers building a one-child tree with
`Shutdown::Infinity`. `Handle`/`ShutdownSignal` stay: a `serve*` caller
never sees the runtime, so it has no `RuntimeHandle` to reach for — but
the poll behind it is gone.
- `examples/crud.rs` is the demonstrator (app-owned tree, store as an
ordered sibling registered under a typed `Name`, shutdown via
`rt.handle().request_shutdown(root_sup)`); it loses its `OnceLock` store
cell and its `SHUTTING_DOWN` flag + 250ms poll. Other examples stay short
on `serve*`.
**Open after this cycle:**
- ~~One unreproduced test failure seen once in ~120 full-suite runs~~
**DIAGNOSED AND FIXED** in v0.8 below — it was the pubsub/channels drain
hang, not the `free_port()` race. It hid because `hammer.sh` builds with
default features while the failure needs `--all-features` load.
- `endpoint()` returns `impl Fn()`, so calling it twice = two endpoints
contending for one `Config.name` (second panics on the clash). Honest
failure, but the type doesn't prevent the mistake; a consume-on-first-use
newtype would. `ChannelHub::new` returning `(hub, children)` in v0.8 is
the same lesson applied: make the type refuse the mistake.
## v0.8 — Bus, hub and session registries are supervised — DONE (2026-08-20)
Same crate version (**0.3.0**); this cycle and v0.7 ship together and are
both breaking.
v0.7 put the endpoint under a supervisor. This puts everything else there
too, because smarm 0.7 left it no choice: commit `415effb` ("lifetime is
the actor's — refs are addresses") removed the rule the pubsub table,
channel hub bus and session registries were built on. A `GenServerRef` no
longer owns its server; the loop holds its own inbox sender and the inbox
never closes when the last ref drops. Nothing commanded these three to
stop, and what still terminated a run was the root-exit sweep racing the
drain.
That is the flake above: `shutdown_with_open_chat_terminates` and
`channels_wire::shutdown_with_open_channel_terminates`, 4 failures in 25
`--all-features` integration runs, all "serve did not return: ... outlived
the drain: Timeout". 0 in 10 full-suite runs after the fix.
**Description separated from instantiation.** The gotchas here all
descended from constructors that spawn — `PubSub::new()`,
`ChannelHub::new()` and `PrefixRouter::channel_session()` all started
actors, which forced in-runtime-only construction, which forced the
`Arc<OnceLock<..>>` lazy-init from the first handler, which forced the
"cell must not be `static`" and "a relay must never hold a `PubSub` clone"
rules. Five documented rules propping up one inverted dependency.
```
your root sup (RestForOne)
├── ChildSpec(Permanent, PubSub::child) <- bus table, named
├── ChildSpec(Permanent, session registry) <- one per channel_session
└── ChildSpec(Permanent, urus::endpoint(..)) <- starts last, drains first
```
- **`PubSub<M>` is a name, not a ref.** `const`-constructible, `Copy`,
spawns nothing, legal outside the runtime and in a `static`. Operations
resolve through the registry per call, so a supervisor-restarted table is
reached transparently. `PubSub::new()` is gone: `PubSub::new(name)` +
`PubSub::child()`.
- **`ChannelHub::new(bus, router) -> (hub, Vec<ChildSpec>)`**, `#[must_use]`.
Both halves come back together because a hub whose children were never
started compiles and fails on the first join. There is no `children()`
method to forget to call.
- **`channel_session` takes a registry name**; each registry is separately
named and supervised.
- **`serve_with` / `serve_with_shutdown` take a `Vec<ChildSpec>`** of app
children and build `RestForOne[..app children, endpoint]`. App children
start before the endpoint and — shutdown being ordered in reverse — stop
after it has drained, so a request still in flight can still reach the
bus. `RestForOne` because a bus crash leaves live sockets addressing a
table that no longer knows them.
- Deleted: the `Arc<OnceLock>` idiom from both examples and both test
pipelines, and every module rule that existed only to hand-manage a
refcount. `ws_chat` lost 21 lines and a struct field; examples net -25.
**Cost accepted:** one registry resolution per operation, including per
broadcast. Caching a `GenServerRef` in the handle would save it and go
stale across exactly the restart the supervisor exists to perform. Bench
before optimising.
**Open after this cycle:**
- `channel_session("session:*", "chat-sessions", f)` puts two unrelated
string literals side by side and nothing catches a transposition. A
`RegistryName` newtype makes it a type error. Not done.
- `PubSub` cannot be protected the way `ChannelHub` was: a `const` is
copied at each use, so it can't be consume-on-first-use. Forgetting
`BUS.child()` compiles and fails at runtime with `PubSubDown`. Judged
worth it — the `const` is what makes the handle pleasant.
- Session actors are still plain-spawned by their registry, not supervised
(they're dynamic, one per key — OTP would want `simple_one_for_one`).
Their exit chain is now commanded rather than refcounted: the registry is
shut down, its state drops, control senders drop, parked sessions wake on
the closed arm. Sound, but it is the last place a lifetime is implied by
a drop rather than stated.
- `hammer.sh` still passes no feature flags, which is why this hid for a
whole cycle. A feature matrix is the obvious fix.
## Known bugs
- ~~**HTTP/1.0 keep-alive: server honors but never advertises**~~ FIXED
(2026-06-12, same session it was found). As-landed decisions:
serialise_response now echoes `connection: keep-alive` on a kept-alive
1.0 response (1.1 keep-alive stays implicit); error statuses (>=400)
on 1.0 force close — no keep-alive advertised on a 4xx/5xx where reuse
is opt-in (conn_actor, alongside the existing 1.0-stream force-close);
HEAD needs nothing special (header echo only, framing untouched);
parse-error responses already carried `connection: close`. Validated
against the original repro: `ab -k -n 5000 -c 50` on GET /users now
reports 5000/5000 keep-alive requests, 0 failed, ~63k req/s
(vs ~18.9k no-keep-alive baseline, same box). Watch out benching:
ab counts a response toward "Non-2xx" silently — a 404'd route still
shows "Failed requests: 0" with great rps; check the route first.
## Later / icebox ## Later / icebox
- **HTTP/2** — spec §2.2 already designs the stream-actor demux (zero-copy - **HTTP/2** — spec §2.2 already designs the stream-actor demux (zero-copy
@@ -182,9 +525,13 @@ PubSub underneath.
- **SSE ergonomics round 2** — auto-reconnect support (`Last-Event-ID`), - **SSE ergonomics round 2** — auto-reconnect support (`Last-Event-ID`),
event-id bookkeeping, backpressure policy on slow consumers. event-id bookkeeping, backpressure policy on slow consumers.
- **TLS** — still deferred; acceptor design still must not preclude it. - **TLS** — still deferred; acceptor design still must not preclude it.
- **`wait_readable_timeout` / fd-in-select upstreaming** — the smarm-side
item that simplifies v0.2-3 and v0.4-3 retroactively.
- **Middleware ecosystem** — compression, static files, auth plugs. - **Middleware ecosystem** — compression, static files, auth plugs.
- **Bench suite** — `urus-bench-spec.md` exists in the artefact store; - **Bench suite** — `urus-bench-spec.md` exists in the artefact store;
wire it up once v0.2 lands (supervision changes the hot path not at all, wire it up once v0.2 lands (supervision changes the hot path not at all,
but prove it). but prove it).
- **Operator introspection via typed names** — partly delivered in v0.7:
the endpoint is a named gen_server (`Config.name`, default `"urus"`),
reachable via `urus::endpoint::whereis(name)`, and answers
`Call::ConnCount`. Still open: per-listener visibility (the pool is
internal and anonymous), and a richer stats call (ws/channel counts,
request rates) — the endpoint is the place to hang them.
+322
View File
@@ -0,0 +1,322 @@
//! Causal-profiling bench (RFC 007): a real urus webserver as the workload,
//! with a planted, known-answer bottleneck.
//!
//! Where smarm's `causal_pipeline` demo is synthetic, this exercises the
//! exact machinery the causal commits touched, for real:
//!
//! - epoll parks (`wait_readable_timeout` / `wait_writable_timeout`)
//! → park-gated resume credit
//! - keep-alive / request / write deadlines → timer-heap virtual time
//! - the experiment controller's fixed windows → wall-anchored timers
//! - mixed CPU (parse/serialize) and IO (socket) → non-trivial ranking
//!
//! Shape: `LOAD_CONNS` plain OS threads hammer `GET /order/:id` over
//! blocking loopback TCP (keep-alive). OS threads on the client side are
//! deliberate: they can't absorb injected virtual delay, so the only
//! delayable code is the server's — the same trick the smarm-side timer
//! tests use. The handler calls a single store actor (crud's pattern:
//! one actor owns the data, serialization is structural) which burns
//! `STORE_US` of calibrated *work* — work-shaped, not timed, because a
//! timed wait absorbs injected delay and flattens every experiment.
//!
//! Known answer: the store is serialized and saturated, so it is the
//! throughput ceiling (1e6/STORE_US rps). A virtual speedup of `store`
//! by p% predicts throughput ×1/(1-p): +33% @25%, +100% @50%. The other
//! five sites (parse, router, pipeline, serialize, socket-write — the
//! urus lib instrumentation) are tens of µs and parallel across
//! connections: predicted impact ≈0. Per the smarm-side validation runs,
//! absolute magnitudes are trustworthy to roughly ±15% while the
//! *ranking* is solid — the verdict thresholds reflect that.
//!
//! Run (needs cores; the verdict is SKIPPED below 4):
//! cargo run --release --example causal_bench --features smarm-causal
//!
//! Prints the summary, writes `profile.coz` next to the CWD, and exits
//! nonzero if the expected separation doesn't hold.
use std::io::{Read, Write};
use std::net::{SocketAddr, TcpListener, TcpStream};
use std::sync::atomic::{AtomicBool, Ordering};
use std::sync::{Arc, OnceLock};
use std::time::{Duration, Instant};
use smarm::{channel, Sender};
use urus::{serve_with_shutdown, shutdown_handle, Config, Conn, Next, Pipeline, Router};
// ---------------------------------------------------------------------------
// Tunables
// ---------------------------------------------------------------------------
/// Planted bottleneck cost: calibrated work per store request, µs.
const STORE_US: u64 = 400;
/// Handler-side render work per request, µs. Runs under the `pipeline`
/// site (no inner guard), parallel across connection actors.
const RENDER_US: u64 = 50;
/// Client connections = OS load threads. Plenty to keep a 400µs store
/// saturated even at a 50% virtual speedup (see saturation note in main).
const LOAD_CONNS: usize = 16;
// ---------------------------------------------------------------------------
// Calibrated work (lifted from smarm's causal_pipeline demo)
// ---------------------------------------------------------------------------
/// LCG-mix `iters` times in dependent sequence (unvectorizable, un-elidable),
/// staying preemptible — and causal-sampleable/delayable — via `check!()`.
fn work_iters(iters: u64) {
let mut acc = 0x2545_f491_4f6c_dd1du64;
let mut i = 0u64;
while i < iters {
let chunk_end = (i + 256).min(iters);
while i < chunk_end {
acc = acc.wrapping_mul(6364136223846793005).wrapping_add(i);
i += 1;
}
std::hint::black_box(acc);
smarm::check!();
}
}
/// Iterations per microsecond on this machine, so stage costs are
/// meaningful in time while staying work-shaped.
fn calibrate_iters_per_us() -> u64 {
let n = 8_000_000u64;
let t = Instant::now();
work_iters(n);
(n / (t.elapsed().as_micros().max(1) as u64)).max(1)
}
static PER_US: OnceLock<u64> = OnceLock::new();
fn work_us(us: u64) {
work_iters(us * PER_US.get().copied().unwrap_or(1));
}
// ---------------------------------------------------------------------------
// Store actor — the planted bottleneck (crud's once-cell pattern)
// ---------------------------------------------------------------------------
enum Req {
Get { id: u64, reply: Sender<u64> },
}
/// Set at teardown; the store polls it (a cross-thread `Sender::send`
/// can't wake a parked actor, so a bare `recv` would hang AllDone —
/// same limitation crud's store works around).
static STORE_STOP: AtomicBool = AtomicBool::new(false);
static STORE_TX: OnceLock<Sender<Req>> = OnceLock::new();
fn store_loop(rx: smarm::Receiver<Req>) {
loop {
let req = match rx.recv_timeout(Duration::from_millis(250)) {
Ok(r) => r,
Err(smarm::channel::RecvTimeoutError::Timeout) => {
if STORE_STOP.load(Ordering::Relaxed) {
return;
}
continue;
}
Err(smarm::channel::RecvTimeoutError::Disconnected) => return,
};
let Req::Get { id, reply } = req;
{
// The known-answer site: serialized by construction (one
// actor, one request at a time), saturated by the load.
let _g = smarm::causal_site!("store");
work_us(STORE_US);
}
let _ = reply.send(id);
}
}
/// Spawned from the first handler that runs — connection actors are
/// inside the runtime, which `spawn` requires; the static then shares
/// the Sender with every later handler.
fn store() -> &'static Sender<Req> {
STORE_TX.get_or_init(|| {
let (tx, rx) = channel::<Req>();
smarm::spawn(move || store_loop(rx));
tx
})
}
// ---------------------------------------------------------------------------
// Handler
// ---------------------------------------------------------------------------
fn order(conn: Conn, _next: Next) -> Conn {
let id: u64 = conn.params.get("id").and_then(|s| s.parse().ok()).unwrap_or(0);
let (tx, rx) = channel::<u64>();
store().send(Req::Get { id, reply: tx }).ok();
let got = rx.recv().expect("store dropped");
// Render work: attributed to the enclosing `pipeline` site.
work_us(RENDER_US);
conn.put_status(200).put_body(got.to_string())
}
// ---------------------------------------------------------------------------
// Load generation — plain OS threads, blocking loopback TCP, keep-alive
// ---------------------------------------------------------------------------
/// Read one HTTP/1.1 response (headers + content-length body) into `buf`,
/// draining exactly what was consumed. Errors mean: reconnect.
fn read_response(s: &mut TcpStream, buf: &mut Vec<u8>) -> std::io::Result<()> {
let mut tmp = [0u8; 4096];
let header_end = loop {
if let Some(i) = buf.windows(4).position(|w| w == b"\r\n\r\n") {
break i + 4;
}
let n = s.read(&mut tmp)?;
if n == 0 {
return Err(std::io::ErrorKind::UnexpectedEof.into());
}
buf.extend_from_slice(&tmp[..n]);
};
let head = String::from_utf8_lossy(&buf[..header_end]).to_ascii_lowercase();
let len: usize = head
.lines()
.find_map(|l| l.strip_prefix("content-length:"))
.and_then(|v| v.trim().parse().ok())
.unwrap_or(0);
while buf.len() < header_end + len {
let n = s.read(&mut tmp)?;
if n == 0 {
return Err(std::io::ErrorKind::UnexpectedEof.into());
}
buf.extend_from_slice(&tmp[..n]);
}
buf.drain(..header_end + len);
Ok(())
}
fn load_loop(port: u16, stop: Arc<AtomicBool>) {
let mut n: u64 = 1;
'outer: while !stop.load(Ordering::Relaxed) {
let mut s = match TcpStream::connect(("127.0.0.1", port)) {
Ok(s) => s,
Err(_) => {
std::thread::sleep(Duration::from_millis(10));
continue;
}
};
s.set_read_timeout(Some(Duration::from_secs(5))).ok();
s.set_nodelay(true).ok();
let mut buf = Vec::with_capacity(4096);
while !stop.load(Ordering::Relaxed) {
let req = format!("GET /order/{n} HTTP/1.1\r\nhost: bench\r\n\r\n");
if s.write_all(req.as_bytes()).is_err() {
continue 'outer;
}
if read_response(&mut s, &mut buf).is_err() {
continue 'outer;
}
n += 1;
}
}
}
// ---------------------------------------------------------------------------
// main — boot, load, sweep, verdict
// ---------------------------------------------------------------------------
fn main() {
let per_us = calibrate_iters_per_us();
PER_US.set(per_us).expect("PER_US set twice");
println!("calibration: {per_us} work iters/µs");
println!(
"planted: store {STORE_US}µs serialized (ceiling ≈{} rps), render {RENDER_US}µs, {LOAD_CONNS} conns",
1_000_000 / STORE_US
);
println!("theory: store +33% @25%, +100% @50%; every other site ≈0");
// Ephemeral port via bind-and-release (integration-test pattern; the
// race window before the server rebinds is far below flake threshold).
let port = {
let l = TcpListener::bind("127.0.0.1:0").expect("bind probe");
l.local_addr().expect("local_addr").port()
};
let addr: SocketAddr = format!("127.0.0.1:{port}").parse().expect("addr");
let (handle, signal) = shutdown_handle();
let server = std::thread::spawn(move || {
let pipe = Pipeline::new().plug(Router::new().get("/order/:id", order));
serve_with_shutdown(Config::new(addr), smarm::Config::default(), pipe, Vec::new(), signal).expect("serve");
});
// Wait until it's accepting.
let up = (0..100).any(|_| {
std::thread::sleep(Duration::from_millis(50));
TcpStream::connect(addr).is_ok()
});
assert!(up, "server didn't come up on {addr}");
let stop = Arc::new(AtomicBool::new(false));
let loaders: Vec<_> = (0..LOAD_CONNS)
.map(|_| {
let stop = stop.clone();
std::thread::spawn(move || load_loop(port, stop))
})
.collect();
// Warm up: queues to steady state, every site + the progress point
// registered by real traffic before the sweep enumerates them.
std::thread::sleep(Duration::from_millis(500));
// The controller runs on this plain OS thread: its windows are wall
// time by construction here; inside a runtime they'd be wall-anchored
// via the efbc254 timer path (the thing the jobrunner sweep confirms).
let results = smarm::causal::run_experiments(&smarm::causal::ExperimentPlan {
speedups_pct: vec![0, 25, 50],
experiment: Duration::from_millis(700),
cooldown: Duration::from_millis(150),
});
stop.store(true, Ordering::Relaxed);
for l in loaders {
l.join().expect("loader panicked");
}
STORE_STOP.store(true, Ordering::Relaxed);
handle.shutdown();
server.join().expect("server panicked");
print!("{}", smarm::causal::render_summary(&results));
match std::fs::write("profile.coz", smarm::causal::render_coz(&results)) {
Ok(()) => println!("\nwrote profile.coz"),
Err(e) => eprintln!("\nfailed to write profile.coz: {e}"),
}
// Verdict — same discipline as the smarm demo: skip where the
// separation physically can't exist.
let cores = std::thread::available_parallelism().map(|n| n.get()).unwrap_or(1);
if cores < 4 {
println!("verdict: SKIPPED ({cores} cores; separation needs real parallelism)");
return;
}
let mut failures: Vec<String> = Vec::new();
let impact =
|site: &str| smarm::causal::impact_pct(&results, site, 25, "responses");
let mut expect = |site: &str, ok: &dyn Fn(f64) -> bool, want: &str| match impact(site) {
Some(p) => {
let verdict = if ok(p) { "ok" } else { "FAIL" };
println!("verdict: {site} @25% -> {p:+.1}% (want {want}) {verdict}");
if !ok(p) {
failures.push(format!("{site}: {p:+.1}% (want {want})"));
}
}
None => {
println!("verdict: {site} @25% -> missing cell FAIL");
failures.push(format!("{site}: missing cell"));
}
};
expect("store", &|p| p > 15.0, "> +15%");
for site in ["parse", "router", "pipeline", "serialize", "socket-write"] {
expect(site, &|p| p < 10.0, "< +10%");
}
if failures.is_empty() {
println!("verdict: PASS — planted bottleneck found, cold sites quiet");
} else {
println!("verdict: FAIL — {}", failures.join("; "));
std::process::exit(1);
}
}
+117
View File
@@ -0,0 +1,117 @@
//! Phoenix-style channels (v0.6): join/leave/event protocol over the
//! v0.4 WebSocket duplex, pubsub underneath, phoenix V2 JSON on the
//! wire.
//!
//! cargo run --example channels_chat --features phoenix
//! websocat ws://127.0.0.1:8080/socket (run two of these)
//! ["1","1","room:lobby","phx_join",{"name":"alice"}]
//! ["1","2","room:lobby","shout",{"text":"hi"}]
//! ["1","3","room:lobby","phx_leave",{}]
//!
//! The shapes this demonstrates:
//!
//! - **One `ChannelHub` for the whole app**, built up front. The hub is
//! a description — a bus name plus the routing table — and spawns
//! nothing; `hub.children()` is the vec of actors it needs (the bus
//! table, plus one registry per session route), handed to
//! `serve_with_shutdown` so the supervisor starts them ahead of the
//! endpoint and stops them after it has drained. Same shape as
//! `examples/ws_chat.rs`, see the docs there.
//!
//! - **A channel per joined topic, not per socket.** `Room` never sees
//! frames, refs, or the transport heartbeat — the conn-side handler
//! answers heartbeats, routes `phx_join`, and the per-join channel
//! actor (linked to the connection) runs the callbacks sequentially.
//!
//! - **`Json<T>` is the phoenix codec's opt-in newtype.** The
//! `"channels"` core is wire-neutral (`Encode`/`Decode`, zero deps);
//! `Json<Value>` here buys phoenix.js V2 interop. Typed payloads work
//! the same way: `Json<MyPayload>` for any serde type.
//!
//! - **`session::*` topics persist across reconnects** (opt-in via
//! `channel_session`): drop a websocat mid-room, reconnect and rejoin
//! within 30s, and the broadcasts you missed drain to you in order —
//! the buffer holds the relay path's own `Arc`s, nothing is copied.
//! The session is keyed by the `"name"` in the join payload.
use serde_json::{json, Value};
use urus::channels::phoenix::Json;
use urus::{
serve_with_shutdown, shutdown_handle, Channel, ChannelHub, ChannelSession, ChannelSocket,
Config, Conn, Next, Pipeline, PrefixRouter, Router, Status,
};
type P = Json<Value>;
/// One chat room. State lives across callbacks (and, for `session:*`
/// topics, across transports).
#[derive(Default)]
struct Room {
name: String,
msgs_seen: u64,
}
impl Channel<P> for Room {
fn join(&mut self, topic: &str, payload: P, _s: &ChannelSocket<P>) -> Result<P, P> {
// Re-entrant by contract for session topics: a rejoin lands
// here again on the same instance — re-auth, then ack.
self.name = payload.0["name"].as_str().unwrap_or("anon").to_string();
Ok(Json(json!({ "joined": topic, "as": self.name, "seen": self.msgs_seen })))
}
fn handle_in(&mut self, event: &str, payload: P, s: &ChannelSocket<P>) {
match event {
"shout" => {
self.msgs_seen += 1;
s.broadcast("shouted", Json(json!({ "from": self.name, "msg": payload.0 })));
}
_ => s.reply(Status::Ok, Json(json!({ "echo": event }))),
}
}
fn terminate(&mut self) {
println!("room channel for {:?} is over", self.name);
}
}
/// Session config for `session:*`: the join payload's `"name"` is the
/// session token. Defaults: 128 buffered broadcasts, 30s TTL.
struct BySessionName;
impl ChannelSession<P> for BySessionName {
type Key = String;
fn session_key(topic: &str, join_payload: &P) -> String {
format!("{topic}/{}", join_payload.0["name"].as_str().unwrap_or("anon"))
}
}
fn main() {
let (hub, children) = ChannelHub::new(
"chat-bus",
PrefixRouter::new()
.channel_default::<Room>("room:*")
.channel_session::<BySessionName>("session:*", "chat-sessions", |_: &str| {
Box::new(Room::default()) as Box<dyn Channel<P>>
}),
);
let router = Router::new().get("/socket", move |c: Conn, _n: Next| hub.upgrade(c));
let (handle, signal) = shutdown_handle();
std::thread::spawn(move || {
println!("channels_chat on ws://127.0.0.1:8080/socket — press Enter to shut down");
let mut line = String::new();
let _ = std::io::stdin().read_line(&mut line);
handle.shutdown();
});
serve_with_shutdown(
Config::new("0.0.0.0:8080".parse().unwrap()),
smarm::Config::default(),
Pipeline::new().plug(router),
children,
signal,
)
.unwrap();
println!("channels_chat: drained, bye");
}
+118 -44
View File
@@ -1,11 +1,18 @@
//! CRUD example: a tiny user database with JSON persistence. //! CRUD example: a tiny user database with JSON persistence.
//! //!
//! Demonstrates urus and the actor model together: //! Demonstrates urus and the actor model together, in the shape a real
//! - The pipeline is shared (Arc) across all connection actors. //! application should use (v0.3):
//! - Handlers do NOT take a lock or share mutable state directly. //! - The APP owns the smarm runtime and the root supervisor. urus is one
//! - A single "store" actor owns the data; handlers send it a request //! child in that tree — `urus::endpoint(...)` — and the store actor is
//! via a channel and block on the reply. Serialization is structural — //! an ordered sibling started BEFORE it, so the supervisor's
//! the store processes one request at a time, no Mutex needed. //! reverse-order shutdown drains HTTP first and only then stops the
//! store. No handler can be mid-request against a store that is gone.
//! - Handlers do NOT take a lock or share mutable state directly. A
//! single "store" actor owns the data; handlers address it by
//! registered name and block on a reply channel. Serialization is
//! structural — one request at a time, no Mutex.
//! - Both actors are supervised: kill the store (or let it panic) and it
//! restarts from the JSON file, with the endpoint left alone.
//! - On every mutating request the store writes the JSON file. Read //! - On every mutating request the store writes the JSON file. Read
//! requests don't touch disk. //! requests don't touch disk.
//! //!
@@ -23,9 +30,8 @@
//! curl -s http://localhost:8080/users/1 //! curl -s http://localhost:8080/users/1
use serde::{Deserialize, Serialize}; use serde::{Deserialize, Serialize};
use smarm::{channel, Sender}; use smarm::{channel, ChildSpec, Name, OneForOne, Restart, Sender};
use std::sync::OnceLock; use urus::{Config, Conn, Next, Pipeline, Router};
use urus::{serve_with, Config, Conn, Next, Pipeline, Router};
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// Domain // Domain
@@ -66,7 +72,12 @@ const DB_PATH: &str = "/tmp/urus-crud.json";
// Store actor body // Store actor body
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
fn store_loop(rx: smarm::Receiver<Request>) { fn store_loop() {
let (tx, rx) = channel::<Request>();
// Self-registration: the name is bound before the first recv, and it
// is re-bound automatically on every restart.
smarm::register(STORE, tx).expect("crud.store name already taken");
// Load on start. Missing file = empty store. Corrupt file = panic; we // Load on start. Missing file = empty store. Corrupt file = panic; we
// don't auto-rebuild because silently losing data is worse than failing // don't auto-rebuild because silently losing data is worse than failing
// loud. // loud.
@@ -78,8 +89,10 @@ fn store_loop(rx: smarm::Receiver<Request>) {
let mut next_id: u64 = users.iter().map(|u| u.id).max().unwrap_or(0) + 1; let mut next_id: u64 = users.iter().map(|u| u.id).max().unwrap_or(0) + 1;
loop { loop {
// A plain park. The supervisor stops this actor at shutdown (after
// the endpoint has drained), so there is nothing to poll for.
let req = match rx.recv() { let req = match rx.recv() {
Ok(r) => r, Ok(r) => r,
Err(_) => return, // all senders dropped Err(_) => return, // all senders dropped
}; };
match req { match req {
@@ -158,19 +171,29 @@ fn persist(users: &[User]) {
// Handler helpers // Handler helpers
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// //
// Once-cell trick: the store actor is spawned the first time a handler // The store is a supervised child that registers its own inbox under a
// runs (smarm requires `spawn` to be called from inside an actor — which // typed name; handlers resolve it per send. That replaces the old
// connection actors are). After that all handlers share the same Sender. // `OnceLock<Sender>` spawn-on-first-use trick — which had no supervisor,
// Simpler than threading the Sender through the pipeline at startup. // no restart, and no defined shutdown point — with a plain actor whose
// lifecycle the tree owns. A restart re-registers the same name, so
// in-flight handlers heal on their next send.
static STORE_TX: OnceLock<Sender<Request>> = OnceLock::new(); const STORE: Name<Request> = Name::new("crud.store");
fn store() -> &'static Sender<Request> { /// Send to the store and wait for its reply. `Err` only if the store is
STORE_TX.get_or_init(|| { /// between incarnations (restarting); the handler turns that into a 503
let (tx, rx) = channel::<Request>(); /// rather than pretending.
smarm::spawn(move || store_loop(rx)); fn ask(make: impl FnOnce(Sender<(u16, Vec<u8>)>) -> Request) -> Option<(u16, Vec<u8>)> {
tx let (tx, rx) = channel::<(u16, Vec<u8>)>();
}) smarm::send(STORE, make(tx)).ok()?;
rx.recv().ok()
}
fn reply(conn: Conn, r: Option<(u16, Vec<u8>)>) -> Conn {
match r {
Some((status, body)) => json(conn, status, body),
None => json(conn, 503, b"{\"error\":\"store unavailable\"}".to_vec()),
}
} }
fn json(conn: Conn, status: u16, body: Vec<u8>) -> Conn { fn json(conn: Conn, status: u16, body: Vec<u8>) -> Conn {
@@ -187,19 +210,34 @@ fn parse_id(s: &str) -> Option<u64> {
// Handlers // Handlers
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
/// SSE demo: `curl -N localhost:8080/ticker` streams a tick every second
/// (with `: keep-alive` comments if it ever goes quiet). The producer
/// exits on SseClosed — client gone, write timeout, or the drain stopping
/// its connection actor. Nothing to flag: the tree's shutdown reaches it.
fn ticker(conn: Conn, _next: Next) -> Conn {
let (conn, events) = conn.sse();
smarm::spawn(move || {
let mut n: u64 = 0;
loop {
if events.send("tick", &n.to_string()).is_err() {
return; // SseClosed
}
n += 1;
smarm::sleep(std::time::Duration::from_secs(1));
}
});
conn
}
fn list(conn: Conn, _next: Next) -> Conn { fn list(conn: Conn, _next: Next) -> Conn {
let (tx, rx) = channel::<(u16, Vec<u8>)>(); let r = ask(|reply| Request::List { reply });
store().send(Request::List { reply: tx }).ok(); reply(conn, r)
let (status, body) = rx.recv().expect("store dropped");
json(conn, status, body)
} }
fn create(conn: Conn, _next: Next) -> Conn { fn create(conn: Conn, _next: Next) -> Conn {
let body = conn.body.as_bytes().to_vec(); let body = conn.body.as_bytes().to_vec();
let (tx, rx) = channel::<(u16, Vec<u8>)>(); let r = ask(|reply| Request::Create { body, reply });
store().send(Request::Create { body, reply: tx }).ok(); reply(conn, r)
let (status, body) = rx.recv().expect("store dropped");
json(conn, status, body)
} }
fn get_one(conn: Conn, _next: Next) -> Conn { fn get_one(conn: Conn, _next: Next) -> Conn {
@@ -207,10 +245,8 @@ fn get_one(conn: Conn, _next: Next) -> Conn {
Some(id) => id, Some(id) => id,
None => return json(conn, 400, b"{\"error\":\"bad id\"}".to_vec()), None => return json(conn, 400, b"{\"error\":\"bad id\"}".to_vec()),
}; };
let (tx, rx) = channel::<(u16, Vec<u8>)>(); let r = ask(|reply| Request::Get { id, reply });
store().send(Request::Get { id, reply: tx }).ok(); reply(conn, r)
let (status, body) = rx.recv().expect("store dropped");
json(conn, status, body)
} }
fn update(conn: Conn, _next: Next) -> Conn { fn update(conn: Conn, _next: Next) -> Conn {
@@ -219,10 +255,8 @@ fn update(conn: Conn, _next: Next) -> Conn {
None => return json(conn, 400, b"{\"error\":\"bad id\"}".to_vec()), None => return json(conn, 400, b"{\"error\":\"bad id\"}".to_vec()),
}; };
let body = conn.body.as_bytes().to_vec(); let body = conn.body.as_bytes().to_vec();
let (tx, rx) = channel::<(u16, Vec<u8>)>(); let r = ask(|reply| Request::Update { id, body, reply });
store().send(Request::Update { id, body, reply: tx }).ok(); reply(conn, r)
let (status, body) = rx.recv().expect("store dropped");
json(conn, status, body)
} }
fn delete(conn: Conn, _next: Next) -> Conn { fn delete(conn: Conn, _next: Next) -> Conn {
@@ -230,10 +264,8 @@ fn delete(conn: Conn, _next: Next) -> Conn {
Some(id) => id, Some(id) => id,
None => return json(conn, 400, b"{\"error\":\"bad id\"}".to_vec()), None => return json(conn, 400, b"{\"error\":\"bad id\"}".to_vec()),
}; };
let (tx, rx) = channel::<(u16, Vec<u8>)>(); let r = ask(|reply| Request::Delete { id, reply });
store().send(Request::Delete { id, reply: tx }).ok(); reply(conn, r)
let (status, body) = rx.recv().expect("store dropped");
json(conn, status, body)
} }
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -261,10 +293,52 @@ fn main() {
.post( "/users", create) .post( "/users", create)
.get( "/users/:id", get_one) .get( "/users/:id", get_one)
.put( "/users/:id", update) .put( "/users/:id", update)
.delete("/users/:id", delete), .delete("/users/:id", delete)
.get( "/ticker", ticker),
); );
let cfg = Config::new("127.0.0.1:8080".parse().unwrap()); let cfg = Config::new("127.0.0.1:8080".parse().unwrap());
// Bind here, on this thread: an address-in-use error is a startup
// failure, not an actor crash. The fds outlive any restart of the
// endpoint child.
let endpoint = urus::endpoint(cfg, pipeline).expect("bind 127.0.0.1:8080");
println!("urus-crud: DB at {DB_PATH}"); println!("urus-crud: DB at {DB_PATH}");
serve_with(cfg, pipeline).unwrap(); println!("urus-crud: listening on 127.0.0.1:8080 — press Enter to shut down");
let rt = smarm::init(smarm::Config::default());
let handle = rt.handle();
let (sup_tx, sup_rx) = std::sync::mpsc::channel();
// Graceful shutdown on stdin-Enter: a plain OS thread blocks on
// read_line and shuts the ROOT SUPERVISOR down. No signal-handling
// crate needed, and no urus-specific shutdown plumbing — this is
// exactly what a SIGTERM handler would do.
std::thread::spawn(move || {
let sup: smarm::Pid = sup_rx.recv().expect("supervisor pid");
let mut line = String::new();
let _ = std::io::stdin().read_line(&mut line);
println!("urus-crud: shutting down (draining in-flight requests)…");
handle.request_shutdown(sup);
});
rt.run(move || {
let sup = smarm::spawn(move || {
OneForOne::new()
// Store FIRST: reverse-order shutdown therefore stops it
// LAST, after the endpoint has finished draining.
.child(ChildSpec::new(Restart::Permanent, store_loop))
.child(
ChildSpec::new(Restart::Permanent, endpoint)
// The endpoint bounds its own drain with
// Config::drain_timeout; a shorter supervisor
// deadline would cut that drain in half.
.shutdown(smarm::supervisor::Shutdown::Infinity),
)
.run()
});
let _ = sup_tx.send(sup.pid());
let _ = sup.join();
});
println!("urus-crud: bye");
} }
+106
View File
@@ -0,0 +1,106 @@
//! Causal load-profiling target (RFC 007): the real urus hot path under
//! EXTERNAL load — no planted bottleneck, no in-process load generation.
//!
//! Where `causal_bench` validates the machinery against a planted,
//! known-answer store, this is the discovery tool: boot a bare server,
//! let an external generator (wrk, pinned to different cores) drive it,
//! sweep virtual speedups over the five lib sites (parse, router,
//! pipeline, serialize, socket-write), and report which one causally
//! limits throughput. External load is deliberate: the generator lives
//! outside the smarm runtime, so it cannot absorb injected virtual
//! delay — the same property `causal_bench` gets from plain OS threads.
//!
//! No pass/fail verdict — there is no known answer here. Trust gates
//! only: traffic actually flowed through the sweep, and the ledger
//! audit (printed) balances.
//!
//! Env:
//! URUS_PORT listen port (default 8080; binds 0.0.0.0)
//! COZ_OUT coz profile output path (default profile.coz)
//! WARMUP_MS settle time after first traffic, ms (default 2000)
//!
//! Orchestration contract (the jobrunner run.sh): start this, wait for
//! it to accept, point wrk at `GET /json/:id` with a duration that
//! outlives the sweep, and stop wrk when this process exits.
//!
//! cargo run --release --example load_profile --features smarm-causal
use std::net::SocketAddr;
use std::time::{Duration, Instant};
use urus::{serve_with_shutdown, shutdown_handle, Config, Conn, Next, Pipeline, Router};
fn env_u64(name: &str, default: u64) -> u64 {
std::env::var(name).ok().and_then(|v| v.parse().ok()).unwrap_or(default)
}
/// The whole handler: param parse + JSON render. Deliberately thin — the
/// subject is the lib path around it, not application work.
fn json_id(conn: Conn, _next: Next) -> Conn {
let id: u64 = conn.params.get("id").and_then(|s| s.parse().ok()).unwrap_or(0);
conn.put_status(200)
.put_header("content-type", "application/json")
.put_body(format!("{{\"id\":{id}}}"))
}
fn main() {
let port = env_u64("URUS_PORT", 8080) as u16;
let warmup = Duration::from_millis(env_u64("WARMUP_MS", 2000));
let coz_out = std::env::var("COZ_OUT").unwrap_or_else(|_| "profile.coz".into());
let addr: SocketAddr = format!("0.0.0.0:{port}").parse().expect("addr");
let (handle, signal) = shutdown_handle();
let server = std::thread::spawn(move || {
let pipe = Pipeline::new().plug(Router::new().get("/json/:id", json_id));
serve_with_shutdown(Config::new(addr), smarm::Config::default(), pipe, Vec::new(), signal).expect("serve");
});
// Readiness is the orchestrator's job (TCP probe); ours is to not
// sweep before real traffic has registered every site and the
// progress point — `run_experiments` enumerates *registered* sites.
// 100 completed responses guarantees full end-to-end coverage.
let responses = |snap: &[(String, u64)]| {
snap.iter().find(|(n, _)| n == "responses").map(|(_, c)| *c).unwrap_or(0)
};
let t0 = Instant::now();
let seen = loop {
let n = responses(&smarm::causal::progress_snapshot());
if n >= 100 {
break n;
}
assert!(
t0.elapsed() < Duration::from_secs(120),
"no load after 120s ({n} responses) — is the generator running?"
);
std::thread::sleep(Duration::from_millis(100));
};
println!("traffic up: {seen} responses; settling {warmup:?}");
std::thread::sleep(warmup);
// Same plan as causal_bench, for comparability across workloads.
let results = smarm::causal::run_experiments(&smarm::causal::ExperimentPlan {
speedups_pct: vec![0, 25, 50],
experiment: Duration::from_millis(700),
cooldown: Duration::from_millis(150),
});
handle.shutdown();
server.join().expect("server panicked");
print!("{}", smarm::causal::render_summary(&results));
print!("{}", smarm::causal::render_ledger_audit(&results));
match std::fs::write(&coz_out, smarm::causal::render_coz(&results)) {
Ok(()) => println!("\nwrote {coz_out}"),
Err(e) => eprintln!("\nfailed to write {coz_out}: {e}"),
}
// Trust gate: the sweep is meaningless if load didn't flow through it.
let total: u64 = results
.iter()
.flat_map(|r| r.deltas.iter())
.filter(|(n, _)| n == "responses")
.map(|(_, c)| c)
.sum();
println!("total responses across windows: {total}");
assert!(total > 0, "sweep saw zero progress");
}
+53
View File
@@ -0,0 +1,53 @@
//! Plain benchmark server: the `load_profile` request path with zero causal
//! machinery — the baseline half of the ka/close A/B matrix (the
//! throughput-inversion chase).
//!
//! Identical route, handler, and `Config` construction to `load_profile`;
//! differs only in having no `smarm-causal` feature, no sweep, and no
//! shutdown path — it serves until killed. Throughput is measured
//! externally (wrk).
//!
//! Env:
//! URUS_PORT listen port (default 8080; binds 0.0.0.0)
//! URUS_SCHED_THREADS smarm scheduler OS threads (default: smarm's own
//! default). Malformed values panic rather than
//! silently falling back — a benchmark knob that
//! quietly reverts to default poisons the cell.
//!
//! cargo run --release --example plain_serve
use std::net::SocketAddr;
use urus::{serve_with, Config, Conn, Next, Pipeline, Router};
/// Mirrors `load_profile`'s handler byte for byte: param parse + JSON render.
fn json_id(conn: Conn, _next: Next) -> Conn {
let id: u64 = conn.params.get("id").and_then(|s| s.parse().ok()).unwrap_or(0);
conn.put_status(200)
.put_header("content-type", "application/json")
.put_body(format!("{{\"id\":{id}}}"))
}
fn main() {
let port: u16 = std::env::var("URUS_PORT")
.ok()
.and_then(|v| v.parse().ok())
.unwrap_or(8080);
let addr: SocketAddr = format!("0.0.0.0:{port}").parse().expect("addr");
let sched_threads: Option<usize> = std::env::var("URUS_SCHED_THREADS").ok().map(|v| {
v.parse()
.unwrap_or_else(|_| panic!("URUS_SCHED_THREADS not a usize: {v:?}"))
});
// Audit line: lands in each bench cell's server.log so the effective
// scheduler count is recorded per cell, same discipline as mode-verify.
eprintln!("plain_serve: scheduler_threads={sched_threads:?}");
// Scheduler threads are a RUNTIME knob, so they live in smarm::Config,
// not urus's — an endpoint placed in someone else's tree could not
// honour them anyway.
let rt_cfg = match sched_threads {
Some(n) => smarm::Config::exact(n),
None => smarm::Config::default(),
};
let pipe = Pipeline::new().plug(Router::new().get("/json/:id", json_id));
serve_with(Config::new(addr), rt_cfg, pipe, Vec::new()).expect("serve");
}
+48
View File
@@ -0,0 +1,48 @@
//! Serve with an optional TOML config overlay.
//!
//! urus is a library and never presumes a config path or reads the
//! environment for one — the binary decides where the file lives and hands
//! the text to `Config::with_toml_str`. Here that's a `--config PATH` flag;
//! with no flag, the compiled defaults are used unchanged.
//!
//! Requires the `config-file` feature:
//! cargo run --example serve_toml --features config-file -- --config urus.toml
//!
//! Example urus.toml (all keys optional, sparse override; seconds):
//! head_timeout_secs = 15
//! body_timeout_secs = 300
//! body_burst_bytes = 4096
//! body_stall_timeout_secs = 20
use std::net::SocketAddr;
use urus::{serve_with, Config, Conn, Next, Pipeline, Router};
fn json_id(conn: Conn, _next: Next) -> Conn {
let id: u64 = conn.params.get("id").and_then(|s| s.parse().ok()).unwrap_or(0);
conn.put_status(200)
.put_header("content-type", "application/json")
.put_body(format!("{{\"id\":{id}}}"))
}
fn main() {
let addr: SocketAddr = "0.0.0.0:8080".parse().expect("addr");
let mut cfg = Config::new(addr);
// Minimal flag scan: `--config PATH`. No presumed default location.
let mut args = std::env::args().skip(1);
while let Some(arg) = args.next() {
if arg == "--config" {
let path = args.next().expect("--config needs a PATH");
let toml = std::fs::read_to_string(&path)
.unwrap_or_else(|e| panic!("reading config {path}: {e}"));
cfg = cfg
.with_toml_str(&toml)
.unwrap_or_else(|e| panic!("invalid config {path}: {e}"));
eprintln!("serve_toml: loaded config from {path}");
}
}
let pipe = Pipeline::new().plug(Router::new().get("/json/:id", json_id));
serve_with(cfg, smarm::Config::default(), pipe, Vec::new()).expect("serve");
}
+115
View File
@@ -0,0 +1,115 @@
//! WebSocket chat rooms (v0.7): `urus::pubsub` wired to the v0.4 duplex.
//!
//! cargo run --example ws_chat
//! websocat ws://127.0.0.1:8080/chat/lobby (run two of these)
//!
//! The shapes this demonstrates:
//!
//! - **One `PubSub<String>` for the whole app**, topics are rooms
//! (`room:{name}`). `BUS` is a `const`: the handle is an address, not
//! the table. The table actor is `BUS.child()`, handed to
//! `serve_with_shutdown` as an app child, so it starts before the
//! endpoint and — shutdown being ordered in reverse — stops after the
//! endpoint has drained.
//!
//! The v0.5 idiom this replaces was a non-static `Arc<OnceLock<PubSub>>`
//! lazily initialised from the first connection, because `PubSub::new()`
//! used to spawn the table and the handle used to own its life. Neither
//! is true any more.
//!
//! - **`on_open` subscribes and spawns the relay** — a listen-only
//! client receives the room without ever sending. The subscription is
//! pinned to the CONNECTION actor (`on_open` runs inside it), so the
//! monitor cleans up exactly when the connection dies.
//!
//! - **The relay holds the `Receiver` and a `WsSender` clone.** It may
//! also hold `BUS` — a handle pins nothing now — but it has no use for
//! one. Conn dies → monitor prunes → sender drops → relay's recv errs
//! → relay exits.
use urus::{
serve_with_shutdown, shutdown_handle, Config, Conn, Message, Next, Pipeline, PubSub, Router,
WsHandler, WsSender,
};
const BUS: PubSub<String> = PubSub::new("chat");
struct ChatHandler {
room: String,
}
impl ChatHandler {
fn topic(&self) -> String {
format!("room:{}", self.room)
}
}
impl WsHandler for ChatHandler {
fn on_open(&mut self, sender: &WsSender) {
let topic = self.topic();
let rx = match BUS.subscribe(&topic) {
Ok(rx) => rx,
Err(_) => {
let _ = sender.close(1011, "chat bus down");
return;
}
};
let _ = BUS.broadcast_from(smarm::self_pid(), &topic, format!("* someone joined {topic}"));
// The relay: room messages -> this socket. Exits when the
// subscription is pruned (conn death / unsubscribe) or the
// socket is gone (WsClosed).
let out = sender.clone();
smarm::spawn(move || {
while let Ok(msg) = rx.recv() {
if out.send(Message::Text((*msg).clone())).is_err() {
break; // client gone or server draining
}
}
});
}
fn on_message(&mut self, msg: Message, sender: &WsSender) {
let Message::Text(text) = msg else {
let _ = sender.close(1003, "text only");
return;
};
// broadcast_from: the sender's own relay is skipped — no echo.
// self_pid() here is the connection actor, the pid on_open
// subscribed as.
let _ = BUS.broadcast_from(smarm::self_pid(), self.topic(), text);
}
fn on_close(&mut self, _code: Option<u16>, _reason: &str) {
let topic = self.topic();
let _ = BUS.broadcast_from(smarm::self_pid(), &topic, format!("* someone left {topic}"));
// No explicit unsubscribe: the connection actor is about to
// exit and the monitor prunes the subscription (which is also
// what stops the relay).
}
}
fn main() {
let pipeline = Pipeline::new().plug(Router::new().get("/chat/:room", |c: Conn, _n: Next| {
let room = c.params.get("room").unwrap_or("lobby").to_string();
c.upgrade(ChatHandler { room })
}));
let (handle, signal) = shutdown_handle();
std::thread::spawn(move || {
println!("ws_chat on ws://127.0.0.1:8080/chat/:room — press Enter to shut down");
let mut line = String::new();
let _ = std::io::stdin().read_line(&mut line);
handle.shutdown();
});
serve_with_shutdown(
Config::new("0.0.0.0:8080".parse().unwrap()),
smarm::Config::default(),
pipeline,
vec![BUS.child()],
signal,
)
.unwrap();
println!("ws_chat: drained, bye");
}
+103
View File
@@ -0,0 +1,103 @@
//! WebSocket echo server (v0.4).
//!
//! cargo run --example ws_echo
//! websocat ws://127.0.0.1:8080/echo
//!
//! Demonstrates the chunk-3 handler model:
//! - `EchoHandler` runs INSIDE the connection actor's select loop — echo
//! is exactly the workload that wants zero extra moving parts.
//! - `/clock` shows the other shape: the handler spawns a producer actor
//! at `on_message("start")` and hands it a `WsSender` clone — the
//! SSE-producer pattern, verbatim. The producer exits when its send
//! fails (`WsClosed`: client gone or server shutting down).
//!
//! Routing happens before the upgrade, so one server can host both
//! endpoints — `Conn::upgrade(handler)` is just another thing a route
//! handler returns.
use std::time::Duration;
use urus::{
serve_with_shutdown, shutdown_handle, Config, Conn, Message, Next, Pipeline, Router,
WsHandler, WsSender,
};
// ---------------------------------------------------------------------------
// /echo — everything comes straight back
// ---------------------------------------------------------------------------
struct EchoHandler;
impl WsHandler for EchoHandler {
fn on_message(&mut self, msg: Message, sender: &WsSender) {
if matches!(&msg, Message::Text(t) if t == "bye") {
let _ = sender.close(1000, "you said bye");
return;
}
let _ = sender.send(msg);
}
fn on_close(&mut self, code: Option<u16>, reason: &str) {
println!("ws_echo: /echo closed (code {code:?}, reason {reason:?})");
}
}
// ---------------------------------------------------------------------------
// /clock — "start" spawns a producer actor ticking once a second
// ---------------------------------------------------------------------------
struct ClockHandler {
started: bool,
}
impl WsHandler for ClockHandler {
fn on_message(&mut self, msg: Message, sender: &WsSender) {
match msg {
Message::Text(t) if t == "start" && !self.started => {
self.started = true;
let out = sender.clone();
smarm::spawn(move || {
let mut n = 0u64;
loop {
if out.text(format!("tick {n}")).is_err() {
return; // WsClosed: connection over, we follow
}
n += 1;
smarm::sleep(Duration::from_secs(1));
}
});
}
_ => {
let _ = sender.text("send \"start\" to begin");
}
}
}
}
// ---------------------------------------------------------------------------
// main
// ---------------------------------------------------------------------------
fn main() {
let pipeline = Pipeline::new().plug(
Router::new()
.get("/echo", |c: Conn, _n: Next| c.upgrade(EchoHandler))
.get("/clock", |c: Conn, _n: Next| {
c.upgrade(ClockHandler { started: false })
}),
);
let cfg = Config::new("127.0.0.1:8080".parse().unwrap());
println!("ws_echo: ws://127.0.0.1:8080/echo and /clock — press Enter to shut down");
let (handle, signal) = shutdown_handle();
std::thread::spawn(move || {
let mut line = String::new();
let _ = std::io::stdin().read_line(&mut line);
println!("ws_echo: shutting down…");
handle.shutdown();
});
serve_with_shutdown(cfg, smarm::Config::default(), pipeline, Vec::new(), signal).unwrap();
println!("ws_echo: bye");
}
+89
View File
@@ -0,0 +1,89 @@
#!/usr/bin/env bash
# E1: does the ka low-concurrency latency floor track idle-scheduler count?
#
# Mechanism under test (smarm): one shared level-triggered wake pipe; every
# completion byte wakes ALL idle schedulers; one drain-lock winner; losers
# stampede the timers/io/queue mutexes and re-sleep; `enqueue` is silent.
# Prediction if right: at c <= 8, ka latency tail shrinks and throughput
# rises as scheduler count drops (fewer idle pollers -> smaller herd);
# close mode (control) stays flat or worsens as threads drop.
#
# Sweeps URUS_SCHED_THREADS x CONNS over the plain ka/close matrix,
# PLAIN_ONLY — causal cells are irrelevant to E1. wrk side is byte-identical
# to the 96b40ad5 conns sweep (THREADS=4, --latency) for comparability.
#
# Knobs: THREADS_SET CONNS_SET DUR REPS OUT (+ everything the matrix takes)
set -euo pipefail
cd "$(dirname "$0")/.."
THREADS_SET="${THREADS_SET:-1 2 4 8}"
CONNS_SET="${CONNS_SET:-4 8}"
DUR="${DUR:-15}"
REPS="${REPS:-2}"
OUT="${OUT:-/workspace/results}"
mkdir -p "$OUT"
say() { echo "[$(date +%H:%M:%S)] $*" | tee -a "$OUT/e1.log"; }
for t in $THREADS_SET; do
for c in $CONNS_SET; do
cell="$OUT/t${t}-c${c}"
say "=== E1 cell: sched_threads=$t conns=$c -> $cell ==="
URUS_SCHED_THREADS="$t" PLAIN_ONLY=1 \
DUR="$DUR" REPS="$REPS" CONNS="$c" OUT="$cell" \
bash scripts/ka-close-matrix.sh
# Cell audit: the knob must have actually reached the server. A cell
# whose server silently ran at default threads poisons the sweep — the
# exact failure class per-cell verification exists to catch.
for m in ka close; do
if ! grep -q "scheduler_threads=Some($t)" "$cell/plain-$m/server.log"; then
say "FATAL: t=$t c=$c mode=$m server.log lacks scheduler_threads=Some($t)"
exit 1
fi
done
say "cell t=$t c=$c audit OK"
done
done
say "=== E1 SUMMARY ==="
python3 - "$OUT" <<'PY' | tee -a "$OUT/e1.log"
import re, sys, pathlib, statistics
out = pathlib.Path(sys.argv[1])
def ms(tok):
m = re.match(r"([\d.]+)(us|ms|s)$", tok)
if not m:
return None
v = float(m.group(1))
return {"us": v / 1000, "ms": v, "s": v * 1000}[m.group(2)]
def cell_stats(d):
reps, pcts = [], {}
for rep in sorted(d.glob("rep*.txt")):
t = rep.read_text()
m = re.search(r"Requests/sec:\s+([\d.]+)", t)
if m:
reps.append(float(m.group(1)))
for p, tok in re.findall(r"^\s+(50|75|90|99)%\s+(\S+)$", t, re.M):
pcts.setdefault(p, []).append(ms(tok))
if not reps:
return None
return (statistics.mean(reps),
{p: statistics.mean([v for v in vs if v is not None])
for p, vs in pcts.items()})
cells = sorted(out.glob("t*-c*"))
hdr = f"{'cell':>10} {'mode':>6} {'req/s':>9} {'p50ms':>7} {'p75ms':>7} {'p90ms':>7} {'p99ms':>7}"
print(hdr); print("-" * len(hdr))
for cell in cells:
for mode in ("ka", "close"):
s = cell_stats(cell / f"plain-{mode}")
if s is None:
continue
rps, p = s
print(f"{cell.name:>10} {mode:>6} {rps:>9.0f} "
f"{p.get('50', float('nan')):>7.3f} {p.get('75', float('nan')):>7.3f} "
f"{p.get('90', float('nan')):>7.3f} {p.get('99', float('nan')):>7.3f}")
PY
say "=== E1 DONE ==="
+59
View File
@@ -0,0 +1,59 @@
#!/usr/bin/env bash
# Flake hammer: the conn-lifecycle test discipline, in one command.
#
# Default (no args) = the full pre-commit ritual for conn-lifecycle changes:
# 1. N x the lifecycle-sensitive integration subset (fail-fast)
# 2. M x the full suite (unit + integration + doc)
#
# Usage:
# scripts/hammer.sh # 35x subset + 3x full
# scripts/hammer.sh -n 50 -f 5 # override counts
# scripts/hammer.sh -n 20 -f 0 -- shutdown sse # custom filter, skip full runs
#
# On failure: stops immediately, output of the failing run is in $LOG.
set -u
cd "$(dirname "$0")/.."
export PATH="$HOME/.cargo/bin:$PATH"
N=35
FULL=3
SUBSET=(shutdown timeout reaped slowloris streaming chunked sse stalled ws_)
while [[ $# -gt 0 ]]; do
case "$1" in
-n) N="$2"; shift 2 ;;
-f) FULL="$2"; shift 2 ;;
--) shift; SUBSET=("$@"); break ;;
*) echo "unknown arg: $1" >&2; exit 2 ;;
esac
done
LOG=target/hammer-fail.log
mkdir -p target
echo "== build once =="
cargo test --no-run --quiet || exit 1
fail() { # $1 = label
echo
echo "FAILED at $1 — output in $LOG"
exit 1
}
echo "== ${N}x subset: ${SUBSET[*]} =="
for ((i = 1; i <= N; i++)); do
printf '\r subset %d/%d ' "$i" "$N"
cargo test --quiet --test integration -- "${SUBSET[@]}" >"$LOG" 2>&1 \
|| fail "subset run $i/$N"
done
echo
echo "== ${FULL}x full suite =="
for ((i = 1; i <= FULL; i++)); do
printf ' full %d/%d\n' "$i" "$FULL"
cargo test --quiet >"$LOG" 2>&1 || fail "full run $i/$FULL"
done
rm -f "$LOG"
echo "hammer green: ${N}x subset + ${FULL}x full"
+181
View File
@@ -0,0 +1,181 @@
#!/usr/bin/env bash
# ka-close-matrix.sh — controlled ka/close A/B for the throughput-inversion
# chase (handoff jar item), with the forgiveness-fix box validation folded
# into the causal cells (handoff v13 PENDING, pull-forward agreed 2026-07-20).
#
# Cells, run in order on fresh ports:
# plain-ka, plain-close examples/plain_serve (no causal feature).
# Metric: wrk Requests/sec, REPS reps after a
# discarded warmup.
# causal-ka, causal-close examples/load_profile (smarm-causal). Metric:
# the sweep summary + ledger audit (forgiveness
# column, books balance); wrk is backdrop load
# whose own numbers are injection-contaminated.
#
# Controls: byte-identical wrk invocation per mode pair except the
# `Connection: close` header; server and wrk pinned to disjoint core sets
# (SMT siblings left idle by default on the 5900X); the negotiated
# connection behavior is verified via curl per cell and logged, so a
# loadgen-config asymmetry can never silently explain a result again.
#
# Knobs (env, defaults for the 5900X box):
# DUR=30 REPS=2 CONNS=64 THREADS=4 PIN=1
# SERVER_CPUS=0-7 WRK_CPUS=8-11
# CAUSAL_WRK_DUR=300 BASE_PORT=8080
# OUT=/workspace/results
# PLAIN_BIN, CAUSAL_BIN binary paths (default: target/release/examples/*)
set -euo pipefail
cd "$(dirname "$0")/.."
DUR="${DUR:-30}"
REPS="${REPS:-2}"
CONNS="${CONNS:-64}"
THREADS="${THREADS:-4}"
PIN="${PIN:-1}"
SERVER_CPUS="${SERVER_CPUS:-0-7}"
WRK_CPUS="${WRK_CPUS:-8-11}"
CAUSAL_WRK_DUR="${CAUSAL_WRK_DUR:-300}"
BASE_PORT="${BASE_PORT:-8080}"
OUT="${OUT:-/workspace/results}"
PLAIN_BIN="${PLAIN_BIN:-target/release/examples/plain_serve}"
CAUSAL_BIN="${CAUSAL_BIN:-target/release/examples/load_profile}"
mkdir -p "$OUT"
say() { echo "[$(date +%H:%M:%S)] $*" | tee -a "$OUT/matrix.log"; }
pin_server=(); pin_wrk=()
if [ "$PIN" = 1 ]; then
pin_server=(taskset -c "$SERVER_CPUS")
pin_wrk=(taskset -c "$WRK_CPUS")
fi
# ---- environment record --------------------------------------------------
{
date -u
uname -a
echo "nproc: $(nproc)"
lscpu -e 2>/dev/null || true
echo "port range: $(cat /proc/sys/net/ipv4/ip_local_port_range 2>/dev/null)"
echo "somaxconn: $(cat /proc/sys/net/core/somaxconn 2>/dev/null)"
echo "wrk: $(wrk --version 2>&1 | head -1 || true)"
echo "PIN=$PIN SERVER_CPUS=$SERVER_CPUS WRK_CPUS=$WRK_CPUS"
echo "DUR=$DUR REPS=$REPS CONNS=$CONNS THREADS=$THREADS CAUSAL_WRK_DUR=$CAUSAL_WRK_DUR"
} > "$OUT/env.txt"
say "env recorded -> env.txt"
# ---- helpers -------------------------------------------------------------
run_wrk() { # $1=mode $2=duration_s $3=port
if [ "$1" = close ]; then
"${pin_wrk[@]}" wrk -t "$THREADS" -c "$CONNS" -d "${2}s" --latency \
-H "Connection: close" "http://127.0.0.1:$3/json/7"
else
"${pin_wrk[@]}" wrk -t "$THREADS" -c "$CONNS" -d "${2}s" --latency \
"http://127.0.0.1:$3/json/7"
fi
}
wait_port() { # $1=port
for _ in $(seq 1 150); do
curl -s -o /dev/null "http://127.0.0.1:$1/json/1" && return 0
sleep 0.2
done
return 1
}
verify_mode() { # $1=mode $2=port — record what actually goes over the wire
echo "--- single request, mode=$1 ---"
if [ "$1" = close ]; then
curl -sv --http1.1 -H "Connection: close" -o /dev/null \
"http://127.0.0.1:$2/json/7" 2>&1 | grep -iE "^(> |< )(GET|HTTP|connection)" || true
else
curl -sv --http1.1 -o /dev/null "http://127.0.0.1:$2/json/7" 2>&1 \
| grep -iE "^(> |< )(GET|HTTP|connection)" || true
fi
echo "--- reuse probe (two requests, one curl) ---"
curl -sv --http1.1 -o /dev/null -o /dev/null \
"http://127.0.0.1:$2/json/1" "http://127.0.0.1:$2/json/2" 2>&1 \
| grep -icE "re-us(ed|ing)" || true
}
# ---- plain cells ---------------------------------------------------------
run_plain() { # $1=mode $2=port
local mode="$1" port="$2" d="$OUT/plain-$1"
mkdir -p "$d"
say "=== plain / $mode (port $port) ==="
URUS_PORT="$port" "${pin_server[@]}" "$PLAIN_BIN" > "$d/server.log" 2>&1 &
local spid=$!
wait_port "$port" || { say "FATAL: plain server never came up"; cat "$d/server.log"; exit 1; }
verify_mode "$mode" "$port" > "$d/mode-verify.txt" 2>&1
ss -s > "$d/ss-before.txt" 2>/dev/null || true
run_wrk "$mode" 5 "$port" > "$d/warmup.txt" 2>&1
for r in $(seq 1 "$REPS"); do
run_wrk "$mode" "$DUR" "$port" > "$d/rep$r.txt" 2>&1
grep -E "Requests/sec|Latency |requests in|Socket errors|Non-2xx" "$d/rep$r.txt" \
| sed "s/^/ [plain-$mode r$r] /" | tee -a "$OUT/matrix.log" || true
done
ss -s > "$d/ss-after.txt" 2>/dev/null || true
awk '{printf "server cpu jiffies (utime+stime): %d\n", $14+$15}' \
"/proc/$spid/stat" > "$d/server-cpu.txt" 2>/dev/null || true
kill "$spid" 2>/dev/null || true
wait "$spid" 2>/dev/null || true
}
# ---- causal cells --------------------------------------------------------
run_causal() { # $1=mode $2=port
local mode="$1" port="$2" d="$OUT/causal-$1"
mkdir -p "$d"
say "=== causal / $mode (port $port) ==="
URUS_PORT="$port" COZ_OUT="$d/profile.coz" WARMUP_MS=2000 \
"${pin_server[@]}" "$CAUSAL_BIN" > "$d/sweep.log" 2>&1 &
local spid=$!
wait_port "$port" || { say "FATAL: causal server never came up"; cat "$d/sweep.log"; exit 1; }
verify_mode "$mode" "$port" > "$d/mode-verify.txt" 2>&1
run_wrk "$mode" "$CAUSAL_WRK_DUR" "$port" > "$d/wrk-backdrop.txt" 2>&1 &
local wpid=$!
local rc=0
wait "$spid" || rc=$?
kill "$wpid" 2>/dev/null || true
wait "$wpid" 2>/dev/null || true
say "causal/$mode server exit=$rc (sweep + audit in sweep.log)"
if [ "$rc" -ne 0 ]; then say "WARNING: causal/$mode exited nonzero"; fi
}
# ---- matrix --------------------------------------------------------------
run_plain ka "$BASE_PORT"
run_plain close "$((BASE_PORT + 1))"
if [ "${PLAIN_ONLY:-0}" != 1 ]; then
run_causal ka "$((BASE_PORT + 2))"
run_causal close "$((BASE_PORT + 3))"
fi
# ---- summary -------------------------------------------------------------
say "=== SUMMARY ==="
python3 - "$OUT" <<'PY' | tee -a "$OUT/matrix.log"
import re, sys, pathlib, statistics
out = pathlib.Path(sys.argv[1])
means = {}
for mode in ("ka", "close"):
vals = []
for rep in sorted((out / f"plain-{mode}").glob("rep*.txt")):
m = re.search(r"Requests/sec:\s+([\d.]+)", rep.read_text())
if m:
vals.append(float(m.group(1)))
if vals:
means[mode] = statistics.mean(vals)
print(f"plain-{mode}: reps={[f'{v:.0f}' for v in vals]} mean={means[mode]:.0f} req/s")
if len(means) == 2:
r = means["close"] / means["ka"]
print(f"close/ka ratio: {r:.3f}", "-> INVERSION PRESENT (close beats ka)" if r > 1.05
else "-> no inversion (ka >= close)" if r < 0.95 else "-> parity")
for mode in ("ka", "close"):
log = out / f"causal-{mode}" / "sweep.log"
if log.exists():
t = log.read_text()
keep = [l for l in t.splitlines()
if re.search(r"forgiv|books|audit|balance|total responses", l, re.I)]
print(f"-- causal-{mode} audit lines --")
for l in keep[:14]:
print(" " + l)
PY
say "=== MATRIX DONE ==="
+1100
View File
File diff suppressed because it is too large Load Diff
+223
View File
@@ -0,0 +1,223 @@
//! Phoenix V2 wire codec — feature `"phoenix"` (implies `"channels"`,
//! spends dependency slot #4 on `serde`/`serde_json`, ratified
//! 2026-06-12).
//!
//! The wire format phoenix.js speaks: one JSON array per frame,
//!
//! ```text
//! [join_ref, ref, topic, event, payload]
//! ```
//!
//! refs are strings or `null`; replies are `event: "phx_reply"` with
//! `payload: {"status": "ok"|"error", "response": <payload>}`.
//!
//! # `Json<T>`, not a blanket impl on `P`
//!
//! The ratified sketch said "JsonCodec blanket impl via serde". A
//! literal `impl<P: Serialize + DeserializeOwned> Encode for P` is
//! coherence poison: features are additive across a dependency graph,
//! so the moment ANY crate enables `"phoenix"`, every hand-written
//! `Encode` impl in every other crate stops compiling (the compiler
//! cannot prove a foreign type will never gain `Serialize`). That
//! breaks the design's own promise that `"channels"`-only users supply
//! their own codecs. The blanket therefore lives one newtype away:
//! [`Json<T>`] is V2-JSON for ANY serde-capable `T` — including
//! `Json<serde_json::Value>` for schemaless payloads — while the
//! `Encode`/`Decode` namespace stays open.
//!
//! # Strictness
//!
//! Decoding is exactly as strict as `T`'s `Deserialize`: a phoenix.js
//! `{}` join payload fails against a `T` with required fields (the
//! socket closes `1002`). Use `#[serde(default)]`, `Option` fields, or
//! `Json<serde_json::Value>` where the client is loose. Encoding a `T`
//! that fails `Serialize` (non-string map keys and friends) panics —
//! that is a bug in the payload type, and the panic rides the
//! documented channel-panic contract (linked teardown).
use super::{
ChannelFrame, ChannelMessage, Decode, DecodeError, Encode, FrameRef, MessageRef, Status,
EV_REPLY,
};
use crate::ws::Message;
use serde::de::DeserializeOwned;
use serde::Serialize;
use serde_json::{json, Map, Value};
/// V2-JSON payload wrapper: `Json<T>` speaks the phoenix wire for any
/// `T: Serialize + DeserializeOwned`. Derefs to `T`.
#[derive(Debug, Clone, PartialEq, Eq, Default)]
pub struct Json<T>(pub T);
impl<T> std::ops::Deref for Json<T> {
type Target = T;
fn deref(&self) -> &T {
&self.0
}
}
impl<T> std::ops::DerefMut for Json<T> {
fn deref_mut(&mut self) -> &mut T {
&mut self.0
}
}
impl<T> From<T> for Json<T> {
fn from(t: T) -> Self {
Json(t)
}
}
fn empty_object() -> Value {
Value::Object(Map::new())
}
fn to_value<T: Serialize>(t: &T) -> Value {
serde_json::to_value(t).expect("phoenix codec: payload failed to serialize")
}
impl<T: Serialize + DeserializeOwned + Send + Sync + 'static> Encode for Json<T> {
fn encode(f: FrameRef<'_, Self>) -> Message {
let (event, payload) = match f.message {
MessageRef::Event { event, payload } => {
(event, payload.map(|p| to_value(&p.0)).unwrap_or_else(empty_object))
}
MessageRef::Reply { status, payload } => (
EV_REPLY,
json!({
"status": match status {
Status::Ok => "ok",
Status::Error => "error",
},
"response": payload.map(|p| to_value(&p.0)).unwrap_or_else(empty_object),
}),
),
};
let arr = json!([f.join_ref, f.reference, f.topic, event, payload]);
Message::Text(arr.to_string())
}
}
impl<T: Serialize + DeserializeOwned + Send + Sync + 'static> Decode for Json<T> {
fn decode(msg: &Message) -> Result<ChannelFrame<Self>, DecodeError> {
let Message::Text(s) = msg else {
return Err(DecodeError("phoenix V2 frames are text".into()));
};
let (join_ref, reference, topic, event, payload): (
Option<String>,
Option<String>,
String,
String,
Value,
) = serde_json::from_str(s).map_err(|e| DecodeError(e.to_string()))?;
let message = if event == EV_REPLY {
// Clients don't reply in practice; decoded for symmetry
// (and ignored upstream).
let status = match payload.get("status").and_then(Value::as_str) {
Some("ok") => Status::Ok,
_ => Status::Error,
};
let response = payload.get("response").cloned().unwrap_or_else(empty_object);
let payload = serde_json::from_value::<T>(response)
.map_err(|e| DecodeError(format!("reply response: {e}")))?;
ChannelMessage::Reply { status, payload: Some(Json(payload)) }
} else {
let payload = serde_json::from_value::<T>(payload)
.map_err(|e| DecodeError(format!("payload for '{event}': {e}")))?;
ChannelMessage::Event { event, payload: Json(payload) }
};
Ok(ChannelFrame { join_ref, reference, topic, message })
}
}
#[cfg(test)]
mod tests {
use super::*;
type V = Json<Value>;
fn parse(m: &Message) -> Value {
let Message::Text(s) = m else { panic!("not text") };
serde_json::from_str(s).unwrap()
}
#[test]
fn decodes_phoenix_js_join_frame() {
let m = Message::Text(r#"["3","3","room:lobby","phx_join",{"token":"t"}]"#.into());
let f = V::decode(&m).unwrap();
assert_eq!(f.join_ref.as_deref(), Some("3"));
assert_eq!(f.reference.as_deref(), Some("3"));
assert_eq!(f.topic, "room:lobby");
let ChannelMessage::Event { event, payload } = f.message else { panic!() };
assert_eq!(event, "phx_join");
assert_eq!(payload.0, json!({"token": "t"}));
}
#[test]
fn decodes_heartbeat_null_refs_allowed() {
let m = Message::Text(r#"[null,"7","phoenix","heartbeat",{}]"#.into());
let f = V::decode(&m).unwrap();
assert_eq!(f.join_ref, None);
assert_eq!(f.topic, "phoenix");
}
#[test]
fn encodes_ok_reply_with_status_response_wrapper() {
let p = Json(json!({"user_count": 2}));
let m = V::encode(FrameRef {
join_ref: Some("3"),
reference: Some("4"),
topic: "room:lobby",
message: MessageRef::Reply { status: Status::Ok, payload: Some(&p) },
});
assert_eq!(
parse(&m),
json!(["3", "4", "room:lobby", "phx_reply",
{"status": "ok", "response": {"user_count": 2}}])
);
}
#[test]
fn encodes_payloadless_reply_and_event_as_empty_object() {
let m = V::encode(FrameRef {
join_ref: None,
reference: Some("9"),
topic: "phoenix",
message: MessageRef::Reply { status: Status::Ok, payload: None },
});
assert_eq!(
parse(&m),
json!([null, "9", "phoenix", "phx_reply", {"status": "ok", "response": {}}])
);
let m = V::encode(FrameRef {
join_ref: Some("3"),
reference: None,
topic: "room:lobby",
message: MessageRef::Event { event: "phx_close", payload: None },
});
assert_eq!(parse(&m), json!(["3", null, "room:lobby", "phx_close", {}]));
}
#[test]
fn typed_payload_round_trip_and_strictness() {
#[derive(serde::Serialize, serde::Deserialize, Debug, PartialEq)]
struct Chat {
body: String,
}
let m = Message::Text(r#"["1","2","room:a","new_msg",{"body":"hi"}]"#.into());
let f = Json::<Chat>::decode(&m).unwrap();
let ChannelMessage::Event { payload, .. } = f.message else { panic!() };
assert_eq!(payload.0, Chat { body: "hi".into() });
// Missing required field: the codec is as strict as T.
let m = Message::Text(r#"["1","2","room:a","new_msg",{}]"#.into());
assert!(Json::<Chat>::decode(&m).is_err());
}
#[test]
fn garbage_and_binary_are_decode_errors() {
assert!(V::decode(&Message::Text("not json".into())).is_err());
assert!(V::decode(&Message::Text(r#"{"a":1}"#.into())).is_err());
assert!(V::decode(&Message::Binary(vec![1, 2, 3])).is_err());
}
}
+780
View File
@@ -0,0 +1,780 @@
//! Opt-in channel session persistence (the ratified v0.6 design).
//!
//! Default channel lifetime is transport-bound: cold start per join,
//! actor dies with the socket. A [`ChannelSession`] impl registered via
//! [`PrefixRouter::channel_session`](super::PrefixRouter::channel_session)
//! changes that: the channel actor outlives its transport, buffering
//! outbound broadcasts (`VecDeque<Arc<Broadcast<P>>>` — the same `Arc`s
//! the pubsub relay path carries, zero re-allocation) until the client
//! rejoins, the buffer fills, or the TTL expires.
//!
//! Shape (one registry gen_server per `channel_session` registration,
//! monomorphic over the session key — no type erasure):
//!
//! - `deploy` on the conn actor computes the session key and casts the
//! whole join handshake to the registry.
//! - The registry owns `key -> (pid, control sender)`. Existing entry:
//! the handshake is forwarded as an `Attach` (reattach). Absent or
//! dead (send failure raced the monitor `Down`): spawn fresh,
//! monitor, insert. Down prunes, guarded by pid match against
//! replaced entries.
//! - The session actor is **not linked** to any connection — outliving
//! the transport is the point. Its exits: explicit leave, rejected
//! (re)join, TTL expiry, buffer cap, pubsub relay death, or the
//! registry's control sender dropping — which is exactly the shutdown
//! chain (the supervisor shuts the registry down after the endpoint
//! has drained -> the registry's state drops -> control senders drop
//! -> every parked session wakes on the closed control arm,
//! terminates, and exits; `AllDone` composes without links).
//!
//! Note this is the last lifetime in urus implied by a drop rather
//! than stated: sessions are dynamic (one per key), so a fixed
//! `ChildSpec` list cannot hold them. The trigger is now a command
//! rather than a refcount, which is what the v0.8 cycle was about, but
//! a dynamic supervisor (OTP's `simple_one_for_one`) is the honest
//! shape if smarm grows one.
//!
//! As-landed decisions (veto by diff):
//! - **Every attach calls `ch.join()` again** on the same instance —
//! the client's `phx_join` needs an ack payload and the channel gets
//! to re-auth. Session channels must treat `join` as re-entrant.
//! - **A rejected rejoin ends the session** (error reply, `terminate`,
//! exit): `Err` means "this transport may not have this channel", and
//! a zombie session held open for an unauthorized client is a
//! liability. Next join is a cold start.
//! - **A second transport evicts the first** (best-effort `phx_close`
//! on the old socket): a session is single-transport by definition;
//! two tabs wanting independent channels should not share a session
//! key.
//! - **TTL is per detach episode** (reset on every disconnect), and the
//! pubsub subscription is made once, on the first successful join, to
//! the first join's topic.
//! - **`ChannelSession::Key: Clone`** beyond the ratified
//! `Eq + Hash + Send` — the registry keeps a `pid -> key` reverse
//! index for monitor-`Down` pruning.
use super::*;
use smarm::gen_server::{self, GenServer, GenServerBuilder, GenServerCtx, GenServerName};
use smarm::{Down, Restart, Watcher};
use std::collections::VecDeque;
use std::hash::Hash;
use std::time::{Duration, Instant};
/// Opt-in session persistence config for a channel type. All methods
/// are static: the config is consulted before any channel instance
/// exists.
pub trait ChannelSession<P>: Send + 'static {
/// What identifies "the same session" across reconnects.
type Key: Eq + Hash + Send + 'static;
/// Derive the session key from a join. Same key, same (live)
/// channel actor; the join payload is the natural carrier for a
/// client-held session token.
fn session_key(topic: &str, join_payload: &P) -> Self::Key;
/// Detached broadcasts buffered before the session tears itself
/// down (next join is a cold start).
fn buffer_cap() -> usize {
128
}
/// How long a detached session waits for a rejoin before tearing
/// itself down. Reset on every detach.
fn ttl() -> Duration {
Duration::from_secs(30)
}
}
// ---------------------------------------------------------------------------
// The session-aware factory (what channel_session registers)
// ---------------------------------------------------------------------------
pub(super) struct SessionFactory<P: Encode + Decode + Send + Sync + 'static, K: SessionKey> {
inner: Arc<dyn ChannelFactory<P>>,
keyfn: fn(&str, &P) -> K,
/// The registry's registered name. Resolved per deploy, so a
/// registry restarted by the supervisor is reached transparently —
/// with its session map empty, which is the honest outcome: the
/// session actors it tracked died with their control senders.
registry: GenServerName<Registry<P, K>>,
/// Everything the registry's `ChildSpec` needs to build it again.
cap: usize,
ttl: Duration,
}
/// The registry's working bounds for a session key.
pub(super) trait SessionKey: Eq + Hash + Clone + Send + 'static {}
impl<K: Eq + Hash + Clone + Send + 'static> SessionKey for K {}
impl<P: Encode + Decode + Send + Sync + 'static, K: SessionKey> SessionFactory<P, K> {
/// Describe the session route. Spawns nothing; the registry actor is
/// started from [`children`](ChannelFactory::children).
pub(super) fn new<S: ChannelSession<P, Key = K>>(
registry: &'static str,
inner: Arc<dyn ChannelFactory<P>>,
) -> Self {
SessionFactory {
inner,
keyfn: S::session_key,
registry: GenServerName::new(registry),
cap: S::buffer_cap(),
ttl: S::ttl(),
}
}
}
impl<P: Encode + Decode + Send + Sync + 'static, K: SessionKey> ChannelFactory<P>
for SessionFactory<P, K>
{
fn create(&self, topic: &str) -> Box<dyn Channel<P>> {
self.inner.create(topic)
}
fn children(&self) -> Vec<ChildSpec> {
let (name, factory, cap, ttl) = (self.registry, self.inner.clone(), self.cap, self.ttl);
vec![ChildSpec::new(Restart::Permanent, move || {
let state = Registry::<P, K> {
factory: factory.clone(),
cap,
ttl,
sessions: HashMap::new(),
pid_key: HashMap::new(),
watcher: None,
};
if let Err(e) = GenServerBuilder::new(state).named(name).run() {
panic!("urus session registry '{}': name already taken: {e:?}", name.as_str());
}
})]
}
fn deploy(&self, hs: JoinHandshake<P>) -> ChannelInbox<P>
where
P: Encode + Decode,
{
let key = (self.keyfn)(&hs.topic, &hs.payload);
let (tx, rx) = smarm::channel::<In<P>>();
// Cast: the actor acks the join straight to the socket; the
// conn actor has nothing to wait for. A dead registry can only
// mean shutdown — the failed join is moot.
let _ = gen_server::cast(self.registry, Join { key, hs, inbound: rx });
ChannelInbox(tx)
}
}
// ---------------------------------------------------------------------------
// The registry gen_server: key -> live session actor
// ---------------------------------------------------------------------------
pub(super) struct Join<P: Send + Sync + 'static, K> {
key: K,
hs: JoinHandshake<P>,
inbound: Receiver<In<P>>,
}
struct Registry<P: Encode + Decode + Send + Sync + 'static, K: SessionKey> {
factory: Arc<dyn ChannelFactory<P>>,
cap: usize,
ttl: Duration,
sessions: HashMap<K, (Pid, Sender<Ctl<P>>)>,
/// Reverse index for `handle_down` — the reason `Key: Clone`.
pid_key: HashMap<Pid, K>,
watcher: Option<Watcher<Registry<P, K>>>,
}
impl<P: Encode + Decode + Send + Sync + 'static, K: SessionKey> GenServer for Registry<P, K> {
type Call = ();
type Reply = ();
type Cast = Join<P, K>;
type Info = ();
type Timer = ();
fn init(&mut self, ctx: &GenServerCtx<Self>) {
self.watcher = Some(ctx.watcher());
}
fn handle_call(&mut self, _request: ()) {}
fn handle_cast(&mut self, Join { key, hs, inbound }: Join<P, K>) {
let attach = Ctl::Attach { hs, inbound };
let attach = match self.sessions.get(&key) {
Some((_, ctl)) => match ctl.send(attach) {
Ok(()) => return, // reattached
// The actor died and its Down hasn't landed yet:
// recover the handshake, replace below.
Err(smarm::channel::SendError(a)) => a,
},
None => attach,
};
let (ctl_tx, ctl_rx) = smarm::channel::<Ctl<P>>();
let (factory, cap, ttl) = (self.factory.clone(), self.cap, self.ttl);
let pid = smarm::spawn(move || run_session(factory, cap, ttl, ctl_rx)).pid();
// Cannot fail: the receiver is owned by the just-spawned actor.
let _ = ctl_tx.send(attach);
self.watcher
.as_ref()
.expect("watcher set in init")
.watch(smarm::monitor(pid));
self.pid_key.insert(pid, key.clone());
self.sessions.insert(key, (pid, ctl_tx));
}
fn handle_down(&mut self, down: Down) {
if let Some(key) = self.pid_key.remove(&down.pid) {
// Pid guard: a stale Down for an entry handle_cast already
// replaced must not evict the replacement.
if self.sessions.get(&key).is_some_and(|(p, _)| *p == down.pid) {
self.sessions.remove(&key);
}
}
}
}
// ---------------------------------------------------------------------------
// The session actor
// ---------------------------------------------------------------------------
enum Ctl<P: Send + Sync + 'static> {
Attach { hs: JoinHandshake<P>, inbound: Receiver<In<P>> },
}
enum Step<P: Send + Sync + 'static> {
Attach { hs: JoinHandshake<P>, inbound: Receiver<In<P>> },
Attached { inbound: Receiver<In<P>> },
Detached,
Exit,
}
fn run_session<P: Encode + Decode + Send + Sync + 'static>(
factory: Arc<dyn ChannelFactory<P>>,
cap: usize,
ttl: Duration,
ctl: Receiver<Ctl<P>>,
) {
// The initial Attach is cast by the registry in the same handler
// that spawned us. Closed instead = registry died first (shutdown
// raced the spawn): nothing exists yet, nothing to clean.
let Ok(Ctl::Attach { hs, inbound }) = ctl.recv() else { return };
// Partial moves: ws/bus/topic/join_ref go into the socket,
// reference/payload stay behind for the first attach below.
let mut socket = ChannelSocket {
ws: hs.ws,
bus: hs.bus,
topic: hs.topic,
join_ref: hs.join_ref,
cur_ref: RefCell::new(None),
};
let mut ch: Option<Box<dyn Channel<P>>> = None;
let mut joined = false;
let mut bus_rx: Option<Receiver<Arc<Broadcast<P>>>> = None;
let mut buffer: VecDeque<Arc<Broadcast<P>>> = VecDeque::new();
let mut step = attach(
&factory,
&mut ch,
&mut joined,
&mut bus_rx,
&mut buffer,
&mut socket,
hs.reference,
hs.payload,
inbound,
);
loop {
step = match step {
Step::Attach { hs, inbound } => {
socket.ws = hs.ws;
socket.join_ref = hs.join_ref;
socket.topic = hs.topic;
attach(
&factory,
&mut ch,
&mut joined,
&mut bus_rx,
&mut buffer,
&mut socket,
hs.reference,
hs.payload,
inbound,
)
}
Step::Attached { inbound } => attached_loop(
ch.as_deref_mut().expect("attached implies created"),
&socket,
&ctl,
&inbound,
bus_rx.as_ref().expect("attached implies subscribed"),
&mut buffer,
),
Step::Detached => detached_loop(
&ctl,
bus_rx.as_ref().expect("detached only after first subscribe"),
&mut buffer,
cap,
ttl,
),
Step::Exit => {
if joined {
if let Some(c) = ch.as_deref_mut() {
c.terminate();
}
}
return;
}
};
}
}
/// One attach handshake: create-once, (re)join, subscribe-once, ack,
/// drain. On a drain failure the undelivered tail stays in `buffer`
/// for the next attach.
#[allow(clippy::too_many_arguments)]
fn attach<P: Encode + Decode + Send + Sync + 'static>(
factory: &Arc<dyn ChannelFactory<P>>,
ch: &mut Option<Box<dyn Channel<P>>>,
joined: &mut bool,
bus_rx: &mut Option<Receiver<Arc<Broadcast<P>>>>,
buffer: &mut VecDeque<Arc<Broadcast<P>>>,
socket: &mut ChannelSocket<P>,
reference: Option<String>,
payload: P,
inbound: Receiver<In<P>>,
) -> Step<P> {
if ch.is_none() {
*ch = Some(factory.create(&socket.topic));
}
let c = ch.as_deref_mut().expect("just created");
socket.set_ref(reference);
let topic = socket.topic.clone();
match c.join(&topic, payload, socket) {
Ok(reply) => {
if bus_rx.is_none() {
// First successful join: subscribe BEFORE the ok reply
// (subscribe is a call — once the client sees its ack,
// membership is a fact, not a race). Once per session:
// the subscription is actor-scoped, survives detach,
// and that survival IS the buffering path.
match socket.bus.subscribe(&topic) {
Ok(rx) => *bus_rx = Some(rx),
Err(_) => {
socket.send_reply(Status::Error, None);
socket.set_ref(None);
return Step::Exit;
}
}
}
*joined = true;
socket.send_reply(Status::Ok, Some(&reply));
socket.set_ref(None);
// Drain with the NEW join generation's ref; buffered pushes
// have no request to correlate to (reference: None).
while let Some(b) = buffer.front() {
let out = P::encode(FrameRef {
join_ref: socket.join_ref.as_deref(),
reference: None,
topic: &socket.topic,
message: MessageRef::Event { event: &b.event, payload: Some(&b.payload) },
});
if socket.ws.send(out).is_ok() {
buffer.pop_front();
} else {
return Step::Detached;
}
}
Step::Attached { inbound }
}
Err(reply) => {
// Rejected (re)join = session over, by decision: error
// reply, terminate (in run_session, iff it ever joined),
// exit. Next join is a cold start.
socket.send_reply(Status::Error, Some(&reply));
socket.set_ref(None);
Step::Exit
}
}
}
fn attached_loop<P: Encode + Decode + Send + Sync + 'static>(
ch: &mut dyn Channel<P>,
socket: &ChannelSocket<P>,
ctl: &Receiver<Ctl<P>>,
inbound: &Receiver<In<P>>,
bus_rx: &Receiver<Arc<Broadcast<P>>>,
buffer: &mut VecDeque<Arc<Broadcast<P>>>,
) -> Step<P> {
// ctl sits at lowest priority: it only ever carries an eviction by
// a second transport (rare) or the closed-arm shutdown signal. A
// closed arm stays ready forever, but every closed observation
// exits the loop immediately — no spin, no starvation window.
loop {
match smarm::select(&[inbound, bus_rx, ctl]) {
0 => match inbound.try_recv() {
Ok(Some(In::Event { reference, event, payload })) => {
socket.set_ref(reference);
ch.handle_in(&event, payload, socket);
socket.set_ref(None);
}
Ok(Some(In::Leave { reference })) => {
// Explicit leave ends the SESSION, not just the
// attachment: ok reply, close event, terminate.
socket.set_ref(reference);
socket.send_reply(Status::Ok, None);
socket.set_ref(None);
close_event(socket);
return Step::Exit;
}
Ok(None) => {}
// Conn actor finished or replaced this join generation:
// the transport is gone, the session is not.
Err(_) => return Step::Detached,
},
1 => match bus_rx.try_recv() {
Ok(Some(b)) => {
let out = P::encode(FrameRef {
join_ref: socket.join_ref.as_deref(),
reference: None,
topic: &socket.topic,
message: MessageRef::Event { event: &b.event, payload: Some(&b.payload) },
});
if socket.ws.send(out).is_err() {
// Transport died under us: this broadcast is
// the buffer's first entry, nothing was lost.
buffer.push_back(b);
return Step::Detached;
}
}
Ok(None) => {}
// The topic relay died: pubsub is gone, session over.
Err(_) => return Step::Exit,
},
2 => match ctl.try_recv() {
Ok(Some(Ctl::Attach { hs, inbound })) => {
// A second transport claims the session key: evict
// this one (best effort — it may already be dead)
// and hand the session over.
close_event(socket);
return Step::Attach { hs, inbound };
}
Ok(None) => {}
// Registry gone = shutdown chain: terminate and exit.
Err(_) => return Step::Exit,
},
_ => unreachable!("three arms"),
}
}
}
fn detached_loop<P: Encode + Decode + Send + Sync + 'static>(
ctl: &Receiver<Ctl<P>>,
bus_rx: &Receiver<Arc<Broadcast<P>>>,
buffer: &mut VecDeque<Arc<Broadcast<P>>>,
cap: usize,
ttl: Duration,
) -> Step<P> {
// TTL is per detach episode: the clock starts now, every time.
let deadline = Instant::now() + ttl;
loop {
// cap.max(1): a cap of 0 means "no buffering tolerated" — the
// first buffered broadcast tears the session down; an empty
// buffer still waits out the TTL.
if buffer.len() >= cap.max(1) {
return Step::Exit;
}
let remaining = deadline.saturating_duration_since(Instant::now());
match smarm::select_timeout(&[ctl, bus_rx], remaining) {
// TTL expired with no rejoin: session over.
None => return Step::Exit,
Some(0) => match ctl.try_recv() {
Ok(Some(Ctl::Attach { hs, inbound })) => return Step::Attach { hs, inbound },
Ok(None) => {}
Err(_) => return Step::Exit,
},
Some(1) => match bus_rx.try_recv() {
Ok(Some(b)) => buffer.push_back(b),
Ok(None) => {}
Err(_) => return Step::Exit,
},
Some(_) => unreachable!("two arms"),
}
}
}
/// Best-effort `phx_close` push for the socket's current join
/// generation.
fn close_event<P: Encode + Decode + Send + Sync + 'static>(socket: &ChannelSocket<P>) {
let _ = socket.ws.send(P::encode(FrameRef {
join_ref: socket.join_ref.as_deref(),
reference: None,
topic: &socket.topic,
message: MessageRef::Event { event: EV_CLOSE, payload: None },
}));
}
// ---------------------------------------------------------------------------
// Tests — runtime-backed, on the toy pipe codec from testkit. The
// session-config types (`Cfg*`) are deliberately separate from the
// channel they configure: `channel_session::<S>` only needs `S` for
// keying/cap/ttl, the factory builds the channel.
// ---------------------------------------------------------------------------
#[cfg(test)]
mod tests {
use super::super::testkit::{eventually, Harness, TP};
use super::super::*;
use std::sync::atomic::{AtomicBool, Ordering};
use std::sync::Mutex;
use std::time::Duration;
/// Counts its joins (cold start replies `joined:1`); flags
/// terminate; broadcasts on "shout".
struct Counter {
joins: u32,
term: Arc<AtomicBool>,
reject_rejoin: bool,
}
impl Channel<TP> for Counter {
fn join(&mut self, _topic: &str, _payload: TP, _s: &ChannelSocket<TP>) -> Result<TP, TP> {
self.joins += 1;
if self.reject_rejoin && self.joins > 1 {
return Err(TP("not again".into()));
}
Ok(TP(format!("joined:{}", self.joins)))
}
fn handle_in(&mut self, event: &str, payload: TP, s: &ChannelSocket<TP>) {
match event {
"shout" => s.broadcast("news", payload),
_ => s.reply(Status::Ok, TP(format!("echo:{}", payload.0))),
}
}
fn terminate(&mut self) {
self.term.store(true, Ordering::SeqCst);
}
}
fn counter_factory(term: &Arc<AtomicBool>, reject_rejoin: bool) -> impl ChannelFactory<TP> {
let term = term.clone();
move |_t: &str| -> Box<dyn Channel<TP>> {
Box::new(Counter { joins: 0, term: term.clone(), reject_rejoin })
}
}
/// Session config: keyed by topic alone, defaults otherwise.
struct Cfg;
impl ChannelSession<TP> for Cfg {
type Key = String;
fn session_key(topic: &str, _p: &TP) -> String {
topic.to_owned()
}
}
/// Tiny TTL (50ms) for the expiry test.
struct CfgTtl;
impl ChannelSession<TP> for CfgTtl {
type Key = String;
fn session_key(topic: &str, _p: &TP) -> String {
topic.to_owned()
}
fn ttl() -> Duration {
Duration::from_millis(50)
}
}
/// Cap of 2 (and a long TTL so only the cap can fire).
struct CfgCap;
impl ChannelSession<TP> for CfgCap {
type Key = String;
fn session_key(topic: &str, _p: &TP) -> String {
topic.to_owned()
}
fn buffer_cap() -> usize {
2
}
fn ttl() -> Duration {
Duration::from_secs(30)
}
}
fn session_hub<S: ChannelSession<TP, Key = String>>(
term: &Arc<AtomicBool>,
reject_rejoin: bool,
) -> (ChannelHub<TP>, Vec<smarm::ChildSpec>) {
ChannelHub::new(
"sess-unit-bus",
PrefixRouter::new().channel_session::<S>(
"room:*",
"sess-unit-registry",
counter_factory(term, reject_rejoin),
),
)
}
#[test]
fn session_buffers_across_reconnect_and_drains_in_order() {
let out = Arc::new(Mutex::new(Vec::<String>::new()));
let (out2, term) = (out.clone(), Arc::new(AtomicBool::new(false)));
smarm::run(move || {
let (hub, children) = session_hub::<Cfg>(&term, false);
let _sup = crate::channels::tests::start_hub(&hub, children);
let mut h1 = Harness::new(&hub);
h1.send("j1|r1|room:a|event|phx_join|u");
out2.lock().unwrap().push(h1.recv());
// Transport dies. Whether the actor has observed the detach
// yet or not, both paths (closed inbound arm, failed ws
// send) land the broadcasts in the buffer.
drop(h1);
hub.broadcast("room:a", "news", TP("one".into()));
hub.broadcast("room:a", "news", TP("two".into()));
let mut h2 = Harness::new(&hub);
h2.send("j2|r2|room:a|event|phx_join|u");
out2.lock().unwrap().push(h2.recv()); // warm rejoin ack
out2.lock().unwrap().push(h2.recv()); // drained, in order
out2.lock().unwrap().push(h2.recv());
});
let out = out.lock().unwrap();
assert_eq!(out[0], "j1|r1|room:a|reply|ok|joined:1");
assert_eq!(out[1], "j2|r2|room:a|reply|ok|joined:2");
assert_eq!(out[2], "j2|-|room:a|event|news|one");
assert_eq!(out[3], "j2|-|room:a|event|news|two");
}
#[test]
fn session_ttl_expiry_tears_down_then_cold_start() {
let out = Arc::new(Mutex::new(Vec::<String>::new()));
let (out2, term) = (out.clone(), Arc::new(AtomicBool::new(false)));
let term2 = term.clone();
smarm::run(move || {
let (hub, children) = session_hub::<CfgTtl>(&term2, false);
let _sup = crate::channels::tests::start_hub(&hub, children);
let mut h1 = Harness::new(&hub);
h1.send("j1|r1|room:a|event|phx_join|u");
out2.lock().unwrap().push(h1.recv());
drop(h1);
assert!(eventually(|| term2.load(Ordering::SeqCst)), "ttl never fired");
let mut h2 = Harness::new(&hub);
h2.send("j2|r2|room:a|event|phx_join|u");
out2.lock().unwrap().push(h2.recv());
});
let out = out.lock().unwrap();
assert_eq!(out[0], "j1|r1|room:a|reply|ok|joined:1");
assert_eq!(out[1], "j2|r2|room:a|reply|ok|joined:1"); // cold
}
#[test]
fn session_buffer_cap_tears_down_then_cold_start() {
let out = Arc::new(Mutex::new(Vec::<String>::new()));
let (out2, term) = (out.clone(), Arc::new(AtomicBool::new(false)));
smarm::run(move || {
let (hub, children) = session_hub::<CfgCap>(&term, false);
let _sup = crate::channels::tests::start_hub(&hub, children);
let mut h1 = Harness::new(&hub);
h1.send("j1|r1|room:a|event|phx_join|u");
out2.lock().unwrap().push(h1.recv());
drop(h1);
// Cap is 2: the second buffered broadcast tears it down
// (ttl is 30s — only the cap can fire here).
hub.broadcast("room:a", "news", TP("one".into()));
hub.broadcast("room:a", "news", TP("two".into()));
assert!(eventually(|| term.load(Ordering::SeqCst)), "cap never fired");
let mut h2 = Harness::new(&hub);
h2.send("j2|r2|room:a|event|phx_join|u");
out2.lock().unwrap().push(h2.recv());
});
let out = out.lock().unwrap();
assert_eq!(out[0], "j1|r1|room:a|reply|ok|joined:1");
assert_eq!(out[1], "j2|r2|room:a|reply|ok|joined:1"); // cold
}
#[test]
fn leave_ends_the_session_not_just_the_attachment() {
let out = Arc::new(Mutex::new(Vec::<String>::new()));
let (out2, term) = (out.clone(), Arc::new(AtomicBool::new(false)));
smarm::run(move || {
let (hub, children) = session_hub::<Cfg>(&term, false);
let _sup = crate::channels::tests::start_hub(&hub, children);
let mut h = Harness::new(&hub);
h.send("j1|r1|room:a|event|phx_join|u");
out2.lock().unwrap().push(h.recv());
h.send("j1|r2|room:a|event|phx_leave|");
out2.lock().unwrap().push(h.recv()); // ok reply
out2.lock().unwrap().push(h.recv()); // phx_close
assert!(eventually(|| term.load(Ordering::SeqCst)), "leave never terminated");
h.send("j2|r3|room:a|event|phx_join|u");
out2.lock().unwrap().push(h.recv());
});
let out = out.lock().unwrap();
assert_eq!(out[0], "j1|r1|room:a|reply|ok|joined:1");
assert_eq!(out[1], "j1|r2|room:a|reply|ok|");
assert_eq!(out[2], "j1|-|room:a|event|phx_close|");
assert_eq!(out[3], "j2|r3|room:a|reply|ok|joined:1"); // cold
}
#[test]
fn second_transport_evicts_first() {
let out = Arc::new(Mutex::new(Vec::<String>::new()));
let (out2, term) = (out.clone(), Arc::new(AtomicBool::new(false)));
smarm::run(move || {
let (hub, children) = session_hub::<Cfg>(&term, false);
let _sup = crate::channels::tests::start_hub(&hub, children);
let mut h1 = Harness::new(&hub);
h1.send("j1|r1|room:a|event|phx_join|u");
out2.lock().unwrap().push(h1.recv());
let mut h2 = Harness::new(&hub);
h2.send("j2|r2|room:a|event|phx_join|u");
out2.lock().unwrap().push(h2.recv()); // warm handover ack
out2.lock().unwrap().push(h1.recv()); // eviction phx_close
hub.broadcast("room:a", "news", TP("x".into()));
out2.lock().unwrap().push(h2.recv()); // only h2 is attached
h1.assert_silence();
// h1's conn-side entry is stale: its inbound receiver was
// dropped on eviction, so the next event self-prunes with
// an error reply.
h1.send("j1|r3|room:a|event|ping|x");
out2.lock().unwrap().push(h1.recv());
});
let out = out.lock().unwrap();
assert_eq!(out[0], "j1|r1|room:a|reply|ok|joined:1");
assert_eq!(out[1], "j2|r2|room:a|reply|ok|joined:2");
assert_eq!(out[2], "j1|-|room:a|event|phx_close|");
assert_eq!(out[3], "j2|-|room:a|event|news|x");
assert_eq!(out[4], "j1|r3|room:a|reply|error|");
}
#[test]
fn rejected_rejoin_ends_the_session() {
let out = Arc::new(Mutex::new(Vec::<String>::new()));
let (out2, term) = (out.clone(), Arc::new(AtomicBool::new(false)));
smarm::run(move || {
let (hub, children) = session_hub::<Cfg>(&term, true);
let _sup = crate::channels::tests::start_hub(&hub, children);
let mut h1 = Harness::new(&hub);
h1.send("j1|r1|room:a|event|phx_join|u");
out2.lock().unwrap().push(h1.recv());
drop(h1);
let mut h2 = Harness::new(&hub);
h2.send("j2|r2|room:a|event|phx_join|u");
out2.lock().unwrap().push(h2.recv()); // rejoin rejected
assert!(eventually(|| term.load(Ordering::SeqCst)), "reject never terminated");
let mut h3 = Harness::new(&hub);
h3.send("j3|r3|room:a|event|phx_join|u");
out2.lock().unwrap().push(h3.recv()); // cold start
});
let out = out.lock().unwrap();
assert_eq!(out[0], "j1|r1|room:a|reply|ok|joined:1");
assert_eq!(out[1], "j2|r2|room:a|reply|error|not again");
assert_eq!(out[2], "j3|r3|room:a|reply|ok|joined:1");
}
}
+121
View File
@@ -0,0 +1,121 @@
//! Shared test scaffolding for the channels module tree: the toy
//! pipe-delimited codec (`TP`) and the conn-side `Harness` that fakes a
//! websocket transport. `cfg(test)` only.
use super::*;
use crate::ws::duplex::Outbound;
use std::time::Duration;
/// Toy payload + codec: `jref|ref|topic|kind|event-or-status|payload`
/// with `-` for absent refs. Enough to drive the machinery.
#[derive(Debug, Clone, PartialEq, Eq)]
pub(crate) struct TP(pub(crate) String);
pub(crate) fn part(o: Option<&str>) -> &str {
o.unwrap_or("-")
}
impl Encode for TP {
fn encode(f: FrameRef<'_, Self>) -> Message {
let s = match f.message {
MessageRef::Event { event, payload } => format!(
"{}|{}|{}|event|{}|{}",
part(f.join_ref),
part(f.reference),
f.topic,
event,
payload.map(|p| p.0.as_str()).unwrap_or("")
),
MessageRef::Reply { status, payload } => format!(
"{}|{}|{}|reply|{}|{}",
part(f.join_ref),
part(f.reference),
f.topic,
match status {
Status::Ok => "ok",
Status::Error => "error",
},
payload.map(|p| p.0.as_str()).unwrap_or("")
),
};
Message::Text(s)
}
}
impl Decode for TP {
fn decode(msg: &Message) -> Result<ChannelFrame<Self>, DecodeError> {
let Message::Text(s) = msg else {
return Err(DecodeError("binary".into()));
};
let p: Vec<&str> = s.splitn(6, '|').collect();
if p.len() != 6 || p[3] != "event" {
return Err(DecodeError(format!("bad frame: {s}")));
}
let opt = |x: &str| (x != "-").then(|| x.to_string());
Ok(ChannelFrame {
join_ref: opt(p[0]),
reference: opt(p[1]),
topic: p[2].to_string(),
message: ChannelMessage::Event { event: p[4].to_string(), payload: TP(p[5].into()) },
})
}
}
/// A fake socket: the conn-side handler plus the raw Outbound channel
/// a real duplex loop would drain. `recv_text` pulls the next encoded
/// frame the "client" would see.
pub(crate) struct Harness {
handler: SocketHandler<TP>,
sender: WsSender,
out_rx: smarm::Receiver<Outbound>,
}
impl Harness {
pub(crate) fn new(hub: &ChannelHub<TP>) -> Self {
let (tx, out_rx) = smarm::channel();
Self {
handler: SocketHandler {
bus: hub.bus,
router: hub.router.clone(),
joined: HashMap::new(),
},
sender: WsSender { tx },
out_rx,
}
}
/// Feed one wire line in, as the duplex loop would.
pub(crate) fn send(&mut self, line: &str) {
self.handler
.on_message(Message::Text(line.to_string()), &self.sender.clone());
}
/// Assert nothing arrives for this client within a short window —
/// for "the broadcast must NOT reach the evicted transport" cases.
pub(crate) fn assert_silence(&self) {
if let Ok(Outbound::Msg(Message::Text(s))) =
self.out_rx.recv_timeout(Duration::from_millis(100))
{
panic!("expected silence, got frame: {s}");
}
}
/// Next encoded frame the client would see (2s budget — channel
/// actors answer asynchronously).
pub(crate) fn recv(&self) -> String {
match self.out_rx.recv_timeout(Duration::from_secs(2)) {
Ok(Outbound::Msg(Message::Text(s))) => s,
other => panic!("expected outbound text frame, got {:?}", other.is_ok()),
}
}
}
pub(crate) fn eventually(mut f: impl FnMut() -> bool) -> bool {
for _ in 0..200 {
if f() {
return true;
}
smarm::sleep(Duration::from_millis(10));
}
false
}
+121 -10
View File
@@ -146,27 +146,77 @@ impl Body {
// RespBody — what plugs put on the wire. // RespBody — what plugs put on the wire.
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// //
// Enum, not a trait, so the connection actor can pattern-match. Leaves room // Enum, not a trait, so the connection actor can pattern-match.
// for future variants like Sse, Stream, Chunked without touching the plug //
// API. // `Stream` is the pull-shaped streaming body (v0.3 design decision): the
// handler hands back a `smarm::Receiver<Vec<u8>>` — typically the read end
// of a channel whose `Sender` lives in a producer actor the handler
// spawned — and the connection actor pumps it onto the wire. The conn
// actor keeps sole ownership of the socket and of write deadlines; the
// producer never touches the fd. End-of-stream is signalled by dropping
// every `Sender` (the conn actor then emits the terminating 0-chunk).
// Empty `Vec`s are skipped by the pump (a zero-length chunk would
// terminate chunked framing early), so they are safe to send but useless.
#[derive(Debug, Clone)] #[derive(Default)]
pub enum RespBody { pub enum RespBody {
#[default]
Empty, Empty,
Bytes(Vec<u8>), Bytes(Vec<u8>),
Stream(StreamBody),
} }
impl RespBody { /// A streaming response body. Construct with [`StreamBody::new`] (or
pub fn len_hint(&self) -> usize { /// `RespBody::from(rx)`) and hand it to [`Conn::put_body`].
///
/// On HTTP/1.1 the connection actor sends it with
/// `Transfer-Encoding: chunked` and the connection stays reusable after
/// the stream completes. On HTTP/1.0 (no chunked framing) the bytes are
/// written raw and the connection closes at end of stream to delimit the
/// body.
pub struct StreamBody {
pub rx: smarm::Receiver<Vec<u8>>,
/// If set, the connection actor's pump waits for each chunk with
/// `recv_timeout(interval)` instead of a bare `recv`, and on expiry
/// writes `payload` as its own chunk and keeps waiting. This is how
/// SSE keep-alive comments work: liveness comes from the ping (a dead
/// client is detected when a heartbeat write stalls past
/// `write_timeout`), not from any request clock.
pub heartbeat: Option<(std::time::Duration, Vec<u8>)>,
}
impl StreamBody {
pub fn new(rx: smarm::Receiver<Vec<u8>>) -> Self {
Self { rx, heartbeat: None }
}
pub fn with_heartbeat(mut self, interval: std::time::Duration, payload: Vec<u8>) -> Self {
self.heartbeat = Some((interval, payload));
self
}
}
impl std::fmt::Debug for RespBody {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
match self { match self {
RespBody::Empty => 0, RespBody::Empty => f.write_str("Empty"),
RespBody::Bytes(b) => b.len(), RespBody::Bytes(b) => f.debug_tuple("Bytes").field(&b.len()).finish(),
RespBody::Stream(_) => f.write_str("Stream(..)"),
} }
} }
} }
impl Default for RespBody { impl RespBody {
fn default() -> Self { RespBody::Empty } /// Body length for fixed bodies; 0 for `Stream` (a stream has no
/// length up front — that is the point — and is never serialised with
/// a `content-length`).
pub fn len_hint(&self) -> usize {
match self {
RespBody::Empty => 0,
RespBody::Bytes(b) => b.len(),
RespBody::Stream(_) => 0,
}
}
} }
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -256,6 +306,11 @@ pub struct Conn {
pub params: Params, pub params: Params,
pub assigns: Assigns, pub assigns: Assigns,
pub halted: bool, pub halted: bool,
/// Set by [`Conn::upgrade`] on an accepted WebSocket handshake; the
/// connection actor sees it (with status 101) and leaves the HTTP
/// loop after writing the response. `None` for plain HTTP.
pub upgrade: Option<crate::ws::WsUpgrade>,
} }
impl Conn { impl Conn {
@@ -277,6 +332,8 @@ impl Conn {
params: Params::new(), params: Params::new(),
assigns: Assigns::new(), assigns: Assigns::new(),
halted: false, halted: false,
upgrade: None,
} }
} }
@@ -306,6 +363,52 @@ impl Conn {
self.halted = true; self.halted = true;
self self
} }
/// Accept a WebSocket upgrade on this request (RFC 6455 §4.2),
/// handing `handler` to the duplex loop that takes the socket over.
///
/// On a valid handshake: status `101`, the `upgrade`/`connection`/
/// `sec-websocket-accept` response headers, the [`WsUpgrade`] payload
/// the connection actor acts on, and the pipeline is halted. After
/// the 101 the connection actor leaves HTTP and drives `handler` —
/// see [`WsHandler`](crate::ws::WsHandler) for the callback contract.
///
/// On an invalid handshake: `handler` is dropped untouched and the
/// appropriate rejection goes out (`400`, or `426` +
/// `sec-websocket-version: 13` for a version mismatch), empty body,
/// pipeline halted; the connection stays plain HTTP with normal
/// keep-alive semantics. Pre-check with
/// [`ws::handshake::is_upgrade_request`](crate::ws::handshake::is_upgrade_request)
/// to branch before committing.
///
/// [`WsUpgrade`]: crate::ws::WsUpgrade
pub fn upgrade(mut self, handler: impl crate::ws::WsHandler) -> Self {
use crate::ws::{handshake, Rejection};
match handshake::validate(&self) {
Ok(accept) => {
self.upgrade = Some(crate::ws::WsUpgrade {
handler: Box::new(handler),
});
self.put_status(101)
.put_header("upgrade", "websocket")
.put_header("connection", "Upgrade")
.put_header("sec-websocket-accept", accept)
.put_body(RespBody::Empty)
.halt()
}
Err(rej) => {
let c = self
.put_status(rej.status())
.put_body(RespBody::Empty)
.halt();
if rej == Rejection::UnsupportedVersion {
c.put_header("sec-websocket-version", "13")
} else {
c
}
}
}
}
} }
impl Default for Conn { impl Default for Conn {
@@ -326,3 +429,11 @@ impl From<&'static str> for RespBody {
impl From<&[u8]> for RespBody { impl From<&[u8]> for RespBody {
fn from(s: &[u8]) -> Self { RespBody::Bytes(s.to_vec()) } fn from(s: &[u8]) -> Self { RespBody::Bytes(s.to_vec()) }
} }
impl From<StreamBody> for RespBody {
fn from(s: StreamBody) -> Self { RespBody::Stream(s) }
}
impl From<smarm::Receiver<Vec<u8>>> for RespBody {
fn from(rx: smarm::Receiver<Vec<u8>>) -> Self {
RespBody::Stream(StreamBody::new(rx))
}
}
+600 -38
View File
@@ -13,14 +13,17 @@
//! `wait_readable` between bytes and `wait_writable` during slow writes; //! `wait_readable` between bytes and `wait_writable` during slow writes;
//! during those parks, other connection actors progress freely. //! during those parks, other connection actors progress freely.
use crate::conn::{Body, Conn, RespBody}; use crate::conn::{Body, Conn, HttpVersion, RespBody, StreamBody};
use crate::endpoint::{Cast, DeregisterGuard, Endpoint};
use crate::net::OwnedFd; use crate::net::OwnedFd;
use crate::parser::{self, ParseError}; use crate::parser::{self, ParseError};
use crate::plug::Pipeline; use crate::plug::Pipeline;
use smarm::GenServerRef;
use std::io::{self, ErrorKind}; use std::io::{self, ErrorKind};
use std::os::fd::RawFd; use std::os::fd::RawFd;
use std::time::Duration; use std::time::{Duration, Instant};
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// Limits // Limits
@@ -38,7 +41,51 @@ pub struct ConnLimits {
/// Hard cap on Content-Length we'll accept. 16 MiB is enough for a CRUD /// Hard cap on Content-Length we'll accept. 16 MiB is enough for a CRUD
/// example; configurable in `Config`. /// example; configurable in `Config`.
pub max_body_bytes: usize, pub max_body_bytes: usize,
/// Idle budget between requests: how long we'll park waiting for the
/// FIRST byte of a request (including the first request on a fresh
/// connection). Expiry closes the connection silently — nothing is
/// owed to a client that isn't talking.
pub keep_alive_timeout: Duration, pub keep_alive_timeout: Duration,
/// Wall-clock budget for reading the request HEAD, measured from the
/// first byte of a request until the head is fully parsed. Expiry
/// mid-head gets a best-effort 408. Kept short: an incomplete head is
/// the classic slowloris, and a legitimate client sends its head in a
/// single burst. The BODY has its own, larger budget (`body_timeout`)
/// so a slow-but-legit upload is not judged by the head clock.
pub head_timeout: Duration,
/// Absolute wall-clock cap on reading the request BODY, measured from
/// the moment the head finished parsing until the body is fully read.
/// Sized for slow links (e.g. a trickling cellular IoT client), so it
/// is much larger than `head_timeout`. Expiry mid-body just closes —
/// nothing is owed to a client this far gone. Pipeline run time is NOT
/// covered (that's the handler's business); the write phase has its own
/// per-write budget (`write_timeout`).
pub body_timeout: Duration,
/// Burst-gated body stall eviction: the bytes that must accumulate
/// since the last advance to count as a "burst" and reset the stall
/// clock. A body that dribbles fewer than this per `body_stall_timeout`
/// window is evicted — the discriminator between a slowloris trickle
/// (near-zero, smooth) and a slow-but-legit client (delivers real
/// bursts). The pair implies an effective floor of
/// body_burst_bytes / body_stall_timeout, enforced in bursts.
pub body_burst_bytes: usize,
/// Max time since the last qualifying burst (`body_burst_bytes`)
/// before a stalled body read is evicted. Must comfortably exceed a
/// legit client's worst quiet gap (e.g. cellular RRC/handover/DRX
/// stalls). The absolute `body_timeout` always backstops it.
pub body_stall_timeout: Duration,
/// Per-write budget for response bytes: every `write_all` (the fixed
/// head+body, and each streamed chunk) must complete within this.
/// A client that stops reading mid-response is dropped when its
/// socket buffer fills and a write stalls past the budget.
pub write_timeout: Duration,
/// WebSocket: hard cap on a single frame's payload, enforced from
/// the frame header BEFORE the payload is buffered. Violation closes
/// with 1009.
pub max_frame_payload: usize,
/// WebSocket: hard cap on a complete (reassembled) message; spans
/// fragments. Violation closes with 1009.
pub max_message_bytes: usize,
} }
impl Default for ConnLimits { impl Default for ConnLimits {
@@ -49,6 +96,13 @@ impl Default for ConnLimits {
max_head_bytes: 64 * 1024, max_head_bytes: 64 * 1024,
max_body_bytes: 16 * 1024 * 1024, max_body_bytes: 16 * 1024 * 1024,
keep_alive_timeout: Duration::from_secs(60), keep_alive_timeout: Duration::from_secs(60),
head_timeout: Duration::from_secs(30),
body_timeout: Duration::from_secs(300),
body_burst_bytes: 4 * 1024,
body_stall_timeout: Duration::from_secs(20),
write_timeout: Duration::from_secs(30),
max_frame_payload: 1024 * 1024,
max_message_bytes: 4 * 1024 * 1024,
} }
} }
} }
@@ -57,49 +111,121 @@ impl Default for ConnLimits {
// run_connection — entry point spawned by the listener actor. // run_connection — entry point spawned by the listener actor.
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
pub fn run_connection(fd: OwnedFd, pipeline: Pipeline, limits: ConnLimits) { pub fn run_connection(
fd: OwnedFd,
pipeline: Pipeline,
limits: ConnLimits,
registry: GenServerRef<Endpoint>,
) {
// The OwnedFd cleans up via Drop on any exit path (panic, error, or // The OwnedFd cleans up via Drop on any exit path (panic, error, or
// normal close). No explicit close calls below. // normal close). No explicit close calls below.
let raw = fd.as_raw(); let raw = fd.as_raw();
let mut buf: Vec<u8> = Vec::with_capacity(limits.initial_read_buf); let mut buf: Vec<u8> = Vec::with_capacity(limits.initial_read_buf);
// Self-register (initially idle: no request head parsed yet) and arm
// the deregistration guard. Both casts come from this actor, so
// Started always precedes Ended in the registry's inbox — see
// endpoint module docs for why the listener must not do this.
let me = smarm::self_pid();
let _ = registry.cast(Cast::ConnStarted(me));
let _guard = DeregisterGuard::new(registry.clone(), me, Cast::ConnEnded);
loop { loop {
// ----- 1. Read until we have a full request head. ----- // ----- 1. Read until we have a full request head. -----
// We are idle until a head parses: stoppable by a draining
// registry while parked here.
let parsed = match read_head(raw, &mut buf, &limits) { let parsed = match read_head(raw, &mut buf, &limits) {
Ok(p) => p, Ok(p) => p,
Err(ReadHeadErr::ClientClosed) => { Err(ReadHeadErr::ClientClosed) => {
// Clean EOF between requests (or before any request). Normal. // Clean EOF between requests (or before any request). Normal.
return; return;
} }
Err(ReadHeadErr::IdleTimeout) => {
// keep_alive_timeout expired waiting for the first byte of
// a request. Nothing is owed; close silently.
return;
}
Err(ReadHeadErr::HeadTimeout) => {
// head_timeout expired mid-head (slowloris and friends).
// Best-effort 408 WITHOUT parking on writability — a client
// that stalls reads must not defeat the timeout by making
// the 408 write park forever.
try_write_once(raw, b"HTTP/1.1 408 Request Timeout\r\ncontent-length: 0\r\nconnection: close\r\n\r\n");
return;
}
Err(ReadHeadErr::Io(_)) => { Err(ReadHeadErr::Io(_)) => {
// Network error or timeout. Best-effort close; we're done. // Network error. Best-effort close; we're done.
return; return;
} }
Err(ReadHeadErr::Parse(e)) => { Err(ReadHeadErr::Parse(e)) => {
emit_error_response(raw, &e); emit_error_response(raw, &e, Instant::now() + limits.write_timeout);
return; return;
} }
}; };
let _ = registry.cast(Cast::ConnBusy(me));
// ----- 2. Read body. ----- // ----- 2. Read body. -----
// The body has its OWN absolute budget, anchored here (head just
// parsed) and independent of the head clock — a slow-but-legit
// upload must not be judged by the short head deadline. Expiry
// closes the connection (nothing owed mid-body).
let body_deadline = Instant::now() + limits.body_timeout;
// Content-Length pre-check only applies to fixed bodies; a chunked
// body is bounded incrementally by the decoder.
let body_len = parsed.content_length.unwrap_or(0); let body_len = parsed.content_length.unwrap_or(0);
if body_len > limits.max_body_bytes { if body_len > limits.max_body_bytes {
let _ = write_all(raw, b"HTTP/1.1 413 Payload Too Large\r\ncontent-length: 0\r\nconnection: close\r\n\r\n"); let _ = write_all(
raw,
b"HTTP/1.1 413 Payload Too Large\r\ncontent-length: 0\r\nconnection: close\r\n\r\n",
Instant::now() + limits.write_timeout,
);
return; return;
} }
// If client sent `Expect: 100-continue`, emit it before reading the // If client sent `Expect: 100-continue`, emit it before reading the
// body. RFC 7231 §5.1.1. We don't gate on app logic here; v1 always // body. RFC 7231 §5.1.1. We don't gate on app logic here; v1 always
// accepts. // accepts.
if parsed.expect_100 { if parsed.expect_100
if write_all(raw, b"HTTP/1.1 100 Continue\r\n\r\n").is_err() { && write_all(raw, b"HTTP/1.1 100 Continue\r\n\r\n", Instant::now() + limits.write_timeout).is_err()
return; {
} return;
} }
let body = match read_body(raw, &mut buf, parsed.head_len, body_len) { // `consumed_past_head`: how many RAW bytes of `buf` past the head
Ok(b) => b, // this request's body occupied — for chunked bodies that is framing
Err(_) => return, // included, NOT the decoded length. The keep-alive drain at the
// bottom of the loop must drop exactly this much to land on the
// next pipelined request.
let (body, consumed_past_head) = if parsed.chunked {
match read_chunked_body(raw, &mut buf, parsed.head_len, &limits, body_deadline) {
Ok(ok) => ok,
Err(ChunkedBodyErr::TooLarge) => {
let _ = write_all(
raw,
b"HTTP/1.1 413 Payload Too Large\r\ncontent-length: 0\r\nconnection: close\r\n\r\n",
Instant::now() + limits.write_timeout,
);
return;
}
Err(ChunkedBodyErr::Malformed) => {
let _ = write_all(
raw,
b"HTTP/1.1 400 Bad Request\r\ncontent-length: 0\r\nconnection: close\r\n\r\n",
Instant::now() + limits.write_timeout,
);
return;
}
// Timeout mid-body (and any other io error) -> just close.
Err(ChunkedBodyErr::Io(_)) => return,
}
} else {
match read_body(raw, &mut buf, parsed.head_len, body_len, &limits, body_deadline) {
Ok(b) => (b, body_len),
// Timeout mid-body (and any other body io error) -> just
// close; there's no point talking HTTP to a client this far
// gone.
Err(_) => return,
}
}; };
let keep_alive = parsed.keep_alive; let keep_alive = parsed.keep_alive;
@@ -111,12 +237,24 @@ pub fn run_connection(fd: OwnedFd, pipeline: Pipeline, limits: ConnLimits) {
// Catch panics at the actor boundary — a panicking handler should // Catch panics at the actor boundary — a panicking handler should
// not take down the whole connection silently with no response. // not take down the whole connection silently with no response.
let result = std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| { let result = std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| {
#[cfg(feature = "smarm-causal")]
let _g = smarm::causal_site!("pipeline");
pipeline.run(conn) pipeline.run(conn)
})); }));
let mut response_conn = match result { let mut response_conn = match result {
Ok(c) => c, Ok(c) => c,
Err(_) => { Err(_) => {
// Distinguish a genuine handler panic from smarm's stop
// sentinel, which is also a panic payload and which this
// catch_unwind would otherwise swallow — turning a
// graceful stop into a 500-and-keep-running. The stop
// flag is persistent (not consumed by raising the
// sentinel), so if we were stopped this re-raises it
// here, outside the catch, and we unwind properly (the
// fd and registry guards clean up).
smarm::preempt::check_cancelled();
// Compose a 500 manually; the original Conn was moved into // Compose a 500 manually; the original Conn was moved into
// the closure. // the closure.
let mut c = Conn::new(); let mut c = Conn::new();
@@ -132,20 +270,80 @@ pub fn run_connection(fd: OwnedFd, pipeline: Pipeline, limits: ConnLimits) {
.put_body(RespBody::Empty); .put_body(RespBody::Empty);
} }
// ----- 4. Write the response. ----- // ----- 3.5 WebSocket upgrade. -----
let bytes = parser::serialise_response(&response_conn, keep_alive); // An accepted handshake (payload + 101) ends HTTP on this socket:
if write_all(raw, &bytes).is_err() { // write the 101 head, hand the fd to the duplex loop with the
// boxed handler and any bytes already read past this request (a
// client may pipeline its first frame behind the handshake — those
// bytes are ws bytes now). The registry entry stays Busy for the
// whole ws lifetime: an open WebSocket is in-flight work, not
// reapable idle HTTP; graceful shutdown force-stops it out of the
// select park at the drain deadline. The status check is
// defensive: a post-handler plug that clobbered the 101 forfeits
// the upgrade and falls through to plain HTTP.
if response_conn.upgrade.is_some() && response_conn.status == Some(101) {
let head = parser::serialise_response(&response_conn, true);
if write_all(raw, &head, Instant::now() + limits.write_timeout).is_err() {
return;
}
let upgrade = response_conn.upgrade.take().expect("checked above");
buf.drain(..head_len + consumed_past_head);
crate::ws::duplex::run_duplex(raw, buf, upgrade.handler, &limits);
return; return;
} }
// ----- 4. Write the response. -----
// A Stream body on HTTP/1.0 has no chunked framing: the body is
// delimited by EOF, so keep-alive is forced off for this response
// (and `connection: close` goes on the wire). Error statuses on
// 1.0 also force close: we don't advertise keep-alive on a
// 4xx/5xx to a protocol generation where reuse is opt-in.
let is_stream = matches!(response_conn.resp_body, RespBody::Stream(_));
let is_error = response_conn.status.unwrap_or(200) >= 400;
let keep_alive = keep_alive
&& !(version == HttpVersion::Http10 && (is_stream || is_error));
let head_bytes = {
#[cfg(feature = "smarm-causal")]
let _g = smarm::causal_site!("serialize");
parser::serialise_response(&response_conn, keep_alive)
};
if write_all(raw, &head_bytes, Instant::now() + limits.write_timeout).is_err() {
return;
}
if let RespBody::Stream(stream) = response_conn.resp_body {
let chunked = version == HttpVersion::Http11;
if pump_stream(raw, stream, chunked, limits.write_timeout).is_err() {
return;
}
if !chunked {
// EOF delimits the HTTP/1.0 stream body.
smarm::progress!("responses");
return;
}
}
// A full response (head + body, streamed or not) is on the wire:
// the unit of useful work for causal profiling (RFC 007).
smarm::progress!("responses");
// ----- 5. Loop or close. ----- // ----- 5. Loop or close. -----
if !keep_alive { if !keep_alive {
return; return;
} }
// Drop the request bytes (head + body) from `buf`; anything past // Response is on the wire; nothing is owed. Going idle here makes
// them is the start of the next pipelined request. // us stoppable by a draining registry while we park for the next
let consumed = head_len + body_len; // keep-alive request. (If the next request is already pipelined in
// `buf`, the very next read_head parses it without parking and we
// go Busy again — a draining registry's request_stop may still
// catch us, which is acceptable: drain means no new work.)
let _ = registry.cast(Cast::ConnIdle(me));
// Drop the request bytes (head + raw body framing) from `buf`;
// anything past them is the start of the next pipelined request.
let consumed = head_len + consumed_past_head;
buf.drain(..consumed); buf.drain(..consumed);
} }
} }
@@ -157,6 +355,12 @@ pub fn run_connection(fd: OwnedFd, pipeline: Pipeline, limits: ConnLimits) {
#[allow(dead_code)] // io::Error is captured for future logging #[allow(dead_code)] // io::Error is captured for future logging
enum ReadHeadErr { enum ReadHeadErr {
ClientClosed, ClientClosed,
/// keep_alive_timeout expired while waiting for the first byte of a
/// request. Close silently.
IdleTimeout,
/// head_timeout expired after the request had started arriving but
/// before the head finished parsing. Best-effort 408.
HeadTimeout,
Io(io::Error), Io(io::Error),
Parse(ParseError), Parse(ParseError),
} }
@@ -164,16 +368,41 @@ enum ReadHeadErr {
/// Read until `parse_head` succeeds or fails definitively. `buf` may already /// Read until `parse_head` succeeds or fails definitively. `buf` may already
/// contain leftover bytes from a previous keep-alive cycle; we try to parse /// contain leftover bytes from a previous keep-alive cycle; we try to parse
/// those before reading more from the socket. /// those before reading more from the socket.
///
/// Two wall-clock budgets govern the waits (each wait uses whichever budget
/// is currently active):
///
/// - while `buf` is empty and nothing has arrived, we are *idle* and the
/// wait is bounded by `keep_alive_timeout`;
/// - the instant the request has started (first byte read, or pipelined
/// bytes already in `buf` at entry), the *head* clock starts: an
/// `Instant` deadline of `head_timeout` from that moment. This budget
/// covers the HEAD only; the body has its own budget (`body_timeout`),
/// which the caller anchors once the head has parsed.
fn read_head( fn read_head(
fd: RawFd, fd: RawFd,
buf: &mut Vec<u8>, buf: &mut Vec<u8>,
limits: &ConnLimits, limits: &ConnLimits,
) -> Result<parser::ParsedHead, ReadHeadErr> { ) -> Result<parser::ParsedHead, ReadHeadErr> {
let entry = Instant::now();
let idle_deadline = entry + limits.keep_alive_timeout;
// Pipelined leftovers count as a started request.
let mut head_deadline: Option<Instant> = if buf.is_empty() {
None
} else {
Some(entry + limits.head_timeout)
};
loop { loop {
// Try to parse what we already have. On the first iteration of a // Try to parse what we already have. On the first iteration of a
// fresh keep-alive cycle, `buf` may already hold the next request. // fresh keep-alive cycle, `buf` may already hold the next request.
if !buf.is_empty() { if !buf.is_empty() {
match parser::parse_head(buf, limits.max_headers) { let head = {
#[cfg(feature = "smarm-causal")]
let _g = smarm::causal_site!("parse");
parser::parse_head(buf, limits.max_headers)
};
match head {
Ok(h) => return Ok(h), Ok(h) => return Ok(h),
Err(ParseError::Incomplete) => {} // need more bytes Err(ParseError::Incomplete) => {} // need more bytes
Err(e) => return Err(ReadHeadErr::Parse(e)), Err(e) => return Err(ReadHeadErr::Parse(e)),
@@ -184,15 +413,81 @@ fn read_head(
return Err(ReadHeadErr::Parse(ParseError::TooManyHeaders)); return Err(ReadHeadErr::Parse(ParseError::TooManyHeaders));
} }
// Read more. // Read more, bounded by whichever budget is active.
match read_some(fd, buf, limits.initial_read_buf) { let deadline = head_deadline.unwrap_or(idle_deadline);
match read_some(fd, buf, limits.initial_read_buf, deadline) {
Ok(0) => return Err(ReadHeadErr::ClientClosed), Ok(0) => return Err(ReadHeadErr::ClientClosed),
Ok(_) => continue, Ok(_) => {
if head_deadline.is_none() {
// First byte(s) of this request: the head clock
// starts now.
head_deadline = Some(Instant::now() + limits.head_timeout);
}
}
Err(e) if e.kind() == ErrorKind::TimedOut => {
return Err(if head_deadline.is_some() {
ReadHeadErr::HeadTimeout
} else {
ReadHeadErr::IdleTimeout
});
}
Err(e) => return Err(ReadHeadErr::Io(e)), Err(e) => return Err(ReadHeadErr::Io(e)),
} }
} }
} }
// ---------------------------------------------------------------------------
// BodyStallGate — burst-gated stall eviction for body reads
// ---------------------------------------------------------------------------
//
// Each body read is bounded by the SOONER of two deadlines: the absolute
// body cap (`body_timeout`, passed in as `cap`) and a sliding stall window
// (`mark + body_stall_timeout`). The stall mark only advances when the
// client delivers a full burst (`body_burst_bytes` accumulated since the
// last advance) — so a steady sub-burst trickle never moves the mark and
// is evicted at ~body_stall_timeout, while a bursty slow-but-legit client
// keeps resetting it and survives up to the absolute cap.
//
// State is two words (`mark`, `since_mark`); the per-read cost is one add
// and one compare. Bytes counted are RAW socket bytes (progress = the
// client is sending *something*), so chunked framing counts too, and a
// burst that the kernel fragments into several reads still accumulates.
struct BodyStallGate {
cap: Instant,
stall_timeout: Duration,
burst_bytes: usize,
mark: Instant,
since_mark: usize,
}
impl BodyStallGate {
fn new(cap: Instant, limits: &ConnLimits, now: Instant) -> Self {
Self {
cap,
stall_timeout: limits.body_stall_timeout,
burst_bytes: limits.body_burst_bytes,
mark: now,
since_mark: 0,
}
}
/// Deadline for the next read: the sooner of the absolute cap and the
/// current stall window.
fn deadline(&self) -> Instant {
(self.mark + self.stall_timeout).min(self.cap)
}
/// Record `n` freshly-read raw body bytes; advance the stall mark if a
/// full burst has accumulated since the last advance.
fn record(&mut self, n: usize, now: Instant) {
self.since_mark += n;
if self.since_mark >= self.burst_bytes {
self.mark = now;
self.since_mark = 0;
}
}
}
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// read_body // read_body
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -202,6 +497,8 @@ fn read_body(
buf: &mut Vec<u8>, buf: &mut Vec<u8>,
head_len: usize, head_len: usize,
body_len: usize, body_len: usize,
limits: &ConnLimits,
cap: Instant,
) -> io::Result<Vec<u8>> { ) -> io::Result<Vec<u8>> {
// Bytes already in `buf` past the head belong to the body. // Bytes already in `buf` past the head belong to the body.
let already = buf.len().saturating_sub(head_len); let already = buf.len().saturating_sub(head_len);
@@ -213,12 +510,17 @@ fn read_body(
return Ok(buf[head_len..head_len + body_len].to_vec()); return Ok(buf[head_len..head_len + body_len].to_vec());
} }
// Read until we have the rest. // Read until we have the rest, bounded by the body cap AND the
// burst-gated stall window (whichever is sooner).
let mut gate = BodyStallGate::new(cap, limits, Instant::now());
let mut total_read = already; let mut total_read = already;
while total_read < body_len { while total_read < body_len {
match read_some(fd, buf, 8 * 1024) { match read_some(fd, buf, 8 * 1024, gate.deadline()) {
Ok(0) => return Err(io::Error::new(ErrorKind::UnexpectedEof, "client closed during body")), Ok(0) => return Err(io::Error::new(ErrorKind::UnexpectedEof, "client closed during body")),
Ok(n) => total_read += n, Ok(n) => {
total_read += n;
gate.record(n, Instant::now());
}
Err(e) => return Err(e), Err(e) => return Err(e),
} }
} }
@@ -226,18 +528,166 @@ fn read_body(
} }
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// read_some — single epoll-park + read loop. // read_chunked_body — incremental chunked transfer-decoding (request side).
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// //
// Appends what it reads onto `buf`. Returns bytes read, 0 for EOF, or the // Decodes `Transfer-Encoding: chunked` from `buf[head_len..]`, reading more
// last io error. // from the socket as needed on the body deadline (anchored by the caller
// when the head finished parsing, independent of the head clock).
// Returns (decoded_body, raw_bytes_consumed_past_head) — the raw
// count includes all framing and the trailer section, so the caller's
// keep-alive drain lands exactly on the next pipelined request.
//
// Bounds: the DECODED size is capped at max_body_bytes (-> TooLarge/413);
// a single size line (incl. chunk extensions, which are ignored) is capped
// at MAX_SIZE_LINE and the trailer section at MAX_TRAILER_BYTES (->
// Malformed/400) so framing spam can't grow `buf` unboundedly. Trailers
// are consumed and discarded — nothing in the pipeline wants them yet.
fn read_some(fd: RawFd, buf: &mut Vec<u8>, chunk: usize) -> io::Result<usize> { const MAX_SIZE_LINE: usize = 128;
const MAX_TRAILER_BYTES: usize = 8 * 1024;
#[allow(dead_code)] // io::Error is captured for future logging
enum ChunkedBodyErr {
Io(io::Error),
Malformed,
TooLarge,
}
fn read_chunked_body(
fd: RawFd,
buf: &mut Vec<u8>,
head_len: usize,
limits: &ConnLimits,
deadline: Instant,
) -> Result<(Vec<u8>, usize), ChunkedBodyErr> {
// Ensure `buf` holds at least `until` bytes, reading under the body
// stall gate (absolute cap AND burst-gated stall window). Io(TimedOut)
// on expiry, UnexpectedEof on early close. All chunked socket reads
// funnel through here, so recording bytes here covers the whole path.
fn fill_to(
fd: RawFd,
buf: &mut Vec<u8>,
until: usize,
gate: &mut BodyStallGate,
) -> Result<(), ChunkedBodyErr> {
while buf.len() < until {
match read_some(fd, buf, 8 * 1024, gate.deadline()) {
Ok(0) => {
return Err(ChunkedBodyErr::Io(io::Error::new(
ErrorKind::UnexpectedEof,
"client closed during chunked body",
)))
}
Ok(n) => gate.record(n, Instant::now()),
Err(e) => return Err(ChunkedBodyErr::Io(e)),
}
}
Ok(())
}
// Find "\r\n" in buf[from..], reading more as needed; the line may be
// at most `max_line` bytes (terminator excluded). Returns the index of
// the '\r'.
fn find_crlf(
fd: RawFd,
buf: &mut Vec<u8>,
from: usize,
max_line: usize,
gate: &mut BodyStallGate,
) -> Result<usize, ChunkedBodyErr> {
let mut scan = from;
loop {
while scan + 1 < buf.len() {
if buf[scan] == b'\r' && buf[scan + 1] == b'\n' {
return Ok(scan);
}
scan += 1;
if scan - from > max_line {
return Err(ChunkedBodyErr::Malformed);
}
}
fill_to(fd, buf, buf.len() + 1, gate)?;
}
}
// `deadline` is the absolute body cap; the gate layers the burst-gated
// stall window under it. All reads below go through find_crlf/fill_to.
let mut gate = BodyStallGate::new(deadline, limits, Instant::now());
let mut pos = head_len;
let mut decoded: Vec<u8> = Vec::new();
loop {
// ----- size line: HEX[;extensions]\r\n -----
let line_end = find_crlf(fd, buf, pos, MAX_SIZE_LINE, &mut gate)?;
let line = &buf[pos..line_end];
let size_str = match line.iter().position(|&b| b == b';') {
Some(i) => &line[..i], // chunk extensions: ignored
None => line,
};
let size_str = std::str::from_utf8(size_str)
.map_err(|_| ChunkedBodyErr::Malformed)?
.trim();
let size = usize::from_str_radix(size_str, 16)
.map_err(|_| ChunkedBodyErr::Malformed)?;
pos = line_end + 2;
if size == 0 {
// ----- trailer section: zero or more header lines, then CRLF -----
let trailer_start = pos;
loop {
let t_end = find_crlf(fd, buf, pos, MAX_SIZE_LINE.max(1024), &mut gate)?;
let empty = t_end == pos;
pos = t_end + 2;
if empty {
return Ok((decoded, pos - head_len));
}
if pos - trailer_start > MAX_TRAILER_BYTES {
return Err(ChunkedBodyErr::Malformed);
}
}
}
if decoded.len() + size > limits.max_body_bytes {
return Err(ChunkedBodyErr::TooLarge);
}
// ----- chunk payload + trailing CRLF -----
fill_to(fd, buf, pos + size + 2, &mut gate)?;
decoded.extend_from_slice(&buf[pos..pos + size]);
if &buf[pos + size..pos + size + 2] != b"\r\n" {
return Err(ChunkedBodyErr::Malformed);
}
pos += size + 2;
}
}
// ---------------------------------------------------------------------------
// read_some — single epoll-park + read loop, bounded by a deadline.
// ---------------------------------------------------------------------------
//
// Appends what it reads onto `buf`. Returns bytes read, 0 for EOF,
// `ErrorKind::TimedOut` when `deadline` passes before the fd turns
// readable, or the last io error.
pub(crate) fn read_some(
fd: RawFd,
buf: &mut Vec<u8>,
chunk: usize,
deadline: Instant,
) -> io::Result<usize> {
// Loop to absorb EAGAIN: a readable wakeup followed by EAGAIN is // Loop to absorb EAGAIN: a readable wakeup followed by EAGAIN is
// possible (signal race, etc). Re-park and retry rather than returning // possible (signal race, etc). Re-park and retry rather than returning
// 0 (which would be confused with EOF by callers). // 0 (which would be confused with EOF by callers). The deadline is an
// Instant, so spurious wakes don't reset the budget.
loop { loop {
smarm::wait_readable(fd)?; let remaining = deadline.saturating_duration_since(Instant::now());
if remaining.is_zero() {
return Err(io::Error::new(ErrorKind::TimedOut, "read deadline elapsed"));
}
if !smarm::wait_readable_timeout(fd, remaining)? {
return Err(io::Error::new(ErrorKind::TimedOut, "read deadline elapsed"));
}
let start = buf.len(); let start = buf.len();
buf.resize(start + chunk, 0); buf.resize(start + chunk, 0);
@@ -262,13 +712,45 @@ fn read_some(fd: RawFd, buf: &mut Vec<u8>, chunk: usize) -> io::Result<usize> {
} }
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// write_all — robust write loop. // try_write_once — single non-parking write attempt, result ignored.
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
//
// For best-effort farewells (the 408) to clients we've decided to drop:
// one non-blocking write syscall, no wait_writable park. A client that
// stalls its read side must not be able to keep this actor alive past its
// own timeout. The socket send buffer almost always has room for a
// one-liner, so in practice the 408 lands.
fn write_all(fd: RawFd, mut buf: &[u8]) -> io::Result<()> { fn try_write_once(fd: RawFd, buf: &[u8]) {
unsafe {
let _ = libc::write(fd, buf.as_ptr() as *const _, buf.len());
}
}
// ---------------------------------------------------------------------------
// write_all — robust write loop, bounded by a deadline.
// ---------------------------------------------------------------------------
//
// Mirrors read_some: each writability park is bounded by the remaining
// budget. `ErrorKind::TimedOut` when the deadline passes before the bytes
// are down — a client that stops reading must not pin this actor in
// wait_writable forever (the write-side twin of slowloris).
pub(crate) fn write_all(fd: RawFd, mut buf: &[u8], deadline: Instant) -> io::Result<()> {
// The whole loop — writability parks included — runs under the
// `socket-write` causal site: a park inside a site is exactly what
// RFC 007's park-gated resume credit exists to attribute.
#[cfg(feature = "smarm-causal")]
let _g = smarm::causal_site!("socket-write");
while !buf.is_empty() { while !buf.is_empty() {
// Park on writability before each syscall. // Park on writability before each syscall, bounded by the budget.
smarm::wait_writable(fd)?; let remaining = deadline.saturating_duration_since(Instant::now());
if remaining.is_zero() {
return Err(io::Error::new(ErrorKind::TimedOut, "write deadline elapsed"));
}
if !smarm::wait_writable_timeout(fd, remaining)? {
return Err(io::Error::new(ErrorKind::TimedOut, "write deadline elapsed"));
}
let n = unsafe { let n = unsafe {
libc::write(fd, buf.as_ptr() as *const _, buf.len()) libc::write(fd, buf.as_ptr() as *const _, buf.len())
@@ -288,11 +770,89 @@ fn write_all(fd: RawFd, mut buf: &[u8]) -> io::Result<()> {
Ok(()) Ok(())
} }
// ---------------------------------------------------------------------------
// pump_stream — drive a RespBody::Stream onto the wire.
// ---------------------------------------------------------------------------
//
// The pull side of the v0.3 streaming design: the handler's producer actor
// owns the Sender; this conn actor owns the socket and every write
// deadline. We park in `recv()` between chunks — that park is stoppable
// (a draining registry's `request_stop` unwinds us out of `park_current`
// via the stop sentinel; fd + registry guards clean up), so an infinite
// stream is force-stoppable at the drain deadline like any other in-flight
// request. End of stream is the channel closing: every Sender dropped.
//
// Each chunk gets a FRESH write_timeout budget — a stream is expected to
// outlive any whole-response clock; what is not tolerated is a single
// write stalling. On write failure we return Err: the conn loop drops the
// Receiver, and the producer's next `send` observes the closed channel and
// should exit (that is the documented producer contract).
//
// `chunked` selects HTTP/1.1 chunked framing (hex-length CRLF payload
// CRLF, terminated by a 0-chunk) vs HTTP/1.0 raw writes (EOF-delimited;
// caller closes). Empty chunks are skipped — a zero-length chunk would
// terminate the framing early.
fn pump_stream(
fd: RawFd,
stream: StreamBody,
chunked: bool,
write_timeout: Duration,
) -> io::Result<()> {
let write_chunk = |payload: &[u8]| -> io::Result<()> {
let deadline = Instant::now() + write_timeout;
if chunked {
let mut framed = Vec::with_capacity(payload.len() + 20);
framed.extend_from_slice(format!("{:x}\r\n", payload.len()).as_bytes());
framed.extend_from_slice(payload);
framed.extend_from_slice(b"\r\n");
write_all(fd, &framed, deadline)
} else {
write_all(fd, payload, deadline)
}
};
loop {
// With a heartbeat configured (SSE), the wait between chunks is the
// heartbeat interval; expiry emits the ping and keeps waiting. A
// ping write that stalls past write_timeout errors out below —
// that is the dead-client detector. Without one, plain recv():
// stoppable by the draining registry either way.
let msg = match &stream.heartbeat {
Some((interval, payload)) => match stream.rx.recv_timeout(*interval) {
Ok(chunk) => Some(chunk),
Err(smarm::RecvTimeoutError::Timeout) => {
write_chunk(payload)?;
continue;
}
Err(smarm::RecvTimeoutError::Disconnected) => None,
},
None => stream.rx.recv().ok(),
};
match msg {
Some(chunk) => {
if chunk.is_empty() {
continue;
}
write_chunk(&chunk)?;
}
None => {
// All senders dropped: end of stream.
if chunked {
write_all(fd, b"0\r\n\r\n", Instant::now() + write_timeout)?;
}
return Ok(());
}
}
}
}
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// Error responses for unparseable / malformed requests. // Error responses for unparseable / malformed requests.
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
fn emit_error_response(fd: RawFd, err: &ParseError) { fn emit_error_response(fd: RawFd, err: &ParseError, deadline: Instant) {
let resp: &[u8] = match err { let resp: &[u8] = match err {
ParseError::TooManyHeaders => ParseError::TooManyHeaders =>
b"HTTP/1.1 431 Request Header Fields Too Large\r\ncontent-length: 0\r\nconnection: close\r\n\r\n", b"HTTP/1.1 431 Request Header Fields Too Large\r\ncontent-length: 0\r\nconnection: close\r\n\r\n",
@@ -300,10 +860,12 @@ fn emit_error_response(fd: RawFd, err: &ParseError) {
b"HTTP/1.1 400 Bad Request\r\ncontent-length: 0\r\nconnection: close\r\n\r\n", b"HTTP/1.1 400 Bad Request\r\ncontent-length: 0\r\nconnection: close\r\n\r\n",
ParseError::Unsupported => ParseError::Unsupported =>
b"HTTP/1.1 411 Length Required\r\ncontent-length: 0\r\nconnection: close\r\n\r\n", b"HTTP/1.1 411 Length Required\r\ncontent-length: 0\r\nconnection: close\r\n\r\n",
ParseError::UnknownTransferCoding =>
b"HTTP/1.1 501 Not Implemented\r\ncontent-length: 0\r\nconnection: close\r\n\r\n",
// Incomplete and Malformed both lead here; Incomplete shouldn't // Incomplete and Malformed both lead here; Incomplete shouldn't
// appear (read_head loops on it). // appear (read_head loops on it).
_ => _ =>
b"HTTP/1.1 400 Bad Request\r\ncontent-length: 0\r\nconnection: close\r\n\r\n", b"HTTP/1.1 400 Bad Request\r\ncontent-length: 0\r\nconnection: close\r\n\r\n",
}; };
let _ = write_all(fd, resp); let _ = write_all(fd, resp, deadline);
} }
+656
View File
@@ -0,0 +1,656 @@
//! The endpoint: one supervised gen_server that owns a listening socket,
//! the pool of listener actors accepting on it, and the set of live
//! connections.
//!
//! # Shape (v0.3)
//!
//! ```text
//! your root supervisor
//! └── ChildSpec(Permanent, urus::endpoint(config, pipeline)?) <- Endpoint gen_server
//! └── listener_sup OneForOne over N listener actors (spawned in init)
//! └── (listeners spawn plain connection actors)
//! ```
//!
//! The endpoint runs *inline as the ChildSpec's actor*
//! (`NamedGenServerBuilder::run`), so your supervisor's ordered shutdown
//! reaches it as [`GenServer::handle_shutdown`] and a restart re-runs the
//! factory on the same still-open listen fds.
//!
//! ## Why the endpoint spawns its own listener supervisor
//!
//! The obvious tree is registry and listener-sup as *siblings* under a
//! `RestForOne`, with listeners resolving the registry by name. It has a
//! boot race: smarm's supervisor starts children with a fire-and-forget
//! spawn, so start *order* is not start *readiness* — a listener can look
//! the name up before the registry actor has run. Rather than paper over
//! that with a retry loop, the registrar spawns its consumers: everything
//! here is program order inside one actor's `init`, not a cross-actor
//! guarantee. (The general gap is filed in smarm's ROADMAP as a
//! supervisor readiness-ack item; when it lands, the sibling shape becomes
//! available, but this one costs nothing and is not waiting on it.)
//!
//! ## Connection accounting
//!
//! Connection actors self-register as their first action and self-
//! deregister via a drop guard (so panic unwinds deregister too); both
//! casts come from the same sender, so `ConnStarted` always precedes
//! `ConnEnded` in the inbox. Were the *listener* to announce the pid
//! instead, a short-lived conn's `Ended` could overtake its `Started` and
//! leak a dead pid into the set forever.
//!
//! ## Shutdown
//!
//! `handle_shutdown` (i.e. a `request_shutdown` from your supervisor, or
//! from anywhere):
//!
//! 1. flip `draining` — from here on, a connection that registers or goes
//! idle is stopped on the spot;
//! 2. `request_shutdown` the listener supervisor, which stops its
//! listeners in reverse order and exits; its monitored death is the
//! "no new connections can ever be accepted" barrier;
//! 3. stop every currently-idle connection;
//! 4. arm one `drain_timeout` timer;
//! 5. return `Continue` — the endpoint exits normally once the listener
//! sup is down *and* the connection set is empty, so "the endpoint
//! actor has exited" is exactly "every connection is gone". With no
//! connections and (once the Down lands) nothing to wait for, that is
//! immediate.
//!
//! At the drain deadline one force sweep `request_stop`s whatever remains
//! (plus, defensively, the listener sup). A single sweep suffices: every
//! path a connection can enter the set on afterwards stops it at the point
//! of entry. Force-stopped conns unwind safely out of their fd waits and
//! close their sockets via `OwnedFd::drop`.
//!
//! Give the endpoint [`Shutdown::Infinity`](smarm::supervisor::Shutdown) in
//! your child spec: it bounds itself with `drain_timeout`, and a
//! supervisor-imposed deadline shorter than that would kill the drain
//! halfway and orphan connections.
//!
//! ## Known residual window
//!
//! A connection *spawned* by a listener in the instant before that
//! listener is stopped, which has not yet run, has not registered. If the
//! set was already empty the endpoint can exit before its `ConnStarted`
//! arrives, and that connection serves on as a forest root until smarm's
//! root-exit sweep collects it. Inherent to self-registration (the
//! alternative loses `Started`/`Ended` ordering, which is worse);
//! documented rather than defended.
use crate::conn_actor::{run_connection, ConnLimits};
use crate::net::{accept_nonblocking, bind_and_listen, OwnedFd};
use crate::plug::Pipeline;
use crate::serve::Config;
use smarm::gen_server::{GenServerName, ShutdownAction, StopHandle, TimerHandle};
use smarm::{
ChildSpec, GenServer, GenServerBuilder, GenServerCtx, GenServerRef, OneForOne, Pid, Restart,
Strategy,
};
use std::collections::HashMap;
use std::io::{self, ErrorKind};
use std::os::fd::RawFd;
use std::sync::atomic::{AtomicU32, Ordering};
use std::sync::Arc;
use std::time::Duration;
// ---------------------------------------------------------------------------
// Messages
// ---------------------------------------------------------------------------
pub enum Cast {
/// A connection actor started (self-registered, initially idle: it has
/// not parsed a request head yet).
ConnStarted(Pid),
/// Parsed a request head; a response is now owed.
ConnBusy(Pid),
/// Response written; parked (or about to park) waiting for the next
/// keep-alive request.
ConnIdle(Pid),
ConnEnded(Pid),
}
pub enum Call {
/// Live connection count — introspection (and the planned hook for
/// ws/channels stats).
ConnCount,
}
pub enum Reply {
ConnCount(usize),
}
// ---------------------------------------------------------------------------
// Endpoint
// ---------------------------------------------------------------------------
#[derive(Clone, Copy, PartialEq, Eq)]
enum ConnState {
Busy,
Idle,
}
/// Everything an endpoint incarnation needs to (re)build its listener
/// pool. Shared behind an `Arc` by the factory closure so a restart reuses
/// the same already-bound fds — no re-bind, no window where the port is
/// unclaimed.
struct Boot {
listener_fds: Vec<Arc<OwnedFd>>,
pipeline: Pipeline,
limits: ConnLimits,
conn_stack_reserve: usize,
name: &'static str,
}
pub struct Endpoint {
boot: Arc<Boot>,
conns: HashMap<Pid, ConnState>,
draining: bool,
drain_timeout: Duration,
/// The listener supervisor, spawned in `init` and monitored. `None`
/// once it is down.
listener_sup: Option<Pid>,
stop: Option<StopHandle<Self>>,
timer: Option<TimerHandle<Self>>,
}
impl Endpoint {
fn name(&self) -> GenServerName<Self> {
GenServerName::new(self.boot.name)
}
/// Exit normally once nothing can arrive and nothing is left: the
/// listener sup is down and the conn set is empty. This exit *is* the
/// barrier a supervisor's ordered shutdown waits on.
fn stop_if_drained(&self) {
if self.draining && self.listener_sup.is_none() && self.conns.is_empty() {
self.stop.as_ref().expect("init ran first").stop();
}
}
}
impl GenServer for Endpoint {
type Call = Call;
type Reply = Reply;
type Cast = Cast;
type Info = ();
/// One meaning: the drain deadline elapsed — force-stop the stragglers.
type Timer = ();
fn init(&mut self, ctx: &GenServerCtx<Self>) {
ctx.trap_exit(); // shutdown arrives as handle_shutdown, not a kill
self.stop = Some(ctx.stop_handle());
self.timer = Some(ctx.timer());
// Our own name is bound (from inside `run`) before `init` runs, so
// this resolves to us. Listeners are handed the resolved ref, not
// the name: no per-connection registry lookup on the accept path.
let me: GenServerRef<Self> =
smarm::gen_server::whereis_server(self.name()).expect("endpoint name bound before init");
let boot = self.boot.clone();
let sup = smarm::spawn(move || {
let mut sup = OneForOne::new().strategy(Strategy::OneForOne);
for lfd in &boot.listener_fds {
let lfd = lfd.clone();
let pipeline = boot.pipeline.clone();
let limits = boot.limits;
let reserve = boot.conn_stack_reserve;
let me = me.clone();
// Permanent: a listener only ever exits by supervisor
// action now (its normal-exit-as-shutdown flag is gone),
// so "exited on its own" always means something broke and
// always deserves a restart.
sup = sup.child(ChildSpec::new(Restart::Permanent, move || {
listener_loop(lfd.clone(), pipeline.clone(), limits, reserve, me.clone());
}));
}
// Default intensity (3 restarts / 5s) applies; a listener
// crash-looping faster than that trips the cap and the pool
// tears down — which our monitor turns into a loud endpoint
// failure rather than a zombie server on a dead port.
sup.run();
});
ctx.watch(smarm::monitor(sup.pid()));
self.listener_sup = Some(sup.pid());
}
fn handle_call(&mut self, request: Call) -> Reply {
match request {
Call::ConnCount => Reply::ConnCount(self.conns.len()),
}
}
fn handle_cast(&mut self, request: Cast) {
match request {
Cast::ConnStarted(pid) => {
self.conns.insert(pid, ConnState::Idle);
if self.draining {
// Accepted just before its listener was stopped.
// Draining means no new work; it leaves the map via
// its guard's ConnEnded.
smarm::request_stop(pid);
}
}
Cast::ConnBusy(pid) => {
if let Some(s) = self.conns.get_mut(&pid) {
*s = ConnState::Busy;
}
}
Cast::ConnIdle(pid) => {
if let Some(s) = self.conns.get_mut(&pid) {
*s = ConnState::Idle;
if self.draining {
smarm::request_stop(pid); // in-flight request finished
}
}
}
Cast::ConnEnded(pid) => {
self.conns.remove(&pid);
self.stop_if_drained();
}
}
}
fn handle_down(&mut self, down: smarm::Down) {
if self.listener_sup != Some(down.pid) {
return;
}
self.listener_sup = None;
if !self.draining {
// The pool died on its own (restart intensity exceeded, or the
// supervisor itself failed). Nothing is listening any more, so
// this endpoint is a zombie: fail loudly and let the caller's
// supervisor decide (restart re-runs init on the same fds).
panic!("urus endpoint '{}': listener pool died: {:?}", self.boot.name, down.reason);
}
// Expected during shutdown: the accept side is now provably gone.
self.stop_if_drained();
}
fn handle_shutdown(&mut self) -> ShutdownAction {
self.draining = true;
if let Some(sup) = self.listener_sup {
// Stops listeners in reverse start order and exits; the
// monitored Down is our "no new connections" barrier.
smarm::request_shutdown(sup);
}
for (pid, state) in &self.conns {
if *state == ConnState::Idle {
smarm::request_stop(*pid);
}
}
// Busy conns get until the deadline; then handle_timer sweeps.
self.timer
.as_ref()
.expect("init ran first")
.arm_after(self.drain_timeout, ());
ShutdownAction::Continue
}
fn handle_timer(&mut self, _deadline: ()) {
// Drain deadline: force-stop everything left. One sweep — late
// registrants are already stopped on arrival (see module docs).
for pid in self.conns.keys() {
smarm::request_stop(*pid);
}
if let Some(sup) = self.listener_sup {
// Defensive: a listener wedged past its supervisor's own grace
// period must not hold the whole shutdown open.
smarm::request_stop(sup);
}
// The map empties via each conn's guard ConnEnded; stop_if_drained
// fires on the last one (or on the sup's Down, whichever is last).
}
}
// ---------------------------------------------------------------------------
// endpoint() — the public constructor
// ---------------------------------------------------------------------------
/// Bind `config.addr` and return the endpoint's supervisable body: pass it
/// to [`ChildSpec::new`] under your own supervisor, alongside your
/// application's other children.
///
/// The bind happens **here**, eagerly, so an address-in-use error surfaces
/// on the caller's thread rather than inside an actor — and the fds
/// outlive any restart of the child.
///
/// ```ignore
/// let endpoint = urus::endpoint(Config::new(addr), pipeline)?;
/// let rt = smarm::init(smarm::Config::default());
/// rt.run(move || {
/// smarm::OneForOne::new()
/// .child(ChildSpec::new(Restart::Permanent, my_app_state))
/// .child(ChildSpec::new(Restart::Permanent, endpoint)
/// .shutdown(Shutdown::Infinity))
/// .run()
/// });
/// ```
///
/// The returned closure is `Clone`, so one endpoint definition can be
/// handed to more than one place; each *invocation* is one running
/// endpoint, and two live at once under the same [`Config::name`] is a
/// name clash (panic on the second).
pub fn endpoint(
config: Config,
pipeline: Pipeline,
) -> io::Result<impl Fn() + Clone + Send + Sync + 'static> {
// One dup'd fd per listener: `accept4` is thread-safe on a single fd,
// but a per-listener RawFd keeps each actor's epoll registration
// distinct in smarm's `waiters: HashMap<RawFd, Pid>`.
let mut listener_fds = Vec::with_capacity(config.listener_pool);
listener_fds.push(Arc::new(bind_and_listen(config.addr)?));
for _ in 1..config.listener_pool {
let dup = dup_fd(listener_fds[0].as_raw())?;
listener_fds.push(Arc::new(dup));
}
let boot = Arc::new(Boot {
listener_fds,
pipeline,
limits: config.to_conn_limits(),
conn_stack_reserve: config.conn_stack_reserve,
name: config.name,
});
let drain_timeout = config.drain_timeout;
Ok(move || {
let state = Endpoint {
boot: boot.clone(),
conns: HashMap::new(),
draining: false,
drain_timeout,
listener_sup: None,
stop: None,
timer: None,
};
let name = GenServerName::<Endpoint>::new(state.boot.name);
// Inline: this actor *is* the server, so a supervisor's shutdown
// reaches handle_shutdown and a restart re-binds the name.
if let Err(e) = GenServerBuilder::new(state).named(name).run() {
panic!("urus endpoint '{}': name already taken: {e:?}", name.as_str());
}
})
}
/// Resolve a running endpoint by [`Config::name`] — for `ConnCount` and
/// other introspection from elsewhere in the app.
pub fn whereis(name: &'static str) -> Option<GenServerRef<Endpoint>> {
smarm::gen_server::whereis_server(GenServerName::<Endpoint>::new(name))
}
// ---------------------------------------------------------------------------
// Listener actors
// ---------------------------------------------------------------------------
fn dup_fd(fd: RawFd) -> io::Result<OwnedFd> {
let new_fd = unsafe { libc::fcntl(fd, libc::F_DUPFD_CLOEXEC, 0) };
if new_fd < 0 {
return Err(io::Error::last_os_error());
}
Ok(OwnedFd::from_raw(new_fd))
}
/// Test-only fault injection. When nonzero, the next accept-loop iteration
/// of whichever listener gets there first decrements this and panics —
/// *before* calling `accept`, so a pending connection stays in the kernel
/// backlog and must be picked up by the restarted listener. Cost when idle
/// is one relaxed load per accept-loop iteration (each of which already
/// pays a syscall). Not public API.
#[doc(hidden)]
pub static INJECT_LISTENER_PANICS: AtomicU32 = AtomicU32::new(0);
/// Accept forever, spawning one connection actor per connection. Exits
/// only by supervisor action: `request_stop` unwinds the untimed
/// `wait_readable` park below (smarm's 06-10 io fix), which is why there
/// is no shutdown flag and no tick — the park is genuinely open-ended and
/// costs nothing while idle.
fn listener_loop(
listener: Arc<OwnedFd>,
pipeline: Pipeline,
limits: ConnLimits,
conn_stack_reserve: usize,
endpoint: GenServerRef<Endpoint>,
) {
let fd = listener.as_raw();
loop {
if INJECT_LISTENER_PANICS
.fetch_update(Ordering::Relaxed, Ordering::Relaxed, |n| n.checked_sub(1))
.is_ok()
{
panic!("urus: injected listener panic (test hook)");
}
match accept_nonblocking(fd) {
Ok(client) => {
// Hand the fd off to a new connection actor. spawn() is
// cheap on smarm — it's a single Vec push under the
// shared lock.
let p = pipeline.clone();
let l = limits;
let e = endpoint.clone();
let opts = smarm::SpawnOpts {
stack_reserve: Some(conn_stack_reserve),
..smarm::SpawnOpts::default()
};
smarm::spawn_with(opts, move || run_connection(client, p, l, e));
}
Err(e) if e.kind() == ErrorKind::WouldBlock => {
if let Err(we) = smarm::wait_readable(fd) {
// epoll registration failed — abnormal, so panic: a
// transient failure (e.g. EMFILE on the epoll set)
// heals by restart instead of silently shrinking the
// pool. smarm catches actor panics in the trampoline;
// this is a Signal::Panic to the supervisor, not
// process noise.
panic!("urus: listener wait_readable failed: {we}");
}
}
Err(e) if e.kind() == ErrorKind::Interrupted => continue,
Err(e) => {
// EMFILE / ENFILE / ECONNABORTED etc. Log and back off
// briefly in case the error is sticky; the system may
// recover.
eprintln!("urus: accept error: {e}");
smarm::sleep(Duration::from_millis(10));
}
}
}
// The Arc clone we were started with drops on unwind, but the
// ChildSpec factory holds another — the fd outlives any one
// incarnation of this listener.
}
// ---------------------------------------------------------------------------
// Deregistration guard
// ---------------------------------------------------------------------------
/// Casts `make(pid)` on drop. Runs on normal return, `request_stop`
/// unwind, and panic unwind alike; the cast is infallible from the
/// guard's perspective (a dead endpoint just returns an ignored Err).
pub struct DeregisterGuard {
endpoint: GenServerRef<Endpoint>,
pid: Pid,
make: fn(Pid) -> Cast,
}
impl DeregisterGuard {
pub fn new(endpoint: GenServerRef<Endpoint>, pid: Pid, make: fn(Pid) -> Cast) -> Self {
Self { endpoint, pid, make }
}
}
impl Drop for DeregisterGuard {
fn drop(&mut self) {
let _ = self.endpoint.cast((self.make)(self.pid));
}
}
// ---------------------------------------------------------------------------
// Tests — the drain protocol, driven purely by request_shutdown.
// ---------------------------------------------------------------------------
#[cfg(test)]
mod tests {
use super::*;
use crate::plug::Pipeline;
use std::time::Instant;
const NAME: &str = "urus.test.endpoint";
/// Spawn a real endpoint (bound to an ephemeral port) as a supervised
/// child, exactly as an application would, and hand back its
/// supervisor's pid.
fn spawn_endpoint(drain: Duration) -> Pid {
let cfg = Config {
listener_pool: 1,
drain_timeout: drain,
name: NAME,
..Config::new("127.0.0.1:0".parse().unwrap())
};
let body = endpoint(cfg, Pipeline::new()).expect("bind");
let sup = smarm::spawn(move || {
OneForOne::new()
.child(
ChildSpec::new(Restart::Permanent, body)
.shutdown(smarm::supervisor::Shutdown::Infinity),
)
.run()
});
await_pred("endpoint up", Duration::from_secs(2), || {
whereis(NAME).is_some()
});
sup.pid()
}
/// A stand-in connection actor: registers with the endpoint, optionally
/// reports busy, then parks forever. Only a `request_stop` ends it; the
/// guard's ConnEnded runs on the unwind.
fn fake_conn(ep: GenServerRef<Endpoint>, busy: bool) {
let me = smarm::self_pid();
let _ = ep.cast(Cast::ConnStarted(me));
let _guard = DeregisterGuard::new(ep.clone(), me, Cast::ConnEnded);
if busy {
let _ = ep.cast(Cast::ConnBusy(me));
}
loop {
smarm::sleep(Duration::from_secs(3600));
}
}
fn spawn_fake_conn(busy: bool) -> Pid {
let ep = whereis(NAME).expect("endpoint live");
smarm::spawn(move || fake_conn(ep, busy)).pid()
}
/// Live conn count, or `Err` once the endpoint is gone.
fn conn_count() -> Result<usize, ()> {
match whereis(NAME) {
Some(ep) => match ep.call(Call::ConnCount) {
Ok(Reply::ConnCount(n)) => Ok(n),
Err(_) => Err(()),
},
None => Err(()),
}
}
fn await_pred(what: &str, deadline: Duration, mut pred: impl FnMut() -> bool) {
let end = Instant::now() + deadline;
while !pred() {
assert!(Instant::now() < end, "timed out waiting for: {what}");
smarm::sleep(Duration::from_millis(5));
}
}
#[test]
fn shutdown_with_no_conns_exits_promptly() {
smarm::run(|| {
let sup = spawn_endpoint(Duration::from_secs(30));
assert_eq!(conn_count(), Ok(0));
let t0 = Instant::now();
smarm::request_shutdown(sup);
// Nothing to drain: gone well inside the (huge) drain window,
// i.e. the idle path does not run the clock out.
await_pred("endpoint exit", Duration::from_secs(3), || {
conn_count().is_err()
});
assert!(t0.elapsed() < Duration::from_secs(3));
});
}
#[test]
fn drain_stops_idle_now_busy_at_deadline_then_exits() {
smarm::run(|| {
let drain = Duration::from_millis(300);
let sup = spawn_endpoint(drain);
spawn_fake_conn(false); // idle
spawn_fake_conn(true); // busy
await_pred("both registered", Duration::from_secs(2), || {
conn_count() == Ok(2)
});
let t0 = Instant::now();
smarm::request_shutdown(sup);
// Idle conn goes promptly, well before the deadline.
await_pred("idle stopped", drain / 2, || conn_count() == Ok(1));
// Busy conn holds until the force sweep, then the endpoint winds up.
await_pred("busy swept + endpoint exit", drain * 6, || {
conn_count().is_err()
});
assert!(
t0.elapsed() >= drain,
"busy conn must not be stopped before the drain deadline"
);
});
}
#[test]
fn conn_finishing_mid_drain_is_stopped_on_idle() {
smarm::run(|| {
// Long deadline: this passes only via the ConnIdle path, not
// the force sweep.
let sup = spawn_endpoint(Duration::from_secs(30));
let conn = spawn_fake_conn(true);
await_pred("registered busy", Duration::from_secs(2), || {
conn_count() == Ok(1)
});
smarm::request_shutdown(sup);
smarm::sleep(Duration::from_millis(50));
assert_eq!(conn_count(), Ok(1), "busy conn survives the idle sweep");
// "Request finishes": the conn reports idle.
let ep = whereis(NAME).expect("endpoint still draining");
let _ = ep.cast(Cast::ConnIdle(conn));
await_pred("stopped on idle + endpoint exit", Duration::from_secs(3), || {
conn_count().is_err()
});
});
}
#[test]
fn conn_registering_mid_drain_is_stopped_on_arrival() {
smarm::run(|| {
let sup = spawn_endpoint(Duration::from_secs(30));
let holder = spawn_fake_conn(true); // keeps the drain open
await_pred("holder registered", Duration::from_secs(2), || {
conn_count() == Ok(1)
});
smarm::request_shutdown(sup);
smarm::sleep(Duration::from_millis(20));
// The listener-death race: a fresh conn registers mid-drain.
spawn_fake_conn(false);
// It is stopped on arrival — the count returns to just the holder.
await_pred("late arrival stopped", Duration::from_secs(2), || {
conn_count() == Ok(1)
});
// Cleanup: release the holder so the run can end.
let ep = whereis(NAME).expect("endpoint still draining");
let _ = ep.cast(Cast::ConnIdle(holder));
});
}
}
+18 -2
View File
@@ -24,10 +24,26 @@ pub mod router;
pub mod parser; pub mod parser;
pub mod net; pub mod net;
pub mod conn_actor; pub mod conn_actor;
pub mod endpoint;
pub mod serve; pub mod serve;
pub mod sse;
pub mod ws;
pub mod pubsub;
#[cfg(feature = "channels")]
pub mod channels;
// Re-exports — what most users want at the crate root. // Re-exports — what most users want at the crate root.
pub use conn::{Assigns, Body, Conn, HeaderMap, HttpVersion, Method, Params, RespBody}; pub use conn::{Assigns, Body, Conn, HeaderMap, HttpVersion, Method, Params, RespBody, StreamBody};
pub use plug::{Next, Pipeline, Plug}; pub use plug::{Next, Pipeline, Plug};
pub use router::Router; pub use router::Router;
pub use serve::{serve, serve_with, Config}; pub use sse::{EventSender, SseClosed};
pub use pubsub::{PubSub, PubSubDown};
#[cfg(feature = "channels")]
pub use channels::{Channel, ChannelHub, ChannelSession, ChannelSocket, PrefixRouter, Status, TopicRouter};
pub use ws::{Message, WsClosed, WsHandler, WsSender};
pub use endpoint::endpoint;
pub use serve::{
serve, serve_with, serve_with_shutdown, shutdown_handle, Config, Handle, ShutdownSignal,
};
#[cfg(feature = "config-file")]
pub use serve::ConfigError;
+7
View File
@@ -52,6 +52,13 @@ impl Drop for OwnedFd {
// becomes the unique closer. // becomes the unique closer.
unsafe impl Send for OwnedFd {} unsafe impl Send for OwnedFd {}
// SAFETY: shared references only expose `as_raw(&self)` — reading an
// integer. No interior mutability; `into_raw` takes `self` by value and is
// unreachable through a shared reference. Needed so a supervisor
// `ChildSpec` factory (an `Fn` closure, hence shared across restarts) can
// own its listener fd via `Arc<OwnedFd>`.
unsafe impl Sync for OwnedFd {}
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// bind_and_listen // bind_and_listen
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
+357 -28
View File
@@ -4,14 +4,15 @@
//! copy, battle-tested). Body framing, keep-alive logic and response writing //! copy, battle-tested). Body framing, keep-alive logic and response writing
//! are ours. //! are ours.
//! //!
//! v1 body framing: //! Body framing:
//! - `Content-Length: N` — read exactly N bytes. //! - `Content-Length: N` — read exactly N bytes.
//! - No body header — empty body. //! - No body header — empty body.
//! - `Transfer-Encoding: chunked` — deferred. Returns ParseError::Unsupported //! - `Transfer-Encoding: chunked` (HTTP/1.1) — flagged in `ParsedHead`;
//! and the connection actor responds 411 Length Required + close. //! the connection actor decodes incrementally (`read_chunked_body`).
//! //! TE is 1.1-only and overrides Content-Length: TE on HTTP/1.0, or TE
//! Keep it stupid simple. Chunked decoding lands when something actually //! together with a Content-Length, is Malformed (400). `chunked` must be
//! requests it. //! the final coding (non-final -> 400); any other coding is unimplemented
//! (-> 501). Only a sole final `chunked` sets the flag (RFC 9112 §6.1/§6.3).
use crate::conn::{Body, Conn, HeaderMap, HttpVersion, Method, RespBody}; use crate::conn::{Body, Conn, HeaderMap, HttpVersion, Method, RespBody};
@@ -30,9 +31,14 @@ pub enum ParseError {
TooManyHeaders, TooManyHeaders,
/// `Content-Length` header could not be parsed as an integer. /// `Content-Length` header could not be parsed as an integer.
BadContentLength, BadContentLength,
/// A wire feature we haven't implemented yet (e.g. chunked encoding). /// A wire feature we haven't implemented. Currently unconstructed
/// Connection actor responds 411 + close. /// (chunked decoding landed in v0.3); kept for future unsupported
/// framings. Connection actor responds 411 + close.
Unsupported, Unsupported,
/// `Transfer-Encoding` names a transfer coding we don't implement
/// (`chunked` is the only one urus decodes). Connection actor responds
/// 501 Not Implemented + close (RFC 9112 §6.1, §7).
UnknownTransferCoding,
} }
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -50,6 +56,10 @@ pub struct ParsedHead {
pub version: HttpVersion, pub version: HttpVersion,
pub headers: HeaderMap, pub headers: HeaderMap,
pub content_length: Option<usize>, pub content_length: Option<usize>,
/// `Transfer-Encoding: chunked` — the body is chunked-framed and the
/// connection actor decodes it (`read_chunked_body`). Mutually
/// exclusive with `content_length` (rejected as Malformed).
pub chunked: bool,
pub keep_alive: bool, pub keep_alive: bool,
pub expect_100: bool, pub expect_100: bool,
} }
@@ -96,6 +106,11 @@ pub fn parse_head(buf: &[u8], max_headers: usize) -> Result<ParsedHead, ParseErr
let mut connection_hdr = None; let mut connection_hdr = None;
let mut chunked = false; let mut chunked = false;
let mut expect_100 = false; let mut expect_100 = false;
let mut host_count = 0usize;
let mut host_ok = true;
let mut cl_count = 0usize;
let mut te_present = false;
let mut te_codings: Vec<String> = Vec::new();
for h in req.headers.iter() { for h in req.headers.iter() {
let name_lower = h.name.to_ascii_lowercase(); let name_lower = h.name.to_ascii_lowercase();
@@ -103,6 +118,10 @@ pub fn parse_head(buf: &[u8], max_headers: usize) -> Result<ParsedHead, ParseErr
match name_lower.as_str() { match name_lower.as_str() {
"content-length" => { "content-length" => {
// Count occurrences; duplicates (even equal) are rejected
// post-loop. A single value must be one decimal integer —
// a comma-list ("5, 5") or non-numeric fails parse here.
cl_count += 1;
content_length = Some( content_length = Some(
value.trim() value.trim()
.parse::<usize>() .parse::<usize>()
@@ -110,18 +129,30 @@ pub fn parse_head(buf: &[u8], max_headers: usize) -> Result<ParsedHead, ParseErr
); );
} }
"transfer-encoding" => { "transfer-encoding" => {
// We only care whether it includes "chunked". Multiple codings // Collect the ordered coding list across any number of TE
// can appear; chunked is the only one we'd need to decode. // headers; finality/known-ness is decided post-loop. Empty
if value.to_ascii_lowercase().split(',').any(|t| t.trim() == "chunked") { // list elements (legacy `#rule`, e.g. a trailing comma) are
chunked = true; // skipped; a wholly empty value leaves te_codings empty and
// is caught below.
te_present = true;
for coding in value.split(',') {
let c = coding.trim().to_ascii_lowercase();
if !c.is_empty() {
te_codings.push(c);
}
} }
} }
"connection" => { "connection" => {
connection_hdr = Some(value.to_ascii_lowercase()); connection_hdr = Some(value.to_ascii_lowercase());
} }
"expect" => { "expect" if value.eq_ignore_ascii_case("100-continue") => {
if value.eq_ignore_ascii_case("100-continue") { expect_100 = true;
expect_100 = true; }
"host" => {
// Presence/uniqueness enforced post-loop; validity here.
host_count += 1;
if !valid_host(value) {
host_ok = false;
} }
} }
_ => {} _ => {}
@@ -129,8 +160,56 @@ pub fn parse_head(buf: &[u8], max_headers: usize) -> Result<ParsedHead, ParseErr
headers.append(&name_lower, value.to_string()); headers.append(&name_lower, value.to_string());
} }
if chunked { // Host (RFC 9112 §3.2): an HTTP/1.1 request MUST carry exactly one valid
return Err(ParseError::Unsupported); // Host; a missing, duplicate, or malformed Host is a 400. HTTP/1.0 may
// omit Host, but a duplicate or invalid one is still rejected on any
// version (ambiguous / malformed authority).
if host_count > 1 || !host_ok {
return Err(ParseError::Malformed);
}
if version == HttpVersion::Http11 && host_count == 0 {
return Err(ParseError::Malformed);
}
// Content-Length (RFC 9112 §6.3): more than one Content-Length is an
// unrecoverable framing ambiguity (CL.CL request smuggling). We are
// strict — reject any duplicate, not only differing values.
if cl_count > 1 {
return Err(ParseError::BadContentLength);
}
// Transfer-Encoding (RFC 9112 §6.1/§6.3). TE is a 1.1 mechanism and
// overrides Content-Length; only `chunked` is implemented here.
if te_present {
// TE on HTTP/1.0 is malformed (no 1.0 chunked).
if version == HttpVersion::Http10 {
return Err(ParseError::Malformed);
}
// TE together with Content-Length is the classic smuggling
// ambiguity; TE overrides CL and we reject rather than forward.
if content_length.is_some() {
return Err(ParseError::Malformed);
}
// A Transfer-Encoding header that carries no coding frames nothing.
if te_codings.is_empty() {
return Err(ParseError::Malformed);
}
let last_is_chunked = te_codings.last().map(String::as_str) == Some("chunked");
let has_chunked = te_codings.iter().any(|c| c == "chunked");
if has_chunked && !last_is_chunked {
// chunked present but not final: body length isn't reliably
// determinable -> 400.
return Err(ParseError::Malformed);
}
if te_codings.iter().any(|c| c != "chunked") {
// Some coding we don't implement (chunked is the only decodable
// one). Whether or not chunked is final, we can't apply it -> 501.
return Err(ParseError::UnknownTransferCoding);
}
// Sole, final `chunked`: the connection actor decodes the body.
chunked = true;
} }
// Keep-alive logic, RFC 7230 §6.3: // Keep-alive logic, RFC 7230 §6.3:
@@ -149,11 +228,38 @@ pub fn parse_head(buf: &[u8], max_headers: usize) -> Result<ParsedHead, ParseErr
version, version,
headers, headers,
content_length, content_length,
chunked,
keep_alive, keep_alive,
expect_100, expect_100,
}) })
} }
/// Conservative RFC 3986 check for a `Host` field-value: non-empty and every
/// byte drawn from the `host[:port]` productions (reg-name / IP-literal
/// brackets / port colon). This is charset-level, not full structural
/// validation (no bracket matching, no pct-encoding well-formedness) — enough
/// to reject the smuggling-relevant garbage (whitespace, controls, `@`, `/`,
/// `?`, `#`) while accepting every legitimate host. Tighter structural checks
/// (bracketed IPv6, single port colon) are a possible follow-up.
fn valid_host(value: &str) -> bool {
!value.is_empty()
&& value.bytes().all(|b| {
b.is_ascii_alphanumeric()
|| matches!(
b,
// unreserved punctuation
b'-' | b'.' | b'_' | b'~'
// sub-delims
| b'!' | b'$' | b'&' | b'\'' | b'(' | b')'
| b'*' | b'+' | b',' | b';' | b'='
// pct-encoded lead
| b'%'
// IP-literal brackets + port separator
| b'[' | b']' | b':'
)
})
}
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// Conn assembly // Conn assembly
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -187,8 +293,9 @@ pub fn build_conn(head: ParsedHead, body: Body) -> Conn {
pub fn serialise_response(conn: &Conn, keep_alive: bool) -> Vec<u8> { pub fn serialise_response(conn: &Conn, keep_alive: bool) -> Vec<u8> {
// Pre-size: status line ~30 + headers ~50/each + body. Good enough. // Pre-size: status line ~30 + headers ~50/each + body. Good enough.
let body_len = conn.resp_body.len_hint(); let body_len = conn.resp_body.len_hint();
let mut out = Vec::with_capacity(64 + conn.resp_headers.len() * 40 + body_len); let is_stream = matches!(conn.resp_body, RespBody::Stream(_));
let mut out = Vec::with_capacity(64 + conn.resp_headers.len() * 40 + body_len);
let status = conn.status.unwrap_or(200); let status = conn.status.unwrap_or(200);
let reason = reason_phrase(status); let reason = reason_phrase(status);
@@ -202,11 +309,17 @@ pub fn serialise_response(conn: &Conn, keep_alive: bool) -> Vec<u8> {
out.extend_from_slice(b"\r\n"); out.extend_from_slice(b"\r\n");
// User headers — written first so subsequent injection can skip them. // User headers — written first so subsequent injection can skip them.
// For Stream bodies WE own the framing: a user `content-length` or
// `transfer-encoding` is dropped rather than emitted (the combination
// of content-length + chunked is a smuggling vector, and a stream has
// no length to promise anyway).
let mut wrote_content_length = false; let mut wrote_content_length = false;
let mut wrote_connection = false; let mut wrote_connection = false;
for (name, value) in conn.resp_headers.iter() { for (name, value) in conn.resp_headers.iter() {
match name { match name {
"content-length" if is_stream => continue,
"transfer-encoding" if is_stream => continue,
"content-length" => wrote_content_length = true, "content-length" => wrote_content_length = true,
"connection" => wrote_connection = true, "connection" => wrote_connection = true,
_ => {} _ => {}
@@ -217,22 +330,44 @@ pub fn serialise_response(conn: &Conn, keep_alive: bool) -> Vec<u8> {
out.extend_from_slice(b"\r\n"); out.extend_from_slice(b"\r\n");
} }
if !wrote_content_length { if is_stream {
// HTTP/1.1: chunked framing, connection reusable afterwards.
// HTTP/1.0: no chunked TE exists; the body is raw bytes delimited
// by EOF — the caller passes keep_alive = false and we emit
// `connection: close` below.
if conn.version == HttpVersion::Http11 {
out.extend_from_slice(b"transfer-encoding: chunked\r\n");
}
} else if !wrote_content_length && !(100..200).contains(&status) {
// 1xx responses have no body by definition (RFC 7230 §3.3.2) —
// injecting `content-length: 0` on the 101 upgrade response is a
// protocol violation some clients reject.
out.extend_from_slice(b"content-length: "); out.extend_from_slice(b"content-length: ");
out.extend_from_slice(body_len.to_string().as_bytes()); out.extend_from_slice(body_len.to_string().as_bytes());
out.extend_from_slice(b"\r\n"); out.extend_from_slice(b"\r\n");
} }
if !wrote_connection && !keep_alive { if !wrote_connection {
out.extend_from_slice(b"connection: close\r\n"); if !keep_alive {
out.extend_from_slice(b"connection: close\r\n");
} else if conn.version == HttpVersion::Http10 {
// HTTP/1.0 defaults to close: a connection we intend to keep
// open MUST be advertised back, or a spec-following client
// waits for an EOF that never comes (`ab -k` deadlocked on
// exactly this). 1.1 keep-alive is the default and stays
// implicit.
out.extend_from_slice(b"connection: keep-alive\r\n");
}
} }
out.extend_from_slice(b"\r\n"); out.extend_from_slice(b"\r\n");
// Body. // Fixed bodies are written inline with the head; a Stream body is
// pumped by the connection actor after this head goes on the wire.
match &conn.resp_body { match &conn.resp_body {
RespBody::Empty => {} RespBody::Empty => {}
RespBody::Bytes(b) => out.extend_from_slice(b), RespBody::Bytes(b) => out.extend_from_slice(b),
RespBody::Stream(_) => {}
} }
out out
@@ -340,11 +475,150 @@ mod tests {
} }
#[test] #[test]
fn parse_chunked_unsupported() { fn parse_chunked_is_flagged() {
let req = b"POST /a HTTP/1.1\r\nHost: x\r\nTransfer-Encoding: chunked\r\n\r\n"; let req = b"POST /a HTTP/1.1\r\nHost: x\r\nTransfer-Encoding: chunked\r\n\r\n";
let head = parse_head(req, 64).unwrap();
assert!(head.chunked);
assert_eq!(head.content_length, None);
}
#[test]
fn parse_chunked_plus_content_length_is_malformed() {
let req = b"POST /a HTTP/1.1\r\nHost: x\r\nContent-Length: 5\r\nTransfer-Encoding: chunked\r\n\r\n";
match parse_head(req, 64) { match parse_head(req, 64) {
Err(ParseError::Unsupported) => {} Err(ParseError::Malformed) => {}
_ => panic!("expected Unsupported for chunked"), _ => panic!("expected Malformed for CL + chunked"),
}
}
#[test]
fn parse_chunked_on_http10_is_malformed() {
let req = b"POST /a HTTP/1.0\r\nHost: x\r\nTransfer-Encoding: chunked\r\n\r\n";
match parse_head(req, 64) {
Err(ParseError::Malformed) => {}
_ => panic!("expected Malformed for chunked on 1.0"),
}
}
// --- Host (RFC 9112 §3.2) -------------------------------------------
#[test]
fn parse_missing_host_http11_is_malformed() {
let req = b"GET / HTTP/1.1\r\n\r\n";
match parse_head(req, 64) {
Err(ParseError::Malformed) => {}
_ => panic!("expected Malformed for missing Host on 1.1"),
}
}
#[test]
fn parse_missing_host_http10_is_allowed() {
// Host is optional in HTTP/1.0.
let req = b"GET / HTTP/1.0\r\n\r\n";
assert!(parse_head(req, 64).is_ok(), "1.0 may omit Host");
}
#[test]
fn parse_duplicate_host_is_malformed() {
let req = b"GET / HTTP/1.1\r\nHost: a\r\nHost: b\r\n\r\n";
match parse_head(req, 64) {
Err(ParseError::Malformed) => {}
_ => panic!("expected Malformed for duplicate Host"),
}
}
#[test]
fn parse_invalid_host_value_is_malformed() {
// Embedded whitespace — invalid in an RFC 3986 authority.
let req = b"GET / HTTP/1.1\r\nHost: bad host\r\n\r\n";
match parse_head(req, 64) {
Err(ParseError::Malformed) => {}
_ => panic!("expected Malformed for invalid Host"),
}
}
#[test]
fn parse_valid_hosts_accepted() {
// Positive controls: reg-name, reg-name:port, and IPv6-literal:port.
for req in [
b"GET / HTTP/1.1\r\nHost: example.com\r\n\r\n".as_slice(),
b"GET / HTTP/1.1\r\nHost: example.com:8080\r\n\r\n".as_slice(),
b"GET / HTTP/1.1\r\nHost: [::1]:443\r\n\r\n".as_slice(),
] {
assert!(parse_head(req, 64).is_ok(), "should accept a valid Host");
}
}
// --- Content-Length (RFC 9112 §6.3) ---------------------------------
#[test]
fn parse_conflicting_content_length_is_rejected() {
// Two differing Content-Length values — classic CL.CL smuggling.
let req = b"POST / HTTP/1.1\r\nHost: x\r\nContent-Length: 5\r\nContent-Length: 7\r\n\r\nhello!!";
match parse_head(req, 64) {
Err(ParseError::BadContentLength) => {}
_ => panic!("expected BadContentLength for conflicting CL"),
}
}
#[test]
fn parse_duplicate_equal_content_length_is_rejected() {
// Strict: even identical duplicates are rejected.
let req = b"POST / HTTP/1.1\r\nHost: x\r\nContent-Length: 5\r\nContent-Length: 5\r\n\r\nhello";
match parse_head(req, 64) {
Err(ParseError::BadContentLength) => {}
_ => panic!("expected BadContentLength for duplicate CL"),
}
}
#[test]
fn parse_single_content_length_still_ok() {
// Regression: the ordinary single-CL path is unchanged.
let req = b"POST / HTTP/1.1\r\nHost: x\r\nContent-Length: 5\r\n\r\nhello";
let head = parse_head(req, 64).unwrap();
assert_eq!(head.content_length, Some(5));
}
// --- Transfer-Encoding (RFC 9112 §6.1/§6.3) -------------------------
#[test]
fn parse_non_final_chunked_is_malformed() {
// chunked must be the FINAL coding.
let req = b"POST / HTTP/1.1\r\nHost: x\r\nTransfer-Encoding: chunked, gzip\r\n\r\n";
match parse_head(req, 64) {
Err(ParseError::Malformed) => {}
_ => panic!("expected Malformed for non-final chunked"),
}
}
#[test]
fn parse_unknown_transfer_coding_is_unimplemented() {
// A coding urus doesn't implement, no chunked at all -> 501.
let req = b"POST / HTTP/1.1\r\nHost: x\r\nTransfer-Encoding: nonsense\r\n\r\n";
match parse_head(req, 64) {
Err(ParseError::UnknownTransferCoding) => {}
_ => panic!("expected UnknownTransferCoding for unknown coding"),
}
}
#[test]
fn parse_gzip_then_chunked_is_unimplemented() {
// chunked IS final, but gzip is still a coding we can't apply -> 501.
let req = b"POST / HTTP/1.1\r\nHost: x\r\nTransfer-Encoding: gzip, chunked\r\n\r\n";
match parse_head(req, 64) {
Err(ParseError::UnknownTransferCoding) => {}
_ => panic!("expected UnknownTransferCoding for gzip,chunked"),
}
}
#[test]
fn parse_te_with_content_length_is_malformed() {
// ANY Transfer-Encoding + Content-Length -> reject (smuggling),
// not only chunked+CL. This closes the old TE:unknown + CL gap.
let req = b"POST / HTTP/1.1\r\nHost: x\r\nTransfer-Encoding: bogus\r\nContent-Length: 5\r\n\r\nhello";
match parse_head(req, 64) {
Err(ParseError::Malformed) => {}
_ => panic!("expected Malformed for TE + CL"),
} }
} }
@@ -366,6 +640,61 @@ mod tests {
assert!(s.contains("connection: close")); assert!(s.contains("connection: close"));
} }
#[test]
fn serialise_stream_http11_is_chunked_no_content_length() {
let (_tx, rx) = smarm::channel::<Vec<u8>>();
let conn = Conn::new().put_status(200).put_body(RespBody::from(rx));
let bytes = serialise_response(&conn, true);
let s = std::str::from_utf8(&bytes).unwrap();
assert!(s.contains("transfer-encoding: chunked"));
assert!(!s.contains("content-length"));
assert!(s.ends_with("\r\n\r\n")); // head only, no body bytes
}
#[test]
fn serialise_stream_strips_user_framing_headers() {
let (_tx, rx) = smarm::channel::<Vec<u8>>();
let conn = Conn::new().put_status(200)
.put_header("content-length", "999")
.put_header("transfer-encoding", "gzip")
.put_body(RespBody::from(rx));
let bytes = serialise_response(&conn, true);
let s = std::str::from_utf8(&bytes).unwrap();
assert!(!s.contains("content-length"));
assert!(!s.contains("gzip"));
assert!(s.contains("transfer-encoding: chunked"));
}
#[test]
fn serialise_stream_http10_no_te_and_closes() {
let (_tx, rx) = smarm::channel::<Vec<u8>>();
let mut conn = Conn::new().put_status(200).put_body(RespBody::from(rx));
conn.version = HttpVersion::Http10;
// The conn actor forces keep_alive=false for a 1.0 stream.
let bytes = serialise_response(&conn, false);
let s = std::str::from_utf8(&bytes).unwrap();
assert!(!s.contains("transfer-encoding"));
assert!(!s.contains("content-length"));
assert!(s.contains("connection: close"));
}
#[test]
fn serialise_http10_keepalive_is_echoed() {
let mut conn = Conn::new().put_status(200).put_body("hi");
conn.version = HttpVersion::Http10;
let bytes = serialise_response(&conn, true);
let s = std::str::from_utf8(&bytes).unwrap();
assert!(s.contains("connection: keep-alive"), "head: {s}");
}
#[test]
fn serialise_http11_keepalive_stays_implicit() {
let conn = Conn::new().put_status(200).put_body("hi");
let bytes = serialise_response(&conn, true);
let s = std::str::from_utf8(&bytes).unwrap();
assert!(!s.contains("connection:"), "head: {s}");
}
#[test] #[test]
fn serialise_user_content_length_is_respected() { fn serialise_user_content_length_is_respected() {
let conn = Conn::new().put_status(200) let conn = Conn::new().put_status(200)
+581
View File
@@ -0,0 +1,581 @@
//! `urus::pubsub` — local-node topic pub/sub, phoenix_pubsub-shaped.
//!
//! Deliberately independent of HTTP: nothing here imports the rest of
//! urus, only smarm. Usable from any smarm app; candidate for crate
//! extraction later.
//!
//! # Shape (v0.5 design, user-ratified)
//!
//! - **Generic per instance** (`PubSub<M>`): zero-cost dispatch, no
//! downcasts. One instance per message domain; heterogeneous events on
//! one bus are an app-side `enum M`.
//! - **Subscribe returns a `Receiver<Arc<M>>`**: the subscriber owns its
//! receive loop and composes with the v0.4 ws outbound pattern (a relay
//! actor pumping the receiver into a `WsSender` clone). One fresh
//! receiver per subscription; fan-in (`subscribe_with(topic, sender)`)
//! is a compatible later addition if wanted.
//! - **Unbounded mailboxes**: `broadcast` never blocks the topic table —
//! a bounded-block policy would head-of-line-block *every* topic behind
//! one slow subscriber, and bounded-drop silently loses messages.
//! Memory risk is a live-but-not-receiving subscriber, bounded by two
//! cleanup paths: monitor `Down` on subscriber death, and pruning on
//! send failure when a receiver was dropped. Bench before sharding.
//! - **User-owned handle**: construct a `PubSub<M>`, clone it around
//! (handler closures already thread state this way). No global named
//! instance — see the module note on `register` below.
//! - **One subscription per `(pid, topic)`**: subscribe is idempotent;
//! re-subscribing replaces the sender (the old receiver's channel
//! closes). Kills the silent double-delivery bug class.
//!
//! # Cleanup mechanics
//!
//! The table monitors each subscriber pid **once** (first subscribe) via
//! the gen_server [`Watcher`]; the `Down` removes the pid from every
//! topic — a full table scan, deliberately: deaths are rare and topic
//! counts small at this stage, a pid→topics reverse index is the
//! documented optimisation if that ever measures hot (same spirit as the
//! roadmap's "don't pre-shard"). A pid that unsubscribes from everything
//! but stays alive keeps its (one-shot, inert) monitor until death;
//! that's a bounded bookkeeping entry, not a leak.
//!
//! # The handle is an address, not an owner (v0.7)
//!
//! [`PubSub::new`] takes a **name** and spawns nothing. It is `const`,
//! costs a `&'static str`, and may be built anywhere — out of the
//! runtime, in a `static`, in a route closure, held by a long-lived
//! relay. Every operation resolves the name through smarm's registry, so
//! a table restarted by its supervisor is reached transparently.
//!
//! The table actor is started by [`PubSub::child`], a `ChildSpec` you put
//! in your supervision tree (or hand to `serve_with*`, which puts it in
//! the root ahead of the endpoint). Its lifetime is the supervisor's:
//! it stops when the supervisor shuts it down, in reverse start order,
//! after the endpoint has drained.
//!
//! ```ignore
//! const BUS: PubSub<Event> = PubSub::new("events");
//! serve_with(cfg, rt_cfg, pipeline, vec![BUS.child()])?;
//! ```
//!
//! This replaces the v0.5–v0.6 idiom where `PubSub::new()` spawned the
//! table and the handle owned its life — a non-static
//! `Arc<OnceLock<PubSub<M>>>` lazily initialised from the first handler,
//! with matching rules that the cell must not be `static` and that
//! relays must never hold a clone. All of that existed to hand-manage a
//! refcount. smarm 0.7 made a server's lifetime its own (refs are
//! addresses; the inbox no longer closes when the last ref drops), which
//! both removed the mechanism those rules relied on and made the named
//! lookup below possible.
//!
//! Cost: one registry resolution per operation, including per broadcast.
//! Caching a `GenServerRef` in the handle would save it and go stale
//! across exactly the restart the supervisor exists to perform. Measure
//! before optimising — see the bench item in `ROADMAP.md`.
use std::collections::{HashMap, HashSet};
use std::fmt;
use std::sync::Arc;
use smarm::gen_server::{self, GenServer, GenServerBuilder, GenServerCtx, GenServerName};
use smarm::{channel, ChildSpec, Down, GenServerRef, Pid, Receiver, Restart, Sender, Watcher};
// ---------------------------------------------------------------------------
// Public handle
// ---------------------------------------------------------------------------
/// The topic table is gone — its gen_server actor exited (panic or runtime
/// shutdown). Every operation on the handle fails with this from then on.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub struct PubSubDown;
impl fmt::Display for PubSubDown {
fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result {
write!(f, "pubsub topic table is down")
}
}
impl std::error::Error for PubSubDown {}
/// The address of one pub/sub instance: a name, resolved per operation.
///
/// Cheap (`Copy`, a `&'static str`), constructible anywhere including
/// `const` context, and owns nothing — the table actor it addresses is
/// started by [`child`](Self::child) under a supervisor. Every operation
/// returns [`PubSubDown`] if no live table currently holds the name.
///
/// Payloads are broadcast as `Arc<M>`: one allocation per broadcast, not
/// per subscriber.
pub struct PubSub<M: Send + Sync + 'static> {
name: GenServerName<Table<M>>,
}
impl<M: Send + Sync + 'static> Clone for PubSub<M> {
fn clone(&self) -> Self {
*self
}
}
impl<M: Send + Sync + 'static> Copy for PubSub<M> {}
impl<M: Send + Sync + 'static> PubSub<M> {
/// Address the topic table registered under `name`. Spawns nothing
/// and never fails: the name is resolved at each use.
pub const fn new(name: &'static str) -> Self {
PubSub { name: GenServerName::new(name) }
}
/// The `ChildSpec` that runs this instance's table actor. Put it in
/// your supervision tree ahead of anything that broadcasts —
/// `serve_with*` takes a `Vec<ChildSpec>` for exactly this.
///
/// `Permanent`: a table that dies is a bug, and its subscribers'
/// receivers died with it, so the restart is only half a repair —
/// pair it with a `RestForOne` parent (as `serve_with*` does) so the
/// endpoint restarts behind it and connections re-subscribe.
pub fn child(&self) -> ChildSpec {
let name = self.name;
ChildSpec::new(Restart::Permanent, move || {
if let Err(e) = GenServerBuilder::new(Table::<M>::new()).named(name).run() {
panic!("urus pubsub '{}': name already taken: {e:?}", name.as_str());
}
})
}
/// The registry key this handle resolves.
pub const fn name(&self) -> &'static str {
self.name.as_str()
}
fn server(&self) -> Result<GenServerRef<Table<M>>, PubSubDown> {
gen_server::whereis_server(self.name).ok_or(PubSubDown)
}
/// Subscribe the **calling actor** to `topic`. Returns the receiving
/// end; messages arrive as `Arc<M>`. Idempotent per `(pid, topic)`:
/// subscribing again replaces the sender, closing the previously
/// returned receiver.
///
/// The subscription is cleaned up when the calling actor dies. If the
/// receive loop runs in a *different* actor than the session that
/// should scope the subscription, use [`subscribe_as`](Self::subscribe_as)
/// to pin cleanup to the right pid (e.g. a ws connection actor
/// subscribing, with a spawned relay doing the receiving).
pub fn subscribe(&self, topic: impl Into<String>) -> Result<Receiver<Arc<M>>, PubSubDown> {
self.subscribe_as(smarm::self_pid(), topic)
}
/// [`subscribe`](Self::subscribe), but the subscription's lifetime is
/// tied to `pid` instead of the calling actor.
pub fn subscribe_as(
&self,
pid: Pid,
topic: impl Into<String>,
) -> Result<Receiver<Arc<M>>, PubSubDown> {
let (tx, rx) = channel();
// A call, not a cast: when this returns the table is updated, so
// a broadcast issued right after by the same caller is seen.
match self.server()?.call(Call::Subscribe { topic: topic.into(), pid, tx }) {
Ok(_) => Ok(rx),
Err(_) => Err(PubSubDown),
}
}
/// Drop the calling actor's subscription to `topic` (its receiver's
/// channel closes). No-op if not subscribed.
pub fn unsubscribe(&self, topic: impl Into<String>) -> Result<(), PubSubDown> {
self.unsubscribe_as(smarm::self_pid(), topic)
}
/// [`unsubscribe`](Self::unsubscribe) for an explicit pid.
pub fn unsubscribe_as(&self, pid: Pid, topic: impl Into<String>) -> Result<(), PubSubDown> {
self.server()?
.cast(Cast::Unsubscribe { topic: topic.into(), pid })
.map_err(|_| PubSubDown)
}
/// Broadcast `msg` to every subscriber of `topic`. Never blocks on
/// subscribers (unbounded mailboxes); returns once the table has the
/// request queued.
pub fn broadcast(&self, topic: impl Into<String>, msg: M) -> Result<(), PubSubDown> {
self.cast_broadcast(topic.into(), msg, None)
}
/// [`broadcast`](Self::broadcast), skipping delivery to `from` —
/// pass `smarm::self_pid()` for the phoenix `broadcast_from` shape
/// ("everyone in the room but me").
pub fn broadcast_from(
&self,
from: Pid,
topic: impl Into<String>,
msg: M,
) -> Result<(), PubSubDown> {
self.cast_broadcast(topic.into(), msg, Some(from))
}
/// Number of live subscriptions on `topic`. Counts entries in the
/// table — a subscriber whose receiver was dropped but hasn't been
/// pruned yet (no broadcast since, still alive) is still counted.
pub fn subscriber_count(&self, topic: impl Into<String>) -> Result<usize, PubSubDown> {
match self.server()?.call(Call::Count { topic: topic.into() }) {
Ok(Reply::Count(n)) => Ok(n),
Ok(Reply::Subscribed) => unreachable!("Count call answered with Subscribed"),
Err(_) => Err(PubSubDown),
}
}
/// The topic table actor's pid, or `None` if no live table holds the
/// name (not started yet, or between a crash and its restart).
pub fn pid(&self) -> Option<Pid> {
self.server().ok().map(|s| s.pid())
}
fn cast_broadcast(&self, topic: String, msg: M, skip: Option<Pid>) -> Result<(), PubSubDown> {
self.server()?
.cast(Cast::Broadcast { topic, msg: Arc::new(msg), skip })
.map_err(|_| PubSubDown)
}
}
// ---------------------------------------------------------------------------
// The table gen_server
// ---------------------------------------------------------------------------
enum Call<M> {
Subscribe { topic: String, pid: Pid, tx: Sender<Arc<M>> },
Count { topic: String },
}
enum Reply {
Subscribed,
Count(usize),
}
enum Cast<M> {
Unsubscribe { topic: String, pid: Pid },
Broadcast { topic: String, msg: Arc<M>, skip: Option<Pid> },
}
struct Table<M: Send + Sync + 'static> {
topics: HashMap<String, HashMap<Pid, Sender<Arc<M>>>>,
/// Pids we already hold a monitor for. Monitors are one-shot and
/// created at most once per live pid (first subscribe); `handle_down`
/// retires the entry so a reused-slot pid (fresh generation) gets a
/// fresh monitor.
monitored: HashSet<Pid>,
watcher: Option<Watcher<Table<M>>>,
}
impl<M: Send + Sync + 'static> Table<M> {
fn new() -> Self {
Table { topics: HashMap::new(), monitored: HashSet::new(), watcher: None }
}
}
impl<M: Send + Sync + 'static> GenServer for Table<M> {
type Call = Call<M>;
type Reply = Reply;
type Cast = Cast<M>;
type Info = ();
type Timer = ();
fn init(&mut self, ctx: &GenServerCtx<Self>) {
self.watcher = Some(ctx.watcher());
}
fn handle_call(&mut self, request: Call<M>) -> Reply {
match request {
Call::Subscribe { topic, pid, tx } => {
if self.monitored.insert(pid) {
let m = smarm::monitor(pid);
self.watcher
.as_ref()
.expect("watcher set in init")
.watch(m);
}
// Insert replaces: idempotent per (pid, topic); the old
// sender drops and the stale receiver's channel closes.
self.topics.entry(topic).or_default().insert(pid, tx);
Reply::Subscribed
}
Call::Count { topic } => {
Reply::Count(self.topics.get(&topic).map_or(0, HashMap::len))
}
}
}
fn handle_cast(&mut self, request: Cast<M>) {
match request {
Cast::Unsubscribe { topic, pid } => {
if let Some(subs) = self.topics.get_mut(&topic) {
subs.remove(&pid);
if subs.is_empty() {
self.topics.remove(&topic);
}
}
}
Cast::Broadcast { topic, msg, skip } => {
let Some(subs) = self.topics.get_mut(&topic) else { return };
// Prune-on-send-failure: a dropped receiver (subscriber
// alive but done listening, e.g. ws conn reaped at
// write_timeout) is removed lazily here; subscriber
// *death* is handled eagerly by the monitor.
subs.retain(|pid, tx| {
if skip == Some(*pid) {
return true;
}
tx.send(Arc::clone(&msg)).is_ok()
});
if subs.is_empty() {
self.topics.remove(&topic);
}
}
}
}
fn handle_down(&mut self, down: Down) {
self.monitored.remove(&down.pid);
// Full scan, by design — see the module docs.
self.topics.retain(|_, subs| {
subs.remove(&down.pid);
!subs.is_empty()
});
}
}
// ---------------------------------------------------------------------------
// Tests (runtime-backed: results collected outside, asserted after run)
// ---------------------------------------------------------------------------
#[cfg(test)]
mod tests {
use super::*;
use std::sync::Mutex;
use std::time::Duration;
/// Poll `f` until it returns true or ~2s elapse. For the
/// asynchronous cleanup paths (monitor Down delivery).
fn eventually(mut f: impl FnMut() -> bool) -> bool {
for _ in 0..200 {
if f() {
return true;
}
smarm::sleep(Duration::from_millis(10));
}
false
}
/// Start `ps`'s table under a supervisor and hand back the sup's pid.
///
/// The readiness poll is smarm's "start order is not start readiness"
/// gap (see smarm ROADMAP): `start_child` spawns and moves on, so the
/// name may not be bound when this returns. Real apps don't hit it —
/// a handler only runs once a connection has been accepted, long
/// after the tree is up — but a test that broadcasts immediately does.
fn start_table<M: Send + Sync + 'static>(ps: PubSub<M>) -> Pid {
let sup = smarm::spawn(move || smarm::OneForOne::new().child(ps.child()).run());
assert!(eventually(|| ps.pid().is_some()), "table never bound its name");
sup.pid()
}
/// Ordered shutdown of the tree `start_table` built.
fn stop_table(sup: Pid) {
smarm::request_shutdown(sup);
}
#[test]
fn broadcast_reaches_all_subscribers_once() {
let got = Arc::new(Mutex::new(Vec::<(u32, String)>::new()));
let got2 = got.clone();
smarm::run(move || {
let ps = PubSub::<String>::new("test-bus");
let sup = start_table(ps);
let mut handles = Vec::new();
for i in 0..2u32 {
let got = got2.clone();
handles.push(smarm::spawn(move || {
let rx = ps.subscribe("room:a").unwrap();
let msg = rx.recv().unwrap();
got.lock().unwrap().push((i, (*msg).clone()));
}));
}
// Subscribes are calls from the spawned actors; wait until
// both are in the table before broadcasting.
assert!(eventually(|| ps.subscriber_count("room:a").unwrap() == 2));
ps.broadcast("room:a", "hello".to_string()).unwrap();
for h in handles {
h.join().unwrap();
}
stop_table(sup);
});
let mut v = got.lock().unwrap().clone();
v.sort();
assert_eq!(v, vec![(0, "hello".into()), (1, "hello".into())]);
}
#[test]
fn arc_payload_is_shared_not_cloned() {
let ptrs = Arc::new(Mutex::new(Vec::<usize>::new()));
let ptrs2 = ptrs.clone();
smarm::run(move || {
let ps = PubSub::<Vec<u8>>::new("test-bus");
let sup = start_table(ps);
let mut handles = Vec::new();
for _ in 0..2 {
let ptrs = ptrs2.clone();
handles.push(smarm::spawn(move || {
let rx = ps.subscribe("t").unwrap();
let msg = rx.recv().unwrap();
ptrs.lock().unwrap().push(Arc::as_ptr(&msg) as usize);
}));
}
assert!(eventually(|| ps.subscriber_count("t").unwrap() == 2));
ps.broadcast("t", vec![1, 2, 3]).unwrap();
for h in handles {
h.join().unwrap();
}
stop_table(sup);
});
let v = ptrs.lock().unwrap();
assert_eq!(v.len(), 2);
assert_eq!(v[0], v[1], "both subscribers should see the same allocation");
}
#[test]
fn broadcast_from_skips_the_sender() {
let got = Arc::new(Mutex::new(Vec::<String>::new()));
let got2 = got.clone();
smarm::run(move || {
let ps = PubSub::<String>::new("test-bus");
let sup = start_table(ps);
let ps_loud = ps;
let got = got2.clone();
let loud = smarm::spawn(move || {
let rx = ps_loud.subscribe("room").unwrap();
ps_loud
.broadcast_from(smarm::self_pid(), "room", "from loud".into())
.unwrap();
// Must receive the OTHER broadcast only.
let msg = rx.recv().unwrap();
got.lock().unwrap().push(format!("loud got {}", *msg));
});
let rx = ps.subscribe("room").unwrap();
assert!(eventually(|| ps.subscriber_count("room").unwrap() == 2));
ps.broadcast("room", "for everyone".into()).unwrap();
let first = rx.recv().unwrap();
let second = rx.recv().unwrap();
got2.lock()
.unwrap()
.push(format!("root got {} then {}", *first, *second));
loud.join().unwrap();
stop_table(sup);
});
let v = got.lock().unwrap();
assert!(v.contains(&"loud got for everyone".to_string()), "{v:?}");
// Root (not the broadcast_from sender) receives both, in order.
assert!(
v.contains(&"root got from loud then for everyone".to_string())
|| v.contains(&"root got for everyone then from loud".to_string()),
"{v:?}"
);
}
#[test]
fn unsubscribe_closes_the_receiver_and_stops_delivery() {
let ok = Arc::new(Mutex::new(false));
let ok2 = ok.clone();
smarm::run(move || {
let ps = PubSub::<u32>::new("test-bus");
let sup = start_table(ps);
let rx = ps.subscribe("t").unwrap();
ps.unsubscribe("t").unwrap();
// Unsubscribe is a cast; the table holds the only sender, so
// once processed the channel closes and recv errs.
assert!(eventually(|| ps.subscriber_count("t").unwrap() == 0));
assert!(rx.recv().is_err());
ps.broadcast("t", 7).unwrap(); // no subscribers: no-op, no panic
*ok2.lock().unwrap() = true;
stop_table(sup);
});
assert!(*ok.lock().unwrap());
}
#[test]
fn resubscribe_replaces_the_old_subscription() {
let got = Arc::new(Mutex::new((0u32, false)));
let got2 = got.clone();
smarm::run(move || {
let ps = PubSub::<u32>::new("test-bus");
let sup = start_table(ps);
let rx_old = ps.subscribe("t").unwrap();
let rx_new = ps.subscribe("t").unwrap();
assert_eq!(ps.subscriber_count("t").unwrap(), 1, "idempotent per (pid, topic)");
ps.broadcast("t", 42).unwrap();
let v = *rx_new.recv().unwrap();
let old_closed = rx_old.recv().is_err();
*got2.lock().unwrap() = (v, old_closed);
stop_table(sup);
});
assert_eq!(*got.lock().unwrap(), (42, true));
}
#[test]
fn dead_subscriber_is_pruned_by_the_monitor() {
let ok = Arc::new(Mutex::new(false));
let ok2 = ok.clone();
smarm::run(move || {
let ps = PubSub::<u32>::new("test-bus");
let sup = start_table(ps);
let ps2 = ps;
let h = smarm::spawn(move || {
let _rx = ps2.subscribe("t").unwrap();
// Exit without ever receiving: rx drops with the stack.
});
h.join().unwrap();
// No broadcast issued — this MUST be the monitor path, not
// prune-on-send-failure.
*ok2.lock().unwrap() = eventually(|| ps.subscriber_count("t").unwrap() == 0);
stop_table(sup);
});
assert!(*ok.lock().unwrap(), "monitor Down never pruned the dead subscriber");
}
#[test]
fn dropped_receiver_is_pruned_on_next_broadcast() {
let counts = Arc::new(Mutex::new((0usize, 0usize)));
let counts2 = counts.clone();
smarm::run(move || {
let ps = PubSub::<u32>::new("test-bus");
let sup = start_table(ps);
let rx = ps.subscribe("t").unwrap();
drop(rx);
let before = ps.subscriber_count("t").unwrap();
ps.broadcast("t", 1).unwrap();
let mut after = ps.subscriber_count("t").unwrap();
// Casts are ordered but count is a call that can overtake
// nothing here — same inbox, so after the broadcast cast is
// handled. Still poll defensively.
if after != 0 {
after = if eventually(|| ps.subscriber_count("t").unwrap() == 0) { 0 } else { after };
}
*counts2.lock().unwrap() = (before, after);
stop_table(sup);
});
assert_eq!(*counts.lock().unwrap(), (1, 0));
}
#[test]
fn topics_are_independent() {
let got = Arc::new(Mutex::new(Vec::<u32>::new()));
let got2 = got.clone();
smarm::run(move || {
let ps = PubSub::<u32>::new("test-bus");
let sup = start_table(ps);
let rx_a = ps.subscribe("a").unwrap();
ps.broadcast("b", 99).unwrap(); // nobody on b; must not reach a
ps.broadcast("a", 1).unwrap();
got2.lock().unwrap().push(*rx_a.recv().unwrap());
stop_table(sup);
});
assert_eq!(*got.lock().unwrap(), vec![1]);
}
}
+19 -6
View File
@@ -135,16 +135,29 @@ impl Plug for Router {
// path matches but a same-path-different-method does, return 405. // path matches but a same-path-different-method does, return 405.
// Otherwise pass through to `next` so outer pipelines can layer a // Otherwise pass through to `next` so outer pipelines can layer a
// 404 handler (or skip and let the connection actor emit a default). // 404 handler (or skip and let the connection actor emit a default).
//
// Matching runs under the `router` causal site (RFC 007); the
// winning handler and the `next` fall-through run outside it, so
// the site measures dispatch, not what it dispatches to.
let mut path_seen = false; let mut path_seen = false;
for route in &self.routes { let mut hit = None;
if let Some(params) = route.pattern.match_path(&conn.path) { {
if route.method == conn.method { #[cfg(feature = "smarm-causal")]
let conn = conn.put_params(params); let _g = smarm::causal_site!("router");
return route.handler.call(conn, next); for (i, route) in self.routes.iter().enumerate() {
if let Some(params) = route.pattern.match_path(&conn.path) {
if route.method == conn.method {
hit = Some((i, params));
break;
}
path_seen = true;
} }
path_seen = true;
} }
} }
if let Some((i, params)) = hit {
let conn = conn.put_params(params);
return self.routes[i].handler.call(conn, next);
}
if path_seen { if path_seen {
// Path is known, method isn't — RFC 7231 §6.5.5. // Path is known, method isn't — RFC 7231 §6.5.5.
conn.put_status(405) conn.put_status(405)
+309 -118
View File
@@ -1,24 +1,19 @@
//! Listener pool and the `serve` entry point. //! [`Config`] and the `serve*` entry points.
//! //!
//! A small fixed pool of listener actors share the same TCP listen fd (via //! The listener pool itself lives in [`crate::endpoint`], which is the
//! `dup`) — each blocks in non-blocking `accept4` + `wait_readable` on its //! real API: an endpoint is a supervisable child you place in your own
//! own copy. When a connection arrives the listener spawns a connection //! tree. `serve*` is the batteries-included path for a process whose only
//! actor with the `OwnedFd` and immediately returns to `accept`. No //! job is serving HTTP — it owns the runtime and wraps one endpoint in a
//! coordination needed; the kernel serialises `accept` calls across the fds. //! one-child supervisor.
//! //!
//! Sharing via `dup` rather than the same fd is deliberate — Linux's use crate::conn_actor::ConnLimits;
//! `accept4` is thread-safe on a single fd, but dup'ing per-listener keeps
//! each actor's epoll registration local to its own RawFd value (so smarm's
//! `waiters: HashMap<RawFd, Pid>` doesn't see collisions between listeners
//! waiting on "the same fd").
use crate::conn_actor::{run_connection, ConnLimits};
use crate::net::{accept_nonblocking, bind_and_listen, OwnedFd};
use crate::plug::Pipeline; use crate::plug::Pipeline;
use smarm::supervisor::Shutdown;
use smarm::{ChildSpec, OneForOne, Restart, Strategy};
use std::io::{self, ErrorKind}; use std::io::{self, ErrorKind};
use std::net::{SocketAddr, ToSocketAddrs}; use std::net::{SocketAddr, ToSocketAddrs};
use std::os::fd::RawFd;
use std::time::Duration; use std::time::Duration;
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -32,14 +27,54 @@ pub struct Config {
pub keep_alive_timeout: Duration, pub keep_alive_timeout: Duration,
pub max_header_count: usize, pub max_header_count: usize,
pub read_buf_size: usize, pub read_buf_size: usize,
pub request_timeout: Duration, /// Wall-clock budget for reading the request HEAD (from first byte to
/// full head parse). Kept short — an incomplete head is the classic
/// slowloris. See `ConnLimits::head_timeout`.
pub head_timeout: Duration,
/// Absolute wall-clock cap on reading the request BODY (from head-parse
/// to full body). Sized for slow links, so much larger than
/// `head_timeout`. See `ConnLimits::body_timeout`.
pub body_timeout: Duration,
/// Burst size that resets the body stall clock. A body dribbling fewer
/// than this per `body_stall_timeout` window is evicted — the slowloris
/// / slow-legit discriminator. See `ConnLimits::body_burst_bytes`.
pub body_burst_bytes: usize,
/// Max time since the last qualifying body burst before eviction;
/// backstopped by `body_timeout`. See `ConnLimits::body_stall_timeout`.
pub body_stall_timeout: Duration,
/// Per-write budget for response bytes (the fixed head+body write, and
/// each streamed chunk). See `ConnLimits::write_timeout`.
pub write_timeout: Duration,
pub max_body_bytes: usize, pub max_body_bytes: usize,
/// Number of smarm scheduler OS threads. `None` means smarm's default /// How long a graceful shutdown waits for in-flight requests before
/// (one per CPU). Set this to a small fixed number in tests so multiple /// force-stopping the remaining connections. Idle keep-alive
/// concurrent test servers don't oversubscribe the host. /// connections are closed immediately on shutdown and do not run the
pub scheduler_threads: Option<usize>, /// clock out.
pub drain_timeout: Duration,
/// WebSocket: cap on a single frame's payload (header-checked
/// before buffering; violation closes 1009).
pub max_frame_payload: usize,
/// WebSocket: cap on a complete reassembled message (spans
/// fragments; violation closes 1009).
pub max_message_bytes: usize,
/// Registry name for this endpoint's gen_server — how it is addressed
/// from elsewhere in the app ([`crate::endpoint::whereis`]), and what
/// must be unique between two endpoints in one process (a public and
/// an admin port, say). Default `"urus"`.
pub name: &'static str,
/// Stack reserve (RFC 019 `smarm::SpawnOpts::stack_reserve`) given to
/// each per-connection actor. Request handlers routinely pull in
/// application code — DB drivers, (de)compression, templating — whose
/// stack needs comfortably exceed smarm's bare-actor default of 64 KiB
/// (the exact shape of bug this exists to head off; see smarm RFC 019).
/// Default: 256 KiB. The reserve is virtual/demand-paged, so raising it
/// costs address space, not RSS, until a handler actually uses it.
pub conn_stack_reserve: usize,
} }
/// Default per-connection actor stack reserve (see [`Config::conn_stack_reserve`]).
pub const DEFAULT_CONN_STACK_RESERVE: usize = 256 * 1024;
impl Config { impl Config {
pub fn new(addr: SocketAddr) -> Self { pub fn new(addr: SocketAddr) -> Self {
let pool = std::thread::available_parallelism() let pool = std::thread::available_parallelism()
@@ -52,137 +87,293 @@ impl Config {
keep_alive_timeout: Duration::from_secs(60), keep_alive_timeout: Duration::from_secs(60),
max_header_count: 64, max_header_count: 64,
read_buf_size: 8 * 1024, read_buf_size: 8 * 1024,
request_timeout: Duration::from_secs(30), head_timeout: Duration::from_secs(30),
body_timeout: Duration::from_secs(300),
body_burst_bytes: 4 * 1024,
body_stall_timeout: Duration::from_secs(20),
write_timeout: Duration::from_secs(30),
max_body_bytes: 16 * 1024 * 1024, max_body_bytes: 16 * 1024 * 1024,
scheduler_threads: None, drain_timeout: Duration::from_secs(30),
max_frame_payload: 1024 * 1024,
max_message_bytes: 4 * 1024 * 1024,
name: "urus",
conn_stack_reserve: DEFAULT_CONN_STACK_RESERVE,
} }
} }
fn to_conn_limits(&self) -> ConnLimits { pub(crate) fn to_conn_limits(&self) -> ConnLimits {
ConnLimits { ConnLimits {
max_headers: self.max_header_count, max_headers: self.max_header_count,
initial_read_buf: self.read_buf_size, initial_read_buf: self.read_buf_size,
max_head_bytes: 64 * 1024, max_head_bytes: 64 * 1024,
max_body_bytes: self.max_body_bytes, max_body_bytes: self.max_body_bytes,
keep_alive_timeout: self.keep_alive_timeout, keep_alive_timeout: self.keep_alive_timeout,
head_timeout: self.head_timeout,
body_timeout: self.body_timeout,
body_burst_bytes: self.body_burst_bytes,
body_stall_timeout: self.body_stall_timeout,
write_timeout: self.write_timeout,
max_frame_payload: self.max_frame_payload,
max_message_bytes: self.max_message_bytes,
} }
} }
} }
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// dup helper // config-file: TOML overlay for tuning knobs
// ---------------------------------------------------------------------------
fn dup_fd(fd: RawFd) -> io::Result<OwnedFd> {
let new_fd = unsafe { libc::fcntl(fd, libc::F_DUPFD_CLOEXEC, 0) };
if new_fd < 0 {
return Err(io::Error::last_os_error());
}
Ok(OwnedFd::from_raw(new_fd))
}
// ---------------------------------------------------------------------------
// listener actor body
// ---------------------------------------------------------------------------
fn listener_loop(listener: OwnedFd, pipeline: Pipeline, limits: ConnLimits) {
let fd = listener.as_raw();
loop {
match accept_nonblocking(fd) {
Ok(client) => {
// Hand the fd off to a new connection actor. spawn() is
// cheap on smarm — it's a single Vec push under the
// shared lock.
let p = pipeline.clone();
let l = limits;
smarm::spawn(move || run_connection(client, p, l));
}
Err(e) if e.kind() == ErrorKind::WouldBlock => {
// No pending connection. Park until the listener is
// readable again, then retry.
if let Err(we) = smarm::wait_readable(fd) {
// epoll registration failed — fatal for this listener.
eprintln!("urus: listener wait_readable failed: {we}");
return;
}
}
Err(e) if e.kind() == ErrorKind::Interrupted => {
continue;
}
Err(e) => {
// EMFILE / ENFILE / ECONNABORTED etc. Log and continue;
// the system may recover.
eprintln!("urus: accept error: {e}");
// Small backoff via smarm's sleep to avoid spinning if
// the error is sticky.
smarm::sleep(Duration::from_millis(10));
}
}
}
// listener OwnedFd drops here, closing the dup'd fd.
}
// ---------------------------------------------------------------------------
// serve_with — main entry. Boots smarm, spawns listeners, blocks.
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// //
// Boots an smarm runtime (one OS thread per CPU by default — see smarm's // urus is a library, so it never presumes a config-file path or reads the
// `Config::default()`) and runs until externally killed. We don't wire a // environment — the embedding binary decides where a file lives and hands
// graceful-shutdown signal in v1; the runtime exits when all actors exit, // the text here. This overlays a sparse TOML document onto an existing
// which they don't (listener loops are infinite). Ctrl-C is your friend. // `Config` (built with an addr the binary chose): only the keys present are
// applied, everything else keeps the compiled default. Durations are
// integer seconds. Unknown keys are a hard error so a typo is loud, not a
// silent no-op.
//
// Scope for now: the slowloris-tuning knobs only. Migrating the rest of the
// Config surface into the file is a separate, additive job (the loader
// mechanism is general — it just extends `TomlOverrides`).
pub fn serve_with(config: Config, pipeline: Pipeline) -> io::Result<()> { /// Error from [`Config::with_toml_str`]: the TOML failed to parse or carried
let listener = bind_and_listen(config.addr)?; /// an unknown/mistyped key.
println!("urus: listening on {}", config.addr); #[cfg(feature = "config-file")]
#[derive(Debug)]
pub enum ConfigError {
Toml(String),
}
// We want one connection-actor-spawning loop per listener pool slot. #[cfg(feature = "config-file")]
// Each gets its own dup'd fd so epoll registrations don't collide. impl std::fmt::Display for ConfigError {
let mut listener_fds = Vec::with_capacity(config.listener_pool); fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
listener_fds.push(listener); // primary keeps the original match self {
ConfigError::Toml(m) => write!(f, "config TOML error: {m}"),
for _ in 1..config.listener_pool { }
let dup = dup_fd(listener_fds[0].as_raw())?;
listener_fds.push(dup);
} }
}
let limits = config.to_conn_limits(); #[cfg(feature = "config-file")]
impl std::error::Error for ConfigError {}
// smarm's runtime API: init(Config) then run(f). The closure is the #[cfg(feature = "config-file")]
// root actor; from there we spawn one listener per fd in the pool. #[derive(serde::Deserialize)]
let smarm_cfg = match config.scheduler_threads { #[serde(deny_unknown_fields)]
Some(n) => smarm::Config::exact(n), struct TomlOverrides {
None => smarm::Config::default(), head_timeout_secs: Option<u64>,
}; body_timeout_secs: Option<u64>,
let rt = smarm::init(smarm_cfg); body_burst_bytes: Option<usize>,
body_stall_timeout_secs: Option<u64>,
}
#[cfg(feature = "config-file")]
impl Config {
/// Overlay a TOML document of tuning knobs onto this config (sparse:
/// only the keys present are applied). Durations are integer seconds.
///
/// Recognized keys: `head_timeout_secs`, `body_timeout_secs`,
/// `body_burst_bytes`, `body_stall_timeout_secs`. Unknown keys error.
/// Other `Config` knobs are not yet file-configurable.
pub fn with_toml_str(mut self, s: &str) -> Result<Self, ConfigError> {
let o: TomlOverrides =
toml::from_str(s).map_err(|e| ConfigError::Toml(e.to_string()))?;
if let Some(v) = o.head_timeout_secs {
self.head_timeout = Duration::from_secs(v);
}
if let Some(v) = o.body_timeout_secs {
self.body_timeout = Duration::from_secs(v);
}
if let Some(v) = o.body_burst_bytes {
self.body_burst_bytes = v;
}
if let Some(v) = o.body_stall_timeout_secs {
self.body_stall_timeout = Duration::from_secs(v);
}
Ok(self)
}
}
// ---------------------------------------------------------------------------
// Handle / ShutdownSignal — graceful shutdown plumbing for the serve* entries.
// ---------------------------------------------------------------------------
/// A clonable trigger for graceful shutdown, usable from any OS thread
/// (e.g. a signal-handling thread) — the shutdown path for callers who let
/// `serve*` own the runtime and so have no [`smarm::RuntimeHandle`] of
/// their own. If you build the tree yourself with [`crate::endpoint`], use
/// `rt.handle().request_shutdown(root_sup)` instead and ignore this.
#[derive(Clone)]
pub struct Handle {
tx: smarm::Sender<()>,
}
impl Handle {
/// Begin graceful shutdown: stop accepting, close idle keep-alive
/// connections, drain in-flight requests up to `Config.drain_timeout`,
/// then force-stop stragglers. `serve*` returns once the runtime has
/// wound down. Idempotent; extra calls are no-ops.
pub fn shutdown(&self) {
let _ = self.tx.send(());
}
}
/// The receiving half consumed by [`serve_with_shutdown`].
pub struct ShutdownSignal {
rx: smarm::Receiver<()>,
}
/// Create a connected [`Handle`]/[`ShutdownSignal`] pair.
pub fn shutdown_handle() -> (Handle, ShutdownSignal) {
let (tx, rx) = smarm::channel();
(Handle { tx }, ShutdownSignal { rx })
}
// ---------------------------------------------------------------------------
// serve* — batteries-included entries for apps whose only job is serving.
// ---------------------------------------------------------------------------
//
// These own the smarm runtime and build a one-child tree around
// `endpoint()`. An app with its own actors should call `endpoint()`
// directly and put it in its own supervision tree — that is the real API;
// everything here is a convenience wrapper over it.
/// Boot a runtime, serve until `signal` fires, then drain and return.
///
/// The tree is `root sup -> [..app_children, endpoint]` on
/// [`Strategy::RestForOne`], with the endpoint on [`Shutdown::Infinity`]
/// so its `drain_timeout` — not a supervisor deadline — bounds the drain.
///
/// `app_children` is your application's actors: a
/// [`PubSub::child`](crate::PubSub::child), a
/// [`ChannelHub::children`](crate::channels::ChannelHub::children), your
/// own state servers. They start **before** the endpoint and, because
/// shutdown is ordered in reverse, stop **after** it has drained — so a
/// request still in flight can still reach the bus. `RestForOne` in that
/// order also means one of them crashing restarts the endpoint behind it,
/// dropping connections whose subscriptions died with it, rather than
/// leaving live sockets addressing a table that no longer knows them.
///
/// The root actor parks on the signal channel; a `Handle::shutdown` from a
/// foreign OS thread wakes it, it shuts the supervisor down and `rt.run`
/// returns when the last actor is gone.
pub fn serve_with_shutdown(
config: Config,
rt_config: smarm::Config,
pipeline: Pipeline,
app_children: Vec<ChildSpec>,
signal: ShutdownSignal,
) -> io::Result<()> {
let addr = config.addr;
let endpoint = crate::endpoint::endpoint(config, pipeline)?;
println!("urus: listening on {addr}");
let rt = smarm::init(rt_config);
rt.run(move || { rt.run(move || {
let n = listener_fds.len(); let sup = smarm::spawn(move || {
let mut handles = Vec::with_capacity(n); let mut sup = OneForOne::new().strategy(Strategy::RestForOne);
for (i, lfd) in listener_fds.into_iter().enumerate() { for child in app_children {
let p = pipeline.clone(); sup = sup.child(child);
let h = smarm::spawn(move || { }
println!("urus: listener {} starting", i); sup.child(ChildSpec::new(Restart::Permanent, endpoint).shutdown(Shutdown::Infinity))
listener_loop(lfd, p, limits); .run()
}); });
handles.push(h);
} // Park until told to shut down. If every Handle was dropped the
// Block forever (until ctrl-C) by joining the listeners. They // channel closes and no shutdown can ever arrive: serve forever,
// never exit on their own in v1. // exactly v1's semantics.
for h in handles { match signal.rx.recv() {
let _ = h.join(); Ok(()) => smarm::request_shutdown(sup.pid()),
Err(_) => {
// Sender side gone. Park indefinitely; the process is
// expected to be killed externally.
loop {
smarm::sleep(Duration::from_secs(3600));
}
}
} }
let _ = sup.join();
}); });
Ok(()) Ok(())
} }
// --------------------------------------------------------------------------- /// [`serve_with_shutdown`] without a shutdown handle: serves until the
// serve — convenience over serve_with. /// process is killed.
// --------------------------------------------------------------------------- pub fn serve_with(
config: Config,
rt_config: smarm::Config,
pipeline: Pipeline,
app_children: Vec<ChildSpec>,
) -> io::Result<()> {
// The Handle is dropped immediately: shutdown can never be signalled.
let (_handle, signal) = shutdown_handle();
serve_with_shutdown(config, rt_config, pipeline, app_children, signal)
}
/// Defaults all round: default [`Config`], default smarm runtime (one
/// scheduler thread per CPU), no app children, serve until killed.
pub fn serve(addr: impl ToSocketAddrs, pipeline: Pipeline) -> io::Result<()> { pub fn serve(addr: impl ToSocketAddrs, pipeline: Pipeline) -> io::Result<()> {
let addr = addr let addr = addr
.to_socket_addrs()? .to_socket_addrs()?
.next() .next()
.ok_or_else(|| io::Error::new(ErrorKind::InvalidInput, "no addresses resolved"))?; .ok_or_else(|| io::Error::new(ErrorKind::InvalidInput, "no addresses resolved"))?;
serve_with(Config::new(addr), pipeline) serve_with(Config::new(addr), smarm::Config::default(), pipeline, Vec::new())
}
#[cfg(all(test, feature = "config-file"))]
mod config_file_tests {
use super::*;
fn base() -> Config {
Config::new("127.0.0.1:0".parse().unwrap())
}
#[test]
fn toml_empty_keeps_defaults() {
let d = base();
let c = base().with_toml_str("").unwrap();
assert_eq!(c.head_timeout, d.head_timeout);
assert_eq!(c.body_timeout, d.body_timeout);
assert_eq!(c.body_burst_bytes, d.body_burst_bytes);
assert_eq!(c.body_stall_timeout, d.body_stall_timeout);
}
#[test]
fn toml_partial_overrides_only_named() {
let d = base();
let c = base().with_toml_str("head_timeout_secs = 5").unwrap();
assert_eq!(c.head_timeout, Duration::from_secs(5)); // overridden
assert_eq!(c.body_timeout, d.body_timeout); // default kept
assert_eq!(c.body_burst_bytes, d.body_burst_bytes); // default kept
assert_eq!(c.body_stall_timeout, d.body_stall_timeout);
}
#[test]
fn toml_full_overrides_all() {
let c = base()
.with_toml_str(
"head_timeout_secs = 10\n\
body_timeout_secs = 120\n\
body_burst_bytes = 8192\n\
body_stall_timeout_secs = 15\n",
)
.unwrap();
assert_eq!(c.head_timeout, Duration::from_secs(10));
assert_eq!(c.body_timeout, Duration::from_secs(120));
assert_eq!(c.body_burst_bytes, 8192);
assert_eq!(c.body_stall_timeout, Duration::from_secs(15));
}
#[test]
fn toml_unknown_key_errors() {
// A mistyped/unknown key is a hard error, not a silent no-op.
let e = base().with_toml_str("body_timeout_sec = 120"); // typo: missing 's'
assert!(e.is_err(), "unknown key should error");
}
#[test]
fn toml_malformed_errors() {
let e = base().with_toml_str("this is not = valid = toml");
assert!(e.is_err(), "malformed TOML should error");
}
} }
+179
View File
@@ -0,0 +1,179 @@
//! Server-Sent Events sugar over `RespBody::Stream`.
//!
//! An SSE response is just a streaming body with the right headers and a
//! heartbeat: `Conn::sse()` wires all of it and hands back an
//! [`EventSender`] for a producer actor to feed.
//!
//! ```no_run
//! use urus::{Conn, Next, Pipeline, Router};
//! use std::time::Duration;
//!
//! let pipeline = Pipeline::new().plug(Router::new().get(
//! "/events",
//! |c: Conn, _n: Next| {
//! let (c, events) = c.sse();
//! smarm::spawn(move || {
//! let mut i = 0;
//! loop {
//! if events.send("tick", &i.to_string()).is_err() {
//! return; // client gone; conn actor dropped the stream
//! }
//! i += 1;
//! smarm::sleep(Duration::from_secs(1));
//! }
//! });
//! c
//! },
//! ));
//! ```
//!
//! Liveness: the connection actor emits a `: keep-alive` comment chunk
//! whenever [`HEARTBEAT_INTERVAL`] passes with no event (see
//! `StreamBody::heartbeat`). A client that went away is detected when a
//! write — event or heartbeat — stalls past `write_timeout`; the conn
//! actor then drops the stream and the producer's next [`EventSender`]
//! call returns `Err(SseClosed)`. There is no request clock on an SSE
//! response: the head/body read budgets cover only the read phase, by
//! design.
use crate::conn::{Conn, RespBody, StreamBody};
use std::time::Duration;
/// Default heartbeat: one comment line per this interval of silence.
pub const HEARTBEAT_INTERVAL: Duration = Duration::from_secs(15);
/// What the heartbeat puts on the wire — an SSE comment, invisible to
/// `EventSource` consumers.
pub(crate) const HEARTBEAT_PAYLOAD: &[u8] = b": keep-alive\n\n";
/// The stream has ended: the connection actor dropped the receiving end
/// (client disconnect, write timeout, or server shutdown). Producers
/// should exit when they see this.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub struct SseClosed;
impl std::fmt::Display for SseClosed {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
f.write_str("SSE stream closed")
}
}
impl std::error::Error for SseClosed {}
/// Producer handle for an SSE response. Clonable — multiple producers may
/// feed one stream; the stream ends when the LAST clone drops.
#[derive(Clone)]
pub struct EventSender {
tx: smarm::Sender<Vec<u8>>,
}
impl EventSender {
/// Send a named event: `event: {event}` + `data:` line(s).
pub fn send(&self, event: &str, data: &str) -> Result<(), SseClosed> {
self.tx
.send(format_event(Some(event), data))
.map_err(|_| SseClosed)
}
/// Send an unnamed event (`data:` line(s) only) — what `EventSource`
/// surfaces as a plain `message`.
pub fn data(&self, data: &str) -> Result<(), SseClosed> {
self.tx
.send(format_event(None, data))
.map_err(|_| SseClosed)
}
/// Send a comment line. Invisible to `EventSource`; useful for
/// application-level pings beyond the built-in heartbeat.
pub fn comment(&self, text: &str) -> Result<(), SseClosed> {
let mut out = Vec::with_capacity(text.len() + 4);
out.extend_from_slice(b": ");
out.extend_from_slice(text.as_bytes());
out.extend_from_slice(b"\n\n");
self.tx.send(out).map_err(|_| SseClosed)
}
}
/// Wire format for one event. Multi-line `data` becomes one `data:` line
/// per line (the SSE framing for embedded newlines); the blank line
/// dispatches the event.
pub(crate) fn format_event(event: Option<&str>, data: &str) -> Vec<u8> {
let mut out = Vec::with_capacity(data.len() + 32);
if let Some(e) = event {
out.extend_from_slice(b"event: ");
out.extend_from_slice(e.as_bytes());
out.push(b'\n');
}
for line in data.split('\n') {
out.extend_from_slice(b"data: ");
out.extend_from_slice(line.as_bytes());
out.push(b'\n');
}
out.push(b'\n');
out
}
impl Conn {
/// Turn this response into a Server-Sent Events stream with the
/// default [`HEARTBEAT_INTERVAL`]. Returns the `Conn` (return it from
/// the handler) and the [`EventSender`] to hand to a producer actor.
pub fn sse(self) -> (Self, EventSender) {
self.sse_with_heartbeat(HEARTBEAT_INTERVAL)
}
/// [`Conn::sse`] with a custom heartbeat interval.
pub fn sse_with_heartbeat(mut self, interval: Duration) -> (Self, EventSender) {
let (tx, rx) = smarm::channel::<Vec<u8>>();
if self.status.is_none() {
self.status = Some(200);
}
self.resp_headers.set("content-type", "text/event-stream");
self.resp_headers.set("cache-control", "no-cache");
self.resp_body = RespBody::Stream(
StreamBody::new(rx).with_heartbeat(interval, HEARTBEAT_PAYLOAD.to_vec()),
);
(self, EventSender { tx })
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn format_named_event() {
assert_eq!(
format_event(Some("tick"), "42"),
b"event: tick\ndata: 42\n\n".to_vec()
);
}
#[test]
fn format_unnamed_event() {
assert_eq!(format_event(None, "hi"), b"data: hi\n\n".to_vec());
}
#[test]
fn format_multiline_data() {
assert_eq!(
format_event(None, "a\nb"),
b"data: a\ndata: b\n\n".to_vec()
);
}
#[test]
fn sse_sets_headers_and_stream_body() {
let (conn, _events) = Conn::new().sse();
assert_eq!(conn.status, Some(200));
assert_eq!(conn.resp_headers.get("content-type"), Some("text/event-stream"));
assert_eq!(conn.resp_headers.get("cache-control"), Some("no-cache"));
match &conn.resp_body {
RespBody::Stream(s) => {
let (interval, payload) = s.heartbeat.as_ref().expect("heartbeat set");
assert_eq!(*interval, HEARTBEAT_INTERVAL);
assert_eq!(payload.as_slice(), HEARTBEAT_PAYLOAD);
}
other => panic!("expected Stream body, got {other:?}"),
}
}
}
+424
View File
@@ -0,0 +1,424 @@
//! The WebSocket duplex loop — v0.4 chunk 3.
//!
//! After an accepted upgrade the connection actor calls [`run_duplex`],
//! which replaces the HTTP request loop for the rest of the socket's
//! life. ONE actor speaks both directions (the v0.4 topology decision):
//! each iteration parks on a two-arm select —
//!
//! try_select(&[&outbound_rx, &FdArm::readable(fd)])
//!
//! — smarm's RFC 008 fd arms composed with a channel arm on one wait
//! epoch; urus is deliberately the first consumer of that machinery.
//! The outbound arm sits at index 0 (select is priority-ordered): owed
//! writes drain before we read more, which is honest backpressure.
//! Accepted gap, by design: while a large outbound frame is mid-write
//! the actor isn't reading, so a ping arriving during that write is
//! answered after it completes.
//!
//! Handler model (the v0.4 chunk-3 decisions, mirroring SSE):
//! - [`WsHandler`] callbacks run IN the connection actor, inside the
//! select loop. A slow `on_message` stops reads — backpressure, not a
//! bug. Handlers needing concurrency spawn their own actor and hand it
//! a [`WsSender`] clone, exactly like an SSE producer.
//! - [`WsSender`] is the clonable outbound handle; every method returns
//! `Err(WsClosed)` once the connection is gone.
//! - Control frames are invisible to the handler in v1: pings are
//! auto-ponged (RFC 6455 §5.5.2 MUST), pongs are absorbed, a peer
//! close is auto-echoed (§5.5.1) and surfaces only as `on_close`.
//! `on_ping`/`on_pong` hooks can land later as default trait methods
//! without breaking anyone.
//!
//! Liveness mirrors SSE: there is no request clock on an open
//! WebSocket. An idle connection parks in the select (stoppable by a
//! draining registry — graceful shutdown force-stops it at the drain
//! deadline); a dead client is caught when a write stalls past
//! `write_timeout`. Reads have no deadline of their own — the select
//! only wakes the read path when bytes are pending.
use crate::conn_actor::{read_some, write_all, ConnLimits};
use crate::ws::frame::{self, Assembler, Frame, FrameError, Message, Opcode};
use smarm::FdArm;
use std::os::fd::RawFd;
use std::time::Instant;
// ---------------------------------------------------------------------------
// Handler trait
// ---------------------------------------------------------------------------
/// Application callbacks for an upgraded WebSocket connection. Both run
/// inside the connection actor's select loop — see the module docs for
/// what that implies (backpressure, spawn-for-concurrency).
///
/// A panic in `on_message` closes the connection with a `1011 internal
/// error` close frame and skips `on_close` (the handler is not re-entered
/// after it panicked).
pub trait WsHandler: Send + 'static {
/// The connection is up: the 101 has been written and the frame loop
/// is about to start. Runs exactly once, first, inside the
/// connection actor. This is the place to subscribe / spawn
/// producers for sessions where the client may never send (a
/// listen-only chat member, a live feed) — `on_message` never fires
/// for those. NOTE the 101 reaches the client *before* `on_open`
/// runs: a client that needs to know its subscriptions are live must
/// use an application-level ack (send something, await the reply —
/// any reply proves `on_open` completed, since callbacks are
/// sequential in the conn actor). A panic here closes `1011` and
/// skips `on_close`, exactly like a panicking `on_message`.
/// Default: no-op.
fn on_open(&mut self, sender: &WsSender) {
let _ = sender;
}
/// A complete data message (fragments already reassembled, text
/// already UTF-8 validated) arrived from the client.
fn on_message(&mut self, msg: Message, sender: &WsSender);
/// The connection is over. `code`/`reason` are what the PEER said:
/// `Some` iff a close frame was received from the client (whether it
/// initiated or echoed ours); `None` for everything wordless — EOF,
/// write failure, or a close handshake that timed out. Called exactly
/// once, last, on every exit path except a handler panic or a
/// force-stop unwind.
fn on_close(&mut self, code: Option<u16>, reason: &str) {
let _ = (code, reason);
}
}
// ---------------------------------------------------------------------------
// WsSender — the clonable outbound handle
// ---------------------------------------------------------------------------
/// The connection is gone (client disconnect, write timeout, close
/// handshake completed, or server shutdown). Producers should exit when
/// they see this — the exact contract of `SseClosed`.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub struct WsClosed;
impl std::fmt::Display for WsClosed {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
f.write_str("WebSocket connection closed")
}
}
impl std::error::Error for WsClosed {}
pub(crate) enum Outbound {
Msg(Message),
Ping,
Close(u16, String),
}
/// Outbound handle for an upgraded connection. Clonable — hand copies to
/// producer actors freely; the connection outlives none of them (sends
/// after it ends return `Err(WsClosed)`).
///
/// v1 sends each message as a single unfragmented frame; outbound
/// fragmentation is not implemented.
#[derive(Clone)]
pub struct WsSender {
pub(crate) tx: smarm::Sender<Outbound>,
}
impl WsSender {
/// Queue a data message for the client.
pub fn send(&self, msg: Message) -> Result<(), WsClosed> {
self.tx.send(Outbound::Msg(msg)).map_err(|_| WsClosed)
}
/// Sugar: queue a text message.
pub fn text(&self, s: impl Into<String>) -> Result<(), WsClosed> {
self.send(Message::Text(s.into()))
}
/// Sugar: queue a binary message.
pub fn binary(&self, b: impl Into<Vec<u8>>) -> Result<(), WsClosed> {
self.send(Message::Binary(b.into()))
}
/// Queue a ping (empty payload). Application-level liveness beyond
/// the built-in write-timeout detection; the pong is absorbed by the
/// loop in v1.
pub fn ping(&self) -> Result<(), WsClosed> {
self.tx.send(Outbound::Ping).map_err(|_| WsClosed)
}
/// Initiate the close handshake: a close frame with `code`/`reason`
/// goes out, the loop then waits (bounded by `write_timeout`) for the
/// peer's echo before dropping the socket. `reason` is truncated to
/// fit the 125-byte control payload on a char boundary.
pub fn close(&self, code: u16, reason: &str) -> Result<(), WsClosed> {
self.tx
.send(Outbound::Close(code, reason.to_string()))
.map_err(|_| WsClosed)
}
}
// ---------------------------------------------------------------------------
// The loop
// ---------------------------------------------------------------------------
/// Run the duplex loop until the connection ends. `buf` carries any bytes
/// already read past the upgrade request's head — a client may pipeline
/// its first frame behind the handshake, and those bytes belong to us
/// now.
pub(crate) fn run_duplex(
raw: RawFd,
mut buf: Vec<u8>,
mut handler: Box<dyn WsHandler>,
limits: &ConnLimits,
) {
let (tx, rx) = smarm::channel::<Outbound>();
// We hold this sender for handler callbacks, so the rx arm can never
// report closed while the loop runs — no closed-arm starvation case.
let sender = WsSender { tx };
let mut asm = Assembler::new(limits.max_message_bytes);
// on_open before anything is decoded — even a pipelined first frame
// is observed by a handler that knows the connection exists. Same
// panic contract as on_message: re-raise a stop sentinel, else 1011.
let r = std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| handler.on_open(&sender)));
if r.is_err() {
smarm::preempt::check_cancelled();
let _ = close_and_drain_code(raw, 1011, "", &mut buf, limits);
return;
}
loop {
// ----- 1. Decode everything already buffered. -----
// Runs before the first select so a pipelined first frame is
// served without waiting for fresh readability.
loop {
let (frame, consumed) =
match frame::decode(&buf, true, limits.max_frame_payload) {
Ok(Some(hit)) => hit,
Ok(None) => break, // incomplete; need more bytes
Err(e) => {
let peer = close_and_drain(raw, e, &mut buf, limits);
finish(&mut handler, peer);
return;
}
};
buf.drain(..consumed);
match frame.opcode {
Opcode::Ping => {
// §5.5.2 MUST pong, payload echoed. Written inline —
// "as soon as practical" beats channel ordering.
let pong = Frame::new(Opcode::Pong, frame.payload);
if write_frame(raw, &pong, limits).is_err() {
finish(&mut handler, None);
return;
}
}
Opcode::Pong => {} // invisible in v1
Opcode::Close => {
// Peer-initiated close: echo its status code (§5.5.1)
// and end. The server side closes the TCP connection
// first by design (§7.1.1 puts that burden on us).
let echo_payload =
frame.payload.get(..2).map(<[u8]>::to_vec).unwrap_or_default();
let peer = match frame::parse_close_payload(&frame.payload) {
Ok(cp) => {
let _ = write_frame(
raw,
&Frame::new(Opcode::Close, echo_payload),
limits,
);
cp
}
Err(e) => {
// Malformed close payload: answer with the
// mapped code instead of an echo.
let _ = write_frame(
raw,
&Frame::new(
Opcode::Close,
frame::close_payload(e.close_code(), ""),
),
limits,
);
None
}
};
finish_with(&mut handler, peer);
return;
}
_ => match asm.push(frame) {
Ok(Some(msg)) => {
// A panicking handler must not swallow smarm's stop
// sentinel — same re-raise dance as the pipeline.
let r = std::panic::catch_unwind(
std::panic::AssertUnwindSafe(|| {
handler.on_message(msg, &sender)
}),
);
if r.is_err() {
smarm::preempt::check_cancelled();
// Genuine handler panic: 1011 out, no further
// handler calls (no on_close).
let _ = close_and_drain_code(
raw, 1011, "", &mut buf, limits,
);
return;
}
}
Ok(None) => {} // mid-fragmentation
Err(e) => {
let peer = close_and_drain(raw, e, &mut buf, limits);
finish(&mut handler, peer);
return;
}
},
}
}
// ----- 2. Park on first-of outbound / fd-readable. -----
// Outbound at index 0: select is priority-ordered, so owed writes
// go out before we read more. try_select over select: an fd-arm
// registration failure (EBADF after a peer RST, EMFILE on the
// epoll set) is a connection event, not a crash. NOTE from the
// RFC 008 review: an `AlreadyExists` here would mean a stale
// waiter survived select's eager cleanup — that is a smarm bug
// candidate; capture it with --features smarm-trace and report,
// do not paper over.
let fd_arm = FdArm::readable(raw);
let idx = match smarm::try_select(&[&rx, &fd_arm]) {
Ok(i) => i,
Err(_) => {
finish(&mut handler, None);
return;
}
};
if idx == 0 {
// Outbound arm won. Single receiver + we hold a sender: a
// ready arm means a message is queued (Empty/closed are
// impossible here, but stay defensive).
let Ok(Some(out)) = rx.try_recv() else { continue };
match out {
Outbound::Msg(msg) => {
let f = match msg {
Message::Text(s) => Frame::new(Opcode::Text, s.into_bytes()),
Message::Binary(b) => Frame::new(Opcode::Binary, b),
};
if write_frame(raw, &f, limits).is_err() {
finish(&mut handler, None);
return;
}
}
Outbound::Ping => {
let f = Frame::new(Opcode::Ping, Vec::new());
if write_frame(raw, &f, limits).is_err() {
finish(&mut handler, None);
return;
}
}
Outbound::Close(code, reason) => {
// We initiate: close out, bounded wait for the echo.
let peer = close_and_drain_code(raw, code, &reason, &mut buf, limits);
finish(&mut handler, peer);
return;
}
}
} else {
// Readable. The deadline is nominal — the fd already reported
// ready, this returns without a fresh park in practice.
match read_some(
raw,
&mut buf,
limits.initial_read_buf,
Instant::now() + limits.write_timeout,
) {
Ok(0) => {
// EOF without a close frame: abnormal but common.
finish(&mut handler, None);
return;
}
Ok(_) => {}
Err(_) => {
finish(&mut handler, None);
return;
}
}
}
}
}
/// `on_close(None, "")` — wordless endings (EOF, write failure, errors).
fn finish(handler: &mut Box<dyn WsHandler>, peer: Option<(u16, String)>) {
finish_with(handler, peer);
}
/// `on_close` with whatever the peer said, panic-shielded like
/// `on_message` (stop sentinel re-raised; a panic here just ends the
/// already-ending connection).
fn finish_with(handler: &mut Box<dyn WsHandler>, peer: Option<(u16, String)>) {
let r = std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| {
match &peer {
Some((code, reason)) => handler.on_close(Some(*code), reason),
None => handler.on_close(None, ""),
}
}));
if r.is_err() {
smarm::preempt::check_cancelled();
}
}
/// One frame onto the wire under a fresh `write_timeout` budget — the
/// streaming-chunk discipline, verbatim.
fn write_frame(raw: RawFd, f: &Frame, limits: &ConnLimits) -> std::io::Result<()> {
write_all(raw, &f.encode(), Instant::now() + limits.write_timeout)
}
/// Server-initiated close for a frame/assembly error: map to the §7.4
/// code and run the handshake tail.
fn close_and_drain(
raw: RawFd,
e: FrameError,
buf: &mut Vec<u8>,
limits: &ConnLimits,
) -> Option<(u16, String)> {
close_and_drain_code(raw, e.close_code(), "", buf, limits)
}
/// Write our close frame, then wait — bounded by ONE `write_timeout`
/// budget — for the peer's close frame, discarding data frames (§1.4: an
/// endpoint that has sent close may discard incoming data). Returns what
/// the peer's close said, `None` if it never arrived (EOF, timeout,
/// garbage). The socket drops at the caller either way.
fn close_and_drain_code(
raw: RawFd,
code: u16,
reason: &str,
buf: &mut Vec<u8>,
limits: &ConnLimits,
) -> Option<(u16, String)> {
let ours = Frame::new(Opcode::Close, frame::close_payload(code, reason));
if write_frame(raw, &ours, limits).is_err() {
return None;
}
let deadline = Instant::now() + limits.write_timeout;
loop {
loop {
match frame::decode(buf, true, limits.max_frame_payload) {
Ok(Some((frame, consumed))) => {
buf.drain(..consumed);
if frame.opcode == Opcode::Close {
return frame::parse_close_payload(&frame.payload)
.ok()
.flatten();
}
// Data/ping mid-handshake: discarded.
}
Ok(None) => break,
Err(_) => return None, // garbage during the tail: give up
}
}
match read_some(raw, buf, limits.initial_read_buf, deadline) {
Ok(0) | Err(_) => return None,
Ok(_) => {}
}
}
}
+637
View File
@@ -0,0 +1,637 @@
//! RFC 6455 §5 — the frame codec. Pure: bytes in, frames out, no io and
//! no actor machinery, so every edge lives in fast unit tests. The
//! connection actor (chunk 3) owns buffering and the socket.
//!
//! Server-side rules enforced here:
//! - client→server frames MUST be masked (§5.1); decode takes
//! `require_masked` so a future client mode can reuse the codec.
//! - server→client frames are NEVER masked; [`Frame::encode`] doesn't
//! offer masking. [`encode_masked`] exists for the client side of
//! tests.
//! - RSV bits must be 0 (we negotiate no extensions), §5.2.
//! - Payload lengths must use the minimal encoding, §5.2.
//! - Control frames: FIN set, payload ≤ 125, §5.5.
//!
//! Errors map to close codes via [`FrameError::close_code`]: protocol
//! violations → 1002, oversize → 1009, bad UTF-8 in text → 1007.
/// §5.2 opcodes. Reserved values are rejected at decode.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum Opcode {
Continuation,
Text,
Binary,
Close,
Ping,
Pong,
}
impl Opcode {
fn from_u4(n: u8) -> Option<Self> {
Some(match n {
0x0 => Opcode::Continuation,
0x1 => Opcode::Text,
0x2 => Opcode::Binary,
0x8 => Opcode::Close,
0x9 => Opcode::Ping,
0xA => Opcode::Pong,
_ => return None,
})
}
fn to_u4(self) -> u8 {
match self {
Opcode::Continuation => 0x0,
Opcode::Text => 0x1,
Opcode::Binary => 0x2,
Opcode::Close => 0x8,
Opcode::Ping => 0x9,
Opcode::Pong => 0xA,
}
}
pub fn is_control(self) -> bool {
matches!(self, Opcode::Close | Opcode::Ping | Opcode::Pong)
}
}
/// One decoded frame; payload already unmasked.
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct Frame {
pub fin: bool,
pub opcode: Opcode,
pub payload: Vec<u8>,
}
impl Frame {
pub fn new(opcode: Opcode, payload: impl Into<Vec<u8>>) -> Self {
Frame { fin: true, opcode, payload: payload.into() }
}
/// Serialise unmasked (server→client, §5.1: a server MUST NOT mask).
pub fn encode(&self) -> Vec<u8> {
let mut out = Vec::with_capacity(self.payload.len() + 10);
encode_head(&mut out, self.fin, self.opcode, self.payload.len(), None);
out.extend_from_slice(&self.payload);
out
}
}
/// Why decoding (or assembly) failed. Fatal for the connection: the ws
/// close handshake should carry [`FrameError::close_code`].
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum FrameError {
/// §5.x violation; the str is a debugging breadcrumb, not protocol.
Protocol(&'static str),
/// Frame or assembled message exceeds the configured cap → 1009.
TooLarge,
/// A complete text message that is not valid UTF-8 → 1007.
BadUtf8,
}
impl FrameError {
pub fn close_code(self) -> u16 {
match self {
FrameError::Protocol(_) => 1002,
FrameError::TooLarge => 1009,
FrameError::BadUtf8 => 1007,
}
}
}
/// Try to decode one frame from the front of `buf`.
///
/// - `Ok(Some((frame, consumed)))` — drop `consumed` bytes off the front.
/// - `Ok(None)` — incomplete; read more. (Header-derived sizes are still
/// bounds-checked first, so a hostile 8-byte length can't make the
/// caller buffer toward it: oversize fails *before* the payload
/// arrives.)
/// - `Err(_)` — fatal; start the close handshake with the mapped code.
pub fn decode(
buf: &[u8],
require_masked: bool,
max_payload: usize,
) -> Result<Option<(Frame, usize)>, FrameError> {
if buf.len() < 2 {
return Ok(None);
}
let b0 = buf[0];
let b1 = buf[1];
let fin = b0 & 0x80 != 0;
if b0 & 0x70 != 0 {
return Err(FrameError::Protocol("RSV bits set without extension"));
}
let opcode = Opcode::from_u4(b0 & 0x0F)
.ok_or(FrameError::Protocol("reserved opcode"))?;
let masked = b1 & 0x80 != 0;
if require_masked && !masked {
return Err(FrameError::Protocol("unmasked client frame"));
}
if opcode.is_control() {
if !fin {
return Err(FrameError::Protocol("fragmented control frame"));
}
if b1 & 0x7F > 125 {
return Err(FrameError::Protocol("control payload > 125"));
}
}
// Payload length: 7-bit, or 126 + u16, or 127 + u64 — minimal form
// required (§5.2).
let (len, mut pos): (u64, usize) = match b1 & 0x7F {
126 => {
if buf.len() < 4 {
return Ok(None);
}
let l = u16::from_be_bytes([buf[2], buf[3]]) as u64;
if l < 126 {
return Err(FrameError::Protocol("non-minimal 16-bit length"));
}
(l, 4)
}
127 => {
if buf.len() < 10 {
return Ok(None);
}
let l = u64::from_be_bytes(buf[2..10].try_into().unwrap());
if l & (1 << 63) != 0 {
return Err(FrameError::Protocol("64-bit length MSB set"));
}
if l < 65536 {
return Err(FrameError::Protocol("non-minimal 64-bit length"));
}
(l, 10)
}
n => (n as u64, 2),
};
// Cap check BEFORE waiting for the payload (see decode docs).
if len > max_payload as u64 {
return Err(FrameError::TooLarge);
}
let len = len as usize;
let mask_key: Option<[u8; 4]> = if masked {
if buf.len() < pos + 4 {
return Ok(None);
}
let k = [buf[pos], buf[pos + 1], buf[pos + 2], buf[pos + 3]];
pos += 4;
Some(k)
} else {
None
};
if buf.len() < pos + len {
return Ok(None);
}
let mut payload = buf[pos..pos + len].to_vec();
if let Some(key) = mask_key {
for (i, b) in payload.iter_mut().enumerate() {
*b ^= key[i & 3];
}
}
Ok(Some((Frame { fin, opcode, payload }, pos + len)))
}
fn encode_head(
out: &mut Vec<u8>,
fin: bool,
opcode: Opcode,
len: usize,
mask: Option<[u8; 4]>,
) {
let mask_bit = if mask.is_some() { 0x80 } else { 0 };
out.push(if fin { 0x80 } else { 0 } | opcode.to_u4());
if len <= 125 {
out.push(mask_bit | len as u8);
} else if len <= u16::MAX as usize {
out.push(mask_bit | 126);
out.extend_from_slice(&(len as u16).to_be_bytes());
} else {
out.push(mask_bit | 127);
out.extend_from_slice(&(len as u64).to_be_bytes());
}
if let Some(k) = mask {
out.extend_from_slice(&k);
}
}
/// Client-side serialisation (masked). Servers never call this in
/// production — it exists so tests (and an eventual client mode) can
/// produce conformant client frames.
pub fn encode_masked(frame: &Frame, key: [u8; 4]) -> Vec<u8> {
let mut out = Vec::with_capacity(frame.payload.len() + 14);
encode_head(&mut out, frame.fin, frame.opcode, frame.payload.len(), Some(key));
out.extend(frame.payload.iter().enumerate().map(|(i, b)| b ^ key[i & 3]));
out
}
// ---------------------------------------------------------------------------
// Close payload (§5.5.1, §7.4)
// ---------------------------------------------------------------------------
/// Parse a close frame payload: empty is fine (`None`); otherwise a
/// 2-byte code + optional UTF-8 reason. A 1-byte payload, an invalid
/// wire code, or a non-UTF-8 reason are protocol errors.
pub fn parse_close_payload(payload: &[u8]) -> Result<Option<(u16, String)>, FrameError> {
match payload.len() {
0 => Ok(None),
1 => Err(FrameError::Protocol("1-byte close payload")),
_ => {
let code = u16::from_be_bytes([payload[0], payload[1]]);
if !close_code_valid_on_wire(code) {
return Err(FrameError::Protocol("invalid close code"));
}
let reason = std::str::from_utf8(&payload[2..])
.map_err(|_| FrameError::Protocol("close reason not UTF-8"))?;
Ok(Some((code, reason.to_string())))
}
}
}
/// §7.4.1/.2: codes an endpoint may put on the wire. 1004 is reserved;
/// 1005/1006/1015 are signalling-only (MUST NOT appear in a close
/// frame); 1000–1011 otherwise fine; 3000–4999 are app/registry space.
fn close_code_valid_on_wire(code: u16) -> bool {
matches!(code, 1000..=1003 | 1007..=1011 | 3000..=4999)
}
/// Build a close frame payload for `code` (+ optional reason, truncated
/// to fit the 125-byte control cap on a UTF-8 boundary).
pub fn close_payload(code: u16, reason: &str) -> Vec<u8> {
let mut p = Vec::with_capacity(2 + reason.len().min(123));
p.extend_from_slice(&code.to_be_bytes());
let mut cut = reason.len().min(123);
while !reason.is_char_boundary(cut) {
cut -= 1;
}
p.extend_from_slice(&reason.as_bytes()[..cut]);
p
}
// ---------------------------------------------------------------------------
// Message assembly (§5.4 fragmentation)
// ---------------------------------------------------------------------------
/// A complete application message.
#[derive(Debug, Clone, PartialEq, Eq)]
pub enum Message {
Text(String),
Binary(Vec<u8>),
}
/// Reassembles data frames into [`Message`]s. Feed it ONLY data frames
/// (Text/Binary/Continuation) — control frames interleave at the caller
/// (§5.4 allows them mid-fragmentation) and never enter the assembler.
///
/// Enforced here: continuation with nothing in progress; a new data
/// frame while a fragmented message is in progress; total message size
/// (1009); text UTF-8 validity, checked once on the complete message
/// (we deliver whole messages, so per-frame incremental validation buys
/// nothing).
#[derive(Debug)]
pub struct Assembler {
buf: Vec<u8>,
in_progress: Option<Opcode>, // Text or Binary
max_message: usize,
}
impl Assembler {
pub fn new(max_message: usize) -> Self {
Assembler { buf: Vec::new(), in_progress: None, max_message }
}
pub fn push(&mut self, frame: Frame) -> Result<Option<Message>, FrameError> {
match (frame.opcode, self.in_progress) {
(Opcode::Text | Opcode::Binary, Some(_)) => {
return Err(FrameError::Protocol("new data frame mid-fragmentation"));
}
(Opcode::Continuation, None) => {
return Err(FrameError::Protocol("continuation without a message"));
}
(Opcode::Text | Opcode::Binary, None) if frame.fin => {
// Unfragmented fast path: no buffer copy.
if frame.payload.len() > self.max_message {
return Err(FrameError::TooLarge);
}
return Self::complete(frame.opcode, frame.payload).map(Some);
}
(Opcode::Text | Opcode::Binary, None) => {
self.in_progress = Some(frame.opcode);
}
(Opcode::Continuation, Some(_)) => {}
(op, _) if op.is_control() => {
return Err(FrameError::Protocol("control frame fed to assembler"));
}
_ => unreachable!(),
}
if self.buf.len() + frame.payload.len() > self.max_message {
self.reset();
return Err(FrameError::TooLarge);
}
self.buf.extend_from_slice(&frame.payload);
if frame.fin {
let kind = self.in_progress.take().expect("checked above");
let payload = std::mem::take(&mut self.buf);
return Self::complete(kind, payload).map(Some);
}
Ok(None)
}
fn complete(kind: Opcode, payload: Vec<u8>) -> Result<Message, FrameError> {
match kind {
Opcode::Binary => Ok(Message::Binary(payload)),
Opcode::Text => String::from_utf8(payload)
.map(Message::Text)
.map_err(|_| FrameError::BadUtf8),
_ => unreachable!(),
}
}
fn reset(&mut self) {
self.buf.clear();
self.in_progress = None;
}
}
#[cfg(test)]
mod tests {
use super::*;
const MAX: usize = 1 << 20;
fn dec(buf: &[u8]) -> Result<Option<(Frame, usize)>, FrameError> {
decode(buf, true, MAX)
}
// ----- §5.7 worked examples. -----
#[test]
fn rfc_masked_hello() {
// "A single-frame masked text message" containing "Hello".
let wire = [0x81, 0x85, 0x37, 0xfa, 0x21, 0x3d, 0x7f, 0x9f, 0x4d, 0x51, 0x58];
let (f, used) = dec(&wire).unwrap().unwrap();
assert_eq!(used, wire.len());
assert!(f.fin);
assert_eq!(f.opcode, Opcode::Text);
assert_eq!(f.payload, b"Hello");
}
#[test]
fn rfc_unmasked_hello_rejected_then_allowed() {
let wire = [0x81, 0x05, b'H', b'e', b'l', b'l', b'o'];
assert_eq!(
dec(&wire),
Err(FrameError::Protocol("unmasked client frame"))
);
// Same bytes are fine when masking isn't required (client mode).
let (f, _) = decode(&wire, false, MAX).unwrap().unwrap();
assert_eq!(f.payload, b"Hello");
}
#[test]
fn rfc_fragmented_hel_lo() {
let f1 = [0x01, 0x03, b'H', b'e', b'l'];
let f2 = [0x80, 0x02, b'l', b'o'];
let (a, _) = decode(&f1, false, MAX).unwrap().unwrap();
let (b, _) = decode(&f2, false, MAX).unwrap().unwrap();
assert!(!a.fin);
assert_eq!(a.opcode, Opcode::Text);
assert_eq!(b.opcode, Opcode::Continuation);
let mut asm = Assembler::new(MAX);
assert_eq!(asm.push(a).unwrap(), None);
assert_eq!(
asm.push(b).unwrap(),
Some(Message::Text("Hello".into()))
);
}
#[test]
fn rfc_ping_pong_hello() {
let ping = [0x89, 0x05, b'H', b'e', b'l', b'l', b'o'];
let (f, _) = decode(&ping, false, MAX).unwrap().unwrap();
assert_eq!(f.opcode, Opcode::Ping);
assert_eq!(f.payload, b"Hello");
// Pong response carries the same payload back, unmasked.
let pong = Frame::new(Opcode::Pong, f.payload.clone()).encode();
assert_eq!(pong, [0x8A, 0x05, b'H', b'e', b'l', b'l', b'o']);
}
#[test]
fn rfc_256_bytes_binary_16bit_len() {
let payload = vec![0xAB; 256];
let wire = Frame::new(Opcode::Binary, payload.clone()).encode();
assert_eq!(&wire[..4], &[0x82, 0x7E, 0x01, 0x00]);
let (f, used) = decode(&wire, false, MAX).unwrap().unwrap();
assert_eq!(used, wire.len());
assert_eq!(f.payload, payload);
}
#[test]
fn rfc_64k_binary_64bit_len() {
let payload = vec![7u8; 65536];
let wire = Frame::new(Opcode::Binary, payload.clone()).encode();
assert_eq!(&wire[..10], &[0x82, 0x7F, 0, 0, 0, 0, 0, 1, 0, 0]);
let (f, _) = decode(&wire, false, MAX).unwrap().unwrap();
assert_eq!(f.payload.len(), 65536);
}
// ----- round trips. -----
#[test]
fn masked_roundtrip_all_lengths() {
// Cross the 125/126 and 65535/65536 encoding boundaries.
for len in [0, 1, 125, 126, 127, 65535, 65536] {
let f = Frame::new(Opcode::Binary, vec![0x5A; len]);
let wire = encode_masked(&f, [0xDE, 0xAD, 0xBE, 0xEF]);
let (g, used) = decode(&wire, true, 1 << 20).unwrap().unwrap();
assert_eq!(used, wire.len(), "len {len}");
assert_eq!(g, f, "len {len}");
}
}
#[test]
fn decode_is_incremental_at_every_prefix() {
let f = Frame::new(Opcode::Text, b"incremental".to_vec());
let wire = encode_masked(&f, [1, 2, 3, 4]);
for cut in 0..wire.len() {
assert_eq!(dec(&wire[..cut]).unwrap(), None, "prefix {cut}");
}
assert!(dec(&wire).unwrap().is_some());
}
#[test]
fn decode_leaves_trailing_bytes() {
let f = Frame::new(Opcode::Text, b"a".to_vec());
let mut wire = encode_masked(&f, [9, 9, 9, 9]);
let first_len = wire.len();
wire.extend_from_slice(&[0x81]); // start of a second frame
let (_, used) = dec(&wire).unwrap().unwrap();
assert_eq!(used, first_len);
}
// ----- protocol violations. -----
#[test]
fn rsv_bits_rejected() {
for rsv in [0x40, 0x20, 0x10] {
let wire = [0x81 | rsv, 0x80, 0, 0, 0, 0];
assert!(matches!(dec(&wire), Err(FrameError::Protocol(_))), "rsv {rsv:#x}");
}
}
#[test]
fn reserved_opcodes_rejected() {
for op in [0x3, 0x4, 0x5, 0x6, 0x7, 0xB, 0xC, 0xD, 0xE, 0xF] {
let wire = [0x80 | op, 0x80, 0, 0, 0, 0];
assert!(matches!(dec(&wire), Err(FrameError::Protocol(_))), "op {op:#x}");
}
}
#[test]
fn fragmented_control_rejected() {
let wire = [0x09, 0x80, 0, 0, 0, 0]; // ping, FIN=0
assert_eq!(dec(&wire), Err(FrameError::Protocol("fragmented control frame")));
}
#[test]
fn oversize_control_rejected() {
let wire = [0x89, 0x80 | 126, 0x00, 0x7E]; // ping, 16-bit len 126
assert_eq!(dec(&wire), Err(FrameError::Protocol("control payload > 125")));
}
#[test]
fn non_minimal_lengths_rejected() {
// 16-bit form carrying 5.
let w16 = [0x81, 0x80 | 126, 0x00, 0x05];
assert_eq!(dec(&w16), Err(FrameError::Protocol("non-minimal 16-bit length")));
// 64-bit form carrying 5.
let w64 = [0x81, 0x80 | 127, 0, 0, 0, 0, 0, 0, 0, 0x05];
assert_eq!(dec(&w64), Err(FrameError::Protocol("non-minimal 64-bit length")));
}
#[test]
fn msb_set_64bit_length_rejected() {
let wire = [0x81, 0x80 | 127, 0x80, 0, 0, 0, 0, 0, 0, 0];
assert_eq!(dec(&wire), Err(FrameError::Protocol("64-bit length MSB set")));
}
#[test]
fn oversize_fails_before_payload_arrives() {
// Header promises 1 MiB + 1 with a 1 MiB cap: must fail with just
// the header in hand, not Ok(None).
let mut wire = vec![0x82, 0x80 | 127];
wire.extend_from_slice(&((MAX as u64) + 1).to_be_bytes());
assert_eq!(dec(&wire), Err(FrameError::TooLarge));
}
// ----- close payloads. -----
#[test]
fn close_payload_parsing() {
assert_eq!(parse_close_payload(&[]), Ok(None));
assert_eq!(
parse_close_payload(&[0x03]),
Err(FrameError::Protocol("1-byte close payload"))
);
let normal = close_payload(1000, "bye");
assert_eq!(parse_close_payload(&normal), Ok(Some((1000, "bye".into()))));
// Signalling-only and reserved codes are invalid on the wire.
for bad in [999u16, 1004, 1005, 1006, 1015, 1100, 2999, 5000] {
assert!(
parse_close_payload(&bad.to_be_bytes()).is_err(),
"code {bad}"
);
}
for good in [1000u16, 1001, 1002, 1003, 1007, 1011, 3000, 4999] {
assert!(
parse_close_payload(&good.to_be_bytes()).is_ok(),
"code {good}"
);
}
}
#[test]
fn close_payload_reason_truncates_on_char_boundary() {
let reason = "é".repeat(100); // 200 bytes of 2-byte chars
let p = close_payload(1000, &reason);
assert!(p.len() <= 125);
assert_eq!(p.len() % 2, 0, "must not split the 2-byte char");
assert!(std::str::from_utf8(&p[2..]).is_ok());
}
// ----- assembler. -----
#[test]
fn assembler_rejects_orphan_continuation() {
let mut a = Assembler::new(MAX);
let f = Frame { fin: true, opcode: Opcode::Continuation, payload: vec![] };
assert_eq!(
a.push(f),
Err(FrameError::Protocol("continuation without a message"))
);
}
#[test]
fn assembler_rejects_interleaved_data_message() {
let mut a = Assembler::new(MAX);
let start = Frame { fin: false, opcode: Opcode::Text, payload: b"a".to_vec() };
a.push(start).unwrap();
let intruder = Frame::new(Opcode::Binary, b"b".to_vec());
assert_eq!(
a.push(intruder),
Err(FrameError::Protocol("new data frame mid-fragmentation"))
);
}
#[test]
fn assembler_message_size_cap_spans_fragments() {
let mut a = Assembler::new(10);
let f1 = Frame { fin: false, opcode: Opcode::Binary, payload: vec![0; 6] };
let f2 = Frame { fin: true, opcode: Opcode::Continuation, payload: vec![0; 6] };
assert_eq!(a.push(f1).unwrap(), None);
assert_eq!(a.push(f2), Err(FrameError::TooLarge));
// And the unfragmented fast path is capped too.
assert_eq!(
a.push(Frame::new(Opcode::Binary, vec![0; 11])),
Err(FrameError::TooLarge)
);
}
#[test]
fn assembler_text_utf8_validated_on_completion() {
let mut a = Assembler::new(MAX);
// A 2-byte UTF-8 char split across the fragment boundary must
// still assemble — validation is whole-message.
let bytes = "héllo".as_bytes();
let f1 = Frame { fin: false, opcode: Opcode::Text, payload: bytes[..2].to_vec() };
let f2 = Frame { fin: true, opcode: Opcode::Continuation, payload: bytes[2..].to_vec() };
assert_eq!(a.push(f1).unwrap(), None);
assert_eq!(a.push(f2).unwrap(), Some(Message::Text("héllo".into())));
// Invalid UTF-8 → BadUtf8 (close 1007).
let bad = Frame::new(Opcode::Text, vec![0xFF, 0xFE]);
assert_eq!(a.push(bad), Err(FrameError::BadUtf8));
assert_eq!(FrameError::BadUtf8.close_code(), 1007);
}
#[test]
fn assembler_reusable_after_message() {
let mut a = Assembler::new(MAX);
for _ in 0..3 {
let f1 = Frame { fin: false, opcode: Opcode::Text, payload: b"He".to_vec() };
let f2 = Frame { fin: true, opcode: Opcode::Continuation, payload: b"llo".to_vec() };
assert_eq!(a.push(f1).unwrap(), None);
assert_eq!(a.push(f2).unwrap(), Some(Message::Text("Hello".into())));
}
}
}
+250
View File
@@ -0,0 +1,250 @@
//! RFC 6455 §4.2 — the opening handshake, server side.
//!
//! Pure functions over an already-parsed request; no io. The connection
//! actor owns the wire.
use crate::conn::{Conn, HttpVersion, Method};
/// RFC 6455 §1.3 — the fixed GUID appended to the client key.
const WS_GUID: &str = "258EAFA5-E914-47DA-95CA-C5AB0DC85B11";
/// Why a handshake was rejected. Maps onto a response via
/// [`Rejection::status`]: everything is a `400` except a version
/// mismatch, which RFC 6455 §4.2.2 answers with `426` + a
/// `sec-websocket-version: 13` header so the client can retry.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum Rejection {
/// Handshake requests must be GET (§4.2.1).
NotGet,
/// HTTP/1.1 or higher required (§4.2.1).
BadHttpVersion,
/// `Host` header missing (§4.2.1 item 2).
MissingHost,
/// `Upgrade` header missing or lacks the `websocket` token.
NotAnUpgrade,
/// `Connection` header missing or lacks the `upgrade` token.
ConnectionNotUpgrade,
/// `Sec-WebSocket-Key` missing or not base64 of exactly 16 bytes.
BadKey,
/// `Sec-WebSocket-Version` is not 13 → 426 Upgrade Required.
UnsupportedVersion,
}
impl Rejection {
pub fn status(self) -> u16 {
match self {
Rejection::UnsupportedVersion => 426,
_ => 400,
}
}
}
/// Cheap routing predicate: does this request *ask* for a WebSocket
/// upgrade? (Token checks only — full validation happens in
/// [`validate`].) Lets a handler branch without committing to the
/// handshake.
pub fn is_upgrade_request(conn: &Conn) -> bool {
has_token(conn.headers.get("upgrade"), "websocket")
&& has_token(conn.headers.get("connection"), "upgrade")
}
/// Full §4.2.1 server-side validation. `Ok` carries the computed
/// `Sec-WebSocket-Accept` value.
pub fn validate(conn: &Conn) -> Result<String, Rejection> {
if conn.method != Method::Get {
return Err(Rejection::NotGet);
}
if conn.version != HttpVersion::Http11 {
return Err(Rejection::BadHttpVersion);
}
if conn.headers.get("host").is_none() {
return Err(Rejection::MissingHost);
}
if !has_token(conn.headers.get("upgrade"), "websocket") {
return Err(Rejection::NotAnUpgrade);
}
if !has_token(conn.headers.get("connection"), "upgrade") {
return Err(Rejection::ConnectionNotUpgrade);
}
match conn.headers.get("sec-websocket-version") {
Some(v) if v.trim() == "13" => {}
_ => return Err(Rejection::UnsupportedVersion),
}
let key = conn
.headers
.get("sec-websocket-key")
.map(str::trim)
.ok_or(Rejection::BadKey)?;
if !key_is_b64_16_bytes(key) {
return Err(Rejection::BadKey);
}
Ok(accept_key(key))
}
/// `base64(SHA1(key ++ GUID))` — §4.2.2 item 5.4.
pub fn accept_key(key: &str) -> String {
let mut h = sha1_smol::Sha1::new();
h.update(key.as_bytes());
h.update(WS_GUID.as_bytes());
b64(&h.digest().bytes())
}
/// Case-insensitive token search in a comma-separated header value.
/// `Connection: keep-alive, Upgrade` must match token `upgrade`.
fn has_token(value: Option<&str>, token: &str) -> bool {
match value {
None => false,
Some(v) => v
.split(',')
.any(|t| t.trim().eq_ignore_ascii_case(token)),
}
}
/// The client key must be base64 of exactly 16 bytes: 24 chars, last two
/// `==`, the rest in the standard alphabet (RFC 4648 §4). We never
/// decode it — §4.2.2 only concatenates the encoded form.
fn key_is_b64_16_bytes(key: &str) -> bool {
let b = key.as_bytes();
b.len() == 24
&& b[22] == b'='
&& b[23] == b'='
&& b[..22]
.iter()
.all(|&c| c.is_ascii_alphanumeric() || c == b'+' || c == b'/')
}
/// RFC 4648 base64 *encoding* (standard alphabet, padded). Encode-only
/// and not cryptographic — kept in-tree by design; do not grow a decoder
/// here without a design conversation (decoding is where the parsing
/// bugs live).
pub(crate) fn b64(input: &[u8]) -> String {
const ALPHA: &[u8; 64] =
b"ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+/";
let mut out = String::with_capacity(input.len().div_ceil(3) * 4);
for chunk in input.chunks(3) {
let b0 = chunk[0] as u32;
let b1 = *chunk.get(1).unwrap_or(&0) as u32;
let b2 = *chunk.get(2).unwrap_or(&0) as u32;
let n = (b0 << 16) | (b1 << 8) | b2;
out.push(ALPHA[(n >> 18) as usize & 63] as char);
out.push(ALPHA[(n >> 12) as usize & 63] as char);
out.push(if chunk.len() > 1 { ALPHA[(n >> 6) as usize & 63] as char } else { '=' });
out.push(if chunk.len() > 2 { ALPHA[n as usize & 63] as char } else { '=' });
}
out
}
#[cfg(test)]
mod tests {
use super::*;
use crate::conn::HeaderMap;
// ----- base64: RFC 4648 §10 test vectors. -----
#[test]
fn b64_rfc4648_vectors() {
assert_eq!(b64(b""), "");
assert_eq!(b64(b"f"), "Zg==");
assert_eq!(b64(b"fo"), "Zm8=");
assert_eq!(b64(b"foo"), "Zm9v");
assert_eq!(b64(b"foob"), "Zm9vYg==");
assert_eq!(b64(b"fooba"), "Zm9vYmE=");
assert_eq!(b64(b"foobar"), "Zm9vYmFy");
}
// ----- accept key: the RFC 6455 §1.3 worked example. -----
#[test]
fn accept_key_rfc6455_example() {
assert_eq!(
accept_key("dGhlIHNhbXBsZSBub25jZQ=="),
"s3pPLMBiTxaQ9kYGzzhZRbK+xOo="
);
}
fn ws_conn() -> Conn {
let mut c = Conn::new();
c.method = Method::Get;
c.version = HttpVersion::Http11;
let mut h = HeaderMap::new();
h.append("host", "example.com");
h.append("upgrade", "websocket");
h.append("connection", "Upgrade");
h.append("sec-websocket-key", "dGhlIHNhbXBsZSBub25jZQ==");
h.append("sec-websocket-version", "13");
c.headers = h;
c
}
#[test]
fn validate_happy_path() {
assert_eq!(
validate(&ws_conn()).unwrap(),
"s3pPLMBiTxaQ9kYGzzhZRbK+xOo="
);
}
#[test]
fn validate_rejects_post() {
let mut c = ws_conn();
c.method = Method::Post;
assert_eq!(validate(&c), Err(Rejection::NotGet));
}
#[test]
fn validate_rejects_http10() {
let mut c = ws_conn();
c.version = HttpVersion::Http10;
assert_eq!(validate(&c), Err(Rejection::BadHttpVersion));
}
#[test]
fn validate_requires_host() {
let mut c = ws_conn();
let mut h = HeaderMap::new();
for (n, v) in c.headers.iter() {
if n != "host" {
h.append(n, v);
}
}
c.headers = h;
assert_eq!(validate(&c), Err(Rejection::MissingHost));
}
#[test]
fn validate_connection_token_in_list() {
let mut c = ws_conn();
c.headers.set("connection", "keep-alive, Upgrade");
assert!(validate(&c).is_ok());
}
#[test]
fn validate_rejects_wrong_ws_version() {
let mut c = ws_conn();
c.headers.set("sec-websocket-version", "8");
assert_eq!(validate(&c), Err(Rejection::UnsupportedVersion));
assert_eq!(Rejection::UnsupportedVersion.status(), 426);
}
#[test]
fn validate_rejects_bad_keys() {
for bad in [
"tooshort==",
"dGhlIHNhbXBsZSBub25jZQ=!", // bad padding char
"dGhlIHNhbXBsZSBub25jZSE=", // single pad = 17 bytes, not 16
"dGhlIHNhbXBsZSBub25jZQQQ", // unpadded 24 chars (18 bytes)
"",
] {
let mut c = ws_conn();
c.headers.set("sec-websocket-key", bad);
assert_eq!(validate(&c), Err(Rejection::BadKey), "key: {bad:?}");
}
}
#[test]
fn is_upgrade_request_predicate() {
assert!(is_upgrade_request(&ws_conn()));
let mut c = ws_conn();
c.headers.set("upgrade", "h2c");
assert!(!is_upgrade_request(&c));
assert!(!is_upgrade_request(&Conn::new()));
}
}
+35
View File
@@ -0,0 +1,35 @@
//! WebSocket support (RFC 6455).
//!
//! v0.4 chunk 1: the upgrade handshake. A handler that wants to speak
//! WebSocket calls [`Conn::upgrade`](crate::Conn::upgrade); on a valid
//! handshake the connection actor writes the `101 Switching Protocols`
//! response and leaves the HTTP request loop. (Until the duplex wiring
//! lands — chunk 3 — leaving the loop closes the socket; the 101 itself
//! is correct and tested.)
//!
//! Dependency note: SHA-1 comes from `sha1_smol` (zero transitive deps),
//! spending urus dependency slot #3 — agreed in the v0.4 design pass over
//! vendoring. Base64 here is the 20-byte *encode* direction only and is
//! not cryptographic, so that stays in-tree (`handshake::b64`).
pub mod duplex;
pub mod frame;
pub mod handshake;
pub use duplex::{WsClosed, WsHandler, WsSender};
pub use frame::{Frame, FrameError, Message, Opcode};
pub use handshake::Rejection;
/// Payload carried on a [`Conn`](crate::Conn) whose handler accepted a
/// WebSocket upgrade: the boxed [`WsHandler`] the duplex loop will drive
/// (the chunk-3 growth this was reserved for). Still opaque outside the
/// crate.
pub struct WsUpgrade {
pub(crate) handler: Box<dyn WsHandler>,
}
impl std::fmt::Debug for WsUpgrade {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
f.write_str("WsUpgrade { handler: .. }")
}
}
+1715 -2
View File
File diff suppressed because it is too large Load Diff