A file-based way to set tuning knobs without recompiling. urus is a
library, so it never presumes a config path or reads the environment — the
binary hands the text in:
- Config::with_toml_str(&str) -> Result<Config, ConfigError>: sparse overlay
onto an existing Config (built with the addr the binary chose). Only keys
present are applied; durations are integer seconds; unknown keys are a
hard error (deny_unknown_fields) so a typo is loud, not a silent no-op.
- Scope: the four slowloris knobs (head_timeout_secs, body_timeout_secs,
body_burst_bytes, body_stall_timeout_secs). Migrating the rest of Config
into the file is a separate, additive job — TomlOverrides just grows.
- Feature `config-file = ["dep:serde", "dep:toml"]`; the optional serde dep
gains the derive feature. The default build is unchanged (deps + code are
all gated).
- examples/serve_toml.rs (required-features = ["config-file"]): a `--config
PATH` demo with no presumed default location. plain_serve and its env
vars are left untouched.
Tests (feature-gated): empty keeps defaults, partial overrides only named,
full overrides all, unknown key errors, malformed errors. 84 lib with the
feature / 79 without; clippy --lib clean both ways; e2e smoke serves 200
from a file and rejects an unknown key loudly.
The body budget from the prior commit is a generous absolute cap; alone it
just hands a body-phase slowloris a bigger window. Add a stall gate under
that cap that distinguishes a slowloris trickle from a slow-but-legit
client by requiring BURSTS, not a mere average rate:
- BodyStallGate: each body read is bounded by min(body cap, mark + stall).
The stall mark advances only when body_burst_bytes accumulate since the
last advance, so a steady sub-burst trickle never moves it and is evicted
at ~body_stall_timeout, while a bursty slow client keeps resetting it.
- Two words of state; one add + one compare per read. Raw socket bytes are
counted, so chunked framing counts and an MSS-fragmented burst still
accumulates. Reuses the existing read_some deadline plumbing.
- Wired into read_body (fixed CL) and read_chunked_body (via fill_to, the
single choke point all chunked reads pass through).
- New knobs body_burst_bytes (4 KiB) + body_stall_timeout (20s); effective
floor ~205 B/s enforced in bursts.
Tests: body_smooth_trickle_evicted_at_stall_timeout (chunked; active
sub-burst trickle evicted at ~stall while the cap is far away) and
bursty_slow_body_survives_stall_gate (fixed CL; real bursts with sub-stall
gaps complete intact). 79 lib + 45 integration green; clippy --lib clean.
The single request_timeout covered head + body under one wall clock, so a
slow-but-legit body upload (e.g. a trickling cellular IoT client) was
judged by the short head deadline and killed mid-body. Split into:
- head_timeout (default 30s): first byte -> full head parse; the classic
slowloris surface, kept short.
- body_timeout (default 300s): head parse -> full body; an absolute cap
sized for slow links, anchored independently once the head has parsed.
read_head no longer returns a shared deadline; run_connection anchors the
body deadline itself. ReadHeadErr::RequestTimeout -> HeadTimeout. Config
and ConnLimits gain body_timeout; request_timeout renamed to head_timeout
(breaking, but this axis is unreleased).
Tests: slow_body_outlives_head_timeout (positive: body survives past the
head clock), fixed_/chunked_body_stall_killed_at_body_timeout (body cap
still bites), slowloris_partial_head_killed_at_head_timeout (head clock
unchanged). 79 lib + 43 integration green; clippy --lib clean.
Config gained a max_actors: Option<usize> (None = smarm's DEFAULT_MAX_ACTORS
of 16_384). serve_with now applies it to the smarm runtime config, so the
per-connection actor slab can be sized to the deployment's peak concurrent
connections. Since each connection is one actor, the slab was the hard cap
on concurrent connections (previously an un-raisable 16_384) regardless of
RAM/fds; slots are ~256 B so raising it is cheap next to per-conn stacks.
Verified on the GPU box: default caps at 16_383 held connections; with
max_actors raised, a paced ramp holds 100_000 concurrent slow-header
connections at 12.4 KB RSS / 2 VMAs each (1.29 GB total) on one pinned core.
The TE arm set chunked whenever the token appeared anywhere in the value,
so 'chunked, gzip' (chunked not final) was accepted and an unknown coding
like 'bogus' was treated as no-body (h1spec #18/#19 -> 404). Collect the
ordered coding list across all TE headers and decide post-loop: TE on
HTTP/1.0 or TE+Content-Length -> 400 (the CL check now covers ANY TE, not
just chunked, closing the old TE:unknown + CL smuggling gap); chunked
present but not final -> 400; any coding other than chunked -> 501 via a
new UnknownTransferCoding variant (emit_error_response gains the 501 arm);
only a sole final chunked sets the flag. Tests cover each branch.
Note: the dead Unsupported/411 variant is left as-is (separate cleanup).
The content-length arm ran content_length = Some(parse) per header, so a
second Content-Length silently overwrote the first with no conflict check
(CL.CL request smuggling; h1spec #21 -> 404 instead of 400). Count
occurrences and reject any duplicate post-loop, strictly (even equal
values), reusing BadContentLength (400). A single value is still required
to be one decimal integer, so a comma-list or non-numeric keeps failing at
parse as before. Tests: differing dup, equal dup, single-CL regression.
parse_head never inspected Host, so a missing (HTTP/1.1), duplicate, or
syntactically invalid Host all passed through to the router (h1spec #8/#9/
#10 -> 404 instead of 400). Add per-header validity (RFC 3986 host[:port]
charset via valid_host) plus a post-loop presence/uniqueness check: 1.1
MUST carry exactly one valid Host; 1.0 may omit it but a duplicate/invalid
one is still 400. Unit matrix mirrors the three h1spec cases with reg-name/
port/IPv6-literal positive controls.
RFC 019 lands upstream: per-actor stacks, park-path shrink, recycle zap,
SIGSEGV diagnostics, introspect surface. No urus code changes required —
E1 interleaved A/B on the box shows every ka cell within +0.3..+2.9% of
the v0.5.0 pin (t8-c4 close-mode control is bistable either side; see
smarm v0.6.0 release notes). Tag must exist upstream before this builds:
push smarm master + v0.6.0 first.