17 KiB
Urus v1 Benchmark Plan
Goal: produce defensible HTTP/1.1 numbers for urus v0.1, framed by hyper (Rust async reference) above and cowboy (BEAM reference) below. Per the urus spec: target is within 2x of hyper; significantly outperform Cowboy 2 on HTTP/1.1.
The benchmark is layered, smallest-first, so we can see what costs what: the plug indirection, the connection-actor context switch, the pipeline middleware stack, and finally the DB layer.
1. Design considerations (read first)
1.1 Don't bench the DB
A "realistic" load that hits a DB on every request will almost certainly bench the DB instead of the HTTP stack. SQLite single-writer contention is the most likely culprit. The plan is:
- Tier the scenarios. The cheapest scenario does no DB work so we get a clean read on the HTTP stack itself. Each subsequent tier adds one source of work and is compared against the previous tier — the delta is the cost of that layer.
- Instrument the handler with a histogram of "handler-inside time"
(microseconds spent between the start of the handler and the moment it
returns the
Conn). If at high concurrency the handler-inside time dominates request latency, we're benching the DB / the in-memory actor / the auth check, not the framework. Report this metric next to RPS. - For the DB tier, use SQLite in WAL mode, with one dedicated writer
actor and N reader connections (or a single-conn reader pool with
BEGIN IMMEDIATEsemantics). This matches urus's actor-style data ownership and is fair to all three frameworks. - Cross-check by also running each scenario with the in-memory store from
examples/crud.rs. If urus's in-memory and SQLite numbers are close at low write ratios, the DB is not the bottleneck; if SQLite is much lower, it is.
1.2 wrk2, not wrk
wrk (Glenn's original) is closed-loop: it sends a new request only after
the previous one returns. Tail latency in that model is meaningless
("coordinated omission" — the load generator stalls when the server stalls,
so it can't measure the stall). We use wrk2 with a fixed rate -R.
For a smoke-test top-end RPS we'll also run plain wrk at maximum offered
load to find each server's saturation point. That number is then used to
pick the target rate for wrk2 — typically 60–80% of saturation, where tail
latency is the actual measurement of interest.
Cross-checks: oha (open-loop, HDR-histogram output) and rewrk.
1.3 Equivalent stacks across servers
The bench is unfair unless the middleware stack is materially the same on all three. Spec:
- Request logger (writes to a bounded ring buffer in memory; no stdout, which would skew everything).
- Request ID middleware (generates a short ID, sets a response header).
- Bearer token auth (constant-time compare against a fixed valid token).
- JSON request/response on the write paths.
- Router with at least one literal route, one parameterised route, and one POST route.
Cowboy's idiomatic equivalent is cowboy_router + handler modules + an
onrequest hook for auth. Hyper's idiomatic equivalent is axum: it
exposes the closest analogue to urus's pipeline (tower::Service + layers
- Router). We bench against axum, and note that "raw hyper" would be 10–20% faster but is not a fair comparison to a routing+middleware stack.
1.4 Two-process discipline
Server and load generator run on the same host but with pinned, disjoint CPU sets to avoid each starving the other. On a 16-core box:
- Server:
taskset -c 0-7 - Loadgen:
taskset -c 8-15
This matters more than it sounds: with both on all cores, you measure the loadgen's scheduling as much as the server's.
1.5 Warm-up
Each run is 30s warm-up + 60s measured. The warm-up is necessary because:
- TCP window scaling needs time to ramp.
- Page-fault costs for the read buffers happen once.
- BEAM in particular has a JIT (BeamAsm) that benefits from warm-up.
Save raw HDR histograms (wrk2 emits one); aggregate p50/p95/p99/p99.9/max.
1.6 What we don't measure (yet)
- TLS: deferred per urus spec (v2+).
- HTTP/2: deferred per urus spec (v2).
- WebSocket: deferred (v3).
- Cold-start / spawn cost: orthogonal to the steady-state question this bench answers.
2. Scenarios
Four scenarios, ordered cheapest to most realistic. Each is run against each framework. Each measures the delta over the previous one.
S1: GET /ping — naked hello world
- Pipeline: a single handler. No logger, no auth, no router.
- Response:
200 OK, bodypong(4 bytes),content-type: text/plain. - Measures: raw protocol throughput. The floor on framework overhead.
- Connection: keep-alive (default in HTTP/1.1).
This is the number to compare against hyper's hello benchmark. The urus
spec target is "within 2x". The plug pipeline isn't really exercised here —
that's deliberate. S1 isolates the connection actor + parser + serialiser.
S2: GET /api/v1/users/:id — router + middleware, no DB
- Pipeline:
logger → request_id → auth → router. - Route:
GET /api/v1/users/:idreturns a hard-codedUserstruct serialised to JSON. The:idis echoed into the response so the serialiser cannot const-fold it away. - Bearer token check: a known good token sent in the
Authorizationheader. Mismatch → 401 + halt. - Measures: the cost of the pipeline (4 plug hops) and the router lookup. This is where urus's spec budget ("one vtable dispatch per plug per request") gets tested.
- Connection: keep-alive.
Expected: lower RPS than S1 by a measurable but non-catastrophic margin (a few % to maybe 20%). If S2 is half of S1, something is allocating per request.
S3: POST /api/v1/users + GET /api/v1/users/:id — in-memory store
- Pipeline: same as S2.
- Routes:
POST /api/v1/userswith a small JSON body ({"name":"alice","email":"a@x"}).GET /api/v1/users/:id.GET /api/v1/users(list, capped at 100).
- Store: in-memory actor (urus: the actor from
examples/crud.rs; hyper:tokio::sync::Mutex<HashMap>; cowboy: agen_server). No persistence. - Traffic mix: 80% GET single, 15% GET list, 5% POST. This is the typical web-app shape and exposes the message-passing path on urus without making writes the bottleneck.
- Measures: actor channel round-trip vs. mutex contention. Urus does a channel send-recv to its store; hyper holds a brief mutex. The bench shows the cost of message passing vs. the cost of contention as concurrency rises.
- Connection: keep-alive.
S4: Same as S3, with SQLite (WAL)
- Same routes and mix as S3.
- Backing store:
- urus: one writer actor owning a write connection; reads go through a
pool of N=
min(cpus, 4)read-only connections (each held by its own actor, requested via a worker pool). AllBEGIN IMMEDIATE. - hyper/axum:
sqlx::SqlitePoolwith WAL + busy_timeout. - cowboy:
esqlitewith the same WAL config; pool viapoolboy.
- urus: one writer actor owning a write connection; reads go through a
pool of N=
- Measures: realism. We expect everyone to be much lower here. If the three framework numbers converge, the DB is the bottleneck and the framework comparison can't be made from this scenario alone — that's the answer S4 might give us, and it's still useful information.
In all four scenarios the handler instruments
handler_microseconds_histogram so we can detect DB domination.
3. The middleware stack (spec for all three implementations)
Each implementation MUST expose these middlewares with materially identical
behaviour. The wire-level outputs (status codes, body shapes, headers) MUST
match — verified by a small curl-based smoke test before each bench run.
3.1 Logger
- Captures:
method,path,status,elapsed_us. - Writes into a bounded ring buffer in memory (e.g.
crossbeam::ArrayQueue<LogEntry>for Rust, an ETS table for cowboy with capped size). - Does not touch stdout, stderr, or the filesystem during the run.
- The buffer is drained on shutdown to a file for inspection.
3.2 Request ID
- 12-byte random ID, base32-encoded.
- Set as response header
x-request-id. - Generated from a thread-local fast RNG (avoid the global
rand::thread_rng()mutex on Rust; use SplitMix64 seeded once per actor/task).
3.3 Auth
- Reads
Authorization: Bearer <token>from request headers. - Compares against a fixed token (configured at startup) using
constant-time compare (
subtle::ConstantTimeEqon Rust;crypto:hash-based on BEAM is fine — the constant-time property is pedagogical here, not a security claim). - On mismatch:
401, body{"error":"unauthorized"}, halt the pipeline. - The bench load generator always sends the valid token (we are not benching the failure path).
3.4 Router
- Methods: GET, POST, PUT, DELETE.
- Patterns:
/api/v1/users,/api/v1/users/:id,/ping. - 404 on no match; 405 on method-mismatch (urus already does this).
3.5 JSON
serde_jsonfor Rust,jsxorjsonefor BEAM.- Bodies are small: well under 1 KiB for the user payload.
4. Queries (load-gen scripts)
wrk2 takes a Lua script. Below are the four scripts. They share a single
common header script (common.lua) for token + content-type.
common.lua
-- Shared header config; included by every script below.
local M = {}
M.token = os.getenv("BEARER") or "test-token-aaaaaaaaaaaaaaaaaaaa"
M.headers = {
["Authorization"] = "Bearer " .. M.token,
["Connection"] = "keep-alive",
["Host"] = os.getenv("HOST") or "127.0.0.1:8080",
}
return M
s1_ping.lua — S1 ping, no auth header (it's a naked endpoint)
wrk.method = "GET"
wrk.path = "/ping"
wrk.headers["Connection"] = "keep-alive"
Invocation:
# Saturation probe (closed-loop wrk):
wrk -t8 -c256 -d60s --latency http://127.0.0.1:8080/ping
# Latency-honest measurement (wrk2, target rate from probe):
wrk2 -t8 -c256 -d60s -R200000 --latency -s s1_ping.lua \
http://127.0.0.1:8080
s2_user_get.lua — single-route GET with middleware
local common = require("common")
wrk.method = "GET"
wrk.headers = common.headers
-- Spread :id across 10_000 values so per-id caches can't trivially win.
math.randomseed(os.time() + os.getpid())
request = function()
local id = math.random(1, 10000)
return wrk.format(nil, "/api/v1/users/" .. id)
end
Invocation:
wrk2 -t8 -c256 -d60s -R150000 --latency -s s2_user_get.lua \
http://127.0.0.1:8080
s3_mixed.lua — 80% GET-one / 15% GET-list / 5% POST
local common = require("common")
wrk.headers = common.headers
wrk.headers["Content-Type"] = "application/json"
math.randomseed(os.time() + os.getpid())
local body_pool = {}
for i = 1, 64 do
body_pool[i] = string.format(
[[{"name":"user%d","email":"u%d@example.test"}]], i, i)
end
request = function()
local r = math.random()
if r < 0.80 then
return wrk.format("GET", "/api/v1/users/" .. math.random(1, 10000))
elseif r < 0.95 then
return wrk.format("GET", "/api/v1/users")
else
local body = body_pool[math.random(1, #body_pool)]
return wrk.format("POST", "/api/v1/users", nil, body)
end
end
Invocation: same shape as s2, lower -R (writes are slower).
wrk2 -t8 -c512 -d60s -R80000 --latency -s s3_mixed.lua \
http://127.0.0.1:8080
s4_mixed_sqlite.lua
Identical to s3_mixed.lua. Only the server's backing store differs.
wrk2 -t8 -c512 -d60s -R20000 --latency -s s4_mixed_sqlite.lua \
http://127.0.0.1:8080
(-R20000 is a guess; the real number comes from a wrk probe of each
framework first.)
5. Implementation plan
Three sibling crates / projects under a urus-bench/ workspace, each
exposing the same routes and middleware:
urus-bench/
├── urus-server/ # Rust, urus
├── axum-server/ # Rust, hyper via axum
└── cowboy-server/ # Erlang/OTP, cowboy 2
Plus:
urus-bench/
├── loadgen/ # wrk2 scripts (common.lua, s1..s4 above)
├── runner.sh # orchestrates pin-cpu, warm-up, measurement, save HDR
└── results/ # one subdir per run: ${date}-${server}-${scenario}/
The runner:
- Boots the target server pinned to cores 0-7.
curl-smoke-tests every route.- Probes saturation with
wrkat a high-c. - Reads back the saturation RPS and chooses 70% of it for wrk2.
- Runs wrk2 with HDR output (
--latency), captures stdout + HDR file. - Greps server process for RSS (
/proc/$pid/status) and CPU (pidstat -p $pid 1) during the run. - Saves everything under
results/${run_id}/. - Kills server, sleeps 5s, moves on.
5.1 urus-server
Port examples/crud.rs with:
- Routes adjusted to
/api/v1/users{,/:id}and/ping. - Logger replaced with a ring-buffer-backed plug.
- New
request_idandauthplugs. - Switchable backing store:
--store=memory(the existing actor) or--store=sqlite(a new actor wrappingrusqlitein WAL).
For SQLite on urus: spawn one writer actor + 4 reader actors, each owning
its own rusqlite::Connection. Handlers route reads to a free reader via
a worker-pool channel, writes to the writer. Pool selection is FIFO over a
single shared Receiver<RequestForReader> — readers race to receive, which
is exactly what smarm's MPSC gives for free.
5.2 axum-server
Standard axum + tokio multi-thread. Same routes. Layers: Logger,
SetRequestIdLayer, RequireAuthorizationLayer. State: sqlx::SqlitePool
for S4, Arc<RwLock<HashMap>> for S3.
5.3 cowboy-server
Standard cowboy 2 with cowboy_router. Handlers as cowboy_handler
modules. Auth as an onrequest-style middleware. Store: a gen_server
for S3, esqlite + poolboy for S4.
Build with rebar3, run with +S 8 +sbt db (8 schedulers, scheduler-thread
binding) under the same taskset -c 0-7.
6. Metrics and reporting
Per (server × scenario) run, capture and report:
| Metric | How |
|---|---|
RPS (saturation, wrk) |
wrk -c256 -d60s — closed-loop peak |
RPS (sustained, wrk2) |
wrk2 -R$target -d60s at 70% of saturation |
| p50 / p95 / p99 / p99.9 / max latency | HDR histogram from wrk2 |
| Server CPU% | pidstat -p $pid 1 60 → mean over run |
| Server RSS | /proc/$pid/status VmRSS sampled at 1 Hz |
| Handler-inside p99 (µs) | server-side histogram, emitted to ring buffer |
| Dropped/errored requests | wrk2's "Non-2xx or 3xx responses" |
The headline table:
| urus | axum | cowboy | |
|---|---|---|---|
| S1 RPS (sat) | |||
| S1 p99 @ 70% | |||
| S2 RPS (sat) | |||
| S2 p99 @ 70% | |||
| S3 RPS (sat) | |||
| S3 p99 @ 70% | |||
| S4 RPS (sat) | |||
| S4 p99 @ 70% | |||
| urus / axum (S1) | — | — | |
| urus / axum (S2) | — | — |
Pass/fail per the urus spec:
- S1: urus RPS ≥ 50% of axum RPS (within 2x).
- S2: urus RPS ≥ 50% of axum RPS.
- All scenarios: urus RPS > cowboy RPS by a clear margin (native expectation).
- All scenarios: urus p99 < 2× axum p99 at the same offered rate.
- Handler-inside p99 ≪ end-to-end p99 in S1/S2 (proves we're benching the framework, not the handler).
7. Order of operations (suggested)
- Build
urus-serverwith S1 (/ping) only. Run S1 against it. Make sure the saturation number is plausible (tens to hundreds of kRPS on a decent box). - Build
axum-serverwith S1. Run S1. Sanity-check the ratio against public hyper hello-world numbers (axum should land near hyper). - Add the middleware stack to both, run S2, capture deltas.
- Add the in-memory store, run S3.
- Bring up
cowboy-serverand bring all three up to S3. - Add SQLite to all three; run S4.
- Write the headline table; produce HDR plots from the saved files (gnuplot or a tiny Python script — wrk2 ships an HDR-plot helper).
If at any point a tier's results are dominated by a non-framework cost (handler-inside p99 ≥ ~50% of end-to-end p99), stop and instrument before moving on. That's the whole point of tiering.
8. Open questions deliberately left for later
- Whether to also bench under
tc-induced packet loss / RTT. The current setup uses loopback, which is a generous environment. A loopback win is necessary but not sufficient for "fast on the real network". - Whether to bench HTTP/1.1 pipelining (multiple in-flight requests per connection without waiting). urus claims it's a non-goal-for-now, and axum/hyper don't pipeline either, so the comparison is moot — but it's worth a note.
- Whether to add a "slow handler" scenario (one route that sleeps 50ms) to demonstrate urus's blocking-handler-without-thread-pool claim. That's more of a qualitative demo than a throughput bench, but it's the kind of thing reviewers ask about.