Files
urus-benches/urus-bench-spec.md
2026-05-26 23:22:32 +02:00

17 KiB
Raw Permalink Blame History

Urus v1 Benchmark Plan

Goal: produce defensible HTTP/1.1 numbers for urus v0.1, framed by hyper (Rust async reference) above and cowboy (BEAM reference) below. Per the urus spec: target is within 2x of hyper; significantly outperform Cowboy 2 on HTTP/1.1.

The benchmark is layered, smallest-first, so we can see what costs what: the plug indirection, the connection-actor context switch, the pipeline middleware stack, and finally the DB layer.


1. Design considerations (read first)

1.1 Don't bench the DB

A "realistic" load that hits a DB on every request will almost certainly bench the DB instead of the HTTP stack. SQLite single-writer contention is the most likely culprit. The plan is:

  • Tier the scenarios. The cheapest scenario does no DB work so we get a clean read on the HTTP stack itself. Each subsequent tier adds one source of work and is compared against the previous tier — the delta is the cost of that layer.
  • Instrument the handler with a histogram of "handler-inside time" (microseconds spent between the start of the handler and the moment it returns the Conn). If at high concurrency the handler-inside time dominates request latency, we're benching the DB / the in-memory actor / the auth check, not the framework. Report this metric next to RPS.
  • For the DB tier, use SQLite in WAL mode, with one dedicated writer actor and N reader connections (or a single-conn reader pool with BEGIN IMMEDIATE semantics). This matches urus's actor-style data ownership and is fair to all three frameworks.
  • Cross-check by also running each scenario with the in-memory store from examples/crud.rs. If urus's in-memory and SQLite numbers are close at low write ratios, the DB is not the bottleneck; if SQLite is much lower, it is.

1.2 wrk2, not wrk

wrk (Glenn's original) is closed-loop: it sends a new request only after the previous one returns. Tail latency in that model is meaningless ("coordinated omission" — the load generator stalls when the server stalls, so it can't measure the stall). We use wrk2 with a fixed rate -R.

For a smoke-test top-end RPS we'll also run plain wrk at maximum offered load to find each server's saturation point. That number is then used to pick the target rate for wrk2 — typically 6080% of saturation, where tail latency is the actual measurement of interest.

Cross-checks: oha (open-loop, HDR-histogram output) and rewrk.

1.3 Equivalent stacks across servers

The bench is unfair unless the middleware stack is materially the same on all three. Spec:

  • Request logger (writes to a bounded ring buffer in memory; no stdout, which would skew everything).
  • Request ID middleware (generates a short ID, sets a response header).
  • Bearer token auth (constant-time compare against a fixed valid token).
  • JSON request/response on the write paths.
  • Router with at least one literal route, one parameterised route, and one POST route.

Cowboy's idiomatic equivalent is cowboy_router + handler modules + an onrequest hook for auth. Hyper's idiomatic equivalent is axum: it exposes the closest analogue to urus's pipeline (tower::Service + layers

  • Router). We bench against axum, and note that "raw hyper" would be 1020% faster but is not a fair comparison to a routing+middleware stack.

1.4 Two-process discipline

Server and load generator run on the same host but with pinned, disjoint CPU sets to avoid each starving the other. On a 16-core box:

  • Server: taskset -c 0-7
  • Loadgen: taskset -c 8-15

This matters more than it sounds: with both on all cores, you measure the loadgen's scheduling as much as the server's.

1.5 Warm-up

Each run is 30s warm-up + 60s measured. The warm-up is necessary because:

  • TCP window scaling needs time to ramp.
  • Page-fault costs for the read buffers happen once.
  • BEAM in particular has a JIT (BeamAsm) that benefits from warm-up.

Save raw HDR histograms (wrk2 emits one); aggregate p50/p95/p99/p99.9/max.

1.6 What we don't measure (yet)

  • TLS: deferred per urus spec (v2+).
  • HTTP/2: deferred per urus spec (v2).
  • WebSocket: deferred (v3).
  • Cold-start / spawn cost: orthogonal to the steady-state question this bench answers.

2. Scenarios

Four scenarios, ordered cheapest to most realistic. Each is run against each framework. Each measures the delta over the previous one.

S1: GET /ping — naked hello world

  • Pipeline: a single handler. No logger, no auth, no router.
  • Response: 200 OK, body pong (4 bytes), content-type: text/plain.
  • Measures: raw protocol throughput. The floor on framework overhead.
  • Connection: keep-alive (default in HTTP/1.1).

This is the number to compare against hyper's hello benchmark. The urus spec target is "within 2x". The plug pipeline isn't really exercised here — that's deliberate. S1 isolates the connection actor + parser + serialiser.

S2: GET /api/v1/users/:id — router + middleware, no DB

  • Pipeline: logger → request_id → auth → router.
  • Route: GET /api/v1/users/:id returns a hard-coded User struct serialised to JSON. The :id is echoed into the response so the serialiser cannot const-fold it away.
  • Bearer token check: a known good token sent in the Authorization header. Mismatch → 401 + halt.
  • Measures: the cost of the pipeline (4 plug hops) and the router lookup. This is where urus's spec budget ("one vtable dispatch per plug per request") gets tested.
  • Connection: keep-alive.

Expected: lower RPS than S1 by a measurable but non-catastrophic margin (a few % to maybe 20%). If S2 is half of S1, something is allocating per request.

S3: POST /api/v1/users + GET /api/v1/users/:id — in-memory store

  • Pipeline: same as S2.
  • Routes:
    • POST /api/v1/users with a small JSON body ({"name":"alice","email":"a@x"}).
    • GET /api/v1/users/:id.
    • GET /api/v1/users (list, capped at 100).
  • Store: in-memory actor (urus: the actor from examples/crud.rs; hyper: tokio::sync::Mutex<HashMap>; cowboy: a gen_server). No persistence.
  • Traffic mix: 80% GET single, 15% GET list, 5% POST. This is the typical web-app shape and exposes the message-passing path on urus without making writes the bottleneck.
  • Measures: actor channel round-trip vs. mutex contention. Urus does a channel send-recv to its store; hyper holds a brief mutex. The bench shows the cost of message passing vs. the cost of contention as concurrency rises.
  • Connection: keep-alive.

S4: Same as S3, with SQLite (WAL)

  • Same routes and mix as S3.
  • Backing store:
    • urus: one writer actor owning a write connection; reads go through a pool of N=min(cpus, 4) read-only connections (each held by its own actor, requested via a worker pool). All BEGIN IMMEDIATE.
    • hyper/axum: sqlx::SqlitePool with WAL + busy_timeout.
    • cowboy: esqlite with the same WAL config; pool via poolboy.
  • Measures: realism. We expect everyone to be much lower here. If the three framework numbers converge, the DB is the bottleneck and the framework comparison can't be made from this scenario alone — that's the answer S4 might give us, and it's still useful information.

In all four scenarios the handler instruments handler_microseconds_histogram so we can detect DB domination.


3. The middleware stack (spec for all three implementations)

Each implementation MUST expose these middlewares with materially identical behaviour. The wire-level outputs (status codes, body shapes, headers) MUST match — verified by a small curl-based smoke test before each bench run.

3.1 Logger

  • Captures: method, path, status, elapsed_us.
  • Writes into a bounded ring buffer in memory (e.g. crossbeam::ArrayQueue<LogEntry> for Rust, an ETS table for cowboy with capped size).
  • Does not touch stdout, stderr, or the filesystem during the run.
  • The buffer is drained on shutdown to a file for inspection.

3.2 Request ID

  • 12-byte random ID, base32-encoded.
  • Set as response header x-request-id.
  • Generated from a thread-local fast RNG (avoid the global rand::thread_rng() mutex on Rust; use SplitMix64 seeded once per actor/task).

3.3 Auth

  • Reads Authorization: Bearer <token> from request headers.
  • Compares against a fixed token (configured at startup) using constant-time compare (subtle::ConstantTimeEq on Rust; crypto:hash-based on BEAM is fine — the constant-time property is pedagogical here, not a security claim).
  • On mismatch: 401, body {"error":"unauthorized"}, halt the pipeline.
  • The bench load generator always sends the valid token (we are not benching the failure path).

3.4 Router

  • Methods: GET, POST, PUT, DELETE.
  • Patterns: /api/v1/users, /api/v1/users/:id, /ping.
  • 404 on no match; 405 on method-mismatch (urus already does this).

3.5 JSON

  • serde_json for Rust, jsx or jsone for BEAM.
  • Bodies are small: well under 1 KiB for the user payload.

4. Queries (load-gen scripts)

wrk2 takes a Lua script. Below are the four scripts. They share a single common header script (common.lua) for token + content-type.

common.lua

-- Shared header config; included by every script below.
local M = {}
M.token = os.getenv("BEARER") or "test-token-aaaaaaaaaaaaaaaaaaaa"
M.headers = {
  ["Authorization"] = "Bearer " .. M.token,
  ["Connection"]    = "keep-alive",
  ["Host"]          = os.getenv("HOST") or "127.0.0.1:8080",
}
return M

s1_ping.lua — S1 ping, no auth header (it's a naked endpoint)

wrk.method = "GET"
wrk.path   = "/ping"
wrk.headers["Connection"] = "keep-alive"

Invocation:

# Saturation probe (closed-loop wrk):
wrk -t8 -c256 -d60s --latency http://127.0.0.1:8080/ping

# Latency-honest measurement (wrk2, target rate from probe):
wrk2 -t8 -c256 -d60s -R200000 --latency -s s1_ping.lua \
  http://127.0.0.1:8080

s2_user_get.lua — single-route GET with middleware

local common = require("common")
wrk.method  = "GET"
wrk.headers = common.headers

-- Spread :id across 10_000 values so per-id caches can't trivially win.
math.randomseed(os.time() + os.getpid())
request = function()
  local id = math.random(1, 10000)
  return wrk.format(nil, "/api/v1/users/" .. id)
end

Invocation:

wrk2 -t8 -c256 -d60s -R150000 --latency -s s2_user_get.lua \
  http://127.0.0.1:8080

s3_mixed.lua — 80% GET-one / 15% GET-list / 5% POST

local common = require("common")
wrk.headers = common.headers
wrk.headers["Content-Type"] = "application/json"

math.randomseed(os.time() + os.getpid())
local body_pool = {}
for i = 1, 64 do
  body_pool[i] = string.format(
    [[{"name":"user%d","email":"u%d@example.test"}]], i, i)
end

request = function()
  local r = math.random()
  if r < 0.80 then
    return wrk.format("GET", "/api/v1/users/" .. math.random(1, 10000))
  elseif r < 0.95 then
    return wrk.format("GET", "/api/v1/users")
  else
    local body = body_pool[math.random(1, #body_pool)]
    return wrk.format("POST", "/api/v1/users", nil, body)
  end
end

Invocation: same shape as s2, lower -R (writes are slower).

wrk2 -t8 -c512 -d60s -R80000 --latency -s s3_mixed.lua \
  http://127.0.0.1:8080

s4_mixed_sqlite.lua

Identical to s3_mixed.lua. Only the server's backing store differs.

wrk2 -t8 -c512 -d60s -R20000 --latency -s s4_mixed_sqlite.lua \
  http://127.0.0.1:8080

(-R20000 is a guess; the real number comes from a wrk probe of each framework first.)


5. Implementation plan

Three sibling crates / projects under a urus-bench/ workspace, each exposing the same routes and middleware:

urus-bench/
├── urus-server/      # Rust, urus
├── axum-server/      # Rust, hyper via axum
└── cowboy-server/    # Erlang/OTP, cowboy 2

Plus:

urus-bench/
├── loadgen/          # wrk2 scripts (common.lua, s1..s4 above)
├── runner.sh         # orchestrates pin-cpu, warm-up, measurement, save HDR
└── results/          # one subdir per run: ${date}-${server}-${scenario}/

The runner:

  1. Boots the target server pinned to cores 0-7.
  2. curl-smoke-tests every route.
  3. Probes saturation with wrk at a high -c.
  4. Reads back the saturation RPS and chooses 70% of it for wrk2.
  5. Runs wrk2 with HDR output (--latency), captures stdout + HDR file.
  6. Greps server process for RSS (/proc/$pid/status) and CPU (pidstat -p $pid 1) during the run.
  7. Saves everything under results/${run_id}/.
  8. Kills server, sleeps 5s, moves on.

5.1 urus-server

Port examples/crud.rs with:

  • Routes adjusted to /api/v1/users{,/:id} and /ping.
  • Logger replaced with a ring-buffer-backed plug.
  • New request_id and auth plugs.
  • Switchable backing store: --store=memory (the existing actor) or --store=sqlite (a new actor wrapping rusqlite in WAL).

For SQLite on urus: spawn one writer actor + 4 reader actors, each owning its own rusqlite::Connection. Handlers route reads to a free reader via a worker-pool channel, writes to the writer. Pool selection is FIFO over a single shared Receiver<RequestForReader> — readers race to receive, which is exactly what smarm's MPSC gives for free.

5.2 axum-server

Standard axum + tokio multi-thread. Same routes. Layers: Logger, SetRequestIdLayer, RequireAuthorizationLayer. State: sqlx::SqlitePool for S4, Arc<RwLock<HashMap>> for S3.

5.3 cowboy-server

Standard cowboy 2 with cowboy_router. Handlers as cowboy_handler modules. Auth as an onrequest-style middleware. Store: a gen_server for S3, esqlite + poolboy for S4.

Build with rebar3, run with +S 8 +sbt db (8 schedulers, scheduler-thread binding) under the same taskset -c 0-7.


6. Metrics and reporting

Per (server × scenario) run, capture and report:

Metric How
RPS (saturation, wrk) wrk -c256 -d60s — closed-loop peak
RPS (sustained, wrk2) wrk2 -R$target -d60s at 70% of saturation
p50 / p95 / p99 / p99.9 / max latency HDR histogram from wrk2
Server CPU% pidstat -p $pid 1 60 → mean over run
Server RSS /proc/$pid/status VmRSS sampled at 1 Hz
Handler-inside p99 (µs) server-side histogram, emitted to ring buffer
Dropped/errored requests wrk2's "Non-2xx or 3xx responses"

The headline table:

urus axum cowboy
S1 RPS (sat)
S1 p99 @ 70%
S2 RPS (sat)
S2 p99 @ 70%
S3 RPS (sat)
S3 p99 @ 70%
S4 RPS (sat)
S4 p99 @ 70%
urus / axum (S1)
urus / axum (S2)

Pass/fail per the urus spec:

  • S1: urus RPS ≥ 50% of axum RPS (within 2x).
  • S2: urus RPS ≥ 50% of axum RPS.
  • All scenarios: urus RPS > cowboy RPS by a clear margin (native expectation).
  • All scenarios: urus p99 < 2× axum p99 at the same offered rate.
  • Handler-inside p99 ≪ end-to-end p99 in S1/S2 (proves we're benching the framework, not the handler).

7. Order of operations (suggested)

  1. Build urus-server with S1 (/ping) only. Run S1 against it. Make sure the saturation number is plausible (tens to hundreds of kRPS on a decent box).
  2. Build axum-server with S1. Run S1. Sanity-check the ratio against public hyper hello-world numbers (axum should land near hyper).
  3. Add the middleware stack to both, run S2, capture deltas.
  4. Add the in-memory store, run S3.
  5. Bring up cowboy-server and bring all three up to S3.
  6. Add SQLite to all three; run S4.
  7. Write the headline table; produce HDR plots from the saved files (gnuplot or a tiny Python script — wrk2 ships an HDR-plot helper).

If at any point a tier's results are dominated by a non-framework cost (handler-inside p99 ≥ ~50% of end-to-end p99), stop and instrument before moving on. That's the whole point of tiering.


8. Open questions deliberately left for later

  • Whether to also bench under tc-induced packet loss / RTT. The current setup uses loopback, which is a generous environment. A loopback win is necessary but not sufficient for "fast on the real network".
  • Whether to bench HTTP/1.1 pipelining (multiple in-flight requests per connection without waiting). urus claims it's a non-goal-for-now, and axum/hyper don't pipeline either, so the comparison is moot — but it's worth a note.
  • Whether to add a "slow handler" scenario (one route that sleeps 50ms) to demonstrate urus's blocking-handler-without-thread-pool claim. That's more of a qualitative demo than a throughput bench, but it's the kind of thing reviewers ask about.