This commit is contained in:
2026-05-26 23:22:32 +02:00
parent 842fca2b72
commit 730de334e0
17 changed files with 5380 additions and 0 deletions
+462
View File
@@ -0,0 +1,462 @@
# Urus v1 Benchmark Plan
> Goal: produce defensible HTTP/1.1 numbers for **urus v0.1**, framed by
> **hyper** (Rust async reference) above and **cowboy** (BEAM reference)
> below. Per the urus spec: target is **within 2x of hyper**; significantly
> outperform Cowboy 2 on HTTP/1.1.
The benchmark is layered, smallest-first, so we can see *what* costs what:
the plug indirection, the connection-actor context switch, the pipeline
middleware stack, and finally the DB layer.
---
## 1. Design considerations (read first)
### 1.1 Don't bench the DB
A "realistic" load that hits a DB on every request will almost certainly
bench the DB instead of the HTTP stack. SQLite single-writer contention is
the most likely culprit. The plan is:
- Tier the scenarios. The cheapest scenario does **no DB work** so we get a
clean read on the HTTP stack itself. Each subsequent tier adds one source
of work and is compared against the previous tier — the *delta* is the
cost of that layer.
- Instrument the handler with a histogram of "handler-inside time"
(microseconds spent between the start of the handler and the moment it
returns the `Conn`). If at high concurrency the handler-inside time
dominates request latency, we're benching the DB / the in-memory actor /
the auth check, not the framework. Report this metric *next to* RPS.
- For the DB tier, use **SQLite in WAL mode**, with one dedicated writer
actor and N reader connections (or a single-conn reader pool with `BEGIN
IMMEDIATE` semantics). This matches urus's actor-style data ownership and
is fair to all three frameworks.
- Cross-check by also running each scenario with the in-memory store from
`examples/crud.rs`. If urus's in-memory and SQLite numbers are close at
low write ratios, the DB is not the bottleneck; if SQLite is much lower,
it is.
### 1.2 wrk2, not wrk
`wrk` (Glenn's original) is closed-loop: it sends a new request only after
the previous one returns. Tail latency in that model is meaningless
("coordinated omission" — the load generator stalls when the server stalls,
so it can't *measure* the stall). We use **wrk2** with a fixed rate `-R`.
For a smoke-test top-end RPS we'll also run plain `wrk` at maximum offered
load to find each server's *saturation point*. That number is then used to
pick the target rate for wrk2 — typically 6080% of saturation, where tail
latency is the actual measurement of interest.
Cross-checks: `oha` (open-loop, HDR-histogram output) and `rewrk`.
### 1.3 Equivalent stacks across servers
The bench is unfair unless the middleware stack is materially the same on
all three. Spec:
- Request logger (writes to a bounded ring buffer in memory; **no stdout**,
which would skew everything).
- Request ID middleware (generates a short ID, sets a response header).
- Bearer token auth (constant-time compare against a fixed valid token).
- JSON request/response on the write paths.
- Router with at least one literal route, one parameterised route, and one
POST route.
Cowboy's idiomatic equivalent is `cowboy_router` + handler modules + an
`onrequest` hook for auth. Hyper's idiomatic equivalent is **axum**: it
exposes the closest analogue to urus's pipeline (`tower::Service` + layers
+ Router). We bench against axum, and note that "raw hyper" would be 1020%
faster but is not a fair comparison to a routing+middleware stack.
### 1.4 Two-process discipline
Server and load generator run on the **same host but with pinned, disjoint
CPU sets** to avoid each starving the other. On a 16-core box:
- Server: `taskset -c 0-7`
- Loadgen: `taskset -c 8-15`
This matters more than it sounds: with both on all cores, you measure the
loadgen's scheduling as much as the server's.
### 1.5 Warm-up
Each run is `30s warm-up + 60s measured`. The warm-up is necessary because:
- TCP window scaling needs time to ramp.
- Page-fault costs for the read buffers happen once.
- BEAM in particular has a JIT (BeamAsm) that benefits from warm-up.
Save raw HDR histograms (wrk2 emits one); aggregate p50/p95/p99/p99.9/max.
### 1.6 What we don't measure (yet)
- TLS: deferred per urus spec (v2+).
- HTTP/2: deferred per urus spec (v2).
- WebSocket: deferred (v3).
- Cold-start / spawn cost: orthogonal to the steady-state question this
bench answers.
---
## 2. Scenarios
Four scenarios, ordered cheapest to most realistic. Each is run against
each framework. Each measures the **delta** over the previous one.
### S1: `GET /ping` — naked hello world
- Pipeline: a single handler. No logger, no auth, no router.
- Response: `200 OK`, body `pong` (4 bytes), `content-type: text/plain`.
- Measures: raw protocol throughput. The floor on framework overhead.
- Connection: keep-alive (default in HTTP/1.1).
This is the number to compare against hyper's `hello` benchmark. The urus
spec target is "within 2x". The plug pipeline isn't really exercised here —
that's deliberate. S1 isolates the connection actor + parser + serialiser.
### S2: `GET /api/v1/users/:id` — router + middleware, no DB
- Pipeline: `logger → request_id → auth → router`.
- Route: `GET /api/v1/users/:id` returns a hard-coded `User` struct
serialised to JSON. The `:id` is echoed into the response so the
serialiser cannot const-fold it away.
- Bearer token check: a known good token sent in the `Authorization`
header. Mismatch → 401 + halt.
- Measures: the cost of the pipeline (4 plug hops) and the router lookup.
This is where urus's spec budget ("one vtable dispatch per plug per
request") gets tested.
- Connection: keep-alive.
Expected: lower RPS than S1 by a measurable but non-catastrophic margin (a
few % to maybe 20%). If S2 is half of S1, something is allocating per
request.
### S3: `POST /api/v1/users` + `GET /api/v1/users/:id` — in-memory store
- Pipeline: same as S2.
- Routes:
- `POST /api/v1/users` with a small JSON body (`{"name":"alice","email":"a@x"}`).
- `GET /api/v1/users/:id`.
- `GET /api/v1/users` (list, capped at 100).
- Store: in-memory actor (urus: the actor from `examples/crud.rs`; hyper:
`tokio::sync::Mutex<HashMap>`; cowboy: a `gen_server`). No persistence.
- Traffic mix: **80% GET single, 15% GET list, 5% POST**. This is the
typical web-app shape and exposes the message-passing path on urus
without making writes the bottleneck.
- Measures: actor channel round-trip vs. mutex contention. Urus does a
channel send-recv to its store; hyper holds a brief mutex. The bench
shows the cost of message passing vs. the cost of contention as
concurrency rises.
- Connection: keep-alive.
### S4: Same as S3, with SQLite (WAL)
- Same routes and mix as S3.
- Backing store:
- urus: one writer actor owning a write connection; reads go through a
pool of N=`min(cpus, 4)` read-only connections (each held by its own
actor, requested via a worker pool). All `BEGIN IMMEDIATE`.
- hyper/axum: `sqlx::SqlitePool` with WAL + busy_timeout.
- cowboy: `esqlite` with the same WAL config; pool via `poolboy`.
- Measures: realism. We expect everyone to be much lower here. If the
three framework numbers *converge*, the DB is the bottleneck and the
framework comparison can't be made from this scenario alone — that's
the answer S4 might give us, and it's still useful information.
In **all four scenarios** the handler instruments
`handler_microseconds_histogram` so we can detect DB domination.
---
## 3. The middleware stack (spec for all three implementations)
Each implementation MUST expose these middlewares with materially identical
behaviour. The wire-level outputs (status codes, body shapes, headers) MUST
match — verified by a small `curl`-based smoke test before each bench run.
### 3.1 Logger
- Captures: `method`, `path`, `status`, `elapsed_us`.
- Writes into a **bounded ring buffer in memory** (e.g. `crossbeam::ArrayQueue<LogEntry>` for Rust, an ETS table for cowboy with capped size).
- Does **not** touch stdout, stderr, or the filesystem during the run.
- The buffer is drained on shutdown to a file for inspection.
### 3.2 Request ID
- 12-byte random ID, base32-encoded.
- Set as response header `x-request-id`.
- Generated from a thread-local fast RNG (avoid the global `rand::thread_rng()` mutex on Rust; use SplitMix64 seeded once per actor/task).
### 3.3 Auth
- Reads `Authorization: Bearer <token>` from request headers.
- Compares against a fixed token (configured at startup) using
**constant-time compare** (`subtle::ConstantTimeEq` on Rust;
`crypto:hash`-based on BEAM is fine — the constant-time property is
pedagogical here, not a security claim).
- On mismatch: `401`, body `{"error":"unauthorized"}`, halt the pipeline.
- The bench load generator always sends the valid token (we are not
benching the failure path).
### 3.4 Router
- Methods: GET, POST, PUT, DELETE.
- Patterns: `/api/v1/users`, `/api/v1/users/:id`, `/ping`.
- 404 on no match; 405 on method-mismatch (urus already does this).
### 3.5 JSON
- `serde_json` for Rust, `jsx` or `jsone` for BEAM.
- Bodies are small: well under 1 KiB for the user payload.
---
## 4. Queries (load-gen scripts)
`wrk2` takes a Lua script. Below are the four scripts. They share a single
common header script (`common.lua`) for token + content-type.
### `common.lua`
```lua
-- Shared header config; included by every script below.
local M = {}
M.token = os.getenv("BEARER") or "test-token-aaaaaaaaaaaaaaaaaaaa"
M.headers = {
["Authorization"] = "Bearer " .. M.token,
["Connection"] = "keep-alive",
["Host"] = os.getenv("HOST") or "127.0.0.1:8080",
}
return M
```
### `s1_ping.lua` — S1 ping, no auth header (it's a naked endpoint)
```lua
wrk.method = "GET"
wrk.path = "/ping"
wrk.headers["Connection"] = "keep-alive"
```
Invocation:
```sh
# Saturation probe (closed-loop wrk):
wrk -t8 -c256 -d60s --latency http://127.0.0.1:8080/ping
# Latency-honest measurement (wrk2, target rate from probe):
wrk2 -t8 -c256 -d60s -R200000 --latency -s s1_ping.lua \
http://127.0.0.1:8080
```
### `s2_user_get.lua` — single-route GET with middleware
```lua
local common = require("common")
wrk.method = "GET"
wrk.headers = common.headers
-- Spread :id across 10_000 values so per-id caches can't trivially win.
math.randomseed(os.time() + os.getpid())
request = function()
local id = math.random(1, 10000)
return wrk.format(nil, "/api/v1/users/" .. id)
end
```
Invocation:
```sh
wrk2 -t8 -c256 -d60s -R150000 --latency -s s2_user_get.lua \
http://127.0.0.1:8080
```
### `s3_mixed.lua` — 80% GET-one / 15% GET-list / 5% POST
```lua
local common = require("common")
wrk.headers = common.headers
wrk.headers["Content-Type"] = "application/json"
math.randomseed(os.time() + os.getpid())
local body_pool = {}
for i = 1, 64 do
body_pool[i] = string.format(
[[{"name":"user%d","email":"u%d@example.test"}]], i, i)
end
request = function()
local r = math.random()
if r < 0.80 then
return wrk.format("GET", "/api/v1/users/" .. math.random(1, 10000))
elseif r < 0.95 then
return wrk.format("GET", "/api/v1/users")
else
local body = body_pool[math.random(1, #body_pool)]
return wrk.format("POST", "/api/v1/users", nil, body)
end
end
```
Invocation: same shape as s2, lower `-R` (writes are slower).
```sh
wrk2 -t8 -c512 -d60s -R80000 --latency -s s3_mixed.lua \
http://127.0.0.1:8080
```
### `s4_mixed_sqlite.lua`
Identical to `s3_mixed.lua`. Only the server's backing store differs.
```sh
wrk2 -t8 -c512 -d60s -R20000 --latency -s s4_mixed_sqlite.lua \
http://127.0.0.1:8080
```
(`-R20000` is a guess; the real number comes from a `wrk` probe of each
framework first.)
---
## 5. Implementation plan
Three sibling crates / projects under a `urus-bench/` workspace, each
exposing the same routes and middleware:
```
urus-bench/
├── urus-server/ # Rust, urus
├── axum-server/ # Rust, hyper via axum
└── cowboy-server/ # Erlang/OTP, cowboy 2
```
Plus:
```
urus-bench/
├── loadgen/ # wrk2 scripts (common.lua, s1..s4 above)
├── runner.sh # orchestrates pin-cpu, warm-up, measurement, save HDR
└── results/ # one subdir per run: ${date}-${server}-${scenario}/
```
The runner:
1. Boots the target server pinned to cores 0-7.
2. `curl`-smoke-tests every route.
3. Probes saturation with `wrk` at a high `-c`.
4. Reads back the saturation RPS and chooses 70% of it for wrk2.
5. Runs wrk2 with HDR output (`--latency`), captures stdout + HDR file.
6. Greps server process for RSS (`/proc/$pid/status`) and CPU (`pidstat -p $pid 1`) during the run.
7. Saves everything under `results/${run_id}/`.
8. Kills server, sleeps 5s, moves on.
### 5.1 urus-server
Port `examples/crud.rs` with:
- Routes adjusted to `/api/v1/users{,/:id}` and `/ping`.
- Logger replaced with a ring-buffer-backed plug.
- New `request_id` and `auth` plugs.
- Switchable backing store: `--store=memory` (the existing actor) or
`--store=sqlite` (a new actor wrapping `rusqlite` in WAL).
For SQLite on urus: spawn one writer actor + 4 reader actors, each owning
its own `rusqlite::Connection`. Handlers route reads to a free reader via
a worker-pool channel, writes to the writer. Pool selection is FIFO over a
single shared `Receiver<RequestForReader>` — readers race to receive, which
is exactly what smarm's MPSC gives for free.
### 5.2 axum-server
Standard axum + tokio multi-thread. Same routes. Layers: `Logger`,
`SetRequestIdLayer`, `RequireAuthorizationLayer`. State: `sqlx::SqlitePool`
for S4, `Arc<RwLock<HashMap>>` for S3.
### 5.3 cowboy-server
Standard cowboy 2 with `cowboy_router`. Handlers as `cowboy_handler`
modules. Auth as an `onrequest`-style middleware. Store: a `gen_server`
for S3, `esqlite` + `poolboy` for S4.
Build with `rebar3`, run with `+S 8 +sbt db` (8 schedulers, scheduler-thread
binding) under the same `taskset -c 0-7`.
---
## 6. Metrics and reporting
Per (server × scenario) run, capture and report:
| Metric | How |
|-----------------------------|------------------------------------------------|
| RPS (saturation, `wrk`) | `wrk -c256 -d60s` — closed-loop peak |
| RPS (sustained, `wrk2`) | `wrk2 -R$target -d60s` at 70% of saturation |
| p50 / p95 / p99 / p99.9 / max latency | HDR histogram from wrk2 |
| Server CPU% | `pidstat -p $pid 1 60` → mean over run |
| Server RSS | `/proc/$pid/status` `VmRSS` sampled at 1 Hz |
| Handler-inside p99 (µs) | server-side histogram, emitted to ring buffer |
| Dropped/errored requests | wrk2's "Non-2xx or 3xx responses" |
The headline table:
| | urus | axum | cowboy |
|--------------------|------|------|--------|
| S1 RPS (sat) | | | |
| S1 p99 @ 70% | | | |
| S2 RPS (sat) | | | |
| S2 p99 @ 70% | | | |
| S3 RPS (sat) | | | |
| S3 p99 @ 70% | | | |
| S4 RPS (sat) | | | |
| S4 p99 @ 70% | | | |
| urus / axum (S1) | — | | — |
| urus / axum (S2) | — | | — |
Pass/fail per the urus spec:
- **S1**: urus RPS ≥ 50% of axum RPS (within 2x).
- **S2**: urus RPS ≥ 50% of axum RPS.
- **All scenarios**: urus RPS > cowboy RPS by a clear margin (native
expectation).
- **All scenarios**: urus p99 < 2× axum p99 at the same offered rate.
- Handler-inside p99 ≪ end-to-end p99 in S1/S2 (proves we're benching the
framework, not the handler).
---
## 7. Order of operations (suggested)
1. Build `urus-server` with S1 (`/ping`) only. Run S1 against it. Make
sure the saturation number is plausible (tens to hundreds of kRPS on
a decent box).
2. Build `axum-server` with S1. Run S1. Sanity-check the ratio against
public hyper hello-world numbers (axum should land near hyper).
3. Add the middleware stack to both, run S2, capture deltas.
4. Add the in-memory store, run S3.
5. Bring up `cowboy-server` and bring all three up to S3.
6. Add SQLite to all three; run S4.
7. Write the headline table; produce HDR plots from the saved files
(gnuplot or a tiny Python script — wrk2 ships an HDR-plot helper).
If at any point a tier's results are dominated by a non-framework cost
(handler-inside p99 ≥ ~50% of end-to-end p99), stop and instrument before
moving on. That's the whole point of tiering.
---
## 8. Open questions deliberately left for later
- Whether to also bench under `tc`-induced packet loss / RTT. The current
setup uses loopback, which is a generous environment. A loopback win is
necessary but not sufficient for "fast on the real network".
- Whether to bench HTTP/1.1 pipelining (multiple in-flight requests per
connection without waiting). urus claims it's a non-goal-for-now, and
axum/hyper don't pipeline either, so the comparison is moot — but it's
worth a note.
- Whether to add a "slow handler" scenario (one route that sleeps 50ms)
to demonstrate urus's blocking-handler-without-thread-pool claim. That's
more of a *qualitative* demo than a throughput bench, but it's the kind
of thing reviewers ask about.