Root cause of the ~100% non-2xx benchmark: the load scripts never actually
ran. Two stacked failures, each of which makes wrk silently fall back to its
built-in default request (GET / with no auth -> 401 -> 100% non-2xx):
1. require('common') failed: wrk runs with cwd at the repo root, where
common.lua (in loadgen/) is not on package.path. Now each script prepends
its own directory (via debug.getinfo source) before requiring.
2. os.getpid() is not in luajit's os library (wrk doesn't add it), so the
randomseed line aborted the chunk once the require was fixed. Guarded with
(os.getpid and os.getpid() or 0). It was per-process anyway (same PID across
wrk threads), so it never varied the per-thread seed regardless.
Verified under luajit 2.1 (same interpreter wrk uses), loading each script as
wrk does from the repo-root cwd and via absolute path: all now load, define
request(), and set the Authorization header. The require always failed first,
so neither bug was ever caught -- the wrk path was never exercised end-to-end
(only manual curl, which worked). Not caused by the gen_server port or the
core/loadgen retune; those stand. Re-run _diag_status.lua to confirm real
status codes, then the A/B is finally measuring store work.
urus-bench
HTTP/1.1 performance benchmarks comparing Urus (actor-based Rust web framework) against Axum (async/await Rust via hyper). Designed to test the Urus spec claim: "within 2× of Axum RPS."
Overview
This benchmark suite measures throughput (RPS) and tail latency (p99) across four scenarios of increasing complexity:
| Scenario | Workload | Framework Test |
|---|---|---|
| S1 | Bare HTTP (GET /ping) |
Protocol + connection handling |
| S2 | Routing + middleware | Logger, request ID, auth, router |
| S3 | In-memory store (mixed) | Message-passing vs. mutex contention |
| S4 | SQLite (mixed) | Database overhead under actor model |
Full specification: see urus-bench-spec.md.
Quick Start
Prerequisites
- Rust 1.70+ (stable)
wrkandwrk2(load generators)pidstat,taskset(process introspection & CPU pinning)curl(sanity checks before benchmarks)
Recommended: Use the provided nix-shell:
nix-shell
This installs all dependencies except Rust (assumed already on your host).
Build
cargo build --release
Both urus-server and axum-server will be built. Runs take ~5–10 minutes total (30s warmup + 60s measure per scenario).
Run a Single Benchmark
./runner.sh urus-mem s1
This runs Urus on scenario S1, pins the server to cores 0–7, the load generator to cores 8–15, and saves results to results/<batch>/<server>-<scenario>/.
Run the Full Matrix
./run_all.sh
Runs 8 benchmarks (Urus + Axum on S1–S3, both on SQLite S4). Outputs a summary table and machine-readable JSON index.
Quick Mode (for development)
./run_all.sh --quick
Uses 15s warmup + 15s measure (instead of 30+60) for faster iteration.
Repository Structure
urus-bench/
├── urus-server/ # Urus implementation (actor-based)
│ ├── src/
│ │ ├── main.rs
│ │ ├── handlers.rs
│ │ ├── middleware.rs
│ │ └── store.rs # In-memory and SQLite store actors
│ └── Cargo.toml
├── axum-server/ # Axum implementation (async/await + hyper)
│ ├── src/
│ │ └── main.rs
│ └── Cargo.toml
├── loadgen/ # wrk2 scripts
│ ├── common.lua # shared headers, bearer token
│ ├── s1_ping.lua # simple GET
│ ├── s2_user_get.lua # router + middleware
│ └── s3_mixed.lua # 80% GET, 15% list, 5% POST
├── runner.sh # Single-run orchestrator
├── run_all.sh # Full-matrix orchestrator
├── urus-bench-spec.md # Detailed specification
├── shell.nix # Nix environment
├── Cargo.toml # Workspace definition
└── results/ # Output (created at runtime)
└── <batch>/
├── summary.md # Headline results
├── index.json # Machine-readable results
└── <server>-<scenario>/
├── result.json # Metrics: RPS, latencies, CPU, RSS
└── histogram.txt # HDR histogram (wrk2 raw output)
Key Scripts
runner.sh
Runs a single (server, scenario) pair. Usage:
./runner.sh [--help] <server> <scenario> [options]
# Examples:
./runner.sh urus-mem s1 # Urus on S1 (defaults)
./runner.sh axum-sqlite s4 # Axum on S4
./runner.sh urus-mem s3 --quick # Urus on S3, short timings
./runner.sh axum-mem s2 --measure-sec=120 # Extended measurement
Server names:
urus-mem,axum-mem(in-memory store)urus-sqlite,axum-sqlite(SQLite store)
Scenario names:
s1,s2,s3,s4(see Specification, below)
Key options:
--quick: Use 15s warmup + 15s measure (fast iteration)--measure-sec=N: Set measurement duration--warmup-sec=N: Set warmup duration--wrk-conns=N: Number of connections (default 256)--wrk-threads=N: Load generator threads (default 8)--batch=ID: Tag for results directory (default: timestamp)
Output: results/<batch>/<server>-<scenario>/result.json + histogram.txt.
run_all.sh
Runs the full benchmark matrix. Usage:
./run_all.sh [options]
# Examples:
./run_all.sh # Full 8-run matrix
./run_all.sh --quick # Full matrix, 15s+15s timings
./run_all.sh --scenarios=s1,s2 # Only S1 and S2
./run_all.sh --servers=urus-mem,axum-mem # Only in-memory
Output:
results/<batch>/summary.md— headline table (text)results/<batch>/index.json— structured results (machine-readable)
Configuration
Environment variables (passed through to runner.sh):
| Variable | Default | Purpose |
|---|---|---|
WARMUP_SEC |
30 | Warm-up duration (TCP window scaling, JIT) |
MEASURE_SEC |
60 | Measurement duration (where latencies are sampled) |
PROBE_SEC |
30 | Saturation probe duration (closed-loop wrk) |
SERVER_CPUS |
0-7 | CPU cores pinned to server process |
LOADGEN_CPUS |
8-15 | CPU cores pinned to load generator |
WRK_CONNS |
256 | Concurrent connections |
WRK_THREADS |
8 | Load generator thread count |
SAT_RATIO |
0.7 | Target RPS = saturation × this ratio |
BEARER |
(required) | Authorization token for auth middleware |
Example:
MEASURE_SEC=120 WARMUP_SEC=60 ./runner.sh urus-mem s3
Interpreting Results
Metrics
Each run produces result.json with:
sustained_rps— requests/second at the target rate (wrk2 measurement)p99_latency_ms— 99th percentile latency (HDR histogram)server_cpu_pct— average CPU utilization (pidstat)server_rss_mb— peak resident set sizehandler_p99_us— server-side handler latency (p99, microseconds)
Reading the Output
========== [1/8] urus-mem s1 ==========
sustained=54640.68 rps p99=2.70ms
This is Urus on S1: 54,640 RPS at sustained load, with p99 latency of 2.70ms.
Comparing Frameworks
The spec target is Urus RPS ≥ 50% of Axum RPS on S1 and S2. Example:
- Urus S1: 54,640 RPS
- Axum S1: 288,330 RPS
- Ratio: 0.19 (54,640 / 288,330) → FAIL (need ≥ 0.50)
Development & Debugging
Running a Single Server in Isolation
cargo run --release --bin urus-server -- --store=memory
# Server listens on 127.0.0.1:8080
curl -H "Authorization: Bearer test-token-aaaaaaaaaaaaaaaaaaaa" \
http://127.0.0.1:8080/ping
# Response: 200 OK, body "pong"
Inspecting SQLite Between Runs
sqlite3 urus.db "SELECT COUNT(*) FROM users;"
SQLite databases are created in the working directory as urus.db (Urus) and axum.db (Axum).
Profiling a Run
Modify runner.sh to include perf record:
perf record -g -F 99 -p $pid -- ... # during the measurement phase
perf report
Viewing Raw Histograms
cat results/<batch>/<server>-<scenario>/histogram.txt
# HDR histogram format (wrk2 output)
Convert to a plot with hdr-plot (wrk2 companion tool) or similar.
Known Limitations
- No TLS: HTTP/1.1 plaintext only. TLS would require additional tooling per framework.
- Loopback only: All networking is on 127.0.0.1 (generous environment). Real-world RTT / packet loss not tested.
- No pipelining: HTTP/1.1 pipelining not exercised.
- No Cowboy: BEAM reference removed; focus is Urus vs. Axum.
Results from 2026-05-26
Summary from the last full run (batch 20260526-220625):
| Urus | Axum | Ratio | |
|---|---|---|---|
| S1 RPS | 54,640 | 288,330 | 0.19 (5.3×) |
| S2 RPS | 53,139 | 315,596 | 0.17 (5.9×) |
| S3 RPS | 53,616 | 316,342 | 0.17 (5.9×) |
| S4 RPS | 52,590 | 320,150 | 0.16 (6.1×) |
Verdict: Urus fails the "within 2×" spec target on throughput. Tail latency (p99) remains sub-3ms under all loads; Axum's p99 scales with concurrency (11–60ms). See benchmark-report.html for full analysis.
References
- Specification:
urus-bench-spec.md— design principles, scenarios, pass/fail criteria - Latest Report:
benchmark-report.html— 2026-05-26 results & interpretation - Urus: smarm/urus
- Axum: tokio-rs/axum
- wrk2: giltene/wrk2 — constant-rate load generator
License
Benchmark code is in the public domain or under your preferred permissive license (MIT/Apache-2.0). See individual crate licenses for dependencies.