Files
2026-05-26 23:25:58 +02:00

8.5 KiB
Raw Permalink Blame History

urus-bench

HTTP/1.1 performance benchmarks comparing Urus (actor-based Rust web framework) against Axum (async/await Rust via hyper). Designed to test the Urus spec claim: "within 2× of Axum RPS."

Overview

This benchmark suite measures throughput (RPS) and tail latency (p99) across four scenarios of increasing complexity:

Scenario Workload Framework Test
S1 Bare HTTP (GET /ping) Protocol + connection handling
S2 Routing + middleware Logger, request ID, auth, router
S3 In-memory store (mixed) Message-passing vs. mutex contention
S4 SQLite (mixed) Database overhead under actor model

Full specification: see urus-bench-spec.md.

Quick Start

Prerequisites

  • Rust 1.70+ (stable)
  • wrk and wrk2 (load generators)
  • pidstat, taskset (process introspection & CPU pinning)
  • curl (sanity checks before benchmarks)

Recommended: Use the provided nix-shell:

nix-shell

This installs all dependencies except Rust (assumed already on your host).

Build

cargo build --release

Both urus-server and axum-server will be built. Runs take ~510 minutes total (30s warmup + 60s measure per scenario).

Run a Single Benchmark

./runner.sh urus-mem s1

This runs Urus on scenario S1, pins the server to cores 07, the load generator to cores 815, and saves results to results/<batch>/<server>-<scenario>/.

Run the Full Matrix

./run_all.sh

Runs 8 benchmarks (Urus + Axum on S1S3, both on SQLite S4). Outputs a summary table and machine-readable JSON index.

Quick Mode (for development)

./run_all.sh --quick

Uses 15s warmup + 15s measure (instead of 30+60) for faster iteration.

Repository Structure

urus-bench/
├── urus-server/          # Urus implementation (actor-based)
│   ├── src/
│   │   ├── main.rs
│   │   ├── handlers.rs
│   │   ├── middleware.rs
│   │   └── store.rs      # In-memory and SQLite store actors
│   └── Cargo.toml
├── axum-server/          # Axum implementation (async/await + hyper)
│   ├── src/
│   │   └── main.rs
│   └── Cargo.toml
├── loadgen/              # wrk2 scripts
│   ├── common.lua        # shared headers, bearer token
│   ├── s1_ping.lua       # simple GET
│   ├── s2_user_get.lua   # router + middleware
│   └── s3_mixed.lua      # 80% GET, 15% list, 5% POST
├── runner.sh             # Single-run orchestrator
├── run_all.sh            # Full-matrix orchestrator
├── urus-bench-spec.md    # Detailed specification
├── shell.nix             # Nix environment
├── Cargo.toml            # Workspace definition
└── results/              # Output (created at runtime)
    └── <batch>/
        ├── summary.md    # Headline results
        ├── index.json    # Machine-readable results
        └── <server>-<scenario>/
            ├── result.json      # Metrics: RPS, latencies, CPU, RSS
            └── histogram.txt    # HDR histogram (wrk2 raw output)

Key Scripts

runner.sh

Runs a single (server, scenario) pair. Usage:

./runner.sh [--help] <server> <scenario> [options]

# Examples:
./runner.sh urus-mem s1               # Urus on S1 (defaults)
./runner.sh axum-sqlite s4            # Axum on S4
./runner.sh urus-mem s3 --quick       # Urus on S3, short timings
./runner.sh axum-mem s2 --measure-sec=120  # Extended measurement

Server names:

  • urus-mem, axum-mem (in-memory store)
  • urus-sqlite, axum-sqlite (SQLite store)

Scenario names:

  • s1, s2, s3, s4 (see Specification, below)

Key options:

  • --quick: Use 15s warmup + 15s measure (fast iteration)
  • --measure-sec=N: Set measurement duration
  • --warmup-sec=N: Set warmup duration
  • --wrk-conns=N: Number of connections (default 256)
  • --wrk-threads=N: Load generator threads (default 8)
  • --batch=ID: Tag for results directory (default: timestamp)

Output: results/<batch>/<server>-<scenario>/result.json + histogram.txt.

run_all.sh

Runs the full benchmark matrix. Usage:

./run_all.sh [options]

# Examples:
./run_all.sh                   # Full 8-run matrix
./run_all.sh --quick           # Full matrix, 15s+15s timings
./run_all.sh --scenarios=s1,s2 # Only S1 and S2
./run_all.sh --servers=urus-mem,axum-mem  # Only in-memory

Output:

  • results/<batch>/summary.md — headline table (text)
  • results/<batch>/index.json — structured results (machine-readable)

Configuration

Environment variables (passed through to runner.sh):

Variable Default Purpose
WARMUP_SEC 30 Warm-up duration (TCP window scaling, JIT)
MEASURE_SEC 60 Measurement duration (where latencies are sampled)
PROBE_SEC 30 Saturation probe duration (closed-loop wrk)
SERVER_CPUS 0-7 CPU cores pinned to server process
LOADGEN_CPUS 8-15 CPU cores pinned to load generator
WRK_CONNS 256 Concurrent connections
WRK_THREADS 8 Load generator thread count
SAT_RATIO 0.7 Target RPS = saturation × this ratio
BEARER (required) Authorization token for auth middleware

Example:

MEASURE_SEC=120 WARMUP_SEC=60 ./runner.sh urus-mem s3

Interpreting Results

Metrics

Each run produces result.json with:

  • sustained_rps — requests/second at the target rate (wrk2 measurement)
  • p99_latency_ms — 99th percentile latency (HDR histogram)
  • server_cpu_pct — average CPU utilization (pidstat)
  • server_rss_mb — peak resident set size
  • handler_p99_us — server-side handler latency (p99, microseconds)

Reading the Output

========== [1/8] urus-mem     s1  ==========
       sustained=54640.68 rps   p99=2.70ms

This is Urus on S1: 54,640 RPS at sustained load, with p99 latency of 2.70ms.

Comparing Frameworks

The spec target is Urus RPS ≥ 50% of Axum RPS on S1 and S2. Example:

  • Urus S1: 54,640 RPS
  • Axum S1: 288,330 RPS
  • Ratio: 0.19 (54,640 / 288,330) → FAIL (need ≥ 0.50)

Development & Debugging

Running a Single Server in Isolation

cargo run --release --bin urus-server -- --store=memory
# Server listens on 127.0.0.1:8080
curl -H "Authorization: Bearer test-token-aaaaaaaaaaaaaaaaaaaa" \
  http://127.0.0.1:8080/ping
# Response: 200 OK, body "pong"

Inspecting SQLite Between Runs

sqlite3 urus.db "SELECT COUNT(*) FROM users;"

SQLite databases are created in the working directory as urus.db (Urus) and axum.db (Axum).

Profiling a Run

Modify runner.sh to include perf record:

perf record -g -F 99 -p $pid -- ...  # during the measurement phase
perf report

Viewing Raw Histograms

cat results/<batch>/<server>-<scenario>/histogram.txt
# HDR histogram format (wrk2 output)

Convert to a plot with hdr-plot (wrk2 companion tool) or similar.

Known Limitations

  • No TLS: HTTP/1.1 plaintext only. TLS would require additional tooling per framework.
  • Loopback only: All networking is on 127.0.0.1 (generous environment). Real-world RTT / packet loss not tested.
  • No pipelining: HTTP/1.1 pipelining not exercised.
  • No Cowboy: BEAM reference removed; focus is Urus vs. Axum.

Results from 2026-05-26

Summary from the last full run (batch 20260526-220625):

Urus Axum Ratio
S1 RPS 54,640 288,330 0.19 (5.3×)
S2 RPS 53,139 315,596 0.17 (5.9×)
S3 RPS 53,616 316,342 0.17 (5.9×)
S4 RPS 52,590 320,150 0.16 (6.1×)

Verdict: Urus fails the "within 2×" spec target on throughput. Tail latency (p99) remains sub-3ms under all loads; Axum's p99 scales with concurrency (1160ms). See benchmark-report.html for full analysis.

References

License

Benchmark code is in the public domain or under your preferred permissive license (MIT/Apache-2.0). See individual crate licenses for dependencies.