sandbox-agent 943080805c loadgen: fix two latent lua bugs that voided every wrk run
Root cause of the ~100% non-2xx benchmark: the load scripts never actually
ran. Two stacked failures, each of which makes wrk silently fall back to its
built-in default request (GET / with no auth -> 401 -> 100% non-2xx):

1. require('common') failed: wrk runs with cwd at the repo root, where
   common.lua (in loadgen/) is not on package.path. Now each script prepends
   its own directory (via debug.getinfo source) before requiring.
2. os.getpid() is not in luajit's os library (wrk doesn't add it), so the
   randomseed line aborted the chunk once the require was fixed. Guarded with
   (os.getpid and os.getpid() or 0). It was per-process anyway (same PID across
   wrk threads), so it never varied the per-thread seed regardless.

Verified under luajit 2.1 (same interpreter wrk uses), loading each script as
wrk does from the repo-root cwd and via absolute path: all now load, define
request(), and set the Authorization header. The require always failed first,
so neither bug was ever caught -- the wrk path was never exercised end-to-end
(only manual curl, which worked). Not caused by the gen_server port or the
core/loadgen retune; those stand. Re-run _diag_status.lua to confirm real
status codes, then the A/B is finally measuring store work.
2026-06-14 22:06:49 +00:00
2026-05-26 23:22:32 +02:00
2026-05-26 23:22:32 +02:00
2026-05-26 23:22:32 +02:00
2026-05-26 23:22:32 +02:00
2026-05-26 23:22:32 +02:00
2026-05-26 21:19:19 +00:00
2026-05-26 23:25:58 +02:00
2026-05-26 23:22:32 +02:00
2026-05-26 23:22:32 +02:00
2026-05-26 23:22:32 +02:00
2026-05-26 23:22:32 +02:00

urus-bench

HTTP/1.1 performance benchmarks comparing Urus (actor-based Rust web framework) against Axum (async/await Rust via hyper). Designed to test the Urus spec claim: "within 2× of Axum RPS."

Overview

This benchmark suite measures throughput (RPS) and tail latency (p99) across four scenarios of increasing complexity:

Scenario Workload Framework Test
S1 Bare HTTP (GET /ping) Protocol + connection handling
S2 Routing + middleware Logger, request ID, auth, router
S3 In-memory store (mixed) Message-passing vs. mutex contention
S4 SQLite (mixed) Database overhead under actor model

Full specification: see urus-bench-spec.md.

Quick Start

Prerequisites

  • Rust 1.70+ (stable)
  • wrk and wrk2 (load generators)
  • pidstat, taskset (process introspection & CPU pinning)
  • curl (sanity checks before benchmarks)

Recommended: Use the provided nix-shell:

nix-shell

This installs all dependencies except Rust (assumed already on your host).

Build

cargo build --release

Both urus-server and axum-server will be built. Runs take ~510 minutes total (30s warmup + 60s measure per scenario).

Run a Single Benchmark

./runner.sh urus-mem s1

This runs Urus on scenario S1, pins the server to cores 07, the load generator to cores 815, and saves results to results/<batch>/<server>-<scenario>/.

Run the Full Matrix

./run_all.sh

Runs 8 benchmarks (Urus + Axum on S1S3, both on SQLite S4). Outputs a summary table and machine-readable JSON index.

Quick Mode (for development)

./run_all.sh --quick

Uses 15s warmup + 15s measure (instead of 30+60) for faster iteration.

Repository Structure

urus-bench/
├── urus-server/          # Urus implementation (actor-based)
│   ├── src/
│   │   ├── main.rs
│   │   ├── handlers.rs
│   │   ├── middleware.rs
│   │   └── store.rs      # In-memory and SQLite store actors
│   └── Cargo.toml
├── axum-server/          # Axum implementation (async/await + hyper)
│   ├── src/
│   │   └── main.rs
│   └── Cargo.toml
├── loadgen/              # wrk2 scripts
│   ├── common.lua        # shared headers, bearer token
│   ├── s1_ping.lua       # simple GET
│   ├── s2_user_get.lua   # router + middleware
│   └── s3_mixed.lua      # 80% GET, 15% list, 5% POST
├── runner.sh             # Single-run orchestrator
├── run_all.sh            # Full-matrix orchestrator
├── urus-bench-spec.md    # Detailed specification
├── shell.nix             # Nix environment
├── Cargo.toml            # Workspace definition
└── results/              # Output (created at runtime)
    └── <batch>/
        ├── summary.md    # Headline results
        ├── index.json    # Machine-readable results
        └── <server>-<scenario>/
            ├── result.json      # Metrics: RPS, latencies, CPU, RSS
            └── histogram.txt    # HDR histogram (wrk2 raw output)

Key Scripts

runner.sh

Runs a single (server, scenario) pair. Usage:

./runner.sh [--help] <server> <scenario> [options]

# Examples:
./runner.sh urus-mem s1               # Urus on S1 (defaults)
./runner.sh axum-sqlite s4            # Axum on S4
./runner.sh urus-mem s3 --quick       # Urus on S3, short timings
./runner.sh axum-mem s2 --measure-sec=120  # Extended measurement

Server names:

  • urus-mem, axum-mem (in-memory store)
  • urus-sqlite, axum-sqlite (SQLite store)

Scenario names:

  • s1, s2, s3, s4 (see Specification, below)

Key options:

  • --quick: Use 15s warmup + 15s measure (fast iteration)
  • --measure-sec=N: Set measurement duration
  • --warmup-sec=N: Set warmup duration
  • --wrk-conns=N: Number of connections (default 256)
  • --wrk-threads=N: Load generator threads (default 8)
  • --batch=ID: Tag for results directory (default: timestamp)

Output: results/<batch>/<server>-<scenario>/result.json + histogram.txt.

run_all.sh

Runs the full benchmark matrix. Usage:

./run_all.sh [options]

# Examples:
./run_all.sh                   # Full 8-run matrix
./run_all.sh --quick           # Full matrix, 15s+15s timings
./run_all.sh --scenarios=s1,s2 # Only S1 and S2
./run_all.sh --servers=urus-mem,axum-mem  # Only in-memory

Output:

  • results/<batch>/summary.md — headline table (text)
  • results/<batch>/index.json — structured results (machine-readable)

Configuration

Environment variables (passed through to runner.sh):

Variable Default Purpose
WARMUP_SEC 30 Warm-up duration (TCP window scaling, JIT)
MEASURE_SEC 60 Measurement duration (where latencies are sampled)
PROBE_SEC 30 Saturation probe duration (closed-loop wrk)
SERVER_CPUS 0-7 CPU cores pinned to server process
LOADGEN_CPUS 8-15 CPU cores pinned to load generator
WRK_CONNS 256 Concurrent connections
WRK_THREADS 8 Load generator thread count
SAT_RATIO 0.7 Target RPS = saturation × this ratio
BEARER (required) Authorization token for auth middleware

Example:

MEASURE_SEC=120 WARMUP_SEC=60 ./runner.sh urus-mem s3

Interpreting Results

Metrics

Each run produces result.json with:

  • sustained_rps — requests/second at the target rate (wrk2 measurement)
  • p99_latency_ms — 99th percentile latency (HDR histogram)
  • server_cpu_pct — average CPU utilization (pidstat)
  • server_rss_mb — peak resident set size
  • handler_p99_us — server-side handler latency (p99, microseconds)

Reading the Output

========== [1/8] urus-mem     s1  ==========
       sustained=54640.68 rps   p99=2.70ms

This is Urus on S1: 54,640 RPS at sustained load, with p99 latency of 2.70ms.

Comparing Frameworks

The spec target is Urus RPS ≥ 50% of Axum RPS on S1 and S2. Example:

  • Urus S1: 54,640 RPS
  • Axum S1: 288,330 RPS
  • Ratio: 0.19 (54,640 / 288,330) → FAIL (need ≥ 0.50)

Development & Debugging

Running a Single Server in Isolation

cargo run --release --bin urus-server -- --store=memory
# Server listens on 127.0.0.1:8080
curl -H "Authorization: Bearer test-token-aaaaaaaaaaaaaaaaaaaa" \
  http://127.0.0.1:8080/ping
# Response: 200 OK, body "pong"

Inspecting SQLite Between Runs

sqlite3 urus.db "SELECT COUNT(*) FROM users;"

SQLite databases are created in the working directory as urus.db (Urus) and axum.db (Axum).

Profiling a Run

Modify runner.sh to include perf record:

perf record -g -F 99 -p $pid -- ...  # during the measurement phase
perf report

Viewing Raw Histograms

cat results/<batch>/<server>-<scenario>/histogram.txt
# HDR histogram format (wrk2 output)

Convert to a plot with hdr-plot (wrk2 companion tool) or similar.

Known Limitations

  • No TLS: HTTP/1.1 plaintext only. TLS would require additional tooling per framework.
  • Loopback only: All networking is on 127.0.0.1 (generous environment). Real-world RTT / packet loss not tested.
  • No pipelining: HTTP/1.1 pipelining not exercised.
  • No Cowboy: BEAM reference removed; focus is Urus vs. Axum.

Results from 2026-05-26

Summary from the last full run (batch 20260526-220625):

Urus Axum Ratio
S1 RPS 54,640 288,330 0.19 (5.3×)
S2 RPS 53,139 315,596 0.17 (5.9×)
S3 RPS 53,616 316,342 0.17 (5.9×)
S4 RPS 52,590 320,150 0.16 (6.1×)

Verdict: Urus fails the "within 2×" spec target on throughput. Tail latency (p99) remains sub-3ms under all loads; Axum's p99 scales with concurrency (1160ms). See benchmark-report.html for full analysis.

References

License

Benchmark code is in the public domain or under your preferred permissive license (MIT/Apache-2.0). See individual crate licenses for dependencies.

S
Description
No description provided
Readme MIT
595 KiB
Languages
Rust 46.7%
Shell 30.7%
HTML 11.7%
Python 4.4%
Lua 4.4%
Other 2.1%