Files
smarm/benches
claude-asm-audit 8fd724a6b6 bench: regenerate baseline.json on the reconciled tree (job 3a1af71f)
Rebased onto upstream v0.7.0, with the Weak-per-park fix (93bc83a) and the
root-sweep lock-free guard (3b0e06e) both in. 20-core box, rq-mpmc default,
wake slot on. Key order kept, so the diff is values only.

Not a flat baseline, and the regress pass in outputs/f21_job_3a1af71f.log
records why. Against 09ad759 the short 1-thread sections carry ~30-45 µs of
residual fixed cost (the root-exit sweep still walks all 16_384 slot words even
though it no longer locks them), and ping_pong_steady is +25% 1T / +23% 20T with
only ~30 µs of that explained — ~28 ns/roundtrip on the channel park chain that
nothing in the upstream diff accounts for. Several 20-thread sections moved
10-18% alongside it. That is finding 20c, open, and regenerating here means
`regress` can no longer see it: the comparison lives in history.md and the job
log instead.

switch_cost 126-128 cyc/roundtrip, unchanged, so the yield path is not involved.
2026-08-21 14:04:19 +00:00
..

Benches

Two families live here: comparison benches (smarm vs tokio, predating v0.5) and the run-queue shootout (v0.5 phase 4). All are plain binaries (harness = false in Cargo.toml), so cargo bench just builds in release and runs main() — no criterion, no magic.

cargo bench --bench <name>           # one bench
cargo bench                          # all of them (slow; rarely what you want)

Catalog

file what it measures
primes.rs Compute fan-out/fan-in: counts primes across W workers. Pure compute throughput + spawn/join/channel cost.
multi_scheduler.rs The original cross-runtime matrix: smarm (1 thread / N threads) vs tokio (current_thread / multi_thread) on compute, ping-pong, and spawn throughput.
general.rs Workloads where neither runtime has a structural edge. Large gaps here mean real per-task/per-yield overhead differences — watch these for regressions.
smarm_favored.rs Workloads the stackful green-thread model is built for. Single-thread numbers isolate per-switch cost from contention.
tokio_favored.rs Workloads tokio's model is built for. Expect to lose; the value is knowing by how much and catching the gap widening.
rq_micro.rs Run-queue structures in isolation (no runtime, no actors): push/pop throughput sweeping thread count × producer:consumer ratio. Covers all three queue types in one binary — the types compile in every build; only the runtime's alias is feature-selected.
rq_runtime.rs The whole scheduler with the compile-time-selected queue: yield-storm (pure queue churn), ping-pong-pairs (park/unpark latency), spawn-storm (slab + free list + queue churn), sweeping scheduler count. Comparing variants requires rebuilding per rq-* feature.

The run-queue shootout

One command; it rebuilds rq_runtime once per queue variant, runs rq_micro once, and aggregates:

./scripts/bench_rq.sh
# on a big box:
SMARM_BENCH_THREADS="1 2 4 8 16 20" ./scripts/bench_rq.sh

Outputs land in bench_results/ (gitignored): one full log per run, plus summary.csv assembled from the machine-readable RQCSV,... lines every config prints alongside the human table.

Manual single-variant runs need the feature dance (features are additive, so the default rq-mutex must be switched off):

cargo bench --bench rq_runtime --no-default-features --features rq-striped

Knobs (env vars, all optional)

var default used by
SMARM_BENCH_THREADS "1 2 4" both — space-separated sweep
SMARM_BENCH_RUNS 5 both — repetitions; the median is reported
SMARM_BENCH_ITEMS 200000 rq_micro — items per measurement
SMARM_BENCH_YIELD_ACTORS / _YIELDS 200 / 500 rq_runtime yield-storm
SMARM_BENCH_PAIRS / _ROUNDTRIPS 32 / 1000 rq_runtime ping-pong
SMARM_BENCH_SPAWNS 5000 rq_runtime spawn-storm

Reading the numbers honestly

  • Core count is the experiment. On a 1-core machine (CI, sandboxes) the sweep only validates the harness and catches gross pathologies — oversubscribed schedulers measure context-switch noise, not contention. Variant decisions come from a many-core box.
  • The striped queue should lose at low thread counts (ticket overhead with no contention to amortize) — that's expected, not a bug.
  • Medians over SMARM_BENCH_RUNS absorb scheduling noise but not thermal / turbo drift; for publishable numbers, pin the CPU governor and run a warmup pass first.
  • spawn-storm batches joins (1024 at a time) to stay well under the slab cap; if you raise SMARM_BENCH_SPAWNS massively, that batching is why it still works.

Adding a bench

  1. benches/<name>.rs with a plain main(); print the house table (see any existing bench) and, if it belongs to a sweep, a greppable CSV line with a distinctive prefix (RQCSV, for the shootout family).
  2. Register it in Cargo.toml:
    [[bench]]
    name = "<name>"
    harness = false
    
  3. Take parameters from SMARM_BENCH_* env vars with modest defaults — the defaults must finish in seconds on one core, the env scales them up on real hardware.
  4. Report medians, and keep one measurement = one fresh runtime (init(Config::exact(t)) inside the measured closure constructor, the run() inside the timed region) so runs don't contaminate each other.