AssetStoreServer runs on a single dedicated smarm actor thread, so a plain HashMap+VecDeque LRU in front of the SQLite lookup needs no locking. Capacity configurable via CCC_CACHE_CAPACITY (default 256). Bench (bench/RPS): +42-47% RPS under a cache-sized/hot-set workload, but a small net loss under adversarial uniform-random access with a cache smaller than the catalog. Real traffic is hot-set skewed, so net win in practice; capacity should be tuned to the expected hot set.
79 lines
3.3 KiB
Plaintext
79 lines
3.3 KiB
Plaintext
# CCC bench: RPS ceiling + LRU cache impact
|
|
|
|
Quick and dirty throughput bench, run locally on a 24-core box. Not
|
|
scientific, just enough to sanity-check the actor-based SQLite store and
|
|
the small LRU cache in front of it.
|
|
|
|
## Harness
|
|
|
|
- `bench/seed.py` fills a fresh `cdn.db` with random packages/versions/
|
|
assets (default: 500 packages x 5 versions = 2500 assets, 512B-8KB
|
|
gzipped JS each) and writes `bench/urls.txt` (one `/assets/...` path per
|
|
line, for every seeded asset).
|
|
- Server pinned to 2 CPUs, load generator (`oha`, via
|
|
`nix-shell -p oha`) pinned to 6 CPUs, both via `taskset`, so client
|
|
capacity is never the bottleneck.
|
|
|
|
```
|
|
nix-shell -p python3 --run "python3 bench/seed.py cdn.db"
|
|
awk '{print "http://127.0.0.1:8333"$0}' bench/urls.txt > /tmp/full_urls.txt
|
|
|
|
CCC_DB_PATH=$(pwd)/cdn.db taskset -c 0,1 ./target/release/CCC serve --port 8333 &
|
|
|
|
nix-shell -p oha --run \
|
|
"taskset -c 2-7 oha -z 8s -c 200 --no-tui --urls-from-file /tmp/full_urls.txt"
|
|
```
|
|
|
|
Skewed/hot-set workload (80% of requests hit the top 50 of 2500 assets,
|
|
i.e. a realistic CDN access pattern) was generated with a short Python
|
|
snippet sampling from `bench/urls.txt` with `random.random() < 0.8` picking
|
|
from the first 50 lines, else uniformly from the rest, written to a
|
|
`urls_hot80_20.txt` file and expanded the same way as above.
|
|
|
|
## Results: no cache (baseline)
|
|
|
|
Single-threaded `smarm` actor (`AssetStoreServer::loop_runner`) serializes
|
|
every asset lookup onto one thread/one SQLite connection, so this is
|
|
inherently CPU-bound on the 2 pinned cores regardless of client
|
|
concurrency.
|
|
|
|
| Server CPUs | Client concurrency | RPS |
|
|
|---|---|---|
|
|
| 2 | 50 | ~56.7k |
|
|
| 2 | 100 | ~60.1k |
|
|
| 2 | 200 | ~62.1k (peak, server at ~176-200% CPU, saturated) |
|
|
| 2 | 400 | ~60.0k (plateaued) |
|
|
| 2 | 800 | ~55.9k (queueing overhead) |
|
|
| 2 (client sends `Accept-Encoding: gzip`, server skips decompression) | 200 | ~60.4k |
|
|
| 1 | 200 | ~30.8k (confirms CPU-bound, scales with cores) |
|
|
|
|
Client (6 CPUs) stayed at ~4% usr / 8% sys throughout - never the
|
|
bottleneck.
|
|
|
|
## Results: with the LRU cache (`CCC_CACHE_CAPACITY`, default 256)
|
|
|
|
The cache lives inside `AssetStoreServer` itself (see `src/main.rs`), so
|
|
it needs no locking - the actor thread is already strictly sequential.
|
|
|
|
| Scenario | Cache capacity | RPS | vs. no-cache baseline (62.1k) |
|
|
|---|---|---|---|
|
|
| Uniform-random over all 2500 assets | 256 (default) | ~55.3k | **-11%** |
|
|
| Uniform-random over all 2500 assets | 3000 (covers full catalog) | ~88.1k | **+42%** |
|
|
| 80/20 hot-set (50 hot assets get 80% of traffic) | 256 (default) | ~91.1k | **+47%** |
|
|
|
|
### Caveat
|
|
|
|
Under a purely uniform-random access pattern with a cache smaller than
|
|
the catalog (low hit rate), the cache is a net loss: every request now
|
|
pays HashMap lookup + insert + eviction bookkeeping on top of the SQLite
|
|
query, for a hit rate too low to earn it back. The `touch()` on hit is
|
|
also an O(n) scan of the recency queue, which doesn't help at low
|
|
capacities.
|
|
|
|
Real CDN traffic is essentially never uniform-random (it's hot-set/
|
|
power-law skewed), so in practice this is a clear win, but `CCC_CACHE_CAPACITY`
|
|
should be sized to the actual hot set rather than left at the arbitrary
|
|
default of 256. A follow-up would swap the O(n) recency scan for a proper
|
|
O(1) LRU (e.g. an intrusive linked-hashmap) to remove the downside case
|
|
entirely.
|