Add causal profiling (RFC 007) behind --features causal
Instrument the four suspect regions from bench/RPS: cache-lookup, cache-insert, sqlite-query, and gzip-decode, plus an asset-served progress point. New `ccc causal` subcommand runs the sweep against live traffic and prints a summary (optionally a .coz file and a ledger audit). Ran it under the same 80/20 hot-set workload as bench/RPS - see bench/CAUSAL.md for the methodology and results. Findings: - sqlite-query is the real bottleneck on a cache miss (+20.5% at a 50% speedup, roughly linear). - cache-lookup/cache-insert are noise-level (0-4%) - the O(n) recency-scan touch() bench/RPS flagged as a possible follow-up is not actually costing anything, so that's off the table. - Switched prepare() -> prepare_cached() on the query as the obvious fix; re-measured and it made no real difference (+20.0% -> +20.5%, within noise). Kept it anyway (strictly not worse), but it shows execution cost (B-tree lookup + BLOB copy) dominates over parse cost in that site. - gzip-decode barely gets exercised since real clients (and oha) negotiate gzip - not worth optimizing further. - Net conclusion: the existing CCC_CACHE_CAPACITY tuning from bench/RPS (+42-47% RPS) is the correct lever, and causal profiling explains why - every cache hit skips the one site that matters. Zero cost when the feature is off: causal_site!/progress! compile to no-ops without smarm-causal.
This commit is contained in:
@@ -3,7 +3,9 @@
|
||||
My super simple CDN built for distributing my own (text) content.
|
||||
|
||||
Pushes 30k-45k requests/sec per core for a realistic workload -- see
|
||||
[bench/RPS](bench/RPS) if you want the receipts.
|
||||
[bench/RPS](bench/RPS) if you want the receipts, and
|
||||
[bench/CAUSAL.md](bench/CAUSAL.md) for causal-profiling which sites
|
||||
actually matter (`cargo build --features causal`).
|
||||
|
||||
## Running in Docker
|
||||
|
||||
|
||||
Reference in New Issue
Block a user