Files
urus/examples/load_profile.rs
T
Claude 8f0da2a806 feat(pubsub,channels)!: handles are addresses — bus, hub and session registries are supervised children
smarm 0.7 (415effb, "lifetime is the actor's — refs are addresses") removed the
rule these three actors were built on: a GenServerRef no longer owns the server,
the loop holds its own inbox sender, and the inbox never closes when the last ref
drops. urus's pubsub table, channel hub bus and session registries were still
governed by that deleted rule — PubSub::new() spawned the table and the handle
owned its life — so nothing commanded them to stop. What still terminated a run
was the root-exit sweep, racing the drain: shutdown_with_open_chat_terminates and
channels_wire::shutdown_with_open_channel_terminates failed 4 times in 25
--all-features runs with "serve did not return: ... outlived the drain: Timeout".
Zero in 10 full-suite runs after this change.

The fix is not a supervisor wrapped around the old shape. Every gotcha in this
area descended from constructors that spawn: PubSub::new(), ChannelHub::new() and
PrefixRouter::channel_session() all started actors, which forced in-runtime-only
construction, which forced the Arc<OnceLock<..>> lazy-init from the first handler,
which forced the "cell must not be static" and "a relay must never hold a PubSub
clone" rules. Five documented rules propping up one inverted dependency. So:
description is separated from instantiation.

- PubSub<M> is a name, not a GenServerRef: const-constructible, Copy, spawns
  nothing, valid outside the runtime and in a static. Operations resolve through
  the registry per call, so a table restarted by its supervisor is reached
  transparently (one lookup per broadcast — bench before caching a ref, which
  would go stale across exactly the restart the supervisor exists to perform).
  PubSub::new() is gone; PubSub::new(name) + PubSub::child() replace it.
- ChannelHub::new(bus, router) returns (hub, Vec<ChildSpec>) — the bus table plus
  one registry per session route. Returning both is the point: a hub whose
  children were never started compiles and fails on the first join, so the vec is
  not left behind a method you can forget to call. #[must_use].
- channel_session gains a registry name; each session registry is separately
  named and separately supervised.
- serve_with/serve_with_shutdown take a Vec<ChildSpec> of app children and build
  the root as RestForOne[..app children, endpoint]. They start before the
  endpoint and, shutdown being ordered in reverse, stop after it has drained, so
  a request still in flight can reach the bus. RestForOne because a bus crash
  leaves live sockets addressing a table that no longer knows them.
- Deleted: the Arc<OnceLock> idiom from both examples and both test pipelines,
  and the module rules that existed only to hand-manage a refcount.

Known cost, not fixed here: channel_session("session:*", "chat-sessions", f) puts
two unrelated string literals side by side and nothing catches a transposition —
a RegistryName newtype is the obvious follow-up.

Tests: 111 lib + 50 integration + 2 doc green, clippy clean, 10/10 full-suite
runs. Unit tests poll for name binding before use — smarm's start-order-is-not-
start-readiness gap; real apps don't hit it, since a handler only runs once a
connection has been accepted.
2026-08-20 14:49:23 +00:00

107 lines
4.5 KiB
Rust

//! Causal load-profiling target (RFC 007): the real urus hot path under
//! EXTERNAL load — no planted bottleneck, no in-process load generation.
//!
//! Where `causal_bench` validates the machinery against a planted,
//! known-answer store, this is the discovery tool: boot a bare server,
//! let an external generator (wrk, pinned to different cores) drive it,
//! sweep virtual speedups over the five lib sites (parse, router,
//! pipeline, serialize, socket-write), and report which one causally
//! limits throughput. External load is deliberate: the generator lives
//! outside the smarm runtime, so it cannot absorb injected virtual
//! delay — the same property `causal_bench` gets from plain OS threads.
//!
//! No pass/fail verdict — there is no known answer here. Trust gates
//! only: traffic actually flowed through the sweep, and the ledger
//! audit (printed) balances.
//!
//! Env:
//! URUS_PORT listen port (default 8080; binds 0.0.0.0)
//! COZ_OUT coz profile output path (default profile.coz)
//! WARMUP_MS settle time after first traffic, ms (default 2000)
//!
//! Orchestration contract (the jobrunner run.sh): start this, wait for
//! it to accept, point wrk at `GET /json/:id` with a duration that
//! outlives the sweep, and stop wrk when this process exits.
//!
//! cargo run --release --example load_profile --features smarm-causal
use std::net::SocketAddr;
use std::time::{Duration, Instant};
use urus::{serve_with_shutdown, shutdown_handle, Config, Conn, Next, Pipeline, Router};
fn env_u64(name: &str, default: u64) -> u64 {
std::env::var(name).ok().and_then(|v| v.parse().ok()).unwrap_or(default)
}
/// The whole handler: param parse + JSON render. Deliberately thin — the
/// subject is the lib path around it, not application work.
fn json_id(conn: Conn, _next: Next) -> Conn {
let id: u64 = conn.params.get("id").and_then(|s| s.parse().ok()).unwrap_or(0);
conn.put_status(200)
.put_header("content-type", "application/json")
.put_body(format!("{{\"id\":{id}}}"))
}
fn main() {
let port = env_u64("URUS_PORT", 8080) as u16;
let warmup = Duration::from_millis(env_u64("WARMUP_MS", 2000));
let coz_out = std::env::var("COZ_OUT").unwrap_or_else(|_| "profile.coz".into());
let addr: SocketAddr = format!("0.0.0.0:{port}").parse().expect("addr");
let (handle, signal) = shutdown_handle();
let server = std::thread::spawn(move || {
let pipe = Pipeline::new().plug(Router::new().get("/json/:id", json_id));
serve_with_shutdown(Config::new(addr), smarm::Config::default(), pipe, Vec::new(), signal).expect("serve");
});
// Readiness is the orchestrator's job (TCP probe); ours is to not
// sweep before real traffic has registered every site and the
// progress point — `run_experiments` enumerates *registered* sites.
// 100 completed responses guarantees full end-to-end coverage.
let responses = |snap: &[(String, u64)]| {
snap.iter().find(|(n, _)| n == "responses").map(|(_, c)| *c).unwrap_or(0)
};
let t0 = Instant::now();
let seen = loop {
let n = responses(&smarm::causal::progress_snapshot());
if n >= 100 {
break n;
}
assert!(
t0.elapsed() < Duration::from_secs(120),
"no load after 120s ({n} responses) — is the generator running?"
);
std::thread::sleep(Duration::from_millis(100));
};
println!("traffic up: {seen} responses; settling {warmup:?}");
std::thread::sleep(warmup);
// Same plan as causal_bench, for comparability across workloads.
let results = smarm::causal::run_experiments(&smarm::causal::ExperimentPlan {
speedups_pct: vec![0, 25, 50],
experiment: Duration::from_millis(700),
cooldown: Duration::from_millis(150),
});
handle.shutdown();
server.join().expect("server panicked");
print!("{}", smarm::causal::render_summary(&results));
print!("{}", smarm::causal::render_ledger_audit(&results));
match std::fs::write(&coz_out, smarm::causal::render_coz(&results)) {
Ok(()) => println!("\nwrote {coz_out}"),
Err(e) => eprintln!("\nfailed to write {coz_out}: {e}"),
}
// Trust gate: the sweep is meaningless if load didn't flow through it.
let total: u64 = results
.iter()
.flat_map(|r| r.deltas.iter())
.filter(|(n, _)| n == "responses")
.map(|(_, c)| c)
.sum();
println!("total responses across windows: {total}");
assert!(total > 0, "sweep saw zero progress");
}