docs(roadmap): record the endpoint milestone (crate v0.3.0)

Full account of the refactor: the tree shape, why the endpoint spawns its
own listener sup (supervisor start order != start readiness), the config
split, what was deleted and the smarm properties re-probed to justify
deleting it. Also:
- PubSub discovery note: the non-static Arc<OnceLock> rule is now
  serve*-only; an app that owns its tree starts the table as a supervised
  sibling and addresses it by name (what crud does).
- Icebox 'operator introspection via typed names': partly delivered — the
  endpoint is a named gen_server answering Call::ConnCount; per-listener
  visibility and richer stats remain open.
- Open after this cycle: the one unreproduced hammer failure, and that
  endpoint() returns impl Fn() so a second invocation clashes on the name.
This commit is contained in:
Claude
2026-08-20 13:27:27 +00:00
parent b37888ec2c
commit 8a568c600c
+80 -6
View File
@@ -252,6 +252,10 @@ smarm has none). All six design questions ratified by the user pre-code
(crud's store shape) hangs `serve_with_shutdown`. Corollary: relays (crud's store shape) hangs `serve_with_shutdown`. Corollary: relays
hold `Receiver` only, never a `PubSub` clone (mutual-keepalive cycle). hold `Receiver` only, never a `PubSub` clone (mutual-keepalive cycle).
Proven by `shutdown_with_open_chat_terminates`. Proven by `shutdown_with_open_chat_terminates`.
**Superseded by v0.7 where the app owns its tree:** start the table as a
supervised sibling of the endpoint and address it by name (what crud now
does with its store). The `OnceLock` rule still holds under `serve*`,
which owns the runtime and offers no in-runtime moment beforehand.
- **`WsHandler::on_open(&mut self, sender)`** added (defaulted, - **`WsHandler::on_open(&mut self, sender)`** added (defaulted,
non-breaking): without it a listen-only ws client can never be non-breaking): without it a listen-only ws client can never be
subscribed (`on_message` never fires). Same panic contract as subscribed (`on_message` never fires). Same panic contract as
@@ -349,6 +353,76 @@ reconnect cycles.
--- ---
## v0.7 — Endpoint as a supervised child — DONE (2026-08-20)
Crate goes 0.2.x -> **0.3.0** (breaking). Unwinds the last deviation from
the spec (§2.1/§6): `serve` owning `rt.run`. The app owns the runtime and
the root supervisor; urus is one ordered child in it.
```
your root sup
└── ChildSpec(Permanent, urus::endpoint(cfg, pipeline)?) <- Endpoint gen_server
└── listener_sup OneForOne over N listeners
└── plain connection actors
```
- `src/conn_registry.rs` -> `src/endpoint.rs`; `ConnRegistry` -> `Endpoint`.
The registry absorbed the listener pool it registers for and runs inline
as the `ChildSpec`'s actor (`NamedGenServerBuilder::run`), so supervisor
shutdown arrives as `handle_shutdown` and a restart re-runs `init` on the
same still-open fds. `endpoint()` binds eagerly on the caller's thread.
- **The endpoint spawns its own listener sup rather than being its
sibling.** smarm's supervisor start *order* is not start *readiness*
(`start_child` spawns and moves on), so a sibling listener could
`whereis` the endpoint name before its actor ran. Registrar-spawns-
consumers makes that program order inside one `init`. Filed in smarm's
ROADMAP as a readiness-ack item; a blocking `spawn` would only shrink the
window ("has begun executing" != "has bound its name") at the cost of a
round-trip per accept.
- Listener sup is monitored: death outside shutdown = panic (the app's
supervisor decides) instead of a zombie on a dead port; death during
shutdown is the "no new connections" barrier.
- Drain is entirely internal and event-driven: `handle_shutdown` flips
draining, shuts the listener sup, stops idle conns, arms one
`drain_timeout` timer; conns that register or go idle mid-drain are
stopped on the spot; one force sweep at the deadline; the endpoint exits
when the sup is down and the set is empty. "Endpoint child stopped" ==
"every connection gone".
- **Deleted:** the `AtomicBool` listener flag, `LISTENER_TICK` (250ms wake
per listener per tick -> untimed `wait_readable` park), `SHUTDOWN_POLL`
(100ms root poll -> real park on the signal channel), `Restart::Transient`
for listeners (now `Permanent`: they only exit by supervisor action),
the root-side drain loop, `Cast::{BeginDrain, ForceStopConns}`.
Re-probed against current smarm: `request_stop` unwinds an untimed
`wait_readable` park, is no longer lossy against a QUEUED actor, and
supervisor shutdown joins in ~200us.
- **Config split**: `scheduler_threads` / `max_actors` are runtime knobs an
endpoint-as-child cannot honour -> `serve_with(cfg, smarm::Config, pipe)`
and `serve_with_shutdown(cfg, smarm::Config, pipe, signal)`. New
`Config.name` (default `"urus"`) is the endpoint's registry name;
`urus::endpoint::whereis(name)` for introspection. Two endpoints in one
process need distinct names.
- `serve*` stay as thin wrappers building a one-child tree with
`Shutdown::Infinity`. `Handle`/`ShutdownSignal` stay: a `serve*` caller
never sees the runtime, so it has no `RuntimeHandle` to reach for — but
the poll behind it is gone.
- `examples/crud.rs` is the demonstrator (app-owned tree, store as an
ordered sibling registered under a typed `Name`, shutdown via
`rt.handle().request_shutdown(root_sup)`); it loses its `OnceLock` store
cell and its `SHUTTING_DOWN` flag + 250ms poll. Other examples stay short
on `serve*`.
**Open after this cycle:**
- One unreproduced test failure seen once in ~120 full-suite runs (output
discarded by the loop that caught it; not reproduced in 60x lib + 20x
integration + 10x concurrent + 15x full since). Candidates: the
`free_port()` bind race the harness already documents, or a timing margin
in the new endpoint tests. Grab the test name next time it fires.
- `endpoint()` returns `impl Fn()`, so calling it twice = two endpoints
contending for one `Config.name` (second panics on the clash). Honest
failure, but the type doesn't prevent the mistake; a consume-on-first-use
newtype would.
## Known bugs ## Known bugs
- ~~**HTTP/1.0 keep-alive: server honors but never advertises**~~ FIXED - ~~**HTTP/1.0 keep-alive: server honors but never advertises**~~ FIXED
@@ -378,9 +452,9 @@ reconnect cycles.
- **Bench suite** — `urus-bench-spec.md` exists in the artefact store; - **Bench suite** — `urus-bench-spec.md` exists in the artefact store;
wire it up once v0.2 lands (supervision changes the hot path not at all, wire it up once v0.2 lands (supervision changes the hot path not at all,
but prove it). but prove it).
- **Operator introspection via typed names** — the v0.2-era - **Operator introspection via typed names** — partly delivered in v0.7:
`urus.server` / `urus.listener.{i}` pid tags were dropped in the the endpoint is a named gen_server (`Config.name`, default `"urus"`),
RFC 014 port (names are messageable endpoints now, self-registered reachable via `urus::endpoint::whereis(name)`, and answers
with a real `Sender<M>`). If wanted back, do it properly: register the `Call::ConnCount`. Still open: per-listener visibility (the pool is
serve loop's shutdown/control channel under a typed `urus.server` internal and anonymous), and a richer stats call (ws/channel counts,
name instead of faking pid tags with unit channels. request rates) — the endpoint is the place to hang them.