feat(conn_registry): registry owns the drain — trapping gen_server, request_shutdown driven

The drain protocol moves wholesale into the registry (smarm >=0.7:
trap_exit + handle_shutdown + gen_server timers + StopHandle):

- handle_shutdown: flip draining, stop idle conns, arm a drain_timeout
  timer, Continue; with no conns, Exit immediately.
- ConnIdle while draining stops the conn (unchanged); ConnStarted while
  draining stops it on arrival (unchanged); ConnEnded that empties the
  set while draining = StopHandle::stop() — the registry's own normal
  exit is now the 'every connection is gone' barrier.
- handle_timer (deadline): one force-stop sweep. The old re-sweep-every-
  10ms loop existed to catch late registrants; stop-on-arrival already
  covers every post-sweep entry path, so one sweep suffices.
- Cast::{BeginDrain,ForceStopConns} deleted (internal now); ConnCount
  stays as the introspection call. start() takes drain_timeout.
- serve.rs: the root's whole drain/poll block collapses to
  registry.shutdown() (graceful, monitors until the server has stopped
  itself). SHUTDOWN_POLL + the listener flag are untouched here; they go
  with the endpoint refactor.
- Known residual window documented in the module docs: a conn spawned
  by a dying listener that has not yet registered can outlive an
  already-empty registry; it is collected by smarm's root-exit sweep.

Tests (in-lib, request_shutdown driven): empty-set immediate exit;
idle-now/busy-at-deadline ordering with exit-after; stop-on-idle
mid-drain; stop-on-arrival mid-drain. 35x hammer subset green.
This commit is contained in:
Claude
2026-08-20 08:57:19 +00:00
parent 8bdec97842
commit 5014b870d6
2 changed files with 276 additions and 87 deletions
+12 -29
View File
@@ -13,7 +13,7 @@
//! waiting on "the same fd").
use crate::conn_actor::{run_connection, ConnLimits};
use crate::conn_registry::{self, Call, Cast, ConnRegistry, Reply};
use crate::conn_registry::{self, ConnRegistry};
use crate::net::{accept_nonblocking, bind_and_listen, OwnedFd};
use crate::plug::Pipeline;
@@ -439,7 +439,7 @@ pub fn serve_with_shutdown(
rt.run(move || {
// Registry first: listeners and conns cast into it from birth.
let registry = conn_registry::start();
let registry = conn_registry::start(drain_timeout);
let mut sup = OneForOne::new().strategy(Strategy::OneForOne);
for (i, lfd) in listener_fds.into_iter().enumerate() {
@@ -488,34 +488,17 @@ pub fn serve_with_shutdown(
shutdown_flag.store(true, Ordering::Relaxed);
let _ = sup_h.join();
// 3 + 4. Drain. Same sweep discipline as listeners on the force-
// stop path: a conn accepted just before its listener died may
// register after the deadline, so keep force-stopping until the
// set is empty (each pass kills everything registered; new
// registrants are a strictly shrinking population once listeners
// are gone).
let _ = registry.cast(Cast::BeginDrain);
let deadline = std::time::Instant::now() + drain_timeout;
let mut force = false;
loop {
match registry.call(Call::ConnCount) {
Ok(Reply::ConnCount(0)) => break,
Ok(_) => {}
Err(_) => break, // registry gone; nothing left to track
}
let now = std::time::Instant::now();
if force || now >= deadline {
force = true;
let _ = registry.cast(Cast::ForceStopConns);
smarm::sleep(Duration::from_millis(10));
} else {
smarm::sleep(Duration::from_millis(50).min(deadline - now));
}
}
// 3 + 4. Drain: the registry owns the whole protocol now (idle
// stopped immediately, busy until its internal drain_timeout
// timer, late registrants stopped on arrival — see
// conn_registry docs). `GenServerRef::shutdown()` delivers the
// request and blocks on a monitor until the registry has
// stopped itself, which it does only once the conn set is
// empty: this line IS the "every connection is gone" barrier.
registry.shutdown();
// 5. Our GenServerRef drops here. The registry's inbox closes once
// the last conn's clone drops with it, and the runtime winds
// down when the last actor exits.
// 5. Root returns; the runtime winds down when the last actor
// exits.
});
Ok(())