feat(cluster): RFC 010 c6a — connection actor + manager subtree

Per-peer connection actor as a single select-loop plain actor owning the
whole FramedConn: one select folds its command inbox and the transport's
readable arm, so reads and control share one execution context — no reader
thread, no read/write split. The handshake is bypassed here (c6b wires it);
the actor is spawned already-established and self-registers with the manager.

Manager gen_server: the peer-name -> conn-pid registry and the uniqueness
source the handshake's NameTaken depends on. It monitors each connection, so
the table self-heals on any exit path. Explicit supervision subtree keeps the
manager up; connections are dynamic and monitored, never restarted (c7
re-dials).

Transport gains an additive Conn::readable_arm -> Option<FdArm> (default None;
TCP returns its fd's arm, loopback stays None). Existing c3 transport tests
unchanged.

Lifecycle test over localhost TCP: up reflected in the table, commanded
shutdown reaps exactly one, peer EOF reaps the other.
This commit is contained in:
Claude
2026-08-14 19:23:39 +00:00
parent 9e49038474
commit bbaaa062e3
6 changed files with 410 additions and 2 deletions
+104
View File
@@ -0,0 +1,104 @@
//! RFC 010 c6a — connection-actor lifecycle against the manager table.
//!
//! The handshake is bypassed here (c6b wires it): each connection is
//! constructed already-established over a real localhost TCP pair, handed a
//! fabricated `Peer`, and spawned. The actor registers with the manager, which
//! monitors it, so the table reflects the connection while it lives and reaps
//! it on any exit path. This proves three things at once: a live connection
//! shows up, a commanded shutdown removes exactly that one, and a peer close
//! (EOF, no command) removes the other.
//!
//! TCP parks the calling actor, so everything runs inside `smarm::run`; the
//! single-threaded runtime is fine because every wait is a cooperative fd park.
#![cfg(feature = "cluster")]
use std::time::Duration;
use smarm::cluster::envelope::NodeMeta;
use smarm::cluster::handshake::Peer;
use smarm::cluster::manager::{Call, Manager, Reply, MANAGER};
use smarm::cluster::spawn_established;
use smarm::cluster::transport::tcp::TcpTransport;
use smarm::cluster::transport::{Conn, FramedConn, Transport};
use smarm::gen_server::{self, GenServerBuilder};
use smarm::pg::Incarnation;
use smarm::{run, sleep};
/// A fabricated post-handshake peer identity. Only `node_name` matters to the
/// manager table; the rest is filler until c7 consumes it.
fn peer(name: &str) -> Peer {
Peer {
node_name: name.to_string(),
incarnation: Incarnation::new(1),
meta: NodeMeta {
role: "test".to_string(),
region: "test".to_string(),
},
}
}
/// One established transport pair over localhost. Relies on TCP backlog so the
/// sequential dial-then-accept needs no concurrent acceptor (same assumption as
/// the c3 conformance suite).
fn pair(t: &dyn Transport) -> (Box<dyn Conn>, Box<dyn Conn>) {
let mut l = t.listen("127.0.0.1:0").unwrap();
let a = t.dial(&l.local_addr()).unwrap();
let b = l.accept().unwrap();
(a, b)
}
/// Poll the manager until its peer set matches `expected` (sorted), or fail.
/// The bound is generous against a sub-millisecond real cost.
fn wait_peers(expected: &[&str]) {
let want: Vec<String> = expected.iter().map(|s| s.to_string()).collect();
for _ in 0..2000 {
if let Ok(Reply::Peers(got)) = gen_server::call(MANAGER, Call::Peers) {
if got == want {
return;
}
}
sleep(Duration::from_millis(1));
}
let got = gen_server::call(MANAGER, Call::Peers);
panic!("timed out waiting for peers == {want:?}; last = {got:?}");
}
#[test]
fn connection_up_commanded_shutdown_and_eof_all_reflected_in_table() {
run(|| {
// The manager, started plainly and reachable at its well-known name.
// (The supervised subtree in `cluster::start` is permanent by design;
// a plainly-started manager lets this test terminate cleanly.)
let mgr = GenServerBuilder::new(Manager::new())
.named(MANAGER)
.start()
.expect("manager name is free");
let t = TcpTransport;
let (a1, b1) = pair(&t);
let (a2, b2) = pair(&t);
// Manage the `a` ends as peers node-b and node-c; keep the `b` far ends
// open so neither socket is closed from the far side yet.
let h1 = spawn_established(FramedConn::new(a1), peer("node-b"));
let _h2 = spawn_established(FramedConn::new(a2), peer("node-c"));
// Up: both connections register and the table shows them.
wait_peers(&["node-b", "node-c"]);
// A commanded shutdown reaps exactly its own connection.
h1.shutdown();
wait_peers(&["node-c"]);
// A peer close (EOF) reaps the other with no command at all.
drop(b2);
wait_peers(&[]);
// node-b's far end stayed open until here, so its removal above was the
// shutdown command and not an EOF.
drop(b1);
// All connection actors have exited; stop the manager so `run` returns.
mgr.shutdown();
});
}