feat(cluster): RFC 010 c6b — handshake on the accept/connect path
Drive the c5 machines as straight-line code on the path (D8): dial_handshake and accept_handshake do the IO on a shared FramedConn, and a connection actor is spawned only after a successful handshake. Rejects, tie-break losses (D7), protocol faults and timeouts are all resolved on the path by closing, so no actor ever exists for a connection that did not establish. The whole FramedConn travels into spawn_established, carrying any read-ahead past the handshake frames. Handshake deadlines land here rather than in c6c: FramedConn::recv_deadline enforces them between reads via the connection's fd arm, so a peer that connects and goes silent cannot wedge the acceptor. Connection lifetime moves to the manager (pulled forward from c7). The path registers each established connection and hands over its ConnHandle; the manager owns it, monitors the actor, and tears the connection down on Disconnect, on peer close, or at manager shutdown. spawn_established returns a Pid, so a connection neither outlives nor dies with whichever actor established it — the ownership that made two-node teardown unorderable. The manager also tracks in-flight dial intents, monitored so a panicking dial cannot wedge the tie-break, and answers HelloCtx for the accept path.
This commit is contained in:
+48
-23
@@ -10,60 +10,85 @@
|
||||
//! so no read/write split is needed.
|
||||
//!
|
||||
//! The handshake completes *before* this actor exists (on the accept/connect
|
||||
//! path — c6b) and produces the [`Peer`]; the actor registers that peer with
|
||||
//! the [`manager`](crate::cluster::manager), which monitors it so any exit
|
||||
//! deregisters the connection. Heartbeat send and fixed-timeout liveness join
|
||||
//! the loop in c6c (the timeout arm of the same `select`).
|
||||
//! path — c6b) and produces the [`Peer`]; the *path* then registers the
|
||||
//! connection with the [`manager`](crate::cluster::manager), which takes
|
||||
//! ownership of its [`ConnHandle`] and monitors the actor, so any exit
|
||||
//! deregisters the connection. The actor itself holds no authority over its
|
||||
//! own lifetime: it runs until the manager drops its handle (deregistration,
|
||||
//! `Disconnect`, or manager shutdown) or the connection ends. Heartbeat send
|
||||
//! and fixed-timeout liveness join the loop in c6c (the timeout arm of the
|
||||
//! same `select`).
|
||||
|
||||
use crate::channel::{channel, Receiver, Selectable, Sender};
|
||||
use crate::cluster::handshake::Peer;
|
||||
use crate::cluster::manager::{Call, Registered, Reply, MANAGER};
|
||||
use crate::cluster::transport::FramedConn;
|
||||
use crate::gen_server;
|
||||
use crate::scheduler::{self, spawn};
|
||||
use crate::pid::Pid;
|
||||
use crate::scheduler::spawn;
|
||||
|
||||
/// Commands to a running connection actor.
|
||||
enum Cmd {
|
||||
Shutdown,
|
||||
}
|
||||
|
||||
/// A handle to a running connection actor.
|
||||
/// The manager's authority over one connection actor: while this handle
|
||||
/// lives the connection lives, and dropping it stops the actor and closes
|
||||
/// the socket. Only the [`manager`](crate::cluster::manager) holds one —
|
||||
/// callers of [`spawn_established`] get a [`Pid`] and no lifetime authority,
|
||||
/// so a connection can never outlive, or die with, whichever actor happened
|
||||
/// to establish it.
|
||||
pub struct ConnHandle {
|
||||
cmd_tx: Sender<Cmd>,
|
||||
}
|
||||
|
||||
impl std::fmt::Debug for ConnHandle {
|
||||
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
|
||||
f.write_str("ConnHandle")
|
||||
}
|
||||
}
|
||||
|
||||
impl ConnHandle {
|
||||
/// Ask the connection to close and exit. Idempotent, and a no-op if the
|
||||
/// actor has already gone.
|
||||
/// actor has already gone. Dropping the handle does the same thing; this
|
||||
/// exists for the manager's explicit `Disconnect` path.
|
||||
pub fn shutdown(&self) {
|
||||
let _ = self.cmd_tx.send(Cmd::Shutdown);
|
||||
}
|
||||
}
|
||||
|
||||
/// Spawn a connection actor for an **already-established** connection: the
|
||||
/// handshake has completed elsewhere and produced `peer`. Returns as soon as
|
||||
/// the actor is spawned; the actor's first act is to register with the manager.
|
||||
pub fn spawn_established(framed: FramedConn, peer: Peer) -> ConnHandle {
|
||||
let (cmd_tx, cmd_rx) = channel();
|
||||
spawn(move || run(framed, peer, cmd_rx));
|
||||
ConnHandle { cmd_tx }
|
||||
}
|
||||
/// The name was already claimed by a live connection, so this one was
|
||||
/// refused; its actor has been stopped and its socket closed.
|
||||
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
|
||||
pub struct RegisterRefused;
|
||||
|
||||
fn run(mut framed: FramedConn, peer: Peer, cmd_rx: Receiver<Cmd>) {
|
||||
let me = scheduler::self_pid();
|
||||
/// Spawn a connection actor for an **already-established** connection (the
|
||||
/// handshake completed on the path and produced `peer`) and register it with
|
||||
/// the manager, synchronously, before returning. The manager takes the
|
||||
/// actor's [`ConnHandle`]; the caller gets only the [`Pid`], because
|
||||
/// connection lifetime belongs to the table and not to the establishing
|
||||
/// actor. A refusal has already stopped the actor and closed the socket.
|
||||
pub fn spawn_established(framed: FramedConn, peer: Peer) -> Result<Pid, RegisterRefused> {
|
||||
let (cmd_tx, cmd_rx) = channel();
|
||||
let name = peer.node_name.clone();
|
||||
let pid = spawn(move || run(framed, peer, cmd_rx)).pid();
|
||||
match gen_server::call(
|
||||
MANAGER,
|
||||
Call::Register {
|
||||
name: peer.node_name.clone(),
|
||||
pid: me,
|
||||
name,
|
||||
pid,
|
||||
handle: ConnHandle { cmd_tx },
|
||||
},
|
||||
) {
|
||||
Ok(Reply::Registered(Registered::Ok)) => {}
|
||||
// Duplicate name, or the manager is unreachable: do not run. The
|
||||
// connection drops (closing the socket) as `framed` falls out of scope.
|
||||
_ => return,
|
||||
Ok(Reply::Registered(Registered::Ok)) => Ok(pid),
|
||||
// Duplicate name, or the manager is unreachable. Either way the
|
||||
// handle went with the call and is dropped there (or never arrived
|
||||
// and dropped with it), which stops the actor and closes the socket.
|
||||
_ => Err(RegisterRefused),
|
||||
}
|
||||
}
|
||||
|
||||
fn run(mut framed: FramedConn, _peer: Peer, cmd_rx: Receiver<Cmd>) {
|
||||
loop {
|
||||
match framed.readable_arm() {
|
||||
Some(arm) => {
|
||||
|
||||
Reference in New Issue
Block a user