Tree snapshot of d9c62a8 (2026-08-18). The 20 source commits between
16ef583 (c9) and d9c62a8 were never pushed and the clone that held them
was lost; this commit carries their combined tree verbatim so the build
history stays auditable from the c1–c9 commits below it. Original
hashes as recorded in the session handoff:
c10 f03e94d pid targeting + auto-serialization (RemotePid, D14 name
on the wire); Phase 3 gate
c11 7ef4bad DownReason::Disconnected, wire tag 5
c12 d124162 remote monitors (Monitor/Demonitor/Down frames)
c13 9de967b connection-loss synthesis (A+B: Monitors::teardown +
unread-command Disconnected); Phase 4 gate
c14 7e822b7 eager pg eviction (reaper actor, ReaperInboxes)
dbe1a22 InboundVerdict::label(), trace::Event::ClusterInbound
31a9877 tests/channel.rs monitor-churn target gated on `go`
653559e Discovery::Withdrawn{name, addr}
c15 b41d76e distributed pg: Sync on NodeUp, Join/Leave broadcast,
NodeDown sweep, members_all; PgMsg wire type
c16 fafa881 pick_any / dispatch_any; Phase 5 complete
Phase 6 Tier A:
195c73e p4 NodeEvent::NodeDown(NodeInfo)
48fd766 p1 connector Candidate{name, addr, state}
ce8cf99 p2+p7 conn.rs select arms as Vec<Arm>; Outbound::Drained
bf24988 p6 RemotePid::from_local -> Option
9ae0380 p3 PeerStanding{Free, Claimed, Dialing}
c7d62a1 p11 cluster::Timing knobs, threaded by value
46f171d p11 cluster_disconnect un-ignored on SMARM_FAST_TIMING
Phase 6 Tier B:
391a9ae p5 cluster::RemoteDownReason{Local, Disconnected};
DownReason::Disconnected removed from core
7ddd908 p9 pg ctl channel unconditional, one cfg seam at spawn
d9c62a8 PeerNameMismatch parks the candidate; ClusterDial trace
Verified at d9c62a8: default 361/0, cluster 448/0, clippy --lib on
default / cluster / cluster+smarm-trace, fmt, 10x flake on
cluster_dial_mismatch, 5x on cluster_pg.
The timeout arm of the connection actor's select: HEARTBEAT_INTERVAL (1s)
paces outbound Frame::Heartbeat (first at spawn, so the peer's window
starts fed) and LIVENESS_TIMEOUT (4s = 4 intervals) declares the peer dead
when no inbound frame arrives inside it — any frame resets the window, so
heartbeats keep an idle connection alive and real traffic (c8+) counts for
free. Fire => close + exit; the manager's monitor reaps the table entry as
on every other exit path. Fixed timeout per RFC v2 §5 (control connection,
heartbeats can't queue behind bulk). Intervals are the one-viable-answer
call flagged for veto at diff review.
The pump was made non-blocking to keep the deadlines honest: a plain recv()
blocks into the socket while the buffer holds a partial frame, parking the
actor past its timers. Two additive FramedConn methods (read_once,
next_buffered): exactly one socket read per level-triggered readable wake
(cannot block, cannot strand — leftovers re-signal), then drain every
complete buffered frame. Liveness resets only on complete frames.
No-fd transports (loopback) still get the command-only loop: no readiness
means no timers, same caveat as recv_deadline.
tests/cluster_conn_liveness.rs 3/0, stable over 5 runs (raw far end over
localhost TCP: heartbeats appear unprompted; mute peer still up at half
the window, gone after it; heartbeat-only peer survives 1.5x the window,
then reaped once silenced). Cluster suites regression-clean; clippy --lib
green both configs; fmt clean.