2 Commits
Author SHA1 Message Date
claude-asm-audit 4c0e42152f feat(cluster): RFC 010 c10–c16, follow-ups and Phase 6 (squash of 16ef583..d9c62a8)
Tree snapshot of d9c62a8 (2026-08-18). The 20 source commits between
16ef583 (c9) and d9c62a8 were never pushed and the clone that held them
was lost; this commit carries their combined tree verbatim so the build
history stays auditable from the c1–c9 commits below it. Original
hashes as recorded in the session handoff:

  c10  f03e94d  pid targeting + auto-serialization (RemotePid, D14 name
                on the wire); Phase 3 gate
  c11  7ef4bad  DownReason::Disconnected, wire tag 5
  c12  d124162  remote monitors (Monitor/Demonitor/Down frames)
  c13  9de967b  connection-loss synthesis (A+B: Monitors::teardown +
                unread-command Disconnected); Phase 4 gate
  c14  7e822b7  eager pg eviction (reaper actor, ReaperInboxes)
       dbe1a22  InboundVerdict::label(), trace::Event::ClusterInbound
       31a9877  tests/channel.rs monitor-churn target gated on `go`
       653559e  Discovery::Withdrawn{name, addr}
  c15  b41d76e  distributed pg: Sync on NodeUp, Join/Leave broadcast,
                NodeDown sweep, members_all; PgMsg wire type
  c16  fafa881  pick_any / dispatch_any; Phase 5 complete
  Phase 6 Tier A:
       195c73e  p4  NodeEvent::NodeDown(NodeInfo)
       48fd766  p1  connector Candidate{name, addr, state}
       ce8cf99  p2+p7 conn.rs select arms as Vec<Arm>; Outbound::Drained
       bf24988  p6  RemotePid::from_local -> Option
       9ae0380  p3  PeerStanding{Free, Claimed, Dialing}
       c7d62a1  p11 cluster::Timing knobs, threaded by value
       46f171d  p11 cluster_disconnect un-ignored on SMARM_FAST_TIMING
  Phase 6 Tier B:
       391a9ae  p5  cluster::RemoteDownReason{Local, Disconnected};
                    DownReason::Disconnected removed from core
       7ddd908  p9  pg ctl channel unconditional, one cfg seam at spawn
       d9c62a8      PeerNameMismatch parks the candidate; ClusterDial trace

Verified at d9c62a8: default 361/0, cluster 448/0, clippy --lib on
default / cluster / cluster+smarm-trace, fmt, 10x flake on
cluster_dial_mismatch, 5x on cluster_pg.
2026-08-18 16:00:00 +02:00
Claude 16ef583455 feat(cluster): RFC 010 c9 — remote Name sends: outbound table + the one inbound seam
Outbound (D13, ratified 2026-08-15): a module-private table node → dedicated
Sender<Frame>, populated/torn down by the manager inside the same serialized
handlers that own connection lifetime (Register / Disconnect / reap /
terminate), on RuntimeInner beside the exposure state. remote::send is one
leaf-lock lookup + one channel send: no gen_server on the data plane (rejected
send-through-manager — worse than the c8 argument, it is the data plane), no
published channel a leaked conn pid could inject raw frames into (rejected
send_dyn-at-conn-pid — module privacy cannot fence a published channel).
Ok(()) = handed to the connection's inbox, local knowledge only (RFC §3);
missing entry / closed channel = NotConnected. The outbound sender is a
SEPARATE channel from cmd_tx on purpose: a clone would let the table hold
lifetime authority (D9 violation); its closure is not a stop signal.
Buffering toward a slow peer is unbounded (BEAM busy_dist_port shape);
backpressure out of scope, documented not silently absent.

Inbound: remote::deliver_named is THE resolution seam (RFC v2) — exposed-set
check (unexposed = unreachable, the safety) → hash check against what the
name was exposed with → registry whereis → c8's decode_deliver. Verdicts are
local-only (InboundVerdict, discarded by the conn actor for now; a trace
hook is the place). The conn actor gains a third select arm (outbound
inbox → wire) and interprets SendNamed; Send/Monitor/Down are consumed for
liveness and ignored until c10/c11.

Public surface: RemoteName<M> (node, Name<M>), RemoteSendError, NotConnected,
send, send_remote_raw (untyped escape hatch so tests can put deliberately
wrong frames on the wire — the typed API cannot express a hash mismatch).

CORE (found the hard way this chunk): publishing a second channel of the
same message type on one live actor silently replaced and CLOSED the first,
so a select/recv on it returned 'closed' immediately forever — a hot loop
starving the single-threaded scheduler. publish_channel now asserts when an
existing same-type channel's receiver is still alive AND the new sender is
not a clone of it (Sender::same_channel via Arc::ptr_eq); replacing a
dead-receiver channel stays silent (an actor re-registering after dropping
its inbox is legit). Message names the sanctioned shapes. Three registry
tests pin panic / cloned-sender-ok / dead-receiver-ok. Full default suite
clean.

Harness: wait_line drains stderr before dumping on timeout/EOF (eprintln!
diagnostics no longer vanish); Node::transcript() accessor for
ordering-proof assertions.

tests/cluster_remote_send.rs 1/0, 10/10 flake runs, subprocess harness:
cross-node name-send delivers; unexposed name unreachable (registered
locally, never delivered); wrong hash never misroutes (unknown-hash and
known-hash-wrong-channel flavours); send to an unconnected node =
NotConnected locally; every frame-bearing send is Ok — 'handed to
transport', asserted and documented at the test. Negatives are proven by
STREAM ORDERING (they precede the positive on one in-order connection),
not by sleeping. All cluster suites regression-clean (envelope 15,
handshake 11, transport 11, lifecycle 1, liveness 3, connect 9, two_node 3,
membership 4, mesh 2, expose 5); clippy --lib green both configs; fmt clean;
default build compiles.
2026-08-15 21:22:56 +00:00