feat(cluster): RFC 010 c9 — remote Name sends: outbound table + the one inbound seam

Outbound (D13, ratified 2026-08-15): a module-private table node → dedicated
Sender<Frame>, populated/torn down by the manager inside the same serialized
handlers that own connection lifetime (Register / Disconnect / reap /
terminate), on RuntimeInner beside the exposure state. remote::send is one
leaf-lock lookup + one channel send: no gen_server on the data plane (rejected
send-through-manager — worse than the c8 argument, it is the data plane), no
published channel a leaked conn pid could inject raw frames into (rejected
send_dyn-at-conn-pid — module privacy cannot fence a published channel).
Ok(()) = handed to the connection's inbox, local knowledge only (RFC §3);
missing entry / closed channel = NotConnected. The outbound sender is a
SEPARATE channel from cmd_tx on purpose: a clone would let the table hold
lifetime authority (D9 violation); its closure is not a stop signal.
Buffering toward a slow peer is unbounded (BEAM busy_dist_port shape);
backpressure out of scope, documented not silently absent.

Inbound: remote::deliver_named is THE resolution seam (RFC v2) — exposed-set
check (unexposed = unreachable, the safety) → hash check against what the
name was exposed with → registry whereis → c8's decode_deliver. Verdicts are
local-only (InboundVerdict, discarded by the conn actor for now; a trace
hook is the place). The conn actor gains a third select arm (outbound
inbox → wire) and interprets SendNamed; Send/Monitor/Down are consumed for
liveness and ignored until c10/c11.

Public surface: RemoteName<M> (node, Name<M>), RemoteSendError, NotConnected,
send, send_remote_raw (untyped escape hatch so tests can put deliberately
wrong frames on the wire — the typed API cannot express a hash mismatch).

CORE (found the hard way this chunk): publishing a second channel of the
same message type on one live actor silently replaced and CLOSED the first,
so a select/recv on it returned 'closed' immediately forever — a hot loop
starving the single-threaded scheduler. publish_channel now asserts when an
existing same-type channel's receiver is still alive AND the new sender is
not a clone of it (Sender::same_channel via Arc::ptr_eq); replacing a
dead-receiver channel stays silent (an actor re-registering after dropping
its inbox is legit). Message names the sanctioned shapes. Three registry
tests pin panic / cloned-sender-ok / dead-receiver-ok. Full default suite
clean.

Harness: wait_line drains stderr before dumping on timeout/EOF (eprintln!
diagnostics no longer vanish); Node::transcript() accessor for
ordering-proof assertions.

tests/cluster_remote_send.rs 1/0, 10/10 flake runs, subprocess harness:
cross-node name-send delivers; unexposed name unreachable (registered
locally, never delivered); wrong hash never misroutes (unknown-hash and
known-hash-wrong-channel flavours); send to an unconnected node =
NotConnected locally; every frame-bearing send is Ok — 'handed to
transport', asserted and documented at the test. Negatives are proven by
STREAM ORDERING (they precede the positive on one in-order connection),
not by sleeping. All cluster suites regression-clean (envelope 15,
handshake 11, transport 11, lifecycle 1, liveness 3, connect 9, two_node 3,
membership 4, mesh 2, expose 5); clippy --lib green both configs; fmt clean;
default build compiles.
This commit is contained in:
Claude
2026-08-15 21:22:56 +00:00
parent a7f98f8d48
commit 16ef583455
10 changed files with 688 additions and 16 deletions
+220
View File
@@ -0,0 +1,220 @@
//! RFC 010 c9 — remote `Name` sends: the outbound seam and the single
//! inbound name-resolution seam, cross-process.
//!
//! Two node processes each run the integrated `cluster::start`. The
//! *receiver* registers a `String` inbox under a name and exposes it (and
//! registers a second name it does NOT expose); the *sender* waits for
//! `node_up`, then sends. Facts cross as stdout lines: `LISTENING <addr>`,
//! `MEMBER-UP <name>`, `GOT <payload>`, `SEND-RESULT <case> <verdict>`.
//! Roles park forever afterwards (retractable-state trap); the parent
//! SIGKILLs via `Drop`.
//!
//! What is asserted at each end (roadmap-binding):
//! - cross-node name-send delivers the payload;
//! - an unexposed name is unreachable — the receiver's inbox stays empty
//! even though the name IS registered locally;
//! - a wrong type hash is a decode failure at the receiver, never a
//! misroute — the `String` inbox does not see a `u64` delivered under a
//! made-up hash, nor a `u64` under `u64`'s hash;
//! - a send to a disconnected (never-connected) node fails locally with
//! `NotConnected`, and `Ok(())` means only "handed to the transport".
//!
//! Timing note for the "stays empty" assertions: they are proven by
//! ORDERING, not by waiting — the sender emits the negative-case frames
//! BEFORE the positive one on the same connection (in-order stream), so when
//! the receiver has seen the positive payload, the negatives have already
//! been processed and refused. No sleep-and-hope.
#![cfg(feature = "cluster")]
mod common;
use common::{maybe_child, spawn_node, Node};
use smarm::cluster::envelope::NodeMeta;
use smarm::cluster::expose::expose;
use smarm::cluster::membership::{subscribe, NodeEvent};
use smarm::cluster::remote::{send_remote_raw, RemoteName, RemoteSendError};
use smarm::cluster::{start, Config, StaticSeeds};
use smarm::{channel, register, Name};
use std::time::Duration;
const ROLES: &[(&str, fn())] = &[("receiver", role_receiver), ("sender", role_sender)];
const INBOX: Name<String> = Name::new("c9.inbox");
const HIDDEN: Name<String> = Name::new("c9.hidden");
fn base_config(name: &str, seeds: Vec<(String, String)>) -> Config {
Config {
node_name: name.to_string(),
meta: NodeMeta {
role: "c9".to_string(),
region: "local".to_string(),
},
listen_addr: "127.0.0.1:0".to_string(),
strategy: Box::new(StaticSeeds::new(seeds)),
}
}
/// Receiver: register + expose INBOX; register HIDDEN unexposed **in a
/// separate actor** (one actor holds one channel per message type — a
/// second `register` of the same `M` on one actor silently replaces the
/// first, closing it); print every payload that lands in either.
fn role_receiver() {
smarm::run(|| {
let cluster = start(base_config("recv", vec![])).expect("listener binds");
println!("LISTENING {}", cluster.local_addr());
// HIDDEN's holder: its own actor, so its String channel does not
// displace INBOX's on the root actor.
let (hidden_ready_tx, hidden_ready_rx) = channel::<()>();
smarm::spawn(move || {
let (hid_tx, hid_rx) = channel::<String>();
register(HIDDEN, hid_tx).unwrap();
hidden_ready_tx.send(()).unwrap();
loop {
match hid_rx.recv() {
Ok(s) => println!("GOT-HIDDEN {s}"),
Err(_) => break,
}
}
});
hidden_ready_rx.recv().unwrap();
let (in_tx, in_rx) = channel::<String>();
register(INBOX, in_tx).unwrap();
expose(INBOX);
println!("READY");
loop {
match in_rx.recv() {
Ok(s) => println!("GOT {s}"),
Err(_) => break,
}
}
loop {
smarm::sleep(Duration::from_secs(3600));
}
});
}
/// Sender: connect to recv, wait for node_up, then in this ORDER on the one
/// connection: hidden-name send, wrong-hash sends (two flavours), then the
/// positive send. Plus a send to a node that is not connected at all.
fn role_sender() {
let recv_addr = std::env::var("SMARM_RECV_ADDR").expect("SMARM_RECV_ADDR");
smarm::run(move || {
let _cluster = start(base_config("send", vec![("recv".to_string(), recv_addr)]))
.expect("listener binds");
let events = subscribe().expect("manager is up");
loop {
match events.rx.recv() {
Ok(NodeEvent::NodeUp(info)) if info.name == "recv" => break,
Ok(_) => continue,
Err(_) => panic!("manager gone"),
}
}
println!("MEMBER-UP recv");
// Not connected: purely local knowledge, no frame leaves.
let ghost: RemoteName<String> = RemoteName::new("nowhere", INBOX);
let r = smarm::cluster::remote::send(ghost, "lost".to_string());
println!(
"SEND-RESULT not-connected {}",
match r {
Err(RemoteSendError::NotConnected(_)) => "NotConnected",
Ok(()) => "Ok",
Err(_) => "OtherErr",
}
);
// Unexposed name at the peer: the frame goes (local knowledge can't
// know the peer's exposed set) and the peer refuses it.
let hidden: RemoteName<String> = RemoteName::new("recv", HIDDEN);
let r = smarm::cluster::remote::send(hidden, "should not land".to_string());
println!(
"SEND-RESULT hidden {}",
if r.is_ok() { "Ok" } else { "Err" }
);
// Wrong hash, two flavours: (a) a u64 payload under a made-up hash
// (unknown type at the peer); (b) a u64 payload under u64's real
// hash against a String-typed name (decoder known, wrong channel).
// Both are raw sends — the typed API cannot express them, by design.
let bogus = 0xdead_beef_u64;
let r = send_remote_raw(
"recv",
"c9.inbox",
bogus,
&smarm::cluster::envelope::encode_payload(&7u64).unwrap(),
);
println!(
"SEND-RESULT wrong-hash-unknown {}",
if r.is_ok() { "Ok" } else { "Err" }
);
let r = send_remote_raw(
"recv",
"c9.inbox",
smarm::cluster::expose::type_hash::<u64>(),
&smarm::cluster::envelope::encode_payload(&7u64).unwrap(),
);
println!(
"SEND-RESULT wrong-hash-known {}",
if r.is_ok() { "Ok" } else { "Err" }
);
// Positive: last on the stream, so its arrival proves the negatives
// were already processed.
let inbox: RemoteName<String> = RemoteName::new("recv", INBOX);
let r = smarm::cluster::remote::send(inbox, "hello from send".to_string());
println!(
"SEND-RESULT positive {}",
if r.is_ok() { "Ok" } else { "Err" }
);
loop {
smarm::sleep(Duration::from_secs(3600));
}
});
}
fn wait_send_result(node: &mut Node, case: &str) -> String {
let prefix = format!("SEND-RESULT {case} ");
let line = node.wait_line(&prefix, |l| l.starts_with(&prefix));
line[prefix.len()..].to_string()
}
#[test]
fn remote_name_send_delivers_and_refusals_never_misroute() {
maybe_child(ROLES);
let mut recv = spawn_node("receiver", &[]);
let addr = recv.wait_listening();
recv.wait_line("READY", |l| l == "READY");
let mut send = spawn_node("sender", &[("SMARM_RECV_ADDR", &addr)]);
send.wait_line("MEMBER-UP recv", |l| l == "MEMBER-UP recv");
// Local-knowledge-only failure for an unknown node.
assert_eq!(wait_send_result(&mut send, "not-connected"), "NotConnected");
// Every frame-bearing send is Ok — Ok means "handed to the transport",
// nothing about what the peer does with it (RFC §3, documented here).
assert_eq!(wait_send_result(&mut send, "hidden"), "Ok");
assert_eq!(wait_send_result(&mut send, "wrong-hash-unknown"), "Ok");
assert_eq!(wait_send_result(&mut send, "wrong-hash-known"), "Ok");
assert_eq!(wait_send_result(&mut send, "positive"), "Ok");
// The positive payload lands...
recv.wait_line("GOT hello from send", |l| l == "GOT hello from send");
// ...and, by stream ordering, every negative before it was refused: no
// GOT for the wrong-hash frames, no GOT-HIDDEN at all. The transcript
// up to this point is the proof.
let transcript = recv.transcript();
let gots: Vec<&str> = transcript
.iter()
.map(|s| s.as_str())
.filter(|l| l.starts_with("GOT"))
.collect();
assert_eq!(
gots,
["GOT hello from send"],
"exactly one delivery, the exposed one"
);
}
+11
View File
@@ -133,6 +133,13 @@ impl Node {
}
}
/// Every stdout/stderr line seen so far, in arrival order. For
/// ordering-proof assertions ("by the time X arrived, Y had not").
#[allow(dead_code)]
pub fn transcript(&self) -> &[String] {
&self.transcript
}
/// Wait until a stdout line satisfies `pred`; return it. Panics with the
/// full transcript after [`WAIT`]. `what` names the expectation in the
/// panic message.
@@ -149,6 +156,9 @@ impl Node {
}
}
Err(RecvTimeoutError::Timeout) => {
// Pull in whatever stderr arrived since the last drain,
// so a role's eprintln! diagnostics survive into the dump.
self.drain_stderr();
panic!(
"node {:?}: timed out waiting for {what} after {WAIT:?}; transcript:\n{}",
self.role,
@@ -156,6 +166,7 @@ impl Node {
);
}
Err(RecvTimeoutError::Disconnected) => {
self.drain_stderr();
panic!(
"node {:?}: output closed while waiting for {what}; transcript:\n{}",
self.role,
+43
View File
@@ -258,3 +258,46 @@ fn send_dyn_to_dead_pid_is_dead() {
assert!(matches!(send_dyn::<u64>(p, 1u64), Err(SendError::Dead(_))));
});
}
// --- one channel per message type per actor -----------------------------------
/// Registering a second name of the same message type on one actor, with a
/// *fresh* channel, would silently replace and close the first — so it
/// panics (found in RFC 010 c9). The sanctioned shapes stay quiet: bind both
/// names to a clone of one sender, or use two actors.
#[test]
#[should_panic(expected = "already publishes a live channel")]
fn second_live_channel_of_same_type_on_one_actor_panics() {
run(|| {
let (tx1, _rx1) = channel::<u64>();
let (tx2, _rx2) = channel::<u64>();
register(Name::<u64>::new("dup-a"), tx1).unwrap();
register(Name::<u64>::new("dup-b"), tx2).unwrap(); // panics
});
}
#[test]
fn two_names_on_one_cloned_sender_is_fine() {
run(|| {
let (tx, rx) = channel::<u64>();
register(Name::<u64>::new("twin-a"), tx.clone()).unwrap();
register(Name::<u64>::new("twin-b"), tx).unwrap();
send(Name::<u64>::new("twin-a"), 1).unwrap();
send(Name::<u64>::new("twin-b"), 2).unwrap();
assert_eq!(rx.recv().unwrap(), 1);
assert_eq!(rx.recv().unwrap(), 2);
});
}
#[test]
fn replacing_a_channel_whose_receiver_is_gone_is_fine() {
run(|| {
let (tx1, rx1) = channel::<u64>();
register(Name::<u64>::new("reborn"), tx1).unwrap();
drop(rx1); // old inbox gone: replacement is the honest thing to do
let (tx2, rx2) = channel::<u64>();
register(Name::<u64>::new("reborn-2"), tx2).unwrap();
send(Name::<u64>::new("reborn-2"), 9).unwrap();
assert_eq!(rx2.recv().unwrap(), 9);
});
}