Commit Graph
95 Commits
Author SHA1 Message Date
Claude a7f98f8d48 feat(cluster): RFC 010 c8 — exposure registry + fixed-seed type hashing
Nothing local is remotely reachable by default (RFC §4). expose(Name<M>)
marks a name remotely addressable and registers M's decoder under
type_hash::<M>(); expose_type::<M>() registers only the decoder (the
reply-to path). exposed_names() is the auditable remote surface.

D3's watchable fold, resolved against the code as it stands (stated in the
module docs): register() ALREADY stamps every named holder watchable ('no
successfully-registered actor can die unflagged', registry.rs), so an
exposed name's holder needs no extra mark — and re-registration after a
holder's death re-stamps the new holder for free, which a per-tenancy mark
taken at expose time could not do. The cluster's own mark_watchable
set-site is therefore the pid crossing the wire (frame serialization, c10)
— the exact analog of the membrane crossing. c8 adds only the name/type
state neither the registry nor slot bits can carry. No new pid registry;
RFC §4 honored.

One-viable calls, flagged:
- State lives on RuntimeInner (the pg pattern: leaf RawMutex field,
  cfg-gated behind cluster, zero-cost-when-off per c1) — c9's inbound
  decode consults it per frame; manager-held state would serialize every
  remote delivery through one gen_server.
- type_hash = FNV-1a 64 (fixed seed: the offset basis) over TypeId: a
  constant of the binary — stable across runs of the same build (the scope
  the build-hash handshake reduces the mesh to), deliberately not across
  builds. Collisions degrade to decode error / refused channel, never a
  misroute (the NoChannel guarantee, RFC §3).
- Decoder = decode-and-deliver-to-pid Arc closure capturing M (the one
  typed site): decode_payload then send_dyn. Wire-name → pid resolution
  stays OUTSIDE — that is c9's single seam, which calls decode_deliver.
  Arc so the call happens with the exposure lock RELEASED: send_dyn takes
  the registry lock, a mutual Leaf (the runtime asserts on nesting — caught
  live by the first test run).
- expose is a name-level fact, valid for an unregistered name (names
  late-bind; c9 resolves per delivery).

tests/cluster_expose.rs 5/0 stable x5, purely local per roadmap:
exposed/unexposed lookup + audit listing; decoder registration and the
delivery contract (happy path into a registered String channel; unknown
hash; corrupt bytes; wrong channel refused — never misrouted); distinct
types distinct hashes; expose/bridge-crossing agreement via the shared
watchable observable (terminal_reason after holder death); hash stability
across runs in the same binary via a c4-harness re-exec. Payload types are
std types — the crate's serde is derive-less by design, user crates bring
their own derive.

All cluster suites regression-clean (envelope 15, handshake 11, transport
11, lifecycle 1, liveness 3, connect 9, two_node 3, membership 4, mesh 2);
clippy --lib green both configs; fmt clean; default build compiles.
2026-08-15 07:25:42 +00:00
Claude 1282c3a08d feat(cluster): RFC 010 c7b — discovery Strategy, static seeds, connector dial loop
Phase 2 gate: 3-node mesh under the subprocess harness, repeatable (10/10).

Strategy (ratified): push-based, spawned as its own actor by the connector —
it emits Discovery events into a channel whenever it learns something and
may run forever; the connector owns all retry/backoff state. StaticSeeds
announces its list once and exits. Discovery is #[non_exhaustive] and
additive-only (candidates announced, never withdrawn) so expiry can land
later without breaking strategies.

One-viable correction to the ratified Discovery shape, flagged: a candidate
is a (name, addr) PAIR, not a bare address. The dial path and the D7
tie-break are keyed by peer name (the dial intent must be registered before
connecting so a crossing inbound Hello sees it), so an anonymous dial would
reintroduce exactly the simultaneous-connect flap D7 exists to prevent.
Discovery mechanisms know names — that is what they discover.

Connector: plain select-loop actor (the c6 shape) folding cmd inbox,
discovery stream, membership stream, and the earliest retry deadline into
one wait. It tracks who is up by SUBSCRIBING TO MEMBERSHIP like any
consumer — first consumer of c7a's snapshot-then-stream surface, no
privileged channel into the manager. Backoff: 250ms doubling to a 5s cap
(the c6c class of one-viable constants), reset on node_up; node_down
schedules a prompt redial with a fresh sequence. A candidate bearing the
local name is parked (that seed is us); every other failure retries — in
particular NameTaken can be our own ghost at the peer, not yet reaped by
its liveness timer, so it must not park. Dials run inline in the loop, the
acceptor's deliberate serialization (each attempt bounded by the connect +
handshake deadlines).

cluster::start(Config {node_name, meta, listen_addr, strategy}) is now the
integrated node start: supervised manager + acceptor + connector. It
completes the node identity: build_hash = cluster::BUILD_HASH (first
consumer, closing the c6d loose end) and incarnation = self_incarnation()
— unix-epoch MILLIS truncated to u32, not seconds: a supervised
crash-and-restart inside one second is routine, and seconds would collide
the ghost with its successor. Cluster handle: local_addr()/local()/
shutdown(); drop stops acceptor+connector loops, manager subtree detaches
(same split as AcceptorHandle alone).

Roadmap-binding, asserted in review: no consumer touches the connection
table — Manager.conns and ConnEntry stay private; the only exposures are
Call::Peers (sorted names, pre-existing) and the membership surface.

tests/cluster_mesh.rs 2/0, 10/10 flake runs: (1) 3-node mesh forms; kill
one (SIGKILL via Drop, per the retractable-state trap: roles park forever)
=> node_down at both survivors; restart same name => new incarnation at
every observer, distinguishable from the ghost; (2) seed unreachable at
start (pre-reserved closed port; accepted micro steal-window, documented)
then arriving later => edge forms via the retry path. All cluster suites
regression-clean (envelope 15, handshake 11, transport 11, lifecycle 1,
liveness 3, connect 9, two_node 3, membership 4); clippy --lib green both
configs; fmt clean; default build compiles.
2026-08-15 07:14:13 +00:00
Claude 160967939b feat(cluster): RFC 010 c7a — membership events + view at the manager
node_up/node_down are derived facts of the manager's own register/remove
events, so the membership state lives in the manager (no cross-actor race
between 'connection exists' and 'node is up'); src/cluster/membership.rs is
the consumer surface: NodeEvent/NodeInfo, subscribe(), view(). The conn
table stays private — no consumer touches it (roadmap-binding).

Ratified semantics: subscribe() is snapshot-then-stream — one NodeUp per
live peer is queued before the subscription joins the list, exact because
gen_server handlers are serialized. Dropped subscribers are pruned on the
next emit (closed channel), no monitor needed.

One-viable call, flagged: NodeId is memoized per (name, incarnation) — a
compact local alias for the wire identity, per pg.rs's framing. A reconnect
blip at the same incarnation keeps its id; a restart (new incarnation) gets
a fresh one, so a ghost and its successor are always distinguishable.
Allocation starts at 1; NodeId(0) stays pg::DEFAULT_NODE_ID (self).

Call::Register now carries the whole handshake Peer (the path already has
it; node_up needs incarnation + meta).

tests/cluster_membership.rs 4/0 stable over 5 runs (live up/down over
localhost TCP with commanded and EOF teardown; late-subscriber snapshot +
view agreement; restart-vs-blip id identity; dead-subscriber pruning). All
cluster suites regression-clean; clippy --lib green both configs; fmt
clean; default build compiles.
2026-08-15 07:07:14 +00:00
Claude ad4958421f feat(cluster): RFC 010 c6c — heartbeat send + fixed-timeout liveness + teardown
The timeout arm of the connection actor's select: HEARTBEAT_INTERVAL (1s)
paces outbound Frame::Heartbeat (first at spawn, so the peer's window
starts fed) and LIVENESS_TIMEOUT (4s = 4 intervals) declares the peer dead
when no inbound frame arrives inside it — any frame resets the window, so
heartbeats keep an idle connection alive and real traffic (c8+) counts for
free. Fire => close + exit; the manager's monitor reaps the table entry as
on every other exit path. Fixed timeout per RFC v2 §5 (control connection,
heartbeats can't queue behind bulk). Intervals are the one-viable-answer
call flagged for veto at diff review.

The pump was made non-blocking to keep the deadlines honest: a plain recv()
blocks into the socket while the buffer holds a partial frame, parking the
actor past its timers. Two additive FramedConn methods (read_once,
next_buffered): exactly one socket read per level-triggered readable wake
(cannot block, cannot strand — leftovers re-signal), then drain every
complete buffered frame. Liveness resets only on complete frames.

No-fd transports (loopback) still get the command-only loop: no readiness
means no timers, same caveat as recv_deadline.

tests/cluster_conn_liveness.rs 3/0, stable over 5 runs (raw far end over
localhost TCP: heartbeats appear unprompted; mute peer still up at half
the window, gone after it; heartbeat-only peer survives 1.5x the window,
then reaped once silenced). Cluster suites regression-clean; clippy --lib
green both configs; fmt clean.
2026-08-15 06:36:07 +00:00
Claude c8ed858e4c feat(cluster): RFC 010 c6b — handshake on the accept/connect path
Drive the c5 machines as straight-line code on the path (D8): dial_handshake
and accept_handshake do the IO on a shared FramedConn, and a connection actor
is spawned only after a successful handshake. Rejects, tie-break losses (D7),
protocol faults and timeouts are all resolved on the path by closing, so no
actor ever exists for a connection that did not establish. The whole
FramedConn travels into spawn_established, carrying any read-ahead past the
handshake frames.

Handshake deadlines land here rather than in c6c: FramedConn::recv_deadline
enforces them between reads via the connection's fd arm, so a peer that
connects and goes silent cannot wedge the acceptor.

Connection lifetime moves to the manager (pulled forward from c7). The path
registers each established connection and hands over its ConnHandle; the
manager owns it, monitors the actor, and tears the connection down on
Disconnect, on peer close, or at manager shutdown. spawn_established returns
a Pid, so a connection neither outlives nor dies with whichever actor
established it — the ownership that made two-node teardown unorderable.

The manager also tracks in-flight dial intents, monitored so a panicking
dial cannot wedge the tie-break, and answers HelloCtx for the accept path.
2026-08-14 21:09:08 +00:00
Claude bbaaa062e3 feat(cluster): RFC 010 c6a — connection actor + manager subtree
Per-peer connection actor as a single select-loop plain actor owning the
whole FramedConn: one select folds its command inbox and the transport's
readable arm, so reads and control share one execution context — no reader
thread, no read/write split. The handshake is bypassed here (c6b wires it);
the actor is spawned already-established and self-registers with the manager.

Manager gen_server: the peer-name -> conn-pid registry and the uniqueness
source the handshake's NameTaken depends on. It monitors each connection, so
the table self-heals on any exit path. Explicit supervision subtree keeps the
manager up; connections are dynamic and monitored, never restarted (c7
re-dials).

Transport gains an additive Conn::readable_arm -> Option<FdArm> (default None;
TCP returns its fd's arm, loopback stays None). Existing c3 transport tests
unchanged.

Lifecycle test over localhost TCP: up reflected in the table, commanded
shutdown reaps exactly one, peer EOF reaps the other.
2026-08-14 19:23:39 +00:00
claude 9e49038474 feat(cluster): RFC 010 c5 — handshake as a pure state machine
Frames in, actions out — no IO, no clocks, no actors; the c6 connection
actor will drive it. Initiator (dial: emit Hello, interpret the single
response) and Responder (accept: judge the first frame) as consuming-self
machines; check order proto -> hash -> name -> tie-break. Driver-supplied
HelloCtx carries the two facts the pure machine cannot know (name claimed,
own dial in flight). Tie-break ratified as a wire fact: the smaller name's
dial survives; the losing inbound closes silently (both ends compute the
same verdict, no reject frame needed). Peer's own name offered => NameTaken.
build_hash is config-supplied; derivation lands with c6.
2026-08-14 17:17:14 +00:00
Claude 8a9e2b81b1 test(cluster): RFC 010 c4 — subprocess two-node harness
The runtime is a process singleton, so multi-node tests mean multiple
processes. tests/common/mod.rs is the reusable harness (precedent: RFC
019 c6 / tests/stack_diag.rs self-re-exec, extended to live tailing):
re-execs the current test binary as named roles, tails stdout/stderr on
reader threads, waits on protocol-visible lines with bounded timeouts
(panic dumps carry the full transcript), and Drop SIGKILLs+reaps so a
panicking test leaves no orphan or zombie. Children run with
--test-threads=1 --quiet --nocapture; the last flag is load-bearing —
libtest's capture would otherwise swallow role output.

Port assignment is race-free by construction: children bind port 0 and
announce the concrete address (LISTENING <addr>); the parent never
pre-picks.

Smoke suite per roadmap: two real nodes, handshake-less TCP connect
through the real framed codec (one Heartbeat across, clean close seen on
both sides, both exit 0), plus reap-on-drop proven via ESRCH and
nonzero-exit surfacing. Flake budget stated in the module doc: 10 s
bound per wait, 10/10 clean at authoring, >1/100 failures = regression.
2026-08-14 14:46:28 +00:00
Claude 39ab92871e feat(cluster): RFC 010 c3 — transport trait, framed codec, TCP + loopback impls
The control-connection abstraction (RFC v2 §5): object-safe Transport/
Listener/Conn over opaque pre-resolved addresses (resolution stays the c9
seam), with FramedConn as the single shared byte->Frame codec feeding
Frame::decode's incremental contract. Nothing forecloses additional
per-peer connections for the jarred bulk plane; the membrane is not a
transport (D2).

TCP parks the calling actor via scheduler fd readiness (MSG_NOSIGNAL
writes, EINPROGRESS dial resolved through SO_ERROR). Loopback is the
shipped in-memory test transport: OS-thread-blocking condvar pipes with
TCP-shaped close semantics, per-instance address registry.

Conformance suite runs the same codec over both impls: roundtrips both
directions, framing across split writes, coalesced frames, peer-close
mid-frame as TruncatedByPeer (not EOF), clean close as Ok(None). Plus
impl-specific establishment/error cases and a 4 MiB cross-buffer TCP
frame under real backpressure.
2026-08-14 14:31:39 +00:00
Claude 3850f6099b feat(cluster): RFC 010 c2 — owned wire envelope
Frame enum per the RFC inventory; hand-rolled encode/decode with u32 LE
length prefix + u8 tag; strings u16-prefixed, payload blobs u32-prefixed;
MAX_FRAME_LEN cap (control plane never carries bulk, §5). Streaming decode:
Ok(None) = need more bytes, every Err = corruption. postcard confined to
encode_payload/decode_payload — the single codec seam (§2); postcard gains
the alloc feature for to_allocvec (still no_std-aligned, no default
features). DownReason/RejectReason travel as single tag bytes; Down tag 5
is reserved for c11's Disconnected.

Tests: per-frame roundtrip, back-to-back frames, golden heartbeat bytes,
zero-length payload, every-prefix incomplete, unknown frame/enum tags,
length prefix lying long (with and without bytes present) and short,
truncation mid-string, adversarial lengths (u32::MAX, cap+1, zero), UTF-8
corruption, payload seam roundtrip through a real Send frame.
2026-08-14 13:23:50 +00:00
Claude (sandbox)andClaude (sandbox) ca1c98336e feat(scheduler,runtime): non-panicking try_spawn for at-capacity load shedding
allocate_slot() panics on a full slab; for a load-shedding caller (an
accept loop spawning one actor per connection) that panic lands in the
spawning actor, which then crash-loops under Restart::Transient into the
still-full slab until its restart budget is spent — and the service stops
accepting entirely. Observed live (urus slowloris scaling, 2026-08-10).
A full slab is a routine overload condition for such callers, not an
invariant violation.

- RuntimeInner::try_allocate_slot() -> Option<u32>: the non-panicking
  core; a single pop under the free-list lock, so the claim is atomic
  (claim-or-report — no check-then-spawn TOCTOU, no headroom margin).
  allocate_slot() is now a thin panicking wrapper over it.
- scheduler::try_spawn / try_spawn_under_with -> Result<JoinHandle,
  SpawnError>: parity with spawn/spawn_under_with except a full slab
  returns Err(SpawnError::AtCapacity) instead of panicking. Minimal
  surface per the agreed strategy; the remaining _with/_addr mirrors are
  trivial wrappers if ever needed.
- Slot-first ordering on the try path (reverse of spawn's stack-first):
  under overload Err is the hot path, and a rejection costs one mutex
  pop — no mmap/pool-pop + init + recycle per shed unit of work. A
  drop-guard returns the claimed slot if stack allocation panics in the
  claim-to-install window (would otherwise leak and trip run()'s
  teardown slot-leak debug_assert).
- SpawnError: non_exhaustive, Display + std::error::Error.
- spawn and every existing call site untouched: the panic remains the
  correct loud invariant check at internal/bounded spawn sites.

tests/try_spawn.rs: parity when slots free; exact slab accounting at
capacity (Err, no panic, repeatable); custom-shape try refuses before
stack allocation; self-heal after slots free; plain spawn still panics
(surfaced via JoinError payload); 4-thread race for the last slots
claims exactly the free count; SpawnError impl checks.

Design doc: smarm-suggestion-try-spawn.md. Downstream consumer change
(canned 503 on AtCapacity in urus's accept loop) is urus scope, not
smarm.

(cherry picked from commit 36de4b36aeaa72b2a5f9f3797b9854652656dcf6)
2026-08-13 15:03:16 +02:00
smarm-agent 95306c7f60 style: cargo fmt sweep under rustc 1.97.1 (toolchain reformat, no semantic change) 2026-08-13 05:56:49 +00:00
smarm-agent 1262cc30e3 monitor: widen stamp eligibility to watchable = named ∪ exported (soak sig 5)
The terminal record existed for watches that raced their target's death, but
e43c673 scoped its stamp to named tenancies — and the pid-identity watch
surface (§4 Slice 3) targets arbitrary actors, including anonymous ones whose
pids cross the boundary in contract replies. The first wild pid-face hit
(width-20 soak, pid_watch_test.exs:47, 1/600 full-suite: a monitor installed
while the child was alive delivered :noproc instead of {:smarm_exit, :panic})
is exactly the residual a0ba9be's commit body deferred.

ever_named becomes `watchable`, with a second set-site: mark_watchable(pid),
which the bridge calls wherever a smarm pid is encoded across the boundary —
BEAM can only watch pids it holds, and can only hold pids that crossed.
Anonymous never-exported churn (holder threads, egress tasks) stays
ineligible, preserving e43c673's LIFO-eviction protection unchanged.

mark_watchable takes the cold lock before the liveness screen: finalize
publishes Done and reads the bit under the same lock, so the mark either
lands before the death stamps or observes the tenancy dead and no-ops —
no lost-stamp window, and marking a corpse cannot invent history (pinned
in the test alongside the mark-while-alive stamp).
2026-08-13 05:56:19 +00:00
smarm-agent 461fe4b768 fix(runtime): only named tenancies stamp the terminal record — anonymous churn must not evict it
Discovered wiring the bridge consult: with an unconditional stamp, the record
for the very death being raced was the shortest-lived data in the runtime.
Every green thread is a slot tenant, the free list is LIFO — so the slot a
named server's death frees is the first one recycled, and the next throwaway
exit (monitor holders, chain-runner work, anything) overwrote the record
before a raced watch could consult it. Deterministic bridge repro: the
corpse resolved fine, terminal_reason read None every time.

register_with now flags the tenancy (ever_named, reset at reclaim) before
the binding lands — set outside the registry lock, so no successfully
registered actor can die unflagged and a failed register's overshoot is
harmless — and finalize stamps only flagged tenancies. Watchable identities
are exactly the named ones (the bridge's pid-identity path deliberately
keeps Erlang's raw :noproc), so nothing consultable is lost.

Contract test updated: the three death modes now self-register; a new
anonymous control pins that unregistered deaths neither stamp nor evict.
2026-08-13 05:56:19 +00:00
smarm-agent b937f1f50f monitor/registry: terminal-outcome record — a raced watch can recover the real down reason (soak sig 4)
A watch installed after its target's death has, until now, only NoProc to
report — but the bridge's proxies install their native watch asynchronously
after acquire returns, so a link established before a crash (from the BEAM's
view) could still lose the panic's translated reason to that blanket NoProc
(width-20 soak signature 4: link_test.exs:26, 1/600 full-suite, 3/2000
link-only, all whereis-miss; deterministic repro in the bridge suite).

Two primitives, no change to monitor()'s own Erlang-faithful stale-pid
semantics — the upgrade is the caller's deliberate act:

- finalize_actor stamps the slot with (generation, DownReason) under the same
  cold-lock block that publishes the outcome. The record survives reclaim,
  registry pruning, and the next tenant's install; only the slot's next death
  overwrites it. terminal_reason(pid) reads it generation-matched.
- resolve_name(name) is whereis with the corpse kept: the dead-holder arm
  returns the stored pid it prunes (NameResolution::Corpse) instead of
  discarding the only evidence of who died — whereis itself prunes on the way
  out, so a whereis-then-lookup consumer would find the evidence already
  destroyed. Live/Unbound match whereis's Some/None; the name heals exactly
  as before.

Contract pinned in tests/terminal_outcome_after_death.rs: one record per way
of dying (Exit/Panic/Stopped), no record while live, corpse capture + heal on
resolve_name, record independence from registry pruning, survival across slot
re-tenancy, overwrite at the next tenancy's death.
2026-08-13 05:56:19 +00:00
Claude (sandbox) 410ba33d82 feat(introspect,runtime): per-actor stack surface on ActorInfo (RFC 019 §8)
- introspect::StackInfo { reserve, guard, depth_high_water,
  parks_since_shrink, shrinks } as ActorInfo.stack; re-exported at crate
  root beside ActorInfo.
- All reads lock-free: geometry from the c6 diag slot atomics, depth =
  top - hwm (the §2 sampled high-water; doc spells out sampled-not-exact
  and that 0 means never-descheduled-at-depth), counters straight off the
  §3 atomics. Coherence for the incarnation rides read_slot's existing
  generation check, same as overruns/messages_received.
- Slot::stack_introspect(): one pub(crate) tuple accessor beside the other
  counter accessors.
- Exact RSS deliberately absent per RFC (mincore = debug tooling only,
  never a runtime path); stack_shape(pid) untouched (cold-lock exact
  variant from c2).
- tests/introspect.rs: defaults surface (64 KiB reserve / 1 MiB guard /
  sampled ~32 KiB depth / gate park counted / zero shrinks) + live shrink
  counters (spike visible pre-shrink; shrinks>=1, cooldown counter reset,
  hwm reset after crossing COOLDOWN) read mid-run -- post-join the slot
  reclaim correctly hides the incarnation, which the first draft of the
  test learned the hard way.

FLAGGED (Claude-solo calls):
- Nested StackInfo struct over five flat ActorInfo fields (grain break;
  the five fields are one concern and ActorInfo is already 12 fields).
- Field names reserve/guard/shrinks (RFC says stack_reserve/stack_guard/
  shrink count; the stack_ prefix is redundant inside StackInfo).
2026-08-08 19:12:46 +00:00
Claude (sandbox) 5fd8aecf55 feat(signal,runtime,stack): SIGSEGV overflow diagnostics + 1 MiB guard default (RFC 019 §7)
- src/signal.rs: process-global SA_SIGINFO|SA_ONSTACK handler installed once
  at runtime::init (before any scheduler thread -> unracing PRIOR save);
  per-scheduler-thread 64 KiB sigaltstack registered at schedule_loop entry
  (a guard hit leaves no stack to handle on). Async-signal-safe throughout:
  classification is plain loads (const-init TLS Cell + slot atomics), print
  is fixed-buffer itoa + one write(2), death is SIG_DFL + refault at the
  same instruction (core-dumpable, correct wait status).
- Two-tier classification (agreed): in-guard = definitive; OVERSHOOT window
  below the guard = 'unprobed (FFI?) frame stepped over it' probable
  attribution -- the RFC's motivating incident (cargo-vendored gz, not
  SQLite as the RFC text says) faults there under a small guard. Pure
  classify() fn, 5 adversarial units incl. saturation at low addresses.
- DEFAULT_STACK_GUARD 64 KiB -> 1 MiB (agreed): kernel stack_guard_gap
  anchor post-Stack-Clash; PROT_NONE is VA-only (no RSS, no page tables,
  no overcommit charge) so width is free at any actor count.
- Unclassified faults reinstate the PRIOR sigaction and refault (agreed):
  std's own OS-thread overflow diagnostics survive our presence.
- Slot: diag_{stack_top,stack_reserve,stack_guard,pid} atomics written in
  install_actor pre-publish; readable without the cold lock (Stack lives
  under it); only consulted while CURRENT_SLOT points at the slot, so
  never stale where read. preempt::current_slot_ptr ungated from
  smarm-causal (now also the classifier's anchor).
- build.rs + cc (agreed Q3): canary/canary.c, 96 KiB local touched low-end
  first, -fno-stack-clash-protection pinned so hardened toolchains don't
  probe the canary into uselessness.
- tests/stack_diag.rs: subprocess x4 -- Rust recursion tier-1; FFI canary
  tier-1 at defaults (1 MiB guard catches the jump); tier-2 at guard=4 KiB
  ('stepped over', reproduces the incident); clean at reserve=256 KiB
  (the §1 knob is the fix, same frame).

FLAGGED (Claude-solo calls):
- OVERSHOOT_SLOP = 1 MiB (matches guard default/kernel gap; beyond it
  attribution would be dishonest).
- Altstack 64 KiB, mmap'd once per OS thread, never freed (bounded by
  thread count; reused across run()s via TLS flag).
- Foreign-fault reinstate permanently deregisters our handler; accepted --
  the process is dying either way.
- Diag geometry as 4 slot atomics (install-time cost only) over a per-switch
  TLS snapshot (hot-path stores).
2026-08-08 18:58:30 +00:00
Claude (sandbox) 7d8b9e0310 feat(stack,runtime): pool-recycle DONTNEED above the retained entry end (RFC 019 §6)
- stack::retain_range: pure checked span fn (retain page-up = zap less;
  None when retain covers the reserve, so the 64 KiB default config never
  pays a syscall) + 6 adversarial units mirroring shrink_range's.
- Stack::recycle_zap: advisory MADV_DONTNEED of [usable_base, top-RETAIN);
  stack is unowned at the call site, synchronous eager zap races nothing.
- recycle_stack: zap OFF-LOCK before pool admission (acquire_stack's
  no-syscall-under-the-pool-lock invariant); rare cap-overflow pays a
  wasted zap ahead of munmap, accepted over a second lock round-trip.
- pub const RECYCLE_RETAIN = 64 KiB beside the shrink knobs, ratified-as-
  constant rationale in doc.
- tests/stack_recycle.rs: mincore-based exact-zero-resident assert over
  the zap span. smaps was tried first and over-counts: a neighboring rw
  anon VMA can merge flush against the stack top (observed once under the
  full-suite run); the PROT_NONE guard pins the usable base exactly.

FLAGGED (Claude-solo calls):
- RFC §6 'above the bottom RETAIN' is direction-ambiguous in address
  terms; implemented as retain the ENTRY end (highest addresses, the
  pages the next actor faults first), zap the cold deep span below.
- Const named RECYCLE_RETAIN (RFC says RETAIN) to sit beside SHRINK_*.
2026-08-08 16:13:53 +00:00
Claude (sandbox) 8225716b11 feat(runtime,stack): sampled stack high-water + MADV_FREE shrink at actor-park (RFC 019 §§2–3)
hwm: AtomicUsize lands beside sp on the slot: the single context-save
site min-updates it (one branch + at most one Relaxed store into the
line the sp store just dirtied), install resets it to the fresh top.
Advisory by construction — correctness never depends on it. The mod-doc
ordering chain gains a line: hwm piggybacks the existing
Relaxed-store-before-Release pattern and adds no edges.

Shrink hook in the YieldIntent::Park arm only, before the park_return
Release transition — the owned window (obligation 1's assert-comment at
the site): after the sp store, before Parked is published, scheduler on
its own stack, actor saved and unstealable. It runs on both arms of the
park_return race (a consumed unpark flag means one wasted-but-harmless
madvise). The preempt/yield path deliberately never checks: §4's
bounded, self-healing leak under saturation, when syscalls are least
affordable.

SHRINK_THRESHOLD = 256 KiB and SHRINK_COOLDOWN = 64 parks are pub
constants with the ratified doc rationale, not Config fields. The freed
span is shrink_range(hwm, sp, page): whole pages of [hwm, sp − 1-page
redzone), rounded inward, checked arithmetic — adversarial inputs
collapse to None (obligation 2). MADV_FREE marks lazily; the kernel's
reclaim-under-pressure IS the hysteresis, cancel-on-write is the safety
net. parks_since_shrink + shrink_count ride the slot for the cooldown
and the future introspect surface.

Tests: 7 adversarial shrink_range units (inverted/empty spans, redzone
underflow, unaligned ends, sp-crossing sweep); integration — 8 MiB
reserve, ~3 MiB spike sampled via yield-at-depth, parks gated on
introspected Parked state past the cooldown, then ≥ 2 MiB LazyFree
asserted inside the stack's smaps range with live data intact; and the
inverse guard — a shallow never-spiking actor ends at exactly 0
LazyFree (also proves the parser isn't vacuously zero via the first
test).
2026-08-08 14:30:32 +00:00
Claude (sandbox) 3cb64eefc2 feat(scheduler,gen_server,gen_statem,introspect): SpawnOpts — per-actor stack shape on every spawn surface (RFC 019 §1)
SpawnOpts { stack_reserve, guard_size } with Option<usize> fields, None
resolving to the Config defaults at spawn time — a deliberate deviation
from the RFC's plain-usize struct so struct-update syntax works without
a runtime handle in scope. Threaded across the five surfaces:
spawn_with, spawn_under_with, spawn_addr_with,
GenServerBuilder::stack_opts (mirrored on NamedGenServerBuilder), and
gen_statem::spawn_with (gen_statem has no builder, so the opts ride a
_with variant — Claude-solo surface call, flagged for review). Existing
spawns forward defaults; no call-site churn.

introspect::stack_shape(pid) pulled forward (agreed) as the first slice
of the RFC 019 introspection surface, giving tests an observable.

Tests (tests/spawn_opts.rs): override/partial-override/rounding on each
surface; obligation 4 from the outside — a dead custom stack is never
handed to the next default spawn (LIFO pool would expose it), and the
reverse (default stacks ARE recycled); 8 MiB reserve behaviorally
permits ~1 MiB recursion. Also: silence unused-Result in the c1
runtime test (join now unwrapped).
2026-08-08 14:22:38 +00:00
Claude (sandbox) 0fe052bc7e feat(stack,runtime): per-shape actor stacks — Stack::new(reserve, guard), Config knobs, pool rule (RFC 019 §1)
Stack takes an explicit (reserve, guard) shape, both page-rounded and
stored; usable_base derives from the stored guard. Guard default raised
4 KiB -> 64 KiB (DEFAULT_STACK_GUARD): probestack makes one page enough
for Rust frames, but an unprobed C frame can leap a page in one sub rsp
— the motivating SQLite segfault. Reserve default stays 64 KiB
(DEFAULT_STACK_RESERVE); ACTOR_STACK_SIZE retired.

Config::{stack_reserve, stack_guard} thread the runtime defaults into
RuntimeInner pre-rounded. All acquisition/recycling now goes through
acquire_stack/recycle_stack carrying the pool rule: only default-shaped
stacks are pooled (pooled ⇒ default-shaped by induction); custom shapes
mmap fresh and munmap at death. Pool lock still dropped before any mmap.

No public spawn API change (SpawnOpts is the next commit).

Tests: shape rounding + accessors, wide-guard faults at both ends
(subprocess), Config::stack_reserve permits >64 KiB recursion that
previously could only segfault.
2026-08-08 14:18:27 +00:00
smarm d4839f1d81 feat(runtime,io): driver-enqueues + park/wake idle path — retire the wake pipe
The swap (RFC 018). Schedulers no longer sleep on a shared level-triggered
wake pipe — the herd source that made the default 8-thread config 7x
slower than 2 threads (E1). They park on per-thread futex parkers via the
coordination layer; IO backends become producers behind a two-call
contract (make runnable, then the enqueue tail wakes exactly one parked
scheduler).

Deleted: the drain lock and the one-winner phase-1 drain; the shared
completions VecDeque; the wake pipe fds, poll_wake, drain_wake_pipe,
wake_scheduler, the FdReady/Blocking Completion enum; the 100us idle nap;
the per-pop io.lock liveness read; io.rs's as_millis timeout truncation.

Added:
- enqueue wake tail (fixes the silent enqueue): wake_one_if_idle, a fence
  + one Relaxed mask load when everyone is busy — the pure-compute hot
  path pays almost nothing.
- driver-enqueues: the pool thread stashes its result in the slot,
  decrements io_outstanding, unparks; the epoll thread removes+DELs the
  waiter under the waiters lock and unparks. Both reach the runtime via a
  Weak (no Arc cycle). The waiters map moves behind its own Arc<Mutex> so
  the epoll thread never takes the runtime io lock (teardown holds it
  while joining that thread).
- io_outstanding / io_fd_waiters atomics: the termination verdict reads
  two atomics instead of taking io.lock on every pop.
- timekeeper idle path: at most one parked scheduler holds the timer
  deadline (an expiry wakes one, not a herd); everyone else parks
  indefinitely and is woken by the enqueue tail.
- busy-path timer due-check (ratified design point (a)): under saturation
  nobody parks and no timekeeper exists, yet due timers must still fire —
  one Relaxed load of the earliest-deadline snapshot per loop, clock read
  only when a timer is armed. Maintained under the timers mutex.
- chain rule: a scheduler that pops with more work queued and a sibling
  parked wakes one, so surplus runs in parallel rather than behind it.

tests/park_wake.rs pins the two new observable properties: timers fire
under full scheduler saturation, and sub-ms sleeps are prompt (the
as_millis truncation regression). Full suite + all loom models green;
clippy --lib clean.
2026-07-24 09:12:12 +02:00
Claude (sandbox) d9addeba5e test(causal): controller test no longer races the sweep's site snapshot
run_experiments snapshots the site registry once at entry. stage-x is
registered lazily (first causal_site! execution in the worker), so on a
1-core box the snapshot deterministically wins whenever a sibling test
has already paid the tsc_hz calibration — stage-x missed the sweep and
the summary assert tripped (silently, pre-propagation; the earlier
cascade attribution was incomplete). The worker now signals after its
first site entry and the test waits on it before starting the sweep.
Jarred separately: the lazy-registration trap is a lib UX hazard worth a
doc note or warm-up guidance.
2026-07-18 21:59:21 +00:00
Claude (sandbox) d5a3ba1934 fix(causal): born current — slot reuse booked phantom park forgiveness
A fresh/reused slot started with causal_delay=0 and causal_parked=true:
the first resume then 'forgave' the entire monotone global backlog, once
per spawn, booked as park forgiveness. Under close-mode conn churn (~95k
spawns/s) that is millions of phantom forgiven ms per 700ms window, even
in 0% cells — the books could not balance under actor churn while
ka-mode stayed plausible (few spawns). Impacts were unaffected: the
first resume always precedes the first check, so nobody ever spun.

reset_counters now installs Coz's new-thread rule: causal_delay starts
at the current global ledger and causal_parked starts false — a newborn
neither owes nor is forgiven the process's history, and delay injected
while it sits spawn-queued (runnable, not blocked) is owed and paid at
its first check, the semantics audit_zero_pct_window_absorbs_leftover_
debt pins. That test (born failing, unmasked by the run() propagation
fix) and the new actor_churn_between_experiments_forgives_nothing
regression test both go green.
2026-07-18 21:55:56 +00:00
Claude (sandbox) 7eae56a296 fix(runtime): a root actor panic escapes run()
The trampoline caught the root's panic, recorded it as Outcome::Panic on
the slot, and run() dropped the initial handle without reading it — every
assert inside run(), the standard test-suite pattern, was silently
vacuous (found live: a failing-first test passed). run() now reads the
root outcome before the handle drop and resume_unwinds the payload after
full teardown, so a caller's catch_unwind leaves the Runtime reusable;
Exit and Stopped return normally. The payload message is printed before
re-raising (the throw-site hook output was suppressed in-actor).

Correct the two tests this unmasked, both born failing and never run:
the select loser-arm test kept a closed arm in the set (the documented
closed-arm rule: a closed arm reports ready forever — observe the
disconnect and drop it); the send_after-to-dead test expected Ok(None)
from a closed+empty channel (documented: Err(RecvError), which proves
nothing-delivered even more strongly).
2026-07-18 21:52:50 +00:00
Claude (sandbox) 0ee3fe7330 feat(causal): offcpu audit bucket — the @50 deficit located (RFC 007)
GPU sweep decomposed the ~23ms/700ms @50 injection deficit: every
ledger bucket is ~zero (drop park 0, discards 0, drop yield ~0.3ms),
books balance at absorbed+forgiven = 4x injected in all 48 cells, and
the 0%-cell contamination signature is absent. The probe pins the
residual: eff 0.933-0.943, a constant 22-27µs missing per site entry
= ~4.9 slice-expiry yields/entry x ~5.6µs runqueue wait. The
"deficit" is runnable off-CPU time inside the site — wall time the
probe's ground truth counts but on-CPU attribution correctly skips
(Coz model: speeding the site's code does not shrink queue-wait).

Measure-only reclassification, no behaviour change: a yield in the
target site stashes (tsc, experiment epoch) on the slot; the next
on_resume counts the gap into OFFCPU_IN_SITE_{CYCLES,N} (would-be
delta terms, MAX_SAMPLE_CYCLES-capped) iff the epoch still matches
and the word is live — a gap straddling end()/a same-word begin()
(live in the probe's 50,50 schedule) is dropped, never a leaked
cooldown. Parks excluded: blocked time is forgiveness territory.
New offcpu column in render_ledger_audit; LedgerCounters and
ExperimentResult grow the two fields. Fidelity footer now states
the on-CPU basis (deliberate wording change to the pinned summary;
the substring pin test still holds). +2 tests (counted gap; epoch
straddle) + render assert.
2026-07-13 12:46:15 +00:00
Claude (sandbox) a3be8f0977 feat(causal): ledger audit — decompose the @50 injection deficit (RFC 007)
Measure-only counters for the deficit hunt (~23ms short per 700ms window
at 50% on the bottleneck site; superlinear vs 25%). Nothing here changes
injection or absorption; the sweep decides the fix.

Buckets, windowed per cell into new ExperimentResult fields (audit
snapshot taken at end() — spin/attribution freeze there, forgiveness
does not):
- spin_absorbed / park_forgiven: where owed delay actually went. Spin
  during a 0% cell is the baseline-contamination signature — leftover
  debt from a prior window being paid in a later one (checks gate on
  the experiment word, so cooldowns pay nothing and debt carries over).
- drop_park / drop_yield (+counts): the deschedule path flushes no
  sample tail — an in-target-site park or yield silently loses
  [last sample -> now]; on_resume re-arms before the actor runs again.
  New on_deschedule hook in all three intent arms (real park; explicit/
  slice-expiry yield; requeued park counts as yield — it never blocked).
  Slice-expiry yields sample at the descheduling checkpoint, so a fat
  yield bucket points at explicit yield_now or requeued parks.
- discard_overmax (+count, in would-be delta terms so columns compare
  against injected_cycles) / discard_unarmed: the attribute() clamps,
  previously silent.

LedgerCounters + ledger_counters() expose cumulative totals (tests,
run-level prints); render_ledger_audit() is the per-cell companion to
render_summary, which stays byte-identical (pinned). ExperimentResult
now derives Default so literals survive future audit-field growth.

Tests: +7 (spin counted, forgiveness counted, in-site park drop, in-site
yield drop, overmax discard, 0%-window leftover absorption — synthesized
deterministically via inject_delay_cycles_for_test with no experiment
active — and the audit render). 22/22 causal.
2026-07-13 12:07:19 +00:00
Claude (sandbox) a2d0b7af18 feat(causal): fidelity footer in render_summary
One unconditional footer line whenever there are results:
"note: impacts are lower bounds — undershoot grows with speedup pct;
rankings unaffected" — surfacing the RFC 007 Validation fidelity statement
where users actually look, instead of only in the RFC. Wording pinned by
the summary test.
2026-07-13 11:11:20 +00:00
Claude (sandbox) a4647f368a feat(causal): wall-anchored send_after — user-facing timer opt-out (RFC 007)
send_after_wall / send_after_named_wall (+ Timers::insert_send_wall) arm a
message-delivery timer that opts out of the RFC 007 virtual-time shift and
fires at its raw deadline regardless of injected delay — the Send-reason
sibling of sleep_wall, closing the jar item whose substrate efbc254 landed.
For deadlines that reflect the outside world (protocol timeouts, wall-clock
schedules) rather than workload pacing. cancel_timer is anchor-agnostic and
unchanged; without the feature the API exists and is identical to send_after.

The gen_server timer layer (send_after_to, RFC 015 §5) deliberately stays
virtual-only — an opt-out there means new options on the gen_server/statem
timeout API, out of scope for now.

Tests: wall send fires at raw deadline while a virtual sibling shifts;
cancel on a wall send with debt outstanding; featureless delivery/cancel
smokes through the public named API. 15/15 causal, 34/34 binaries both
feature configs, lib clippy clean both.
2026-07-13 11:10:11 +00:00
Claude (sandbox) efbc254634 feat(causal): wall-anchored timers — controller windows keep fixed wall length (RFC 007)
New timer anchor: insert_sleep_wall / scheduler::sleep_wall (exported) opt a
Sleep entry out of the RFC 007 virtual-time shift, so it fires at its raw
deadline regardless of injected delay. Featureless config is unchanged (the
API exists but is identical to sleep).

The causal controller's window/cooldown sleeps and the tsc_hz calibration
sleep use it on the actor path (the OS-thread path was already wall). This
fixes the controller's own sleeps dilating under its own injection —
experiment windows stretched ~2x at 50% speedup (337ms -> 646ms injected/
window). Deltas were rate-normalized so results were unbiased; this fixes
sweep cost, not bias. The general wall-anchored-timer-semantics jar item
(user-facing opt-out) remains open; this lands the substrate.

Test: wall_timer_ignores_injected_delay — wall entry fires at raw deadline
while a virtual sibling in the same heap shifts. 13/13 causal, 34/34
binaries both feature configs.
2026-07-13 07:44:51 +00:00
Claude (sandbox) 04dbac1f4b feat(causal): timer-heap virtual time — deadlines chase injected delay (RFC 007)
Injected delays dilate virtual time for the workload, but timer deadlines
stayed wall-anchored: a sleep or receive-timeout fired early in virtual
terms, so timeout/retry behaviour sped up relative to the dilated world
(v1 known gap #1).

Every heap Entry now carries delay_stamp — the global delay ledger at
(re-)queue time, cfg-gated on smarm-causal. pop_due converts any debt
accrued since the stamp to wall time (tsc_hz) and shifts the effective
deadline; a not-yet-due entry is re-queued at the shifted deadline with a
fresh stamp, so it keeps chasing delay injected while it waits. seq is
preserved across re-queues, keeping send_after cancellation identity
intact (cancelled entries are discarded before any shift). Zero debt is
byte-identical to the old path; peek_deadline may under-report, costing
one spurious scheduler wake per injected chunk (documented).

This also makes the park-gated resume credit *correct* rather than
forgiving for sleepers: a sleeping actor now physically pays its debt by
sleeping longer, so the on_resume fast-forward reflects real payment
(sleeping_actor_pays_injected_delay pins this end-to-end through the
runtime).

New test hooks: inject_delay_cycles_for_test (deterministic ledger
driver, eagerly TSC-calibrating so conversion never stalls a scheduler
loop) and cycles_to_duration. The ledger is process-global, so the
delta-sensitive causal tests now serialize on a shared test mutex — they
were racy under the parallel test harness before this, in principle.
2026-07-13 07:31:08 +00:00
Claude (sandbox) d496914d40 fix(causal): flush target-site samples at guard boundaries (RFC 007)
Samples were taken only when maybe_preempt's cold block happened to fire
in-site, so the interval between the last check and SiteGuard drop was
discarded on every site entry. Measured live on the 24-core box:
22-29us lost per entry, a constant attribution efficiency of ~0.93-0.94,
which under-reported every impact (+83.5% where theory says +100%; the
observed shortfall fits 1/(1-pct*eff)-1 at both 25% and 50%).

SiteGuard enter/drop now call site_transition(): leaving the target site
flushes the pending interval into the ledger (sample-only, never spins,
so safe under no-preempt regions); entering the target site re-arms the
sample clock so pre-site time is never attributed (the symmetric
over-attribution). Winner attribution is factored into attribute(),
shared by the cold check and the flush, with the same interval clamps.

Adds examples/causal_attrib_probe.rs (ground-truth in-site time vs
ledger attribution, the probe that confirmed the leak) and the
site_boundaries_flush_tail regression test (a site entry that never
hits a cold check must still be attributed). Also gates causal_probe
on smarm-causal in Cargo.toml - it never was, so featureless builds
of the examples were broken.
2026-07-12 19:35:06 +00:00
Claude (sandbox) 2668f4018f feat(causal): native causal profiling behind smarm-causal (RFC 007 v1)
causal_site! scoped site guards per actor slot, progress! throughput
points, and a Coz-style virtual-speedup engine hooked into
maybe_preempt's amortized cold block: target-site samples grow a global
delay ledger; bystanders spin-absorb their debt at the next causal
check, with timeslice extension so injected delay is not charged
against the slice.

Resume credit (Coz's blocked-thread rule) is gated on a causal_parked
slot bit set only by a real park: crediting on every resume made any
yield-cadence actor delay-immune and every experiment inert (found
live on a 24-core run — dead-flat deltas across all sites).

Report normalization uses a measured TSC frequency (~50ms calibration
on first use) instead of the crate-wide 3 GHz assumption, which
uniformly inflated impact numbers on a 3.7 GHz box. impact_pct() is
the machine-readable form of the summary for programmatic checks.

examples/causal_pipeline.rs burns fixed *work* (calibrated LCG loop),
not fixed wall time — a timed busy-wait absorbs injected delay into
its own budget and reads as a no-op. Self-checking: exits nonzero if
causal separation fails; skips the verdict below 4 cores. Validated
on a 24-core box: reserve (true bottleneck) +29.3%@25/+83.5%@50;
serialize and background-compaction ~0%.

Known v1 gaps (jar): timer-heap deadlines unshifted, no-check!/no-alloc
actors undelayable, multi-scheduler coherence best-effort Relaxed,
off-CPU blame punted, Instant::now() uncorrected.

Zero-cost with the feature off; clippy -D warnings clean both ways;
full suite green with and without smarm-causal.
2026-07-12 19:10:04 +00:00
smarm-agent 1c90a4ef5e fix(registry): by_name stores the full holder Pid — a dead name heals under slot reuse
Root cause of soak20 signature 2 (refcount_test.exs 'watcher crash',
110235x fast {:error, :server_down} probes over the full await window):
by_name mapped name -> slot *index*, so a name whose holder died (no stop
path unregisters; prune is lazy) and whose slot was then re-tenanted read
as live-held: register failed NameTaken{holder: <unrelated tenant>} (which
the bridge macro's generated start() swallows -> start_server/1 reports :ok
for a server that never came up), while name resolution reached the
tenant's mailbox, missed on the message TypeId and failed fast WITHOUT
pruning — the wedge self-sustained for the tenant's lifetime. Name-
addressed send additionally judged liveness on the slot's *current*
mailbox pid, so a same-typed tenant would have received the message
(misdelivery) and a differently typed one a misleading NoChannel.

Fix: by_name: HashMap<&'static str, Pid> — every reader judges the
*stored* holder with the generation-checked live(), so a recycled slot's
tenant no longer impersonates a dead holder, and every touch (register /
whereis / resolve / send) prunes and heals a stale name. prune(index)
becomes prune_holder(pid): names bound to the holder go; the mailbox goes
only while still the holder's own (a tenant's replacement mailbox is left
untouched). Introspection matches names to mailboxes by full pid, so a
stale name never annotates a slot's new tenant.

Deterministic regression test added first and shown to fail pre-fix
(tests/stale_name_slot_reuse.rs: tiny slab forces re-tenanting; old slot
(1,0) died, tenant (1,1) took the index; register -> NameTaken pre-fix).
Post-fix it asserts the healed contract: whereis -> None (pruned), call ->
ServerDown, re-register -> Ok. Suite 33 ok-binaries, clippy gate clean.

In the wild the window opened at every splice_test teardown:
Splice.terminate -> exit_server('subtree') left the name bound; width 20
raised the re-tenant probability. Downstream (smarm_beam): install_child's
unregister-before-register workaround becomes dead code (removed there);
the #[smarm_server] macro's swallowed register error becomes truthful
idempotency (a NameTaken now really is a live holder).
2026-07-12 07:23:30 +00:00
smarm-agent 6c2b7e91cf channel: drop queued messages when the Receiver drops
A queued Envelope::Call was stranded until the last Sender dropped, so a
caller parked in gen_server::call was never released with ServerDown when a
*named* server was request_stop'd — the registry's inbox Sender clone (lazy
prune) kept the channel Arc, and the queued reply_tx, alive indefinitely.

Receiver::Drop now drains the queue (items dropped after releasing the lock,
since a reply_tx drop reaches a different channel's lock + the scheduler),
restoring the documented ServerDown guarantee on every teardown path.

Adds tests/stop_with_queued_call.rs: deterministic pure-smarm reproducer.
2026-06-24 20:53:27 +00:00
smarm-agent 531571bfa5 gen_statem: postpone events for replay after a transition
A `=> postpone` row (cast/call/info) defers the current event untouched,
to be replayed after the next real transition. `handle` is now two-phase:
a borrow-only postpone pre-pass that hands the event back as
`Step::Postponed(ev)`, then the existing consuming `match (state, event)`.
The loop owns a FIFO queue, drained in the new state ahead of further
intake; a replayed event may postpone again. A postponed `call` keeps its
Reply, so a later state answers it.

`handle` returns `Step` (Postponed / Transitioned / Stayed) so the loop
can see both deferral and transition without reading the state cell.
`Resolution::Postpone` is removed: postpone is a pre-dispatch routing
decision, not a consuming-dispatch outcome.
2026-06-20 13:02:15 +00:00
smarm-agent acf67fef06 gen_statem: state and named timeouts, info events
Add the two timeout flavours, both surfacing as ordinary events matched
in on-state arms:

- cx.state_timeout(d): fires a state_timeout event after d in the current
  state, auto-reset by the loop on every real transition.
- cx.timeout(name, d): fires a timeout(name) event after d, surviving state
  changes, keyed by name, with cx.cancel_timeout(name).

Both ride the existing timer min-heap via send_after_to onto a new per-loop
system channel, selected above the inbox so a fire can't be starved by inbox
traffic. A local-id stamp on each fire lets a reset/cancel that loses the race
discard a stale fire. The macro grows an info: clause and folds three internal
Ev variants (Info, StateTimeout, Timeout) alongside cast/call, with new row
keywords info / state_timeout / timeout. Unmatched info silently drops (the
gen_server default); state/named timeouts have no default, so a state that can
see one must handle it or the match is non-exhaustive.

Rename the hand-written expansion-target example fused -> expanded, retire the
deprecated Switch demo machine (its round-trip / enter / panic-down coverage
moves onto the timer machine), and refresh the macro docs to the door machine.

cargo build --all-targets warning-free; cargo test green.
2026-06-20 12:22:45 +00:00
smarm-agent 0cf6b80396 gen_statem: GenStatem* type prefix + cleanup
- rename StatemRef/StatemCallError/StatemSendError -> GenStatem*
- move the inline unit test out of src; consolidate the Switch coverage
  onto a single macro-driven harness in tests/gen_statem.rs
- drop the redundant hand-written Switch test machine and the two
  untracked rejected-direction probes (succ_enums, typed_edges)
- rename examples statem_{fused,macro}.rs -> gen_statem_{fused,macro}.rs
- strip RFC/chunk/spike provenance and fix the mislabeled "throwaway"
  example header and dead cross-references
2026-06-20 10:48:33 +00:00
smarm-agent 3e316066c3 gen_server: prefix public types with Gen (GenServerRef, GenServerCtx, GenServerName, GenServerBuilder, NamedGenServerBuilder) 2026-06-20 10:48:33 +00:00
smarm f646c5cd72 cleanup some LLM crud 2026-06-20 12:22:33 +02:00
smarm-agent acc37c5fc9 RFC 017 chunk 1: gen_statem primitives (no macro yet)
Runtime support layer for gen_statem, built against existing public API
(channel + scheduler::spawn), sibling to gen_server. No macro: per review,
build the primitives first and hand-write the Switch example to evaluate
whether a statem! macro earns its place before committing to one.

- src/statem.rs: Machine trait (on_start/handle), Resolution<S> with
  From<S>, Cx (on_unhandled), Reply<T> move-only reply handle, StatemRef
  (send/call), spawn + inbox loop. Real time only; Postpone + Cx timeout
  arming are in the type surface but not yet acted on (chunks 2-3).
- examples/statem_switch.rs: the RFC Switch machine hand-written against the
  primitives, tagged USER vs MACRO to mark what a macro would generate.
  Asserts the RFC end-state (flips=1, enters=3).
- tests/statem.rs: call/cast round-trip, enter-on-start/transition-not-stay,
  panicking-handler -> Down.

Reply<T> included (the call helper needs a handle type; keeps the example
true to the RFC surface) but isolated and trivially removable if we drop it.
2026-06-19 19:52:02 +00:00
smarm-agent 6df4cd4a0b RFC 016 Chunk 4: observer gen_server (feature-gated)
A thin GenServer consumer of the Chunk-1 read primitive — the live
observer process (D12). ObserverRequest/ObserverReply are the wire
contract (D11); the version rides along on the snapshot/tree payloads,
which already carry SNAPSHOT_FORMAT_VERSION (D1). Behind the new
`observer` Cargo feature, off by default (D10): the primitive stays
always-on, only the transport is gated. Cast is Infallible, so the
server takes no async traffic and handle_cast is statically unreachable.

Gated integration test proves each verb relays exactly what the
corresponding primitive returns (snapshot/tree/actor_info), incl. a
forged-pid None and a live Parked classification.
2026-06-19 10:18:54 +00:00
smarm-agent 48d47c45c9 RFC 016 Chunk 2c: approximate per-actor time-budget (reductions-like)
Accumulate on-CPU cycles per actor as ActorInfo.budget_cycles, behind
the off-by-default budget-accounting feature (D6) — a reductions-style
work metric for relative comparison across runs.

- Approximate by design (per Mark): charge now - slice-start once at the
  yield point, reusing the timestamp reset_timeslice already sets, so
  one RDTSC per resume not two. Wake-slot resumes inherit the slice and
  so slightly over-attribute the chain's time to the woken actor — noise
  that averages out; we trade exactness for half the hot-path cost.
- Field/ActorInfo member are unconditional (keeps the snapshot shape
  stable across the feature flag, D1); only the accumulation is gated,
  so default builds are byte-identical and pay nothing. Reads return 0
  when off. Single-writer Relaxed like the other counters; reset in
  reset_counters (D7).

Matrix: default + feature-on + rq-mpmc + rq-striped + release + trace all
green; loom unaffected (feature off under --cfg loom).
2026-06-19 08:47:58 +00:00
smarm-agent e93b3120ec RFC 016 Chunk 2b: per-actor messages-received counter
Tally each dequeued message against the receiving actor, surfaced as
ActorInfo.messages_received (the 'is this actor draining slower than its
mailbox fills' signal).

- messages-received, not sent (D4): the receiver counts on its own
  thread, so it's a single-writer Relaxed load+store on a hot Slot
  AtomicU64, no atomic RMW (D5); reuses the stashed *const Slot from 2a.
- Incremented at all six channel dequeue-success sites (recv,
  recv_timeout x2, recv_match, try_recv_match, try_recv) via
  preempt::note_message_received; no-op outside an actor (null slot).
- Resets with overruns in reset_counters across the three lifecycle
  sites (D7).

Matrix: debug default + rq-mpmc + rq-striped + release green; loom
slot_state/run_queue models pass (the 3 send_after_to loom failures are
pre-existing at fc014c4 — plain #[test]s under --cfg loom, not Chunk 2).
2026-06-19 07:34:32 +00:00
smarm-agent 354eef9f88 RFC 016 Chunk 2a: per-actor timeslice overrun counter
Tally overruns at the slice-expiry site in preempt.rs (the RFC 006
signal), surfaced as ActorInfo.overruns.

- Counter is a hot-region AtomicU64 on Slot, single-writer Relaxed
  load+store (no atomic RMW), read Relaxed by the snapshot (D5).
- Reached from the rare expiry branch via a stashed *const Slot in a
  preempt thread-local, set/cleared on the same resume/return boundary
  as CURRENT_STOP — one TLS load, no runtime lookup. Shared infra for
  the messages-received counter next.
- Reset in all three slot-lifecycle sites: Slot::vacant, reclaim_slot,
  install_actor (standing invariant, D7) — per-incarnation counts.
2026-06-19 07:18:03 +00:00
smarm-agent 7ef915c81e RFC 016 Chunk 3: parentage tree view
tree() / tree_from() fold a Chunk-1 snapshot into a forest by grouping
each actor under its parent pid — one O(n) pass, no new reads. tree_from
is public so a held (or synthetic) snapshot can be folded without a
second scan.

- D8: actors parented at ROOT_PID are genuine roots; an actor whose
  recorded parent is absent from the snapshot is re-rooted under the
  sentinel and flagged orphaned, so the forest stays total. take()-on-
  place doubles as a guard against re-entering a node.
- D9: the edge is parent/spawned-by, documented as not necessarily
  supervision.
2026-06-19 06:33:21 +00:00
smarm-agent c66691943d RFC 016 Chunk 1: runtime introspection read primitive
snapshot() / actor_info() return owned ActorInfo over the slab: pid,
names, fine-grained scheduling state, parent edge, trap flag, mailbox
depth, and monitor/link/joiner counts. Pure reads, no hot-path change.

- ActorState maps the packed slot word (no new storage); introspect.rs.
- D1: RuntimeSnapshot carries SNAPSHOT_FORMAT_VERSION from day one.
- D2: per-slot tearing (ps semantics); actor_info coherent per actor.
- D3: mailbox depth included. Registry channels now stored behind an
  ErasedSender trait (downcast for clone_sender + type-erased
  queued_len); depth summed over published channels under the registry
  Leaf (Leaf -> Channel), kept out of the cold-lock pass so no two
  Leaves are ever held at once. Depth covers published channels only.
- Done slots surface as root-less tombstones.
2026-06-19 06:31:10 +00:00
smarm-agent 1df85e2384 gen_server: drop-guard drains live timers + debug_assert no leak (RFC 015 §4.7) 2026-06-18 19:19:15 +00:00
smarm-agent 8c8af55928 gen_server: idle/receive timeout via select_timeout/recv_timeout + handle_idle (RFC 015 §4.4) 2026-06-18 19:17:20 +00:00
smarm-agent c0cfa01f37 gen_server: tick_every periodic sugar — loop-driven re-arm via factory + Sys::Tick (RFC 015 §4.3) 2026-06-18 19:13:52 +00:00