perf(runtime): skip take_closure's locked swap after first resume

Every resume paid an unconditional AtomicPtr::swap (lock xchg, full
barrier) to check for a first-resume closure that is null on all
resumes after the first. A Relaxed null-load fast path is sound:
store_closure runs only before publish_queued, whose Release pairing
with try_claim's Acquire orders it before this call, so no writer can
race the load within an occupancy.

Measured on switch_cost (1-core sandbox, rq-mutex, cycles): mean
roundtrip 350-355 -> 324-328, ~7.5%. All lib + scheduler/channel/
supervisor tests pass.
This commit is contained in:
claude-asm-audit
2026-08-21 12:22:45 +00:00
parent d7082eb266
commit 6172f4231d
+9
View File
@@ -873,6 +873,15 @@ impl Slot {
}
fn take_closure(&self) -> Option<Closure> {
// Fast path: every resume after the first (the overwhelming case)
// finds null. A plain load suffices to prove it — `store_closure`
// runs only before `publish_queued`, whose Release/Acquire pairing
// with the claimer's `try_claim` orders it before this call, so no
// writer can race the load within an occupancy. This keeps the
// locked RMW (full barrier, ~20+ cycles) off the per-resume path.
if self.closure.load(Ordering::Relaxed).is_null() {
return None;
}
let raw = self.closure.swap(std::ptr::null_mut(), Ordering::Acquire);
if raw.is_null() {
None