fix(tls): actor-side thread-local accessors are #[inline(never)] + fence — LLVM caches TLS addresses across context switches (finding 19)

Without LTO (any downstream `cargo build --release`), LLVM keeps `%fs:0` in a
callee-saved register across `switch_to_scheduler`; after the actor migrates,
`ACTOR_DONE`/`PREEMPTION_ENABLED`/`CURRENT_PID` etc. hit the OLD thread's TLS.
8 multi-thread tests aborted with "scheduler resumed a done actor" (error 5).
Thin LTO only worked because it picked the local-exec model.

- context.rs: module doc "Thread-locals and migration" (the rule), tls_fence()
- every TLS accessor reachable from an actor stack is out of line:
  actor::{current_pid, publish_outcome (new)}, preempt::{maybe_preempt,
  check_cancelled, current_slot_ptr, note_*, preemption_swap/enabled (new;
  replaces direct PREEMPTION_ENABLED.with in NoPreempt/RawMutex/with_runtime/
  trace/debug asserts)}, context::{get,set}_scheduler_sp, runtime::{sched_slot,
  slot_push, set_yield_intent}, causal, trace
- raw_mutex order checks: out of line only under debug_assertions
- Cargo.toml: [profile.reltest] = release without LTO; 41 test bins link in
  ~1.5 s each instead of ~12 s (suite 11 min -> 1.5 min) AND it is the
  regression oracle for this bug: `cargo test --profile reltest`
Both profiles: 41/41 + doctests green.
This commit is contained in:
claude-asm-audit
2026-08-21 12:23:46 +00:00
parent d300a9d536
commit a9341e2d82
10 changed files with 143 additions and 34 deletions
+8
View File
@@ -60,6 +60,14 @@ panic = "unwind"
lto = "thin"
codegen-units = 1
# `cargo test --profile reltest`: release codegen for the crate (same opt-level,
# same panic strategy) but no LTO at the final link. Thin LTO is what makes each
# of the ~40 test binaries cost ~12 s to link instead of ~2 s; the tests don't
# need cross-crate LTO, the benches do (they keep using `release`).
[profile.reltest]
inherits = "release"
lto = false
[[bench]]
name = "primes"
harness = false