fix(tls): actor-side thread-local accessors are #[inline(never)] + fence — LLVM caches TLS addresses across context switches (finding 19)
Without LTO (any downstream `cargo build --release`), LLVM keeps `%fs:0` in a
callee-saved register across `switch_to_scheduler`; after the actor migrates,
`ACTOR_DONE`/`PREEMPTION_ENABLED`/`CURRENT_PID` etc. hit the OLD thread's TLS.
8 multi-thread tests aborted with "scheduler resumed a done actor" (error 5).
Thin LTO only worked because it picked the local-exec model.
- context.rs: module doc "Thread-locals and migration" (the rule), tls_fence()
- every TLS accessor reachable from an actor stack is out of line:
actor::{current_pid, publish_outcome (new)}, preempt::{maybe_preempt,
check_cancelled, current_slot_ptr, note_*, preemption_swap/enabled (new;
replaces direct PREEMPTION_ENABLED.with in NoPreempt/RawMutex/with_runtime/
trace/debug asserts)}, context::{get,set}_scheduler_sp, runtime::{sched_slot,
slot_push, set_yield_intent}, causal, trace
- raw_mutex order checks: out of line only under debug_assertions
- Cargo.toml: [profile.reltest] = release without LTO; 41 test bins link in
~1.5 s each instead of ~12 s (suite 11 min -> 1.5 min) AND it is the
regression oracle for this bug: `cargo test --profile reltest`
Both profiles: 41/41 + doctests green.
This commit is contained in:
@@ -60,6 +60,14 @@ panic = "unwind"
|
||||
lto = "thin"
|
||||
codegen-units = 1
|
||||
|
||||
# `cargo test --profile reltest`: release codegen for the crate (same opt-level,
|
||||
# same panic strategy) but no LTO at the final link. Thin LTO is what makes each
|
||||
# of the ~40 test binaries cost ~12 s to link instead of ~2 s; the tests don't
|
||||
# need cross-crate LTO, the benches do (they keep using `release`).
|
||||
[profile.reltest]
|
||||
inherits = "release"
|
||||
lto = false
|
||||
|
||||
[[bench]]
|
||||
name = "primes"
|
||||
harness = false
|
||||
|
||||
Reference in New Issue
Block a user