Evidence (history.md session 5, findings 10-11, 20-core jobrunner run): rq-mutex collapses with thread count on queue-heavy load (yield-storm 6.9x slower than striped at 20T; attributed cause of the baseline's multi-thread regression on yield_many/chained_spawn), while rq-mpmc is best-or-close everywhere: best 1-thread, best slot-on ping-pong (1237us vs mutex 6754us at 20T), -23%/-41 cyc per yield roundtrip on real hardware (single-core switch_cost, interleaved). rq-striped remains the churn-heavy many-core option, selectable per build. Docs + compile_error hints in run_queue.rs updated to name rq-mpmc as the default. slot_state.rs untouched, no loom-relevant changes; lib (60) + scheduler/channel/supervisor/park_wake/wake_slot/preempt (41) pass under the new default.
smarm
SMARM: Smarm, Marks Actor Runtime Machinery. A proof-of-concept green-thread actor runtime for Rust.
SMARM is my attempt to implement the erlang/OTP philosophy in the Rust programming language. This has yielded a fault-tolerant, fast, and scalable runtime. This runtime allows the creation of asynchronous applications in Rust without the function coloring associated with the async/await system. It encourages the writing of simple, synchronous code, and largely elides the need for lifetime annotations.
Overview
SMARM implements green-thread actors on a shared heap, communicating only by Send messages. By sharing the heap, SMARM avoids the copying overhead of Erlang, which is safe to do due to Rust's borrow checker.
On top of the core runtime mechanics, SMARM also provides a library of primitives for making applications closely inspired by erlang/OTP. This includes generic servers (gen_servers), generic state machines (gen_statem), and supervision trees.
Supervision trees are the core primitive to allow your application to survive an unexpected panic. Supervisors are processes dedicated to monitoring other processes, which can restart these should they fail. This means that when set up properly an application may 'self-heal' when encountering unforeseen circumstances.
SMARM is not cooperatively scheduled; it uses preemption. This means a heavy task will not starve out other lighter tasks. Everything will make steady progress, which translates to very beneficial behaviour under (over)load: average latency goes up, but tail latency does not blow up.
To help diagnose these unforeseen circumstances, smarm may be compiled with its tracing feature, which emits a full trace using Perfetto.
Should you want to optimize your application, SMARM is unusually well poised to help. As the runtime functionally controls time, SMARM comes with a built in causal profiler, under the causal feature.
I also built a Phoenix-Framework inspired HTTP 1.1 library on top of SMARM called URUS, which implements Pub/Sub, Channels, and basic amenities like Websockets and Server Sent Events.
Limitations
This runtime requires naked assembly to function, and has thus far only been implemented for x86-64 assembly. It expects an operating system that supports virtual address space, and is therefore not (yet) suited for embedded targets. The IO implementation is currently based around the Linux kernel's epoll mechanism, meaning it requires a (GNU+)Linux distribution to run.
The preemption mechanism works by wrapping the memory allocator and checking how many CPU cycles you have used compared to your timeslice budget. This allows preemption to fire in most normal code, but tight zero-allocation loops do not get caught and require manual insertion of check!() if you want preemption to function.
This library is still in its early stages, and while I try my best with loom and tests, stable operation cannot be guaranteed. Therefore it is not (yet) recommended for production use.
At this moment, stack memory for each green thread is capped. Uncapping this may lead to performance benefits for deeply recursive algorithms that in a traditional async runtime might require pointer-chases through the heap. This is as yet unrealised.
At this stage, the codebase is largely LLM-generated, which is obvious if you start to read through the internals. While I did the design, and I keep the LLM under tight rein, the codebase is not in a state that I am very happy with. This also goes for the documentation.
Quick taste
use smarm::{run, spawn, channel};
run(|| {
let (tx, rx) = channel::<i64>();
let h = spawn(move || {
for _ in 0..3 {
let v = rx.recv().unwrap();
println!("got {v}");
}
});
for v in 1..=3i64 {
tx.send(v).unwrap();
}
h.join().unwrap();
});
Stopping actors
Two strengths, as in OTP. request_stop(pid) is exit(Pid, kill): a cooperative
hard stop, unwinding at the actor's next observation point. request_shutdown(pid)
is exit(Pid, shutdown): an actor that traps exits (trap_exit(), or
ctx.trap_exit() in a gen_server / cx.trap_exit() in a gen_statem) receives it
as a signal — handle_shutdown / a shutdown row — and may drain before stopping
itself; one that does not trap is stopped outright. Supervisors trap:
request_shutdown(sup) tears the tree down top-down, each child per its
ChildSpec Shutdown policy (Timeout(d), Infinity, BrutalKill). The run's
root actor returning means "the program is done": every top-level actor gets a
request_shutdown, and run() returns when they are gone. From outside the
runtime (a signal thread), Runtime::handle().request_shutdown(pid) does the
same. examples/graceful_shutdown.rs shows all of it.
A gen_server or gen_statem lives until it stops, is shut down, or is killed; its
refs are addresses — dropping them never ends it (a forgotten one is swept at
root exit; with --features smarm-trace each such sweep is a root_sweep
trace line). The supervised shape is GenServerBuilder::named(N).run() /
gen_statem::run_named(N, m): the server runs inline as the ChildSpec child
itself, so the supervisor's shutdown reaches it directly, a restart re-binds the
name, and the program addresses it by name.
Layout
src/
stack.rs context.rs preempt.rs pid.rs actor.rs
scheduler.rs channel.rs mutex.rs timer.rs io.rs
supervisor.rs monitor.rs link.rs runtime.rs
gen_server.rs lib.rs
tests/
per-module integration tests
benches/
primes.rs fan-out/fan-in compute, vs tokio current_thread
Building and running
Standard Cargo. Requires Rust 1.95 or newer (the #[naked] attribute went stable in 1.88; we use a few unrelated post-1.88 features). I have worked hard to keep this library as dependency-free as possible. master is x86-64 Linux only. An experimental, untested aarch64 context-switch backend lives on the arm-port branch (extracted into a target_arch-gated src/arch/); it has not been validated on hardware yet. macOS remains on the deferred list because of the epoll dependency.
cargo test # all tests
cargo test --test mutex # one module
cargo bench # primes benchmark vs tokio
Docs
| Document | What it covers |
|---|---|
Architecture.md |
Design intent, runtime model, and deferred work |
smarm - Deep Dive.html |
Generated walkthrough of the system; good starting point if you want to learn about the internals |
BENCHMARKS_AND_TUNING.md |
Where smarm wins and loses vs tokio, preemption knob recommendations |
benchmarks.md |
Raw benchmark results, methodology, and tuning experiment log |
Coming up
Clustering: clustering multiple SMARM nodes together is in the pipeline. SMARM-BEAM Interop: Running SMARM as a supervised node under the BEAM via a Rustler NIF works, including message passing and supervision trees that span the runtimes. However, this library is still too unstable to release. SMARM is an interesting platform for implementing a 'dataflow' library, but work on this has not yet started.
Contributing
This started as a personal proof-of-concept, but it is starting to outgrow that name. If you want to contribute, please get in contact to discuss what you want to work on. Code without prior communication is not welcome.
A note on open source
An open source project is a gift, and by giving it, it is no longer mine. I highly enourage you to fork it, to make it your own. This repository, however, is still mine.