12 KiB
LLM Debugger — Vision & Spec
Why
Debuggers are built for humans at a terminal. They optimize for interactive, ad-hoc inspection: type a command, read output, think, repeat. LLMs operating as coding agents have fundamentally different needs:
- Predictable, terse output — token efficiency matters; verbose JSON or gdb's human-readable output wastes context window
- Scriptable observation — the agent defines what it wants to see per step, not just what to do
- Stateful sessions — the debugger holds state across many tool calls so the LLM's context window doesn't have to
- Live human visibility — a human watching the agent work needs a real-time view without being in the loop
Existing tools don't fit. gdb/lldb optimize for humans. gdb's MI protocol is unstable and incomplete. DAP is well-designed but IDE-shaped, not agent-shaped. There is no debugger built from first principles for an LLM agent as the primary user.
What
A stateful debugger REPL process that:
- Attaches to or spawns a Linux process via ptrace
- Exposes a rich, lazy
stateobject representing the current stopped state - Embeds a QuickJS interpreter — the LLM sends raw JavaScript, stdout comes back
- Hooks into debugger events (step, breakpoint, watchpoint) by letting the LLM register JS event listeners
- Communicates over stdin/stdout — making it trivially usable as a subprocess from any agent harness, and easily testable in a sandbox
The LLM is the primary author of scripts. A human may observe via a browser visualization (deferred — see Not Yet).
The Core Loop
agent harness
│ raw JS (stdin)
▼
debugger REPL process
│ evaluates JS against live state
│ fires registered listeners on events
▼
stdout (terse output — tab/newline delimited)
No custom protocol. No structured command parser. The language is the interface.
What the LLM Can Do
Anything expressible in JS against the state object. Examples ranging from trivial to sophisticated:
// One-liner inspection
state.pc
// Register a step hook
addEventListener('step', e => {
let dist = e.state.sp - e.state.stack.usable_base;
print(`pc=${hex(e.state.pc)} sp_dist=${dist}`);
});
// Stateful tracking across steps via ctx
addEventListener('step', e => {
ctx.min_sp = Math.min(ctx.min_sp ?? e.state.sp, e.state.sp);
print(`pc=${hex(e.state.pc)} lowest_sp=${hex(ctx.min_sp)}`);
});
// DOM mutation for human visualization (future)
addEventListener('step', e => {
dom.id('pc').textContent = hex(e.state.pc);
print(hex(e.state.pc));
});
How
Stack
| Layer | Technology | Notes |
|---|---|---|
| Debugger backend | ptrace (Linux) |
Via nix crate |
| Host process | Rust | CLI binary |
| Scripting | QuickJS | Via rquickjs crate |
| Disassembly | iced-x86 |
x86-64 |
| DWARF / unwinding | gimli |
Frame unwinding, debug info |
| ELF / symbols | object |
Symbol table, TLS segments |
Process Model
The REPL is a CLI daemon — launched by the agent harness as a subprocess, communicating over stdin/stdout. One REPL process per debug session. State (breakpoints, watchpoints, ctx, registered listeners) persists for the lifetime of the process.
On attach, the REPL:
- Parses the ELF binary (symbols, TLS layout, section map)
- Attaches via ptrace
- Enters the read-eval-print loop
On detach, the target process is allowed to continue (Linux ptrace default).
The state Object
Exposed to every JS evaluation. Access is lazy — each property triggers a ptrace/procfs read on demand. Accessed fields are snapshotted after each evaluation, making them available as prev on the next event.
state.pc // current instruction pointer
state.sp // stack pointer
state.registers.{rax, rbx, ...} // all GP registers
state.flags.{z, c, n, v, ...} // condition flags
state.instruction.addr // address of current instruction
state.instruction.bytes // raw bytes (hex string)
state.instruction.mnemonic // e.g. "mov rsp, rax"
state.instruction.operands // array of operand strings
state.memory(addr).u8() // typed memory accessors
state.memory(addr).u32()
state.memory(addr).u64()
state.memory(addr).i64()
state.memory(addr).bytes(n) // raw hex string, n bytes
state.memory(addr).str() // null-terminated C string
state.disasm(n) // disassemble n instructions from pc
state.symbol_at(addr) // → string | null
state.thread_local(name) // → value (resolved via fs_base + ELF TLS offset)
state.stack.top // highest valid stack address
state.stack.usable_base // lowest usable address (above guard)
state.stack.guard_size // guard page size in bytes
state.unwind() // → [{frame, pc, sp, symbol}]
prev.pc // same shape, previous stop's snapshot
prev.sp
// ... etc
The ctx Object
A plain mutable JS object, persistent across all evaluations for the lifetime of the REPL process. The LLM initialises it however it wants. Resets only on REPL restart.
ctx.my_counter = (ctx.my_counter ?? 0) + 1;
Event Listeners
The REPL drives the QuickJS event loop explicitly — no libuv, no async. On each debugger stop, the REPL:
- Determines stop reason (step, breakpoint, watchpoint, signal)
- Looks up registered listeners for that event type
- Calls each with an event object
{ state, prev, stop_reason } - Collects printed output, returns it to stdout
addEventListener('step', handler)
addEventListener('breakpoint', handler) // fires with e.breakpoint_id
addEventListener('watchpoint', handler) // fires with e.watchpoint_id, e.access_type
addEventListener('signal', handler) // fires with e.signal (SIGSEGV etc.)
Execution Control (Built-in Globals)
Control flow and setup are plain JS functions available as globals — no special syntax, no parser. They compose naturally with the rest of the scripting layer.
step(n?) // step n instructions (default 1)
continue() // resume execution
attach(pid) // attach to running process
spawn(binary, ...args) // spawn and attach
detach() // detach, target continues
breakpoint.set(addr_or_symbol, name?) // → id
breakpoint.remove(id_or_name)
breakpoint.list() // → [{id, addr, name, hits}]
watchpoint.set(addr, 'r'|'w'|'rw', name?)
watchpoint.remove(id_or_name)
watchpoint.list()
thread.list() // → [{tid, state, ...}]
thread.select(tid)
symbols() // dump loaded symbol table
ctx.reset() // clear ctx object
Everything is JS. A breakpoint can be set inside a conditional, step() can be called in a loop, return values are inspectable. No escape hatch needed.
Output Format
Terse, tab-and-newline delimited. No JSON. Designed for token efficiency:
pc 0x8020
sp 0x7ffe0000
instr mov rsp, rax
dist 4096
The LLM author of the script controls output entirely via print(). The format above is a convention, not enforced.
Error Handling
- JS exceptions → clean error string on stdout, no panic
- ptrace failures mid-script (bad address, process died) → error string on stdout with context
- Unknown
:command→ error string
Symbol & TLS Resolution
On attach, the REPL parses the ELF binary and any loaded shared libraries (via /proc/pid/maps) to build:
- Address → symbol name map
- Symbol name → address map
- TLS variable name →
fs_baseoffset map (from ELF TLS segments)
state.thread_local("SCHEDULER_SP") resolves to *(fs_base + offset) — a direct memory read, no runtime lookup overhead.
Open Questions
Threading model during step/continue When multiple threads are running and a non-selected thread hits a breakpoint, does the REPL surface that as an event, silently stop that thread, or let it continue? The right answer probably depends on use case but needs a concrete default.
Attach vs spawn lifecycle differences On spawn, the REPL owns the process and killing the REPL kills the target. On attach, detach should leave the target running. Are there other meaningful differences in what state is available?
Batch stepping efficiency
:step 1000 fires the step listener 1000 times. For large step counts this may be slow. A future --summary mode (first, last, interesting events) would help, but the right shape of "interesting" is unclear without real usage data.
Memory accessor return types
state.memory(addr).u64() returns a JS number, which loses precision above 2^53. For addresses and large register values this matters. Options: BigInt (awkward), hex string (loses arithmetic), or a thin wrapper type. Needs a decision before implementation.
Reverse execution ptrace doesn't support it natively. Would require record-and-replay (rr-style). Desirable but the implementation cost is high.
What Not
gdb/lldb as backend gdb MI is unstable across versions, incomplete, and not designed for programmatic use. The wrapper code is always a mess. lldb's C API is clean but pulls in all of LLVM. ptrace is 30 syscalls we own completely — less total work for a cleaner result.
DAP (Debug Adapter Protocol) Well-designed, but IDE-shaped. Adds significant protocol surface for a benefit (editor integration) that isn't the goal. The LLM doesn't need a DAP client.
Python as scripting layer
Best LLM familiarity, but pyo3 embedding is heavy, startup is slow, and multi-line string ergonomics in a single tool call are worse than JS. QuickJS wins on the embedded use case.
Lua as scripting layer
Lightweight and excellent for embedding. Loses to JS specifically because DOM manipulation is the other major scripting use case, and dom.id('x').textContent = y is native JS idiom. LLMs have also seen vastly more JS than Lua.
Emulator/sandbox backend (Unicorn, qemu-user) Deterministic and safe, but no real syscalls, no real threads, no signals. The goal is debugging real programs in production-like conditions.
Structured command protocol A JSON or XML command/response protocol between harness and REPL adds parsing complexity on both ends for no benefit. stdin/stdout with raw JS in and terse text out is simpler and more flexible.
Eager state snapshots Snapshotting all ptrace-readable state on every stop wastes syscalls when scripts touch 3–5 fields. Lazy access with snapshot-on-read is the right default.
What Not Yet (MVP Scope)
Browser visualization / DOM API
The dom object and SSE transport to a browser (Datastar or otherwise) are architecturally clean to add later — the JS scripting layer is already there. Not in MVP.
Rich browser inspection surfaces
The hook author (LLM or human) knows the semantic meaning of memory that the compiler forgot. This opens up a class of visualization far beyond register tables and SP gauges. A pixel buffer becomes a canvas image (putImageData). A mesh becomes a Three.js render. A ring buffer becomes a waveform or spectrogram via Web Audio. An actor graph becomes a D3 force layout. The DOM API design should be rich enough to support canvas, WebGL, and Web Audio — not just textContent mutations — so this possibility isn't accidentally closed off. Implementation deferred until the basic visualization layer exists.
DWARF locals and source-level debugging
gimli can give us source lines and local variable names. High value, significant implementation work. Frame unwinding is in MVP; locals are not.
Conditional breakpoints
:break <addr> --condition "state.sp < 0x1000" is a natural extension. Defer until the core hook system is proven.
Multi-threaded event routing Selecting threads and stepping them is in MVP. Simultaneous multi-thread event dispatch (all threads running, any can hit a breakpoint) is a harder scheduling problem — defer.
Reverse execution / time-travel Requires record-and-replay infrastructure. Significant project in itself.
Remote debugging gdbserver-style attach over a socket. The stdin/stdout model works locally; network transport is a later concern.
Shared library / dynamic linker awareness
Symbols from dynamically loaded libraries (via dlopen) won't be in the initial ELF parse. Watching ld.so events to update the symbol table at runtime is deferred.