# LLM Debugger — Vision & Spec ## Why Debuggers are built for humans at a terminal. They optimize for interactive, ad-hoc inspection: type a command, read output, think, repeat. LLMs operating as coding agents have fundamentally different needs: - **Predictable, terse output** — token efficiency matters; verbose JSON or gdb's human-readable output wastes context window - **Scriptable observation** — the agent defines *what it wants to see* per step, not just *what to do* - **Stateful sessions** — the debugger holds state across many tool calls so the LLM's context window doesn't have to - **Live human visibility** — a human watching the agent work needs a real-time view without being in the loop Existing tools don't fit. gdb/lldb optimize for humans. gdb's MI protocol is unstable and incomplete. DAP is well-designed but IDE-shaped, not agent-shaped. There is no debugger built from first principles for an LLM agent as the primary user. --- ## What A **stateful debugger REPL process** that: - Attaches to or spawns a Linux process via ptrace - Exposes a rich, lazy `state` object representing the current stopped state - Embeds a QuickJS interpreter — the LLM sends raw JavaScript, stdout comes back - Hooks into debugger events (step, breakpoint, watchpoint) by letting the LLM register JS event listeners - Communicates over stdin/stdout — making it trivially usable as a subprocess from any agent harness, and easily testable in a sandbox The LLM is the primary author of scripts. A human may observe via a browser visualization (deferred — see *Not Yet*). ### The Core Loop ``` agent harness │ raw JS (stdin) ▼ debugger REPL process │ evaluates JS against live state │ fires registered listeners on events ▼ stdout (terse output — tab/newline delimited) ``` No custom protocol. No structured command parser. The language is the interface. ### What the LLM Can Do Anything expressible in JS against the `state` object. Examples ranging from trivial to sophisticated: ```js // One-liner inspection state.pc ``` ```js // Register a step hook addEventListener('step', e => { let dist = e.state.sp - e.state.stack.usable_base; print(`pc=${hex(e.state.pc)} sp_dist=${dist}`); }); ``` ```js // Stateful tracking across steps via ctx addEventListener('step', e => { ctx.min_sp = Math.min(ctx.min_sp ?? e.state.sp, e.state.sp); print(`pc=${hex(e.state.pc)} lowest_sp=${hex(ctx.min_sp)}`); }); ``` ```js // DOM mutation for human visualization (future) addEventListener('step', e => { dom.id('pc').textContent = hex(e.state.pc); print(hex(e.state.pc)); }); ``` --- ## How ### Stack | Layer | Technology | Notes | |---|---|---| | Debugger backend | `ptrace` (Linux) | Via `nix` crate | | Host process | Rust | CLI binary | | Scripting | QuickJS | Via `rquickjs` crate | | Disassembly | `iced-x86` | x86-64 | | DWARF / unwinding | `gimli` | Frame unwinding, debug info | | ELF / symbols | `object` | Symbol table, TLS segments | ### Process Model The REPL is a **CLI daemon** — launched by the agent harness as a subprocess, communicating over stdin/stdout. One REPL process per debug session. State (breakpoints, watchpoints, `ctx`, registered listeners) persists for the lifetime of the process. On attach, the REPL: 1. Parses the ELF binary (symbols, TLS layout, section map) 2. Attaches via ptrace 3. Enters the read-eval-print loop On detach, the target process is allowed to continue (Linux ptrace default). ### The `state` Object Exposed to every JS evaluation. Access is **lazy** — each property triggers a ptrace/procfs read on demand. Accessed fields are **snapshotted** after each evaluation, making them available as `prev` on the next event. ```js state.pc // current instruction pointer state.sp // stack pointer state.registers.{rax, rbx, ...} // all GP registers state.flags.{z, c, n, v, ...} // condition flags state.instruction.addr // address of current instruction state.instruction.bytes // raw bytes (hex string) state.instruction.mnemonic // e.g. "mov rsp, rax" state.instruction.operands // array of operand strings state.memory(addr).u8() // typed memory accessors state.memory(addr).u32() state.memory(addr).u64() state.memory(addr).i64() state.memory(addr).bytes(n) // raw hex string, n bytes state.memory(addr).str() // null-terminated C string state.disasm(n) // disassemble n instructions from pc state.symbol_at(addr) // → string | null state.thread_local(name) // → value (resolved via fs_base + ELF TLS offset) state.stack.top // highest valid stack address state.stack.usable_base // lowest usable address (above guard) state.stack.guard_size // guard page size in bytes state.unwind() // → [{frame, pc, sp, symbol}] prev.pc // same shape, previous stop's snapshot prev.sp // ... etc ``` ### The `ctx` Object A plain mutable JS object, persistent across all evaluations for the lifetime of the REPL process. The LLM initialises it however it wants. Resets only on REPL restart. ```js ctx.my_counter = (ctx.my_counter ?? 0) + 1; ``` ### Event Listeners The REPL drives the QuickJS event loop explicitly — no libuv, no async. On each debugger stop, the REPL: 1. Determines stop reason (step, breakpoint, watchpoint, signal) 2. Looks up registered listeners for that event type 3. Calls each with an event object `{ state, prev, stop_reason }` 4. Collects printed output, returns it to stdout ```js addEventListener('step', handler) addEventListener('breakpoint', handler) // fires with e.breakpoint_id addEventListener('watchpoint', handler) // fires with e.watchpoint_id, e.access_type addEventListener('signal', handler) // fires with e.signal (SIGSEGV etc.) ``` ### Execution Control (Built-in Globals) Control flow and setup are plain JS functions available as globals — no special syntax, no parser. They compose naturally with the rest of the scripting layer. ```js step(n?) // step n instructions (default 1) continue() // resume execution attach(pid) // attach to running process spawn(binary, ...args) // spawn and attach detach() // detach, target continues breakpoint.set(addr_or_symbol, name?) // → id breakpoint.remove(id_or_name) breakpoint.list() // → [{id, addr, name, hits}] watchpoint.set(addr, 'r'|'w'|'rw', name?) watchpoint.remove(id_or_name) watchpoint.list() thread.list() // → [{tid, state, ...}] thread.select(tid) symbols() // dump loaded symbol table ctx.reset() // clear ctx object ``` Everything is JS. A breakpoint can be set inside a conditional, `step()` can be called in a loop, return values are inspectable. No escape hatch needed. ### Output Format Terse, tab-and-newline delimited. No JSON. Designed for token efficiency: ``` pc 0x8020 sp 0x7ffe0000 instr mov rsp, rax dist 4096 ``` The LLM author of the script controls output entirely via `print()`. The format above is a convention, not enforced. ### Error Handling - JS exceptions → clean error string on stdout, no panic - ptrace failures mid-script (bad address, process died) → error string on stdout with context - Unknown `:command` → error string ### Symbol & TLS Resolution On attach, the REPL parses the ELF binary and any loaded shared libraries (via `/proc/pid/maps`) to build: - Address → symbol name map - Symbol name → address map - TLS variable name → `fs_base` offset map (from ELF TLS segments) `state.thread_local("SCHEDULER_SP")` resolves to `*(fs_base + offset)` — a direct memory read, no runtime lookup overhead. --- ## Open Questions **Threading model during step/continue** When multiple threads are running and a non-selected thread hits a breakpoint, does the REPL surface that as an event, silently stop that thread, or let it continue? The right answer probably depends on use case but needs a concrete default. **Attach vs spawn lifecycle differences** On spawn, the REPL owns the process and killing the REPL kills the target. On attach, detach should leave the target running. Are there other meaningful differences in what state is available? **Batch stepping efficiency** `:step 1000` fires the step listener 1000 times. For large step counts this may be slow. A future `--summary` mode (first, last, interesting events) would help, but the right shape of "interesting" is unclear without real usage data. **Memory accessor return types** `state.memory(addr).u64()` returns a JS number, which loses precision above 2^53. For addresses and large register values this matters. Options: BigInt (awkward), hex string (loses arithmetic), or a thin wrapper type. Needs a decision before implementation. **Reverse execution** ptrace doesn't support it natively. Would require record-and-replay (rr-style). Desirable but the implementation cost is high. --- ## What Not **gdb/lldb as backend** gdb MI is unstable across versions, incomplete, and not designed for programmatic use. The wrapper code is always a mess. lldb's C API is clean but pulls in all of LLVM. ptrace is 30 syscalls we own completely — less total work for a cleaner result. **DAP (Debug Adapter Protocol)** Well-designed, but IDE-shaped. Adds significant protocol surface for a benefit (editor integration) that isn't the goal. The LLM doesn't need a DAP client. **Python as scripting layer** Best LLM familiarity, but `pyo3` embedding is heavy, startup is slow, and multi-line string ergonomics in a single tool call are worse than JS. QuickJS wins on the embedded use case. **Lua as scripting layer** Lightweight and excellent for embedding. Loses to JS specifically because DOM manipulation is the other major scripting use case, and `dom.id('x').textContent = y` is native JS idiom. LLMs have also seen vastly more JS than Lua. **Emulator/sandbox backend (Unicorn, qemu-user)** Deterministic and safe, but no real syscalls, no real threads, no signals. The goal is debugging real programs in production-like conditions. **Structured command protocol** A JSON or XML command/response protocol between harness and REPL adds parsing complexity on both ends for no benefit. stdin/stdout with raw JS in and terse text out is simpler and more flexible. **Eager state snapshots** Snapshotting all ptrace-readable state on every stop wastes syscalls when scripts touch 3–5 fields. Lazy access with snapshot-on-read is the right default. --- ## What Not Yet (MVP Scope) **Browser visualization / DOM API** The `dom` object and SSE transport to a browser (Datastar or otherwise) are architecturally clean to add later — the JS scripting layer is already there. Not in MVP. **Rich browser inspection surfaces** The hook author (LLM or human) knows the semantic meaning of memory that the compiler forgot. This opens up a class of visualization far beyond register tables and SP gauges. A pixel buffer becomes a canvas image (`putImageData`). A mesh becomes a Three.js render. A ring buffer becomes a waveform or spectrogram via Web Audio. An actor graph becomes a D3 force layout. The DOM API design should be rich enough to support canvas, WebGL, and Web Audio — not just `textContent` mutations — so this possibility isn't accidentally closed off. Implementation deferred until the basic visualization layer exists. **DWARF locals and source-level debugging** `gimli` can give us source lines and local variable names. High value, significant implementation work. Frame unwinding is in MVP; locals are not. **Conditional breakpoints** `:break --condition "state.sp < 0x1000"` is a natural extension. Defer until the core hook system is proven. **Multi-threaded event routing** Selecting threads and stepping them is in MVP. Simultaneous multi-thread event dispatch (all threads running, any can hit a breakpoint) is a harder scheduling problem — defer. **Reverse execution / time-travel** Requires record-and-replay infrastructure. Significant project in itself. **Remote debugging** gdbserver-style attach over a socket. The stdin/stdout model works locally; network transport is a later concern. **Shared library / dynamic linker awareness** Symbols from dynamically loaded libraries (via `dlopen`) won't be in the initial ELF parse. Watching `ld.so` events to update the symbol table at runtime is deferred.