Files

12 KiB
Raw Permalink Blame History

LLM Debugger — Vision & Spec

Why

Debuggers are built for humans at a terminal. They optimize for interactive, ad-hoc inspection: type a command, read output, think, repeat. LLMs operating as coding agents have fundamentally different needs:

  • Predictable, terse output — token efficiency matters; verbose JSON or gdb's human-readable output wastes context window
  • Scriptable observation — the agent defines what it wants to see per step, not just what to do
  • Stateful sessions — the debugger holds state across many tool calls so the LLM's context window doesn't have to
  • Live human visibility — a human watching the agent work needs a real-time view without being in the loop

Existing tools don't fit. gdb/lldb optimize for humans. gdb's MI protocol is unstable and incomplete. DAP is well-designed but IDE-shaped, not agent-shaped. There is no debugger built from first principles for an LLM agent as the primary user.


What

A stateful debugger REPL process that:

  • Attaches to or spawns a Linux process via ptrace
  • Exposes a rich, lazy state object representing the current stopped state
  • Embeds a QuickJS interpreter — the LLM sends raw JavaScript, stdout comes back
  • Hooks into debugger events (step, breakpoint, watchpoint) by letting the LLM register JS event listeners
  • Communicates over stdin/stdout — making it trivially usable as a subprocess from any agent harness, and easily testable in a sandbox

The LLM is the primary author of scripts. A human may observe via a browser visualization (deferred — see Not Yet).

The Core Loop

agent harness
    │  raw JS (stdin)
    ▼
debugger REPL process
    │  evaluates JS against live state
    │  fires registered listeners on events
    ▼
stdout (terse output — tab/newline delimited)

No custom protocol. No structured command parser. The language is the interface.

What the LLM Can Do

Anything expressible in JS against the state object. Examples ranging from trivial to sophisticated:

// One-liner inspection
state.pc
// Register a step hook
addEventListener('step', e => {
  let dist = e.state.sp - e.state.stack.usable_base;
  print(`pc=${hex(e.state.pc)} sp_dist=${dist}`);
});
// Stateful tracking across steps via ctx
addEventListener('step', e => {
  ctx.min_sp = Math.min(ctx.min_sp ?? e.state.sp, e.state.sp);
  print(`pc=${hex(e.state.pc)} lowest_sp=${hex(ctx.min_sp)}`);
});
// DOM mutation for human visualization (future)
addEventListener('step', e => {
  dom.id('pc').textContent = hex(e.state.pc);
  print(hex(e.state.pc));
});

How

Stack

Layer Technology Notes
Debugger backend ptrace (Linux) Via nix crate
Host process Rust CLI binary
Scripting QuickJS Via rquickjs crate
Disassembly iced-x86 x86-64
DWARF / unwinding gimli Frame unwinding, debug info
ELF / symbols object Symbol table, TLS segments

Process Model

The REPL is a CLI daemon — launched by the agent harness as a subprocess, communicating over stdin/stdout. One REPL process per debug session. State (breakpoints, watchpoints, ctx, registered listeners) persists for the lifetime of the process.

On attach, the REPL:

  1. Parses the ELF binary (symbols, TLS layout, section map)
  2. Attaches via ptrace
  3. Enters the read-eval-print loop

On detach, the target process is allowed to continue (Linux ptrace default).

The state Object

Exposed to every JS evaluation. Access is lazy — each property triggers a ptrace/procfs read on demand. Accessed fields are snapshotted after each evaluation, making them available as prev on the next event.

state.pc                          // current instruction pointer
state.sp                          // stack pointer
state.registers.{rax, rbx, ...}   // all GP registers
state.flags.{z, c, n, v, ...}     // condition flags
state.instruction.addr            // address of current instruction
state.instruction.bytes           // raw bytes (hex string)
state.instruction.mnemonic        // e.g. "mov rsp, rax"
state.instruction.operands        // array of operand strings

state.memory(addr).u8()           // typed memory accessors
state.memory(addr).u32()
state.memory(addr).u64()
state.memory(addr).i64()
state.memory(addr).bytes(n)       // raw hex string, n bytes
state.memory(addr).str()          // null-terminated C string

state.disasm(n)                   // disassemble n instructions from pc
state.symbol_at(addr)             // → string | null
state.thread_local(name)          // → value (resolved via fs_base + ELF TLS offset)

state.stack.top                   // highest valid stack address
state.stack.usable_base           // lowest usable address (above guard)
state.stack.guard_size            // guard page size in bytes

state.unwind()                    // → [{frame, pc, sp, symbol}]

prev.pc                           // same shape, previous stop's snapshot
prev.sp
// ... etc

The ctx Object

A plain mutable JS object, persistent across all evaluations for the lifetime of the REPL process. The LLM initialises it however it wants. Resets only on REPL restart.

ctx.my_counter = (ctx.my_counter ?? 0) + 1;

Event Listeners

The REPL drives the QuickJS event loop explicitly — no libuv, no async. On each debugger stop, the REPL:

  1. Determines stop reason (step, breakpoint, watchpoint, signal)
  2. Looks up registered listeners for that event type
  3. Calls each with an event object { state, prev, stop_reason }
  4. Collects printed output, returns it to stdout
addEventListener('step', handler)
addEventListener('breakpoint', handler)   // fires with e.breakpoint_id
addEventListener('watchpoint', handler)   // fires with e.watchpoint_id, e.access_type
addEventListener('signal', handler)       // fires with e.signal (SIGSEGV etc.)

Execution Control (Built-in Globals)

Control flow and setup are plain JS functions available as globals — no special syntax, no parser. They compose naturally with the rest of the scripting layer.

step(n?)                              // step n instructions (default 1)
continue()                            // resume execution
attach(pid)                           // attach to running process
spawn(binary, ...args)                // spawn and attach
detach()                              // detach, target continues

breakpoint.set(addr_or_symbol, name?) // → id
breakpoint.remove(id_or_name)
breakpoint.list()                     // → [{id, addr, name, hits}]

watchpoint.set(addr, 'r'|'w'|'rw', name?)
watchpoint.remove(id_or_name)
watchpoint.list()

thread.list()                         // → [{tid, state, ...}]
thread.select(tid)

symbols()                             // dump loaded symbol table
ctx.reset()                           // clear ctx object

Everything is JS. A breakpoint can be set inside a conditional, step() can be called in a loop, return values are inspectable. No escape hatch needed.

Output Format

Terse, tab-and-newline delimited. No JSON. Designed for token efficiency:

pc	0x8020
sp	0x7ffe0000
instr	mov rsp, rax
dist	4096

The LLM author of the script controls output entirely via print(). The format above is a convention, not enforced.

Error Handling

  • JS exceptions → clean error string on stdout, no panic
  • ptrace failures mid-script (bad address, process died) → error string on stdout with context
  • Unknown :command → error string

Symbol & TLS Resolution

On attach, the REPL parses the ELF binary and any loaded shared libraries (via /proc/pid/maps) to build:

  • Address → symbol name map
  • Symbol name → address map
  • TLS variable name → fs_base offset map (from ELF TLS segments)

state.thread_local("SCHEDULER_SP") resolves to *(fs_base + offset) — a direct memory read, no runtime lookup overhead.


Open Questions

Threading model during step/continue When multiple threads are running and a non-selected thread hits a breakpoint, does the REPL surface that as an event, silently stop that thread, or let it continue? The right answer probably depends on use case but needs a concrete default.

Attach vs spawn lifecycle differences On spawn, the REPL owns the process and killing the REPL kills the target. On attach, detach should leave the target running. Are there other meaningful differences in what state is available?

Batch stepping efficiency :step 1000 fires the step listener 1000 times. For large step counts this may be slow. A future --summary mode (first, last, interesting events) would help, but the right shape of "interesting" is unclear without real usage data.

Memory accessor return types state.memory(addr).u64() returns a JS number, which loses precision above 2^53. For addresses and large register values this matters. Options: BigInt (awkward), hex string (loses arithmetic), or a thin wrapper type. Needs a decision before implementation.

Reverse execution ptrace doesn't support it natively. Would require record-and-replay (rr-style). Desirable but the implementation cost is high.


What Not

gdb/lldb as backend gdb MI is unstable across versions, incomplete, and not designed for programmatic use. The wrapper code is always a mess. lldb's C API is clean but pulls in all of LLVM. ptrace is 30 syscalls we own completely — less total work for a cleaner result.

DAP (Debug Adapter Protocol) Well-designed, but IDE-shaped. Adds significant protocol surface for a benefit (editor integration) that isn't the goal. The LLM doesn't need a DAP client.

Python as scripting layer Best LLM familiarity, but pyo3 embedding is heavy, startup is slow, and multi-line string ergonomics in a single tool call are worse than JS. QuickJS wins on the embedded use case.

Lua as scripting layer Lightweight and excellent for embedding. Loses to JS specifically because DOM manipulation is the other major scripting use case, and dom.id('x').textContent = y is native JS idiom. LLMs have also seen vastly more JS than Lua.

Emulator/sandbox backend (Unicorn, qemu-user) Deterministic and safe, but no real syscalls, no real threads, no signals. The goal is debugging real programs in production-like conditions.

Structured command protocol A JSON or XML command/response protocol between harness and REPL adds parsing complexity on both ends for no benefit. stdin/stdout with raw JS in and terse text out is simpler and more flexible.

Eager state snapshots Snapshotting all ptrace-readable state on every stop wastes syscalls when scripts touch 35 fields. Lazy access with snapshot-on-read is the right default.


What Not Yet (MVP Scope)

Browser visualization / DOM API The dom object and SSE transport to a browser (Datastar or otherwise) are architecturally clean to add later — the JS scripting layer is already there. Not in MVP.

Rich browser inspection surfaces The hook author (LLM or human) knows the semantic meaning of memory that the compiler forgot. This opens up a class of visualization far beyond register tables and SP gauges. A pixel buffer becomes a canvas image (putImageData). A mesh becomes a Three.js render. A ring buffer becomes a waveform or spectrogram via Web Audio. An actor graph becomes a D3 force layout. The DOM API design should be rich enough to support canvas, WebGL, and Web Audio — not just textContent mutations — so this possibility isn't accidentally closed off. Implementation deferred until the basic visualization layer exists.

DWARF locals and source-level debugging gimli can give us source lines and local variable names. High value, significant implementation work. Frame unwinding is in MVP; locals are not.

Conditional breakpoints :break <addr> --condition "state.sp < 0x1000" is a natural extension. Defer until the core hook system is proven.

Multi-threaded event routing Selecting threads and stepping them is in MVP. Simultaneous multi-thread event dispatch (all threads running, any can hit a breakpoint) is a harder scheduling problem — defer.

Reverse execution / time-travel Requires record-and-replay infrastructure. Significant project in itself.

Remote debugging gdbserver-style attach over a socket. The stdin/stdout model works locally; network transport is a later concern.

Shared library / dynamic linker awareness Symbols from dynamically loaded libraries (via dlopen) won't be in the initial ELF parse. Watching ld.so events to update the symbol table at runtime is deferred.