docs: add SPEC.md (original vision) and README.md

This commit is contained in:
claude
2026-05-27 21:31:36 +00:00
parent 2f64981729
commit 98e3347b4c
2 changed files with 527 additions and 0 deletions
+288
View File
@@ -0,0 +1,288 @@
# LLM Debugger — Vision & Spec
## Why
Debuggers are built for humans at a terminal. They optimize for interactive, ad-hoc inspection: type a command, read output, think, repeat. LLMs operating as coding agents have fundamentally different needs:
- **Predictable, terse output** — token efficiency matters; verbose JSON or gdb's human-readable output wastes context window
- **Scriptable observation** — the agent defines *what it wants to see* per step, not just *what to do*
- **Stateful sessions** — the debugger holds state across many tool calls so the LLM's context window doesn't have to
- **Live human visibility** — a human watching the agent work needs a real-time view without being in the loop
Existing tools don't fit. gdb/lldb optimize for humans. gdb's MI protocol is unstable and incomplete. DAP is well-designed but IDE-shaped, not agent-shaped. There is no debugger built from first principles for an LLM agent as the primary user.
---
## What
A **stateful debugger REPL process** that:
- Attaches to or spawns a Linux process via ptrace
- Exposes a rich, lazy `state` object representing the current stopped state
- Embeds a QuickJS interpreter — the LLM sends raw JavaScript, stdout comes back
- Hooks into debugger events (step, breakpoint, watchpoint) by letting the LLM register JS event listeners
- Communicates over stdin/stdout — making it trivially usable as a subprocess from any agent harness, and easily testable in a sandbox
The LLM is the primary author of scripts. A human may observe via a browser visualization (deferred — see *Not Yet*).
### The Core Loop
```
agent harness
│ raw JS (stdin)
debugger REPL process
│ evaluates JS against live state
│ fires registered listeners on events
stdout (terse output — tab/newline delimited)
```
No custom protocol. No structured command parser. The language is the interface.
### What the LLM Can Do
Anything expressible in JS against the `state` object. Examples ranging from trivial to sophisticated:
```js
// One-liner inspection
state.pc
```
```js
// Register a step hook
addEventListener('step', e => {
let dist = e.state.sp - e.state.stack.usable_base;
print(`pc=${hex(e.state.pc)} sp_dist=${dist}`);
});
```
```js
// Stateful tracking across steps via ctx
addEventListener('step', e => {
ctx.min_sp = Math.min(ctx.min_sp ?? e.state.sp, e.state.sp);
print(`pc=${hex(e.state.pc)} lowest_sp=${hex(ctx.min_sp)}`);
});
```
```js
// DOM mutation for human visualization (future)
addEventListener('step', e => {
dom.id('pc').textContent = hex(e.state.pc);
print(hex(e.state.pc));
});
```
---
## How
### Stack
| Layer | Technology | Notes |
|---|---|---|
| Debugger backend | `ptrace` (Linux) | Via `nix` crate |
| Host process | Rust | CLI binary |
| Scripting | QuickJS | Via `rquickjs` crate |
| Disassembly | `iced-x86` | x86-64 |
| DWARF / unwinding | `gimli` | Frame unwinding, debug info |
| ELF / symbols | `object` | Symbol table, TLS segments |
### Process Model
The REPL is a **CLI daemon** — launched by the agent harness as a subprocess, communicating over stdin/stdout. One REPL process per debug session. State (breakpoints, watchpoints, `ctx`, registered listeners) persists for the lifetime of the process.
On attach, the REPL:
1. Parses the ELF binary (symbols, TLS layout, section map)
2. Attaches via ptrace
3. Enters the read-eval-print loop
On detach, the target process is allowed to continue (Linux ptrace default).
### The `state` Object
Exposed to every JS evaluation. Access is **lazy** — each property triggers a ptrace/procfs read on demand. Accessed fields are **snapshotted** after each evaluation, making them available as `prev` on the next event.
```js
state.pc // current instruction pointer
state.sp // stack pointer
state.registers.{rax, rbx, ...} // all GP registers
state.flags.{z, c, n, v, ...} // condition flags
state.instruction.addr // address of current instruction
state.instruction.bytes // raw bytes (hex string)
state.instruction.mnemonic // e.g. "mov rsp, rax"
state.instruction.operands // array of operand strings
state.memory(addr).u8() // typed memory accessors
state.memory(addr).u32()
state.memory(addr).u64()
state.memory(addr).i64()
state.memory(addr).bytes(n) // raw hex string, n bytes
state.memory(addr).str() // null-terminated C string
state.disasm(n) // disassemble n instructions from pc
state.symbol_at(addr) // → string | null
state.thread_local(name) // → value (resolved via fs_base + ELF TLS offset)
state.stack.top // highest valid stack address
state.stack.usable_base // lowest usable address (above guard)
state.stack.guard_size // guard page size in bytes
state.unwind() // → [{frame, pc, sp, symbol}]
prev.pc // same shape, previous stop's snapshot
prev.sp
// ... etc
```
### The `ctx` Object
A plain mutable JS object, persistent across all evaluations for the lifetime of the REPL process. The LLM initialises it however it wants. Resets only on REPL restart.
```js
ctx.my_counter = (ctx.my_counter ?? 0) + 1;
```
### Event Listeners
The REPL drives the QuickJS event loop explicitly — no libuv, no async. On each debugger stop, the REPL:
1. Determines stop reason (step, breakpoint, watchpoint, signal)
2. Looks up registered listeners for that event type
3. Calls each with an event object `{ state, prev, stop_reason }`
4. Collects printed output, returns it to stdout
```js
addEventListener('step', handler)
addEventListener('breakpoint', handler) // fires with e.breakpoint_id
addEventListener('watchpoint', handler) // fires with e.watchpoint_id, e.access_type
addEventListener('signal', handler) // fires with e.signal (SIGSEGV etc.)
```
### Execution Control (Built-in Globals)
Control flow and setup are plain JS functions available as globals — no special syntax, no parser. They compose naturally with the rest of the scripting layer.
```js
step(n?) // step n instructions (default 1)
continue() // resume execution
attach(pid) // attach to running process
spawn(binary, ...args) // spawn and attach
detach() // detach, target continues
breakpoint.set(addr_or_symbol, name?) // → id
breakpoint.remove(id_or_name)
breakpoint.list() // → [{id, addr, name, hits}]
watchpoint.set(addr, 'r'|'w'|'rw', name?)
watchpoint.remove(id_or_name)
watchpoint.list()
thread.list() // → [{tid, state, ...}]
thread.select(tid)
symbols() // dump loaded symbol table
ctx.reset() // clear ctx object
```
Everything is JS. A breakpoint can be set inside a conditional, `step()` can be called in a loop, return values are inspectable. No escape hatch needed.
### Output Format
Terse, tab-and-newline delimited. No JSON. Designed for token efficiency:
```
pc 0x8020
sp 0x7ffe0000
instr mov rsp, rax
dist 4096
```
The LLM author of the script controls output entirely via `print()`. The format above is a convention, not enforced.
### Error Handling
- JS exceptions → clean error string on stdout, no panic
- ptrace failures mid-script (bad address, process died) → error string on stdout with context
- Unknown `:command` → error string
### Symbol & TLS Resolution
On attach, the REPL parses the ELF binary and any loaded shared libraries (via `/proc/pid/maps`) to build:
- Address → symbol name map
- Symbol name → address map
- TLS variable name → `fs_base` offset map (from ELF TLS segments)
`state.thread_local("SCHEDULER_SP")` resolves to `*(fs_base + offset)` — a direct memory read, no runtime lookup overhead.
---
## Open Questions
**Threading model during step/continue**
When multiple threads are running and a non-selected thread hits a breakpoint, does the REPL surface that as an event, silently stop that thread, or let it continue? The right answer probably depends on use case but needs a concrete default.
**Attach vs spawn lifecycle differences**
On spawn, the REPL owns the process and killing the REPL kills the target. On attach, detach should leave the target running. Are there other meaningful differences in what state is available?
**Batch stepping efficiency**
`:step 1000` fires the step listener 1000 times. For large step counts this may be slow. A future `--summary` mode (first, last, interesting events) would help, but the right shape of "interesting" is unclear without real usage data.
**Memory accessor return types**
`state.memory(addr).u64()` returns a JS number, which loses precision above 2^53. For addresses and large register values this matters. Options: BigInt (awkward), hex string (loses arithmetic), or a thin wrapper type. Needs a decision before implementation.
**Reverse execution**
ptrace doesn't support it natively. Would require record-and-replay (rr-style). Desirable but the implementation cost is high.
---
## What Not
**gdb/lldb as backend**
gdb MI is unstable across versions, incomplete, and not designed for programmatic use. The wrapper code is always a mess. lldb's C API is clean but pulls in all of LLVM. ptrace is 30 syscalls we own completely — less total work for a cleaner result.
**DAP (Debug Adapter Protocol)**
Well-designed, but IDE-shaped. Adds significant protocol surface for a benefit (editor integration) that isn't the goal. The LLM doesn't need a DAP client.
**Python as scripting layer**
Best LLM familiarity, but `pyo3` embedding is heavy, startup is slow, and multi-line string ergonomics in a single tool call are worse than JS. QuickJS wins on the embedded use case.
**Lua as scripting layer**
Lightweight and excellent for embedding. Loses to JS specifically because DOM manipulation is the other major scripting use case, and `dom.id('x').textContent = y` is native JS idiom. LLMs have also seen vastly more JS than Lua.
**Emulator/sandbox backend (Unicorn, qemu-user)**
Deterministic and safe, but no real syscalls, no real threads, no signals. The goal is debugging real programs in production-like conditions.
**Structured command protocol**
A JSON or XML command/response protocol between harness and REPL adds parsing complexity on both ends for no benefit. stdin/stdout with raw JS in and terse text out is simpler and more flexible.
**Eager state snapshots**
Snapshotting all ptrace-readable state on every stop wastes syscalls when scripts touch 35 fields. Lazy access with snapshot-on-read is the right default.
---
## What Not Yet (MVP Scope)
**Browser visualization / DOM API**
The `dom` object and SSE transport to a browser (Datastar or otherwise) are architecturally clean to add later — the JS scripting layer is already there. Not in MVP.
**Rich browser inspection surfaces**
The hook author (LLM or human) knows the semantic meaning of memory that the compiler forgot. This opens up a class of visualization far beyond register tables and SP gauges. A pixel buffer becomes a canvas image (`putImageData`). A mesh becomes a Three.js render. A ring buffer becomes a waveform or spectrogram via Web Audio. An actor graph becomes a D3 force layout. The DOM API design should be rich enough to support canvas, WebGL, and Web Audio — not just `textContent` mutations — so this possibility isn't accidentally closed off. Implementation deferred until the basic visualization layer exists.
**DWARF locals and source-level debugging**
`gimli` can give us source lines and local variable names. High value, significant implementation work. Frame unwinding is in MVP; locals are not.
**Conditional breakpoints**
`:break <addr> --condition "state.sp < 0x1000"` is a natural extension. Defer until the core hook system is proven.
**Multi-threaded event routing**
Selecting threads and stepping them is in MVP. Simultaneous multi-thread event dispatch (all threads running, any can hit a breakpoint) is a harder scheduling problem — defer.
**Reverse execution / time-travel**
Requires record-and-replay infrastructure. Significant project in itself.
**Remote debugging**
gdbserver-style attach over a socket. The stdin/stdout model works locally; network transport is a later concern.
**Shared library / dynamic linker awareness**
Symbols from dynamically loaded libraries (via `dlopen`) won't be in the initial ELF parse. Watching `ld.so` events to update the symbol table at runtime is deferred.