Static Ghidra decompilation hit a wall on the alt-tab-while-not-paused reset bug - state machine confirmed correct, but the actual render function was never located and the cause couldn't be determined from code alone. This is the trigger for building runtime-observation tooling: a Windows-hosted debug agent (extending tools/dbg_path.py's proven ctypes Win32 debug-loop approach) plus a thin MCP server translating tool calls to it, reachable over the LAN from Caspian (the machine that actually runs the game).
9.5 KiB
Havoc Runtime-Observation MCP -- Design Spec
Date: 2026-07-03 Status: approved by user, ready for implementation planning
Why this exists
Static Ghidra decompilation of HAVOC_NOCD.EXE hit a wall investigating the alt-tab-while-
not-paused "resets to menu" bug (see docs/FUNCTIONS.md "Alt-tab / focus-loss investigation").
The state machine was confirmed correct by reading code; a candidate stale-object-pointer
hypothesis was ruled out by reading code; but the actual frame-render function was never
located, and the bug's real cause could not be determined from decompiled code alone. This is
the trigger condition noted in project memory for building runtime-observation tooling: when
static analysis can't answer a question, the next tool is watching the program actually run.
This spec covers one subsystem: an MCP server that lets an agent (this Claude Code
session, or a future one) attach a debugger to the running HAVOC_NOCD.EXE process, set
breakpoints/watchpoints, read/write memory and registers, and capture screenshots -- turning
"we can't tell what this does without running it" into an actual, drivable capability.
Two related ideas were raised and explicitly deferred, not designed here:
- An MCP/interaction layer for a future Godot remaster port (blocked: the Godot project is still empty).
- A local-LLM agent on the muninn/sif cf-orch GPU nodes that plays the game autonomously (a later automation layer that would sit on top of whichever MCP exists by then).
Environment
- Caspian (
10.1.10.228) is the Windows machine that actually runs the game, via dgVoodoo2. It is the user's personal machine, not a CircuitForge/cf-orch homelab node. - Caspian and the Linux RE box share the same NFS/SMB-mounted storage already used all
session for Ghidra work (
Z:\Development\devl\Havocon Windows,/Library/Development/ devl/Havocon Linux) -- this already-proven bridge is why a new network/file-sharing mechanism isn't needed for anything except the MCP protocol traffic itself. - Claude Code can also run natively on Caspian via PowerShell (with SMB access to the same share). This means the debug agent and MCP server can be built and iterated on directly by an agent running on Caspian, rather than authored blind from Linux and round-tripped through manual testing.
- There is existing precedent for exactly this kind of debugging:
tools/dbg_path.py, a purectypesWin32 debugger written during the original boot-crash investigation. It already proves outCreateProcesswithDEBUG_ONLY_THIS_PROCESS,WaitForDebugEvent, and reading WOW64 (x86-on-x64) thread context -- including aWOW64_CONTEXTstruct that already definesDr0-Dr7(the x86 hardware debug registers), which data breakpoints need. No admin rights or third-party tooling (WinDbg/cdb) were required. - Future direction (informs the design, not required now): the user intends to eventually
port Havoc to run under Wine on Heimdall (a CircuitForge homelab node), eliminating the need
for a separate Windows box entirely. Wine implements the same Win32 debug APIs
dbg_path.pyalready uses, so the debug agent's core logic should not need rewriting for that move -- only its host changes (Wine-on-Linux instead of native Windows).
Architecture
Two separate processes on Caspian, both started manually alongside the game (no auto-launch):
Claude Code ──MCP over HTTP──▶ MCP server (0.0.0.0:<mcp_port>, LAN-reachable)
(this or a │
future session) │ localhost only
▼
Debug agent (127.0.0.1:<agent_port>)
(extends tools/dbg_path.py)
│ ctypes Win32 debug APIs
▼
HAVOC_NOCD.EXE (via dgVoodoo2)
- Debug agent owns all Win32-specific logic: the
ctypesdebug loop, breakpoints, watchpoints, memory/register access, screenshot capture. Exposes a local JSON-over-TCP API on127.0.0.1only -- never LAN-exposed directly. - MCP server is a thin translation layer: implements the MCP protocol (Streamable HTTP transport) and turns each MCP tool call into a call to the local debug agent. No debugging logic or state of its own beyond what it forwards.
This split is deliberate for the future Wine-on-Heimdall move: only the debug agent needs to run inside Wine (the only piece touching real Win32 debug APIs against the game process). The MCP server could then run as a normal native Linux process on Heimdall, talking to the Wine-hosted agent over the same local API -- no MCP/HTTP stack needs to work inside Wine at all. Only the debug agent's host changes; its design and the MCP server's design do not.
Components and tool surface
Debug agent (tools/havoc_debug_agent.py, extends tools/dbg_path.py):
| Operation | Purpose |
|---|---|
launch(path) |
Start the game under the debugger (or attach if already running) |
set_breakpoint(addr, condition=None) |
Software (INT3) breakpoint at a VA. Optional condition is a small declarative check evaluated on hit -- e.g. {register: "eax", op: "!=", value: "0x47"} or {memory: "0x47eb48+0x18", size: 4, op: "!=", value: "0x47"}. If false, silently resume; if true, actually stop and report |
set_watchpoint(addr, size, mode) |
Hardware data breakpoint (read/write/readwrite) using DR0-DR3/DR7. Hard limit: 4 concurrent (x86 has exactly 4 hardware debug address registers) -- callers must be told this limit explicitly, not discover it via a silent failure |
remove_breakpoint(addr) / remove_watchpoint(addr) |
|
list_breakpoints() / list_watchpoints() |
|
continue_execution(timeout) |
Resume until a breakpoint/watchpoint hits, the process exits, or timeout elapses (returns "still running" in that case -- never blocks forever) |
read_memory(addr, len) / write_memory(addr, bytes) |
|
get_registers() |
Full WOW64 register set |
screenshot() |
Capture the game window |
get_status() |
running / stopped-at-breakpoint (which one, why) / exited (with exit code) |
MCP server (tools/havoc_mcp_server.py): one MCP tool per debug-agent operation above,
plus list_breakpoints/list_watchpoints exposed directly. Binds 0.0.0.0:<mcp_port> on
Caspian. Requires a shared-secret bearer token (see Security) on every request.
Data flow (worked example: the alt-tab bug)
- Agent calls
launch(or the game is already running and this attaches to it) - Agent calls
set_watchpointon whatever memory holds current game/menu state, mode=write - User alt-tabs away and back on Caspian
- The watchpoint trips the moment anything writes that memory; the debug agent reports which instruction did it
- Agent calls
get_registers()andread_memory()around that address, plusscreenshot()to correlate with what's visually on screen at that instant - Agent calls
continue_execution()to keep running, or sets a new breakpoint based on what was just learned, repeating until the actual cause is found
Error handling
continue_executionrequires a timeout; never blocks indefinitely- Reads/writes to unmapped memory return a clear error, not a debug-agent crash
- Process exit (
EXIT_PROCESS_DEBUG_EVENT) is detected and reported with exit code, not treated as a hang or silent disconnect - All tool calls are serialized (single lock) -- one game process, no concurrent-call races
- Requesting a 5th concurrent watchpoint returns a clear "limit is 4" error
Security
The MCP server is LAN-reachable and exposes real remote influence over a running process (memory read/write, execution control). Even on a trusted home LAN, this needs a lightweight shared-secret bearer token (env var on Caspian, matching value in the MCP client config) so an arbitrary other device on the LAN can't drive it. Nothing heavier than that is warranted (no TLS, no rate limiting, no user accounts) -- this is a personal single-user dev tool, not a production or multi-tenant system, and adding more would be overengineering for what it is.
Testing
No meaningful way to unit-test real OS debug APIs against a live GUI game in isolation, so validation is a manual acceptance checklist instead of an automated test suite:
- Breakpoint at a known-good, already-understood address (e.g.
busywait_frame_limiter_60hz@0x40C8E0) -- confirm register/memory values at the stop match what Ghidra's decompile predicts. - Conditional breakpoint that's initially false (game running normally) then becomes true (trigger the condition) -- confirm it only stops on the true case.
- Watchpoint on
DAT_0047c48c(the delta-time baseline, already fully understood from the FPS-coupling investigation) -- confirm it fires exactly whencompute_frame_delta_ticks_60hz_min1_clampwrites it, and not otherwise. - Screenshot capture during a real gameplay session, confirmed to produce a viewable image.
- End-to-end: reproduce the alt-tab bug live using the worked data-flow example above, and confirm the tooling actually surfaces new information the static analysis couldn't.
Build approach
Since Claude Code can run natively on Caspian via PowerShell with SMB access to the same
share, the recommended build path is to have an agent running there write and iterate on both
havoc_debug_agent.py and havoc_mcp_server.py directly against the real environment, rather
than authoring them blind from Linux and relying on manual round-trip testing by the user.