Skip to main content
Glama
maci0

winedbg-mcp

by maci0
README.md
# winedbg-mcp

An MCP server for interacting with `winedbg` (the Wine debugger). It wraps the interactive debugger so an LLM can drive it through MCP tools.

## Status

The server is implemented and tested. `src/` holds the MCP entry point, the
winedbg session state machine, the environment parsing, the command-line
parsing and the tool-argument validation; `build/` is the compiled output of
`bun run build`; `tests/` covers all of those, and CI (`.github/workflows/ci.yml`)
runs the install, `bun run check` (Biome, both type-check passes and the
suite), the build and the artifact check. The session tests drive
`WinedbgSession` against a stand-in that speaks the same `Wine-dbg>` prompt
protocol, so the suite needs no Wine. No test here has been run against a real
`winedbg`: the debugger is the one thing the fixtures replace, so the suite
proves the prompt protocol and the tool argument handling, not Wine itself. See
[`docs/THREAT_MODEL.md`](docs/THREAT_MODEL.md) for the attack surface, the
trust boundaries and the mitigations, read off the source rather than off this
README.
[`CHANGELOG.md`](CHANGELOG.md) records what changed in each release.

Source layout, one concern per module:

| Path | Concern |
| --- | --- |
| `src/index.ts` | entrypoint: config load, transport, process lifecycle |
| `src/cli.ts` | the command line: `--help`, `--version`, argument errors |
| `src/tools.ts` | the MCP tool list and the call dispatch |
| `src/validate.ts` | validation of untyped tool arguments |
| `src/session.ts` | the winedbg child process and its prompt protocol |
| `src/runtime.ts` | the process, the clock and the process-group probe the session reaches the outside world through |
| `src/logger.ts` | the stderr log line format and the level filter |
| `src/config.ts` | reading and validating the environment |
| `src/constants.ts` | defaults and limits shared across the above |
| `src/version.ts` | the version reported to MCP clients, read from `package.json` |
| `scripts/verify-artifact.sh` | asserts the built entry point runs and ships nothing but compiled JavaScript |

## Prerequisites

- [Bun](https://bun.sh/), the package manager, build tool and test runner
- [Wine](https://www.winehq.org/), which includes `winedbg`

Supported platforms are Linux and macOS. `winedbg_stop` signals a process group
rather than a single process, so the debuggee `winedbg` launched is stopped with
it, and that needs POSIX process groups: neither the `detached` child group nor
the negative-pid `kill` exists on Windows, where the same call would leave the
program under debug running. CI runs the whole gate on both, so the claim is
tested rather than inferred from POSIX.

Node.js 18 or higher is only needed if you run the built `build/index.js`
with `node` instead of `bun`, or to run `scripts/verify-artifact.sh`, which
checks the artifact under both hosts. `package.json` declares the same floor
under `engines`, so a client on an older host is warned at install time.

## Installation

1. Clone the repository.
2. Install dependencies:
   ```bash
   bun install
   ```
   CI runs the same install as `bun install --frozen-lockfile`, so `bun.lock` is
   the only dependency set any build is allowed to resolve. A plain `bun install`
   updates it, which is how a dependency change enters the tree.
3. Build the project:
   ```bash
   bun run build
   ```

The build writes `build/index.js`. To run the server straight from source
without a build step, use `bun run dev` (`bun run start` runs the built file).
The build clears `build/` first, so a module deleted from `src/` cannot linger
in the artifact, and the output is byte-identical wherever the checkout sits,
whatever the locale or timezone, and at any wall-clock time: `tsc` emits no
timestamp, no absolute path and no source map.

## Configuration

To use this with an MCP client (like Claude Desktop or Gemini), configure the MCP server to point to the built `index.js`.

```json
{
  "mcpServers": {
    "winedbg": {
      "command": "bun",
      "args": ["/path/to/winedbg-mcp/build/index.js"],
      "env": {
        "WINEDBG_MCP_BINARY": "/opt/wine-staging/bin/winedbg",
        "WINEDBG_MCP_READY_TIMEOUT_MS": "60000"
      }
    }
  }
}
```

### Environment variables

All four are optional and read once at startup. There are no secrets and no
config file: the environment is the only place to set these.

That says what this server reads, not what its process holds. The environment an
MCP client launches a server with is the agent's own: API keys, registry tokens
and cloud credentials sit in it, and the program under debug is whoever supplied
the target. `winedbg` is therefore started with an allowlist rather than the
whole environment: `PATH`, `HOME`, `DISPLAY`, the `WINE*`, `XDG_*`, `LC_*`,
`SDL_*`, `MESA_*` and graphics-driver families, and the Windows-path variables a
program under Wine reads. Everything else stays here. Name a variable winedbg
turns out to need in `WINEDBG_MCP_PASSTHROUGH_ENV`. `docs/THREAT_MODEL.md`
records this and the rest of the attack surface.

| Variable | Default | Valid values |
| --- | --- | --- |
| `WINEDBG_MCP_BINARY` | `winedbg` (found on `PATH`) | A non-empty command name or path, with no NUL byte in it |
| `WINEDBG_MCP_READY_TIMEOUT_MS` | `10000` | Whole milliseconds, 1 to 600000. How long `winedbg_start` waits for the first prompt. Raise it for a cold wineprefix, which takes far longer than a warm one |
| `WINEDBG_MCP_COMMAND_TIMEOUT_MS` | `30000` | Whole milliseconds, 1 to 600000. How long `winedbg_execute` waits for a reply when the call names no `timeout` of its own. A `cont` on a busy process is slower than 30s on some hosts, and the per-call `timeout` argument still overrides this one |
| `WINEDBG_MCP_LOG_LEVEL` | `info` | `debug`, `info`, `warn` or `error`. Below this level a line is never written |
| `WINEDBG_MCP_PASSTHROUGH_ENV` | empty | Comma-separated variable names to forward to `winedbg` beside the allowlist, e.g. `COREPACK_ENABLE_STRICT` |

A value the server cannot use stops it at startup with the variable named,
rather than failing later as a spawn error or a start timeout. That includes a
variable set to the empty string, a name that is not a variable name, and a
misspelled `WINEDBG_MCP_*` name, which would otherwise be ignored while the
deployment ran on defaults. The startup line
on stderr reports the values in effect:

```json
{"time":"2026-09-27T10:00:00.000Z","level":"info","message":"winedbg MCP server running on stdio","version":"1.0.0","config":"WINEDBG_MCP_BINARY=winedbg WINEDBG_MCP_READY_TIMEOUT_MS=10000 WINEDBG_MCP_COMMAND_TIMEOUT_MS=30000 WINEDBG_MCP_LOG_LEVEL=info"}
```

### Logging

stdout carries JSON-RPC and nothing else, so every diagnostic goes to stderr as
one JSON object per line: an ISO-8601 `time`, a `level`, a fixed `message` and
flat named fields. A multiline `winedbg` reply therefore cannot break a line
parse, and a log aggregator can filter on a field instead of on a phrase.

The fields an operator pivots on:

| Field | Where | Answers |
| --- | --- | --- |
| `callId` | one tool call, and every session line that call produced | ties the start, the outcome, the duration and the debugger lines of a call together, e.g. `call-7` |
| `tool` | one tool call | which of the three tools ran |
| `durationMs` | tool call result | how long the call took, success or failure |
| `error` | a failure | why: the same text the client got back as the tool result |
| `readyMs`, `lifetimeMs`, `pid` | session start and exit | how long the debugger took to answer, and how it ended |
| `command`, `timeoutMs` | a command that timed out | which command, and the bound it hit |
| `droppedChars` | a reply past the 1M buffer limit | that the reply was shortened, and by how much |
| `kind`, `stack`, `version` | a crash | an uncaught exception or a rejection nobody awaited, with the build it came from |

One tool call is one path through the log: a `winedbg_start` that hangs shows
`tool call started` and the `winedbg spawned, waiting for its first prompt` line
under one `callId`, then either the first prompt with its `readyMs` or the ready
timeout, then the tool call's own outcome with its `durationMs`. A `callId` is
absent from a line that belongs to the process rather than to a call: startup,
shutdown and a signal are the server's own.

A session that ends because the debugger died is logged at `error`; one that
ends because `winedbg_stop`, `SIGTERM` or the client hanging up asked for it is
logged at `info`, so a quiet log holds no shutdown noise. `debug` adds the
command text of each `winedbg_execute`. A crash is recorded before the process
leaves, on the same one-object-per-line surface, so it carries a level, a
timestamp and the version rather than node's plain-text output. There are no
metrics, traces or alerts to configure: the server is one process per MCP client
with nothing to scrape, so the log is the whole surface.

## Tools Available

- **`winedbg_start`**: Start or attach to `winedbg`. Use this before running any commands. Optional `args` are passed to `winedbg` unchanged, so anything it accepts works: the program to launch (e.g. `{"args": ["myapp.exe"]}`) or a PID to attach to (`{"args": ["1234"]}`). `args` must be an array of strings; a bare string is rejected rather than split into one argument per character. At most 64 entries, each at most 4096 characters, none carrying a NUL: an argv entry is cut at the first NUL by the C runtime, and an unbounded array is a caller filling the process table rather than a debugging session. Repeating the call with the same `args` is safe and launches nothing: it answers with the session already at its prompt, whether the first call is still waiting for that prompt or has been there for a while, and says that it launched nothing. A client that re-sends a start it never saw answered, or a model that repeats a call, therefore gets one debugger and one debuggee rather than a second launch of a program that has effects of its own. `args` naming something else is a different request and is refused.
- **`winedbg_execute`**: Execute one command in the active `winedbg` session (e.g., `{"command": "bt"}`).
  Takes an optional `timeout` in milliseconds (minimum 1, maximum 600000), defaulting to `WINEDBG_MCP_COMMAND_TIMEOUT_MS` so a deployment can set it once rather than on every call. The command is at most 4096 characters and carries no line break or NUL, for the reason under command framing below. This is the one tool that is not replay-safe: a repeated call runs the command again, so a `step`, a `cont` or anything that writes debugger state advances the program a second time.
- **`winedbg_stop`**: Stop the active `winedbg` session. A start that follows a stop waits for the stopped debugger to be gone, so alternating the two cannot leave one detached process group, each holding a debuggee, per cycle. A stop returns once the signal is sent, since a tool call should not be held open for a grace period; the wait happens where nothing is waiting on it, in the next `winedbg_start` and in the server's own exit on SIGINT, SIGTERM or the end of stdin. Calling it again, or with nothing running, signals nothing a second time and still reports success.

### Audit log

Every tool call is recorded on stderr, which is the operator's log and not the
reply stream, since the program under debug writes to that stream and could
otherwise forge a record of what it did. A start logs its arguments, a command
that is sent logs the command at `debug`, a command that times out logs the
command and its timeout at `error`, a stop shows up as the session's exit, and a
failure logs the tool and the message.

Each line is a JSON object, so a control character in a command is escaped
rather than able to break the one-line parse, and a recorded field is at most as
long as the argument limits above (4096 characters). Nothing bounds how many
records a session writes, and no reply is logged, so the log is a record of what
was asked for and not of what came back. A debugger line records the tool call
that reached it, so what a session did is read under the call that asked for it
rather than by matching text by hand.

### Command framing

`winedbg` has no way to label which output belongs to which command; the only
marker in the stream is the `Wine-dbg>` prompt. Two consequences are visible
through the tools:

- One command per `winedbg_execute` call. A command carrying a line terminator is
  rejected, because every one of them draws its own prompt and puts every later reply
  one command behind. The rejected set is `\n`, `\r`, vertical tab, form feed, NEL
  (U+0085), the Unicode line and paragraph separators (U+2028, U+2029), and NUL. A
  stream reader splits on `\n`, `\r` and vertical tab, and readers disagree on the
  rest, so a command is one line only if it is one line under every one of them.
  NUL is not a line break, but it truncates the line for most C readers, which
  desynchronises the reply stream the same way a second line would. The check
  runs both at the tool-argument boundary and in the session, so the rule holds
  for a caller that reaches the session directly.
- After a command times out, further commands are refused until the debugger
  prints its prompt again. A debugger that has not returned to its prompt is not
  reading commands, and whatever it prints next belongs to the command that timed
  out. If it never comes back (a `cont` into a program that does not stop), call
  `winedbg_stop` and start again.

A single reply is buffered up to 1M UTF-16 code units of decoded text, so a
BMP character costs one unit and an astral one costs two, whatever the target
prints. The child's output is decoded as UTF-8, and a byte sequence that is not
valid UTF-8 becomes U+FFFD rather than being passed through. A multi-byte
character split across two reads of the pipe decodes as the one character it
is, not as two replacement characters. Past the cap the oldest output is
dropped to keep the last three quarters, and the reply then opens with
`[N characters of earlier output dropped: buffer limit]`, counting code points
rather than the units the cap is measured in. The cut moves back off the low
half of a surrogate pair, so it never lands inside one and no reply starts
with an unpaired surrogate.

## Usage Example

1. Call `winedbg_start` with `{"args": ["myapp.exe"]}`.
2. Call `winedbg_execute` with `{"command": "break main"}`.
3. Call `winedbg_execute` with `{"command": "run"}`.
4. Call `winedbg_execute` with `{"command": "bt"}` to get a backtrace.
5. Call `winedbg_stop` when finished.

The debugger gets a process group of its own, and stopping signals the whole
group, so the program under debug does not outlive the debugger that owns it. A
`winedbg_stop` asks with `SIGTERM` and escalates to `SIGKILL` two seconds later.
When the server itself is exiting there is no time left to ask, so it signals
`SIGKILL` to the group outright rather than leave a debuggee behind.

## Command line

The server takes no positional arguments and no options other than the two
below. Everything else is the environment, and a flag does not exist to
override it, so a client configuration cannot pass a misspelled flag and have
the server start on the defaults anyway.

```
Usage: winedbg-mcp [OPTION]

Options:
  -h, --help       Print this help and exit
  -V, --version    Print the version and exit
```

`--` ends the options. The server takes no operands, so anything after it is
refused the same way an unknown flag is.

| Invocation | Stream | Exit |
| --- | --- | --- |
| `winedbg-mcp` | Serves JSON-RPC on stdin/stdout | 0 on SIGINT, SIGTERM, or end of stdin, after winedbg and the debuggee it started are waited for (up to 6s) |
| `winedbg-mcp --help` | Help on stdout | 0 |
| `winedbg-mcp --version` | The `package.json` version on stdout | 0 |
| `winedbg-mcp --anything-else` | The offending argument, the one-line usage, and a pointer to `--help`, on stderr | 2 |
| `winedbg-mcp` with an unusable environment value | The reason and the variable, on stderr | 1 |

stdout carries protocol traffic and nothing else, so `winedbg-mcp --help | less`
and `winedbg-mcp --version` both behave, and every diagnostic goes to stderr.
The command line is resolved before the environment, so `--help` and
`--version` still work in a deployment whose `WINEDBG_MCP_*` value the server
would otherwise refuse to start on.

## Tests

`bun run check` is the gate for this tree: Biome, then shellcheck over
`scripts/`, then the two type-check passes, then the suite. `bun run typecheck`
and `bun test` are the pieces it runs, for iterating on one of them at a time.
CI runs the build as well, so a
tree that type-checks but does not emit is red there rather than at release.
The suite runs against `src/`; the compiled layout gets its own check, run by
CI after `bun run build`.

```bash
bun run check          # what CI runs
bun run lint           # Biome formatting and lint rules, then shellcheck
bun run format         # Biome autofix: formatting, imports and every safe rule fix
bun run typecheck      # tsc on src/, then on src/ + tests/
bun run build          # tsc, then the executable build/index.js
bun test
scripts/verify-artifact.sh
```

The suite is one `bun test` over `tests/`, so the loop while editing is one file
or one test rather than the lot:

```bash
bun test tests/session.test.ts        # one file
bun test -t "rejects a NUL"           # every test whose name matches, in any file
```

`tests/session.test.ts` drives `WinedbgSession` against `tests/fake-winedbg.js`,
a stand-in that speaks the same `Wine-dbg>` prompt protocol, so the suite runs
without Wine installed. It spawns a real child process and drives its stdio, so
a failure is a real spawn, stream or lifecycle failure rather than a mock
disagreeing, and it waits on real time, which is what covers the parts a
simulator cannot: a pipe the debugger stops reading, a signal with no exit code,
a debuggee that outlives its debugger.

`tests/simulation.test.ts` covers the same state machine with no process and no
host clock. It supplies its own `SessionRuntime` (see `src/runtime.ts`): a
virtual clock the session also measures against, a debugger in memory, and the
process-group probe the wait after a stop makes, answered from the fake's own
state rather than from the host's process table. One seed chooses every reply
delay, chunking pattern, crash and kill outcome, so a run is reproducible and a
failure prints the seed that produced it:

```bash
WINEDBG_MCP_SIM_SEED=1014 bun test tests/simulation.test.ts
```

`tests/concurrency.test.ts` drives `callTool` the way the server does, with
several tool calls in flight at once on one session: starts racing each other,
two commands racing for the single command slot, a stop racing an in-flight
command, and a burst where every answer names the command that asked for it.
The server hands tool calls in from the event loop without serializing them, so
that is the case worth pinning.

`tests/validate.test.ts` covers the tool-argument boundary and
`tests/config.test.ts` the environment parsing described above.
`tests/cli.test.ts` spawns the entry point to pin the exit codes and which
stream each message lands on, and `tests/server.test.ts` spawns it to speak
JSON-RPC over its stdio: the handshake, the tool list on the wire, a
start/execute/stop round trip, both kinds of refused call, the `callId` each
outcome is logged under, and the exit 0 a client hanging up gets. Between them
the two files reach `src/index.ts`, the one module no unit test can cover.

CI runs the same `bun run check` and `bun run build` on every push and pull
request, so a green local run is a green remote run. `bun run typecheck` covers
`src/` on its own, which is the set that ships as `build/`, and then re-checks it
together with `tests/` so a mistyped test helper fails the build rather than the
suite.

Formatting is Biome's, and the line width is 120 columns, the width the tree was
already written to. `src/index.ts` keeps one scoped `noConsole` suppression:
stdout carries the MCP JSON-RPC stream, so the configuration failure has to go to
stderr. Everything else it writes goes through the logger, which writes to
stderr by construction (`src/logger.ts:45-48`).

## Troubleshooting

Every tool failure comes back as `Error: <message>` with the message the code
produced, so the text names the state to fix:

| Message | Cause |
| --- | --- |
| `winedbg is not running. Please start it first.` | A command ran with no live session, including one whose `winedbg` exited |
| `winedbg is already running. Please stop it first.` | `winedbg_start` was called with arguments other than the running session's. Repeat it with the same `args` to get the running session instead, or call `winedbg_stop` first |
| `Another command is already in progress: ...` | One command per call, and the previous one has not answered yet. The message names that command and tells the caller to wait for its reply |
| `Configuration error: ...` on stderr at startup | An environment value the server cannot use, named in the message. The server exits with status 1 instead of starting on defaults |
| `Timeout waiting for <binary> to print its first prompt (Nms)` | No prompt within `WINEDBG_MCP_READY_TIMEOUT_MS`; the child is killed. Raise the variable for a cold wineprefix |
| `winedbg stopped before it was ready`, `winedbg was not started: a newer winedbg_start call replaced this one` | A start that was waiting for the previous debugger to be gone was cancelled, by `winedbg_stop` or by a newer start. Start again |
| `command must be a single line: ...` | The command carried a line break or a NUL. One command per call, one line per command |
| `args[1] contains a NUL byte, which no argument can carry` | An argv entry carried a NUL, which the C runtime cuts at. It names the index so the offending argument can be found |
| `args must have at most 64 entries`, `args entries must be at most 4096 characters`, `command must be at most 4096 characters` | An argument past its bound. Shorten it, or split the work across calls |

## License

ISC, declared in `package.json`. This tree ships no `LICENSE` file, so the full
terms are stated here only.