Skip to main content
Glama
README.md
# log-probe-mcp

An agent-agnostic MCP server for hypothesis-driven, runtime-data-backed debugging.

Point any MCP-compatible coding agent (Claude Code, Cursor, etc.) at it and say something like:

> Debug this using the log-probe mcp — checkout returns the wrong cart for some users.

log-probe-mcp never edits your code. Instead it runs a local HTTP log-ingestion server, correlates
incoming logs to specific executions (a script run, a test invocation, a request), and tracks
hypotheses — so the calling agent can drive the classic debugging loop (form hypotheses →
instrument → reproduce → analyze → converge) with **real runtime data** instead of guessing from
static code, and without a human manually copy-pasting console output back into the chat.

## How it works

1. The agent calls `probe_server_start` to spin up a local HTTP server that accepts structured log
   events (`POST /ingest`).
2. It records a debugging session and a few falsifiable hypotheses (`debug_session_create`,
   `hypothesis_create`).
3. It fetches the logging contract (`instrumentation_get_contract`) and inserts a few log calls at
   the decision points that would distinguish between the hypotheses, using its own file-editing
   tools — log-probe-mcp only tells it *what* to send and *where* to send it, it never touches your
   source files itself.
4. It reproduces the bug, either by running a script/test itself (`execution_run`, which also
   captures stdout/stderr automatically) or by minting an execution for something already running
   (`execution_create`) and asking a human to trigger it.
5. It reads the real data back (`execution_get_logs`, `execution_compare` for flaky/intermittent
   bugs) and marks each hypothesis confirmed/refuted with evidence (`hypothesis_update`).
6. Once resolved, it applies the actual fix itself, removes the temporary instrumentation, and
   optionally writes a durable record (`knowledge_base_export`, `debug_session_resolve`).

Call `debug_workflow_guide` (or use the `debug` MCP prompt, on clients that support prompts) for
the full step-by-step guidance an agent needs to run this loop well.

## Installation / client config

```json
{
  "mcpServers": {
    "log-probe": {
      "command": "npx",
      "args": ["-y", "log-probe-mcp"]
    }
  }
}
```

For local development against a checkout of this repo, build it and point a client directly at
`dist/bin.js`:

```json
{
  "mcpServers": {
    "log-probe": {
      "command": "node",
      "args": ["/absolute/path/to/log-probe-mcp/dist/bin.js"]
    }
  }
}
```

Data (the SQLite store and any exported knowledge-base files) lives under `.log-probe/` in the
project the agent is working in. Because MCP clients don't consistently launch servers with `cwd`
set to the project root, the authoritative source is, in priority order: the `dataDir` argument to
`probe_server_start`, the `LOG_PROBE_DATA_DIR` environment variable, then the server process's own
`cwd`.

## Tool surface

| Tool | Purpose |
|---|---|
| `probe_server_start` / `probe_server_stop` / `probe_server_status` | Ingestion server lifecycle |
| `debug_session_create` / `_list` / `_get` / `_resolve` | Track a debugging investigation |
| `hypothesis_create` / `_update` / `_list` | Track falsifiable hypotheses and their evidence |
| `execution_create` / `_run` / `_end` / `_list` / `_get_logs` / `_compare` | Mint/run/query correlated executions |
| `instrumentation_get_contract` | The ingestion HTTP contract + ready-to-paste snippets per language |
| `debug_workflow_guide` | The hypothesis-driven methodology, full guide or per-stage |
| `knowledge_base_export` | Writes a durable markdown record of a session |

Plus a `debug` MCP prompt for clients that support the prompts primitive — a thin wrapper around
`debug_workflow_guide`'s content, so guidance is reachable via tools everywhere regardless of
prompt support.

## Ingestion contract

Instrumented code sends a `POST` to `<ingestion url>/ingest` with a JSON body (single event, or an
array of up to 500 for batching):

```json
{
  "executionId": "exec_...",
  "hypothesisId": "hyp_...",
  "level": "info",
  "message": "cache key computed",
  "data": { "key": "route:/x" },
  "source": "checkout.ts:88"
}
```

`executionId` must already exist (minted via `execution_create` or `execution_run`) — this is what
correlation is built on. Instrumentation should always be fire-and-forget with a short timeout; see
`instrumentation_get_contract` for language-specific snippets that already do this correctly.

## Known limitations

- The ingestion server binds `127.0.0.1` only and has no auth token — acceptable for a local dev
  tool, but don't run it anywhere multi-tenant or expose the port.
- One MCP server process serves one data directory for its lifetime; to point at a different
  project, restart/reconnect the client rather than changing `dataDir` mid-session.

## Example

`examples/buggy-node-service/` is a small, intentionally-buggy HTTP server for trying the full
workflow end to end — see its README.

## Development

```
npm install
npm run build       # compiles to dist/ and copies the SQL migration
npm run dev          # tsx watch, for iterating
npm run typecheck
npm test
npm run inspect       # build + launch the MCP inspector against dist/bin.js
```

TDQS

A3.9/5.0

Scored across 19 tools

Disambiguation4/5

Most tools have clearly distinct purposes across server lifecycle, sessions, executions, hypotheses, and knowledge base. The only minor overlap is between debug_session_get (which includes execution summaries) and execution_list, and execution_create vs execution_run require careful reading to distinguish, but descriptions resolve this.

Naming Consistency5/5

All tool names follow a consistent snake_case convention with a resource_action pattern (e.g., probe_server_start, debug_session_create, execution_run, hypothesis_update). The convention is uniformly applied across all 19 tools, even if the order is noun-verb rather than the more common verb-noun.

Tool Count4/5

19 tools is on the heavier end of the typical range, but each tool maps to a distinct phase of the debugging workflow: server management, session lifecycle, execution management, hypothesis management, instrumentation contract, guide, and export. The count feels justified by the comprehensive scope.

Completeness5/5

The toolset covers the full debug lifecycle without obvious gaps: start/stop/status the server, create/list/get/resolve sessions, create/run/end/list/log/compare executions, manage hypotheses, get instrumentation contract, read the guide, and export the knowledge base. No essential operation appears missing.

Maintenance

ActivityStale
ResponsivenessNo issues