Skip to main content
Glama

Tuning Fork

A step debugger and experiment lab for agent skills running on local LLMs.

Status: pre-alpha, 0.1.0. Recording, divergence localization, resampling from a checkpoint, wording A/B tests, live breakpoints, workspace snapshots, restarting a recorded run from any step to measure it end to end, a web UI that starts all of it, token-level analysis — how sure the model was of each token, and which token decided a step — all of it as tools for an agent over MCP — and refining a skill at a step, with edits proposed by you, an agent or a model, each screened against the original, validated with full runs on held-out tasks, and written to the skill only when a person approves, all work, with Hermes Agent and with Claude Code.

New here? examples/README.md walks through everything below, command by command: first a five-minute tour that needs no GPU, then the real experiment. examples/hermes-design-md/ does it on a real skill, Hermes Agent's own design-md, from the failing step to an approved fix and the runs after.

The problem

You write a skill (an agentskills.io-style SKILL.md) for a local agent. It works three times out of five. So you reword it, run the task five more times, and squint at the results.

That loop is broken in two ways:

  • Small samples are noise. 3/5 is consistent with a true success rate anywhere from 23% to 88% (95% Wilson interval). Separating a 60% skill from a 90% one takes roughly 30+ runs per variant. Almost nobody does that, so wording decisions get made on noise.

  • Wording sensitivity is real and model-specific. Open models swing hard on meaning-preserving format changes, and the best format does not transfer between models (Sclar et al., ICLR 2024). Every model swap or quantization change means partial retuning.

Related MCP server: vibe-debug

The idea

Treat an agent run as a program execution and the skill as its source code, then supply the missing debugger:

  • Record every model call at the proxy, with the exact context the model saw.

  • Break and inspect — pause before or after any call, read and edit the full context.

  • Fork and resample — re-run step k N times from a checkpoint and cluster the outcomes.

  • Localize — find the first step where successful and failed runs diverge.

  • A/B the wording there, with history held fixed, and report the difference with a confidence interval and a significance test instead of a bare percentage.

Pausing and editing LLM traffic is table stakes; several tools do it. The point here is measurement and localization for instruction text.

How it works

tuningfork is a recording proxy that sits at the model-server seam. It speaks both OpenAI chat-completions and Anthropic Messages:

agent harness ──(OpenAI or Anthropic HTTP)──▶ tuningfork ──▶ inference server
                                                  │
                                        SQLite + content-addressed blobs
                                                  │
                              Python API / CLI / web UI / MCP server

Model requests are stateless, so every request carries the model's complete context. The proxy sees the full program state at every step without touching harness code, and works with any harness that accepts a custom base URL.

Start llama-server with --swa-full

Strongly recommended for any model with sliding-window attention layers — Gemma 2, 3 and 4 among them. Without the flag, llama-server keeps only part of those layers' cache and reuses saved states between requests. The same request can then behave differently depending on what the server processed before it.

We measured this on Gemma 4 26B-A4B. One step's request gave its target behaviour 1 time in 40 right after a restart, and 22 times in 40 once a similar prompt had run. A screen samples each proposal right after the original, so on such a server it favours the proposals. With --swa-full that swing was gone in our repeated tests.

The flag costs GPU memory. The 26B with two 64k slots and a q4_0 cache needed 27 GiB with it and 22 GiB without. The dense 31B at a 64k context did not fit a 32 GB card with it. llama-server does not report the setting, so tuningfork cannot check it for you.

Even with the flag, a decision that hangs on a near-tie can still change with how the server happened to process the prompt. So tuningfork asks llama.cpp to process each experiment sample's prompt whole (cache_prompt: false), where it can: on chat completions, not on its Messages endpoint, which Claude Code uses.

And with --no-cache-prompt

Strongly recommended too. By default llama-server answers a prompt from what it cached of an earlier one, and its Messages endpoint, which Claude Code uses, ignores a request asking otherwise. Then a Claude Code step's samples all share the first one's processing of the prompt, and forks whose prompts match — every fork, in the sandbox below — share one cached state. We measured one resample give a behaviour 16 times in 20; with the flag, three resamples of the same step gave it 1 time in 10 each.

The flag makes every step process its whole prompt: about 1.5 s per 14,000 tokens of Gemma 4 26B-A4B on one RTX 5090, so runs take a fifth to a third longer. tuningfork connections check measures whether the server reuses prompts, and a batch's plan warns when it does.

Install bubblewrap (Linux)

Recommended wherever tuningfork launches a harness. With bubblewrap (sudo apt install bubblewrap), every run, fork and held-out run is launched in a sandbox that looks the same each time: its home is /home/user, it works in /home/user/project, and the harness keeps its files where it installs them. The model reads those paths. Without the sandbox, each launch sees its own folders under the recordings folder, and a few digits of a fork's folder name were seen to decide whether the model wrote a file or only said it did: resampled at the same step, two forks whose requests differed only in that number wrote it 5 times in 20 and 20 in 20.

Inside, the rest of the machine is read-only, /tmp and the process list are the launch's own, and the recordings folder and the task's spec folder are empty: the harness can read neither earlier runs nor the check. tuningfork connections show says whether launches run in the sandbox, and a batch's plan warns when they cannot. TUNINGFORK_SANDBOX=off turns it off.

It keeps runs alike; it is not a security boundary. Inside, the rest of the machine stays readable, your home included, and the network is reachable (SECURITY.md).

What works today

A skill that never passes. Where do runs of it stop agreeing?

$ tuningfork batch examples/tally-report.evals.json -n 8 --model <model-id>
4 runs of tally-report/1 — agreed for 0 step(s), diverged at step 0
A  3/4  75%  [30%-95%]  read_file(path='data.csv')
B  1/4  25%  [5%-70%]   execute_code(code='<ast:a62cb33a>')

Resample the step where the skill's instructions bite — twenty times, with the history held fixed — once as written and once with an edit swapped in:

$ tuningfork ab <run> 1 --variant examples/variants/tally-report/SKILL.v2.md --target-match row_count
   behaviour                                        original              variant
A  write_file(content='{"amount_total":61.5,…       0/20   0% [0%-16%]    20/20 100% [84%-100%]
B  write_file(content='{"count":4,"total":61.5}'…  20/20 100% [84%-100%]   0/20   0% [0%-16%]

target (declared before sampling): behaviour matching /row_count/
  Fisher exact p = 1.5e-11 → significant at α=0.05, variant higher

Four minutes of sampling, one model call per sample. Full runs agree — tuningfork batch ... --skill examples/variants/tally-report/SKILL.v2.md puts the edit in every run, and it passes 8/8, against 0/5 before. tuningfork tasks keeps the two apart: a task run under two wordings has a rate for each, never one pooled across them.

The same edit tested at step 0, where the model decides how to read the file, came back not significant (p = 0.748) — correctly, since the edit is about output keys. A checkpoint A/B tells you which decision an edit changed, not just whether the pass rate moved.

From that step to the finish line

A checkpoint A/B measures one step. --to-end makes each sample a fork: the harness is relaunched on the same task in a fresh workspace, answered from the recording up to the step — it re-runs every recorded tool call for real, but the model is not asked — and left to run from there to the end, where the task's check decides:

$ tuningfork ab <run> Write --variant examples/variants/tally-report/SKILL.v2.md --to-end -n 20
target: the task's check passes, end to end from step 2
  original  0/20   0% [0%-16%]
  variant   20/20 100% [84%-100%]
  Fisher exact p = 1.5e-11 → significant at α=0.05, variant higher

Forty forks through Claude Code in eleven minutes, each answered by the recording exactly as it was up to the step — a fork whose harness asks for something the recording never answered is left out, and says where. tuningfork fork <run> <step> runs a single fork with breakpoints armed: a restart from any step of a recorded run. And every call of a run tuningfork launched finds its workspace snapshotted: tuningfork files <run> <step> shows it as that step's call found it.

Which token decided it

A step's answer can be read as the model's own tokens. tuningfork tokens <run> <k> --score sends the step's exact context once more with the answer forced, one call, and reads how sure the model was of each token: a heatmap in the terminal and on the run's page. At the tally-report write step it says count was never in doubt — p = 1.00 — so no amount of resampling would find row_count: the wording has to change.

Where a step does split, tuningfork forking-tokens finds the token that splits it. After Bigelow et al. (ICLR 2025): at the tokens whose runner-up the step's own sampler would most likely have drawn, force each likely alternative, let the model go on, and test which change what it does. A Hermes run that searched for data.csv before reading it:

$ tuningfork forking-tokens <run> 0 --in reasoning
[134] reasoning  …, I'll check if `data.csv` exists⟦ in⟧
  as run  ' in' 19%                10/10 100% [72%-100%]
  forced  ' and' 72%               0/10   0% [0%-28%]   p=1.1e-05 → Holm 7.6e-05  flips away from the target
  forced  '.' 9%                   10/10 100% [72%-100%]   p=1.000 → Holm 1.000

1 of 7 forced tokens changed what the model did (Holm-corrected over the 7 planned before sampling; α=0.05).

"Exists in the working directory" leads to a search; "exists and…", which the model thought likelier, to reading the file. Every alternative is compared with the run's own token at the same position, and the family of tests is fixed before sampling. This needs llama.cpp's server and a model whose raw output tuningfork reads (Gemma 4 today); anything else is refused with the reason.

Stopping a run where it matters

Breakpoints stop a live call before the model sees its context, or after the model has decided and before the harness acts. Here the flaky step is caught as it happens:

$ tuningfork break add --tool write_file --path '*.json'
breakpoint 181d8b0f — after the model: write_file on *.json

$ tuningfork hold wait
hold 7f8094a1 — after the model, run 33f256ea step 1, held 0.4s
the model answered
  tool_call write_file({"content":"{\n  \"count\": 4,\n  \"total\": 61.50\n}\n", ...})

The harness is simply waiting on a slow response. Fix the keys and let it carry on:

$ tuningfork hold show 7f8094a1 --json > held.json     # edit the arguments in held.json
$ tuningfork hold continue 7f8094a1 --edited held.json

The harness writes the corrected file and the run passes its check. The run is still not counted in any pass rate, because its verdict now measures the edit as much as the skill. The recorded step keeps what the model actually produced, alongside what the harness received. hold step walks a run one stop at a time instead, and hold drop fails a call.

All of it from the browser

Everything above is also in a local web UI, served by the proxy itself at /_tuningfork/ui/: a run tree with pass rates and their intervals; each run as one conversation, top to bottom, with each step's answer — reasoning included — where it happened, and the skill marked inside it and mapped to its file line by line; each answer as a heatmap of its tokens, a click showing what else the model weighed there; a step's exact context on a page of its own; divergence and diff views, experiment results, and a Live page that edits held calls in place. A tool call the model made without any reasoning — in a run that reasons elsewhere, often right after a tool call failed — is marked wherever a step is shown, and an experiment says how many of a behaviour's samples reasoned where some did not; break add --no-reasoning stops there live.

The front page says what is set up and where you left off, draws how the pieces fit — one loop, from a skill through its runs, the experiments at a step and a refinement, to the skill's next version — and lists the questions each way of looking at a step answers, with the clicks and the terminal command for each, beside a glossary of tuningfork's words (tuningfork glossary in a terminal). The Skills page lists every skill the recordings know: each version its runs were given, with its pass rates, what its text is known as (the file now, a saved variant, a proposal, a version an approval replaced) and its SKILL.md rendered; any two of its texts compared line by line; and the changelog an approval keeps beside the skill's folder (tuningfork skills, tuningfork skill <name>).

And the page can start the work, not only show it. From a step on a run's conversation, or from Start… on the Jobs page: resample it, or A/B a wording edited right there against the skill as the run was given it, at that step or with every sample run to the end; fork the run there and stop it at the step; score its answer, or select tokens on its heatmap and find which of them decide it — saying what counts as the same behaviour in words (a tool it calls, a phrase it mentions), set against everything the step has already done before anything is sampled. A task's check can be a shell command run in the workspace — a longer one a script beside the spec, which $TUNINGFORK_SPEC_DIR names, out of the agent's sight — or a check of the agent's final reply, which needs none — and it can be tried before it is trusted: on a recorded run of the task, whose verdict it is set beside, or on a reply and files written for the purpose, with nothing recorded. A form shows the plan the proxy would run before its Start button does anything, and an A/B asks what it is about — a behaviour, a pattern, or openly exploratory — before a sample is drawn. Scenarios keep tasks in the recordings folder, edited as the exact evals.json they save to, with their input files, and run as batches. A job's page follows it as it goes; a call held at a breakpoint can be sampled where it stands (evidence, filed under the step it becomes) or asked about (the model's own account, labelled a hypothesis). The Connections page shows the server and what was measured of it, the harnesses it found and their versions, and switches the server while nothing is using it.

The page changes nothing without the proxy's access token, which tuningfork serve prints in the address to open; tuningfork token prints it again. tuningfork ui serves the same pages, read-only, over existing recordings with no proxy running. Runs, experiments, jobs, scenarios and refinements can be moved to a trash — out of every list and pass rate until restored — and deleted for good from there, with everything tied to them: a refinement with its screens, its validations' runs and its proposals' saved texts, never the changelog an approval wrote.

Refining a skill at a step

Once the failing step is found, tuningfork can gather fixes and test each one there. A refinement names the step and what the model should do at it, before anything is proposed:

$ tuningfork refine start 586cf615 write_file --target-rules \
    '[{"if":"calls","value":"write_file"},{"if":"contains","value":"row_count"},{"if":"contains","value":"amount_total"}]'

Edits come from you (--file, your own wording), from an agent over MCP, or from a built-in proposer: a model — the local one by default, through the proxy and never recorded; a cloud API only once you allow it — shown the evidence as data. That is the step as the model saw it and where the skill sits in it, the answers seen there that did and did not do what the target says (here none did, and it says so), how the task's runs with this harness, model and skill version went there, the task's own criteria, and what was tried before. A proposal is at most three exact edits, kept with who proposed it, why, and what they were shown:

$ tuningfork refine propose e811ff0d --builtin
...
  answered in 75s, 3332 tokens · kept in proposer/20261002T204545-e4aee8
  proposal f2b5c51b: 1 line changed
    why  The agent currently uses 'count' and 'total' as keys in the JSON report, likely because
the skill's final line mentions 'just the count and the total'. ...
  proposal fa63747b: 1 line changed
...
  proposal 95ebcd3c: 1 line changed
...

Each is a hypothesis until a screen tests it: a fresh arm of the original wording first, then one arm per proposal, each compared with the original by Fisher's exact test, Holm-corrected over the proposals named before the first sample. Before it samples, a screen says what an arm would need to survive, and it refuses to run when nothing could:

$ tuningfork refine screen e811ff0d -n 20
  survives  with at least 7 of 20 — the step did it 0 times in 20 in resample 118becf9
            Holm over 3 proposals, α=0.05
...
          target                 p        Holm
original  0/20   0% [0%-16%]
f2b5c51b  20/20 100% [84%-100%]  1.5e-11  4.4e-11  survives
fa63747b  20/20 100% [84%-100%]  1.5e-11  4.4e-11  survives
95ebcd3c  20/20 100% [84%-100%]  1.5e-11  4.4e-11  survives

The web UI does the same from a step's Refine the skill here…, and an agent with tuningfork mcp --tools refine. A screen at one step is evidence, not proof: its arms inherit the history the original wording wrote, and a wording can reach the target without fixing the task. In a later run, an edit that told the model to skip the file and write the training task's answers survived its screen exactly as the fix did.

So a survivor is validated with full runs, each arm against a fresh arm of the original: forks from the step to the end, runs of the task with the edit from the first step, and runs of each task the spec holds out ("held_out": true on an eval). A held-out task's runs are recorded in a temporary folder outside the recordings folder and deleted when they end. Only their verdicts are kept, and no agent is shown those:

$ tuningfork refine validate 044cbd20 10c1b095
...
  1568ae72  done  proposal 10c1b095 — meets the gate
    forward   edit 5/5 (100%, 95% CI 57–100%); original 0/5 (0%, 95% CI 0–43%)
    training  edit 5/5 (100%, 95% CI 57–100%); original 0/5 (0%, 95% CI 0–43%)
    held out 2: edit 5/5 (100%, 95% CI 57–100%); original 0/5 (0%, 95% CI 0–43%)
    G0 met     survived a screen: it survived screen b28f6a16
    G1 met     a held-out task was validated: 1 held-out task(s) validated
    G2 met     finished, with enough runs: every arm ran its runs on one server
    G3 met     its own task succeeds from the target forward: with the edit 5/5 forks passed, without it 0/5
    G4 met     the training task's full runs are not worse (low-powered): significantly better (Fisher p = 0.00794)
    G5 met     held-out task 2 is not worse (low-powered): significantly better (Fisher p = 0.00794)
    G6 met     the skill's file is still the one refined: it holds the text the refinement started from (line endings aside)
    G7 met     the refinement is open, and nothing in it is approved: open
    G8 met     held-out task 2 improves where the original fails: with the edit 5/5, without it 0/5

Each check keeps its counts and the name of the rule that judged it; at five runs an arm, "not worse" says it can see only a large drop. The cheat passed every other check: its held-out task failed with and without it, so it was no worse. G8 refuses it: a held-out task the original fails must pass more often with the edit, and an edit that writes in the training task's answers leaves it failing.

Approving is the person's. It shows the file it will write, as WSL and Windows name it, what it keeps, the gate and the edit, and writes nothing until the code it shows is typed:

$ tuningfork refine approve 044cbd20 10c1b095
approve proposal 10c1b095 of refinement 044cbd20, on validation 1568ae72
  writes     /mnt/c/Users/me/tf/tally-report/SKILL.md
             C:\Users\me\tf\tally-report\SKILL.md
  keeps      CRLF line endings
  file before → /mnt/c/Users/me/tf/tally-report.history/
  changelog   → /mnt/c/Users/me/tf/tally-report.changelog.md
  the gate   every check met:
...
  code  6da89c69
Type the code to write it (anything else writes nothing): 6da89c69
written: /mnt/c/Users/me/tf/tally-report/SKILL.md
  the file before  /mnt/c/Users/me/tf/tally-report.history/SKILL.20261005T021037Z-c52cbb98.md
  changelog        /mnt/c/Users/me/tf/tally-report.changelog.md

The file keeps its line endings and byte-order mark. The old file and a changelog entry go beside the skill folder, never into it, since every run copies that folder. The refinement keeps the approval with every piece of evidence it stood on, whatever is deleted later. No MCP tool approves, and an MCP client's token is refused where approving happens.

For an agent: the same tools over MCP

tuningfork mcp serves tuningfork's operations to an agent — Claude Code, Hermes, an OpenJarvis teacher — over the Model Context Protocol. An agent can do what a person does: find where a task's runs split, read that step's context and the skill lines in it, measure how often each behaviour happens there, and A/B-test a wording it wrote.

$ claude mcp add --transport stdio tuningfork -- tuningfork mcp --home ~/recordings

Reading recordings needs no proxy. Work that calls the model — a resample, an A/B, scoring an answer, a forking-token probe — runs as a job inside tuningfork serve on the same recordings, through its gate, and shows on its Jobs page as "from MCP · ". Each answer is compact JSON opened by sentences that say it statistics first: every rate with its 95% interval, an A/B's declared target with Fisher's p and its threshold, or that the comparison was exploratory. Text from the recordings is data to the agent, never instructions, and the server's instructions say so.

By default an agent is offered only work that calls the model; --tools all adds work that runs a harness or a task's check on this machine. --tools refine offers the refinement loop: start a refinement, read its evidence, propose edits of its own, and screen them. No agent is offered the built-in proposer — it sends recordings to an endpoint, which is the person's decision — and nothing offered approves an edit. A fork an agent starts never stops at a breakpoint, and an agent's check never runs a shell command. A job may make at most 100 checkpoint samples or 20 harness runs (--limit resample=200 changes one), agents may have 4 jobs queued or running at once, the same work asked for twice is given back rather than started again, and an agent cancels only its own jobs. Live debugging, the trash and switching the server stay with the person.

A client that connects rather than starts a server uses the running proxy's endpoint, http://127.0.0.1:8787/_tuningfork/mcp, with the proxy's MCP token as a bearer token: tuningfork token --mcp prints it, TUNINGFORK_MCP_TOKEN pins it across restarts, and tuningfork token --header-json prints the header for a client that runs a command for its headers, such as Claude Code's headersHelper. The MCP token opens that endpoint and nothing else, so an agent's client never holds a key to what its tools leave out, such as approving an edit; the access token, in turn, is refused there. Hermes lists the server under mcp_servers in its config.yaml (its mcp extra installed); OpenJarvis under [mcp] servers in its config.toml, by command or by URL and token. From Windows with tuningfork in WSL, the command is wsl.exe -d <distro> -e <venv>/bin/tuningfork mcp --home <recordings>: the server reads the token from the recordings folder itself, so nothing secret goes in a Windows config. Point the agent's own model at the inference server, not at the proxy it studies, or its calls are recorded beside the runs it is reading.

Recording is transparent: point any harness that speaks OpenAI chat-completions or Anthropic Messages at the proxy and it behaves as it did before. If a model server on your machine serves something that matters, protect it, and tuningfork sends it nothing — not a resample, not a probe, not a harness's call: list it in ~/.config/tuningfork/config.json ({"protected": ["localhost:1234"]}), in TUNINGFORK_PROTECTED, or in a recordings folder's config.json. Every message is content-addressed, so the large constant part of an agent's context is stored once rather than once per step. Rates are never reported without the interval they actually imply, and an A/B's target is declared before sampling — without one, the comparison is labelled exploratory and corrected for multiple comparisons.

Portability

Harness and server specifics live behind two seams — harness/ adapters and backends/ capability profiles. Optional features (logprobs, server-side n, KV-cache reuse) are enabled by a runtime capability probe, not by assuming a particular server. Running tuningfork probe against your own stack produces a profile you can contribute back.

First-class targets today: Hermes Agent and Claude Code as harnesses, and LM Studio / llama.cpp as servers. Claude Code runs against a local model with no router, because llama.cpp serves Anthropic's Messages API itself. These are the first implementations, not the assumptions.

Development

Requires Python 3.12+ and uv. The full test suite runs against a scriptable mock server, with no GPU and no model:

uv run pytest

The web UI is plain HTML, CSS and JavaScript modules with no build step. Its behaviour in a real browser is covered by an opt-in test in tests/browser/. What the token-level features stand on is measured against a real llama.cpp server by another opt-in set, skipped unless you name one: TUNINGFORK_REAL_SERVER=http://127.0.0.1:8081 uv run pytest -m real_server.

Tested on Linux — Ubuntu, natively and under WSL2 on Windows — with Python 3.12, 3.13 and 3.14. On Windows, run it inside WSL, with the checkout and the recordings on the Linux file system rather than under /mnt/c, where SQLite's locking does not work. macOS is untested, and the sandbox needs Linux. How to contribute: CONTRIBUTING.md; how to report a vulnerability, and what tuningfork does and does not guard: SECURITY.md.

License

Apache-2.0: see LICENSE, and NOTICE for the work it follows.

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    D
    maintenance
    Enables AI assistants to perform interactive Python debugging with breakpoints, step execution, and variable inspection using the Debug Adapter Protocol (DAP) through an MCP server interface.
    8
    1
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables coding agents to use a real debugger (Python via debugpy) for launching, attaching, setting breakpoints, stepping through code, inspecting stack frames, and evaluating expressions through MCP tools.
    27
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables AI agents to inspect debug state, control execution, and set breakpoints in VS Code by exposing the Debug Adapter Protocol as an MCP server.
    Apache 2.0
  • A
    license
    A
    quality
    C
    maintenance
    Enables MCP-compatible coding agents to debug applications using runtime log data, by providing tools to start a local log-ingestion server, track debugging sessions and hypotheses, and correlate logs to specific executions.
    19
    5 npm
    MIT