cold-run
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@cold-runaudit my agent skills and tell me which ones actually help"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
cold-run
I tested 47 agent skills. I deleted 23 of them. Three were making Claude's code worse.
Then I built the tool that does it, pointed it at the 24 survivors, and it graded three of the first five NEGATIVE.
That is not a bug. It is the most interesting thing I found, and I am publishing it rather than waiting for you to find it: runs/bare-baseline/.
npx cold-runThe finding
A skill that does nothing is the boring failure. The common one is worse:
A skill spends the agent's attention. It buys a specific check, and it pays for that check out of everything else the task needed.
Asked to build a cached lookup, the agent following amplification-check added jitter and
single-flight — exactly what the skill asks for. The agent that had never seen the skill added
those and an LRU cap, stale-while-revalidate, and negative caching. The skilled version's cache
grows without bound.
Asked to add retries inside a Lambda, the skilled agent produced a tidy retry-budget table and a
blanket except that retries a definitive 404 as eagerly as a timeout.
You cannot see this by reading a skill. Both of those skills are well written. You only see it by running the task twice and diffing the artifacts.
What it does
Finds your skills, picks an adversarial test task for each one, runs a cold/skilled pair in parallel, and tells you which of your skills change nothing.
cold-run v0.1 — 30 skills in ~/.claude/skills
REAL retry-guard added a breaker + cap; cold declined one
NONE be-thorough cold output was identical
NONE check-edge-cases cold found two more edge cases
NEGATIVE refactor-first skilled dropped a null check cold kept
...
11 change nothing.
2 make the output worse.
17 earn their place.
Deleting the 13 frees ~7,400 tokens of always-loaded
context, every session.
full report -> cold-run-report.md
every transcript -> .cold-run/It never deletes anything. It writes the report and every transcript, and the decision is yours.
Cost is real and it asks first. Four model calls per skill. On a library over 12 it tests a
spread of five so your first run is cheap; --all does the rest.
npx cold-run # auto-finds your skills
npx cold-run path/to/skills # or point it somewhere
npx cold-run --all --yes # everything, no promptRuns through your local claude CLI if you have one, otherwise ANTHROPIC_API_KEY.
Zero dependencies — npx starts instantly.
As an MCP server
Same audit, driven by whatever agent you already talk to — Claude Code, Claude Desktop, Cursor. It uses your agent's own subagents, so there's no second API key and no hidden spend.
claude mcp add cold-run -- npx -y --package=cold-run cold-run-mcp{
"mcpServers": {
"cold-run": {
"command": "npx",
"args": ["-y", "--package=cold-run", "cold-run-mcp"]
}
}
}Then just ask: "which of my skills actually do anything?"
The five verdicts
REAL | The artifact differs materially and the skilled run is better. |
MARGINAL | A small real addition on top of an output the cold run already got right. |
NONE | Indistinguishable — or the skilled run only talked about the principle. Talking is not changing. |
NEGATIVE | The cold output was better. This is not rare. Three of my 47 landed here. |
INVALID | One side never produced an artifact, so the pair is void. Reported, never counted. A verdict from a run that did not happen is worse than no verdict. |
The 24 that survived, and the asterisk on them
The library that produced this tool is in skills/. Every one has a recorded cold/skilled pair behind it — see EVIDENCE.md and runs/ledger.md.
Read this before you install them. That audit ran both agents under my machine's global
operating contracts — house rules that force stop-and-ask behaviour and a fixed reporting format.
cold-run deliberately strips all of that. Under the bare baseline, three of the first five
survivors grade NEGATIVE. Both results are published, in full, with transcripts:
runs/bare-baseline/.
Which number applies to you depends on what your own CLAUDE.md already forces. Neither has been
run at full scale under the bare baseline. I would rather ship the contradiction than hide it.
The 23 that didn't make it, and why none could be repaired, are in runs/dropped.md.
A few of the survivors, so you can judge the bar:
yagni-audit— "implement a rate limiter." Cold wrote ~700 lines: sync, async, keyed, Redis, Lua, decorators. Skilled wrote 15.trust-source-check— cold authenticated services with a shared static env-var token. Skilled required a JWKS-verified OIDC token with audience pinning.reverse-path-proof— skilled ran a real Postgres container and found the rollback aborts on a 255-character name, wiping every name in the table. Cold had the same bug.pit-of-success— cold shipped immediate hard delete as the default. Skilled made it reversible for 30 days and put purge behind a confirmation echoing the user id.
Install them the normal way, or don't. The tool is the point.
Skills are graded on two gates here, not one
Cold-run delta — does it change the artifact? 23 skills failed this. → EVIDENCE.md
Routing — does the description actually get the skill loaded? A perfect skill with a description that never fires is dead weight that still costs tokens. Tested as a classifier against 360 written requests; two rounds failed before all 24 passed. → tests/router/
A skill passing one and failing the other is still broken.
vs. skill-doctor
skill-doctor is a linter. It reads your skills and checks frontmatter validity, broken resource links, and whether a trigger description looks weak, then scores 0-100. It never runs anything.
cold-run is the other half. It executes the skill and an agent that cannot see the skill, on the same task, and compares the artifacts.
skill-doctor | cold-run | |
Reads your skills | yes | yes |
Runs them | no | yes |
Catches a broken link or bad frontmatter | yes | no |
Catches a skill that is well-formed and changes nothing | no | yes |
A skill can score 100/100 on structure and still be a no-op. That is the failure this catches, and it is the common one — 23 of my 47 were structurally perfect.
Run both. They fail on different things.
Prior art
obra/superpowers is the workflow layer — brainstorm,
plan, execute, verify. There is no overlap with this: zero workflow skills here, zero domain
checks there. Its writing-skills prescribes the same subagent testing this tool automates.
License
MIT.
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Verifies AI agent work end to end: real artifacts and outcomes checked, not self-reported success.
Deterministic AI code review, with an audit record. Governance inside the agent loop.
Git-backed platform for skills, tools, and context for AI agents
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/sina-heidariaan/cold-run'
If you have feedback or need assistance with the MCP directory API, please join our Discord server