crapkit
This server exposes read-side crapkit tooling: it ranks functions by complexity and uncovered risk, tracks debt over runs, and checks whether an edit would pass the commit gate.
Queue and worklist –
get_next_itemreturns the worst function(s) by crap score;list_worklistranks every admitted function by ccn × churn risk.Run history and trends –
list_runsshows scored runs;get_trendgives per-run totals and per-scope grades.Per-function deep dives –
get_function_briefreturns a start-editing packet (source, coverage, uncovered lines, commands, coupling, duplication twins, ratchet mark);get_function_historyshows a function's score across runs, optional commit/test history.Health and config –
check_configvalidates crapkit.toml against the repo, reports problems, versions, and resource policy.Repo-wide analyses –
list_coupled_filesfinds files that change together;list_duplicate_functionsfinds near-duplicate function pairs.Debt and ratchet –
get_ratchet_reportgives burn-down of marked debt, repayments, and policy violations.Gate preview –
check_gatejudges an edited file against the complexity ceiling, pardoning ratchet-marked debt.Claims –
list_claimsshows open queue claims to coordinate work.
Analyzes Git-tracked repositories, using git history to rank functions by churn and change coupling, and providing a commit gate that blocks changes that add complexity debt.
crapkit

crapkit scores every function in your repo on complexity times uncovered risk, ranks the worst ones by how often the file changes, and blocks commits that add more. It reads Python, TypeScript, TSX, JavaScript, Swift, Go, Rust, shell, PowerShell, C and C++, Objective-C, Vue, Java and Zig through lizard, and joins per-function branch coverage from the istanbul or coverage.py artifact your own test command already writes. JSON commands use sorted keys and a versioned schema for scripts, coding agents and the optional MCP server.
CRAP = ccn^2 * (1 - cov)^3 + ccnThe name is not ours: C.R.A.P. (Change Risk Anti-Patterns) was coined for crap4j by Alberto Savoia and Bob Evans in 2007.
ccn is the smaller of standard and modified cyclomatic complexity, both read off one
lizard pass. cov is branch coverage inside the function's span; with no branches it
falls back to statement coverage, and with no statements to invoked-or-not, so a
half-executed straight-line function never reads as fully covered.
Above the ceiling, coverage cannot save you. Decompose. At the default target of 6, a function at ccn 7 with 100% coverage still scores 7 and still fails the gate. The only move that clears it is splitting the function.
Why 6 and not 30. crap4j's conventional threshold of 30 is a CRAP score: it lets an
untested ccn 5 through (25 + 5 = 30) and a fully covered ccn 30 too. crapkit's default
is a complexity ceiling, because coverage can at best collapse CRAP to ccn, and a
function you cannot cover past ccn 6 is one you decompose. Set target = 30 in
crapkit.toml if you want the crap4j number. A repo with existing debt does not need to:
ratchet seed marks today's over-ceiling functions at today's score, the gate then judges
only the functions a change touches, and marks may only fall, so adoption never starts with
a wall of red. Next to crap4py, radon, xenon, wily and SonarQube:
docs/comparison.md.
crapkit scores git-tracked files only. Source you have not git added is invisible to
it.
Start with | When |
You want the first score in an existing Git repository. | |
Python or TypeScript quickstart | You want a worked example from setup through a passing verify. |
You need to choose scopes, wire tests or introduce a ratchet to existing debt. | |
You already have saved runs, ratchet marks or an installed plugin. | |
You are scripting commands or connecting a coding agent. |
The 60-second start
pip install crapkit
cd your-repo
crapkit init # write crapkit.toml and ignore measurement output
crapkit doctor # check scopes, test commands and coverage dependencies
crapkit coverage # runs the lane, joins coverage, stores a scored run
crapkit worklist # the ranked risk map
crapkit ratchet seed
git add crapkit.toml crapkit-ratchet.tsv .gitignoreNot a Python repo? uvx crapkit init runs the same commands and adds nothing to your
manifest: see A repo that is not Python.
init detects pytest, Vitest and Jest from the repository's own files. Review the
generated config before running its commands. When detection leaves a commented
lane, fill it in using the lane recipes.
Commit the adoption files, then run crapkit verify to establish a passing verdict.
Install the commit gate when the config and ratchet are ready.
coverage scores, worklist ranks:
$ crapkit coverage
run 1 @ fae4db93108: 2 functions scored: 2 measured, 1 over ceiling 6, CRAP load 41.0, grade F
-> next: crapkit worklist
$ crapkit worklist
worklist @ fae4db93108 (run 1, floor ccn>=5, churn 12mo) - 1 of 1 active (worklist_top 50), 0 dormant
risk 14.0 ccn 14 crap 38.5 cov 50% 1c/1a calc/grade.py:7 classify( score , attempts , late , bonus )risk 14.0 is ccn times a churn weight of one: a one-commit repo has no spread of commits
to weight, so each commit counts once and the ranking is complexity order until the
history grows (Risk). crap 38.5 and cov 50% are the
score and the coverage behind it.
ratchet seed signs today's debt at today's score. From then on marks only ever fall, so
the repo can get better and never worse while you burn it down.
One thing stops most first runs: the coverage plugin. init writes a lane that shells
out to your own test runner, and the runner needs its coverage package installed:
pytest-cov for pytest, @vitest/coverage-v8 (pinned to your vitest major) for vitest.
Without it the lane produces no artifact and coverage exits 5 quoting the runner's own
error. For pytest, init probes the python its lane will run and prints the install
command when pytest_cov is missing; pip install "crapkit[py]" pulls the plugin
alongside crapkit when the two share a venv. On a Windows PATH holding only the py
launcher it writes py, not a python3 the lane could never run, and when cmd.exe cannot
start the interpreter at all (exit 9009, the Store alias) it names that instead of guessing
at pytest-cov. A repo that pins no lockfile and carries its own .venv gets that venv's
interpreter in the lane, when that interpreter can import pytest, rather than whichever
python the shell answers with. The two quickstarts below walk a real repo end to end.
On Windows a lane command is read by cmd.exe, the shell that will run it, not by sh. Double quotes are the portable quoting. A single-quoted value is refused at config load with exit 3, because cmd.exe would hand pytest five words and the lane would write no artifact:
# the lane in crapkit.toml
command = "python -m pytest -m 'not live and not perf' --cov=calc --cov-branch --cov-report=json:.crapkit/cov/py.json"
$ crapkit doctor
crapkit: lane 'py': positional argument 'live' narrows a full-suite coverage run; drop it, attach it to the flag it belongs to (-n8, --numprocesses=8), or set full_suite = false deliberately (cmd.exe does not treat ' as a quote: write the value in double quotes); a suite whose testpaths cannot be collected in one process needs one lane per testpath, each with full_suite = false and its own artifactWrite it -m "not live and not perf". Carets, && and | segments, redirections and
empty quoted arguments all read the way the shell reads them, so a chained lane
(cd tests && python -m pytest --cov ...) is checked one segment at a time. doctor reads
a lane the same way, and FAILs one whose runner will not start.
Related MCP server: PhpCodeArcheology
Install
pip install crapkitThat is the release on PyPI. For the unreleased tip
of main, or from a local clone (run at the clone root):
pip install git+https://github.com/JeanFrancoisGagne/crapkit.git
pip install .A repo that is not Python
crapkit is a command-line tool, never a dependency of the code it scores. A TypeScript,
Go or Rust repo adds nothing to its own manifest. With uv
on the machine, uvx fetches crapkit into a cache of its own and runs it:
$ uvx crapkit init
wrote crapkit.toml with 1 scope(s): src
detected 1 lane(s) from this repo's own files: js - next: run `crapkit coverage`
added to .gitignore: .crapkit/
$ uvx crapkit coverage
$ uvx crapkit worklistThe lane still runs your own test runner, so Vitest or Jest and its coverage package come
from the repo's node_modules as they do today. uv tool install crapkit or
pipx install crapkit puts a crapkit command on PATH once, which is what the
commit gate and the Claude Code plugin call. uv brings its own Python when
the machine has none.
Requires Python 3.11 or newer and Git on PATH. The CLI has one runtime dependency,
lizard>=1.24.0; a package mirror needs both distributions. Install into the environment
you intend to use, then check crapkit --version. The pip install -e ".[dev]" under
Development is a different thing: it adds the test extra, for people
changing crapkit.
Python projects can install pip install "crapkit[py]" in their test environment to
include pytest-cov and subprocess-capable coverage.py. A separate tool installation
still needs the coverage plugin in the environment that runs the suite.
Analysis and scoring run locally and send no telemetry. Configured lane, mutation and alert commands run with your permissions and can contact services or change files. Review those commands before running Crapkit in a repository you do not trust (SECURITY.md).
$ crapkit --version
crapkit 0.7.6python -m crapkit works identically to the console script and is what to use from a
source checkout. Every subcommand accepts --repo PATH (default: the nearest crapkit.toml
at or above the current directory, so a monorepo workspace finds the root's), and with it
you never have to cd into the repo you are scoring; Subcommands shows
where the flag goes.
Upgrading
Keep the CLI and plugin versions aligned, measure fresh coverage after upgrading, and review any ratchet identity refusal before reseeding. The current reader is analysis version 10; older JavaScript and TypeScript callback marks can require a reviewed mapping. Follow the upgrade guide for saved state, portable records and Windows launcher locks.
Upgrading from 0.4.4
This historical example describes the 0.4.4 to 0.4.5 transition, from analysis version 7 to 8. It is retained to explain older refusal messages:
$ crapkit verify
crapkit: ratchet marks were recorded under [crapkit-analysis=7 lizard=1.24.0] but this run measures [crapkit-analysis=8 lizard=1.24.0] — CRAP scores are not comparable across metric versions; re-baseline with `crapkit ratchet seed`That transition changed cognitive complexity, not ccn or the CRAP formula.
Later reader changes also affect function identity. Use the current upgrade guide
when moving from any older release to today's reader.
The exe lock on Windows
An active MCP server can hold crapkit.exe open and make an upgrade fail with
Windows error 32. Stop that server or its agent session, rerun the upgrade with
the same installer, then restart the client. See the
Windows upgrade procedure.
The Claude Code plugin
claude plugin marketplace add JeanFrancoisGagne/crapkit
claude plugin install crapkit@crapkitTwo commands, installed once per user, and every repo on the machine gets it. The plugin
ships three skills, the read-side MCP server, and one advisory PostToolUse hook that names
any function an edit pushed over its ceiling. Claude reaches two of the skills by itself,
crapkit and crapkit-recover; the third you type, as /crapkit:crapkit-onboard, because
wiring a repo up happens once and its description has no business in every turn's window.
It adds no files to your repo, and it needs the crapkit CLI on PATH.
A repo with no crapkit.toml costs a silent sub-50 ms no-op per edit. After upgrading
the CLI, refresh the marketplace before updating the installed plugin:
claude plugin marketplace update crapkit
claude plugin update crapkit@crapkit --scope user
crapkit doctor --plugin-rootRestart existing Claude Code sessions to apply the plugin update. The check above compares installed files with the CLI on PATH; it does not reload a running session.
The hook registers on Edit|Write, which is every write that names a file. An agent that
writes its source through a shell heredoc names none, so a Bash event is judged off the
working tree instead. That half is yours to register, because it costs two
git spawns per shell call. Add a second PostToolUse entry to your own settings, same
command, matcher Bash:
{
"hooks": {
"PostToolUse": [
{
"matcher": "Bash",
"hooks": [
{ "type": "command", "command": "crapkit claude-hook --protocol 1", "timeout": 20 }
]
}
]
}
}The cost is one git rev-parse --show-toplevel and one git status --porcelain -z -uall
per shell call in any git repo, whether or not crapkit measures it: about 30 ms together
on crapkit's own checkout, and more on a bigger tree. What comes back is the dirty or
untracked *.py files written in the last 12 seconds, 25 at most, each judged the way an
edit is. Python only, so a TypeScript or Go repo pays the two spawns and hears nothing.
Codex
Codex can install the same marketplace's plugin through its own manager:
codex plugin marketplace add https://github.com/JeanFrancoisGagne/crapkit.git
codex plugin add crapkit@crapkitUse the three skills and MCP server in Codex. The advisory hook instructions above configure Claude Code's PostToolUse event. To refresh an existing Codex installation:
codex plugin marketplace upgrade crapkit
codex plugin add crapkit@crapkit
codex plugin list --marketplace crapkit --jsonCheck the installed Codex plugin with an explicit crapkit doctor --plugin-root PATH.
See plugin upgrades
for choosing that path and starting a fresh MCP session. A runtime with a skills
directory but no compatible marketplace can copy plugin/skills/* instead; other
MCP clients use the stdio setup.
Languages
14 languages, two coverage parsers. Coverage joins where a parser exists; everything else scores on complexity alone.
Language | Files | Coverage |
Python |
| coverage.py |
TypeScript |
| istanbul |
TSX |
| istanbul |
JavaScript |
| istanbul |
Vue |
| istanbul, when your vitest run reports on |
Swift |
| none: cc-only |
Go |
| none: cc-only |
Rust |
| none: cc-only |
shell |
| none: cc-only |
PowerShell |
| none: cc-only |
C and C++ |
| none: cc-only |
Objective-C |
| none: cc-only |
Java |
| none: cc-only |
Zig |
| none: cc-only |
A cc-only scope declares coverage_optional = true, scores crap = ccn, and needs no
lane. Nothing about it is provisional: the ceiling still binds and the gate still refuses
a function over it. Add a coverage lane the day a parser exists and the same scope starts
joining coverage.
crapkit init writes that key itself, on every scope whose languages all lack a parser,
and leaves it off any scope a lane could still measure. So the 60-second start above runs
unchanged on a Go, Rust or shell repo: crapkit coverage scores it with no lane at all,
and that run is the baseline worklist, next-item, ratchet seed and verify read.
Three readers are crapkit's own. lizard ships none for shell or PowerShell, so crapkit
counts their functions itself. Its Rust reader scores a 7-arm match as ccn 2 (filed as
lizard #494), so crapkit counts each non-wildcard arm like a C case and retires the
override the day upstream fixes it. The cognitive column charges that same block once,
the way Sonar charges a switch.
Expression arrows in arrays and argument lists are measured separately. In TypeScript,
wrap an arrow body in parentheses when it contains < before a comma, such as
x => (pair<T,U>(x)) or x => (x < 0). Without that delimiter, analysis refuses
the file because this reader cannot distinguish type arguments from an expression
separator. Generic arrow parameter declarations remain supported.
Functions on the same line have separate occurrence identifiers. Existing ratchet marks with ambiguous old identities require a reviewed mapping; see same-line function identity.
The gate
Use the advisory while editing, the gate when committing, and verify for the
full verdict. The preview and hooks differ in what their available evidence can prove:
Surface | Fires | Power |
| after an agent's edit lands | advisory. Names the breach on stderr. Blocks nothing, because PostToolUse runs after the write |
| when you ask, after the first coverage run | preview. The commit gate's verdict on demand, sub-second, before you stage. With no run behind it, exit 1 and |
|
| blocks. The hook exits 6; git reports 1. Staged blobs only, so it costs the size of the commit and needs no coverage |
| before you push, and in CI | the verdict. Gate, ratchet, new test failures, diff coverage, against the trusted baseline |
Both hooks exempt a function the committed ratchet already carries a mark for, so touching
signed debt never refuses a commit. verify is what fails a mark that rises. Since 0.4.5
its gate exempts a touched function whose fresh CRAP sits at or under its mark, the
rule rescore --gate already applied; push it past the mark and the gate fires again. The
pre-commit hook still exempts on the mark's existence alone, on purpose: a staged blob has
no coverage, so there is no fresh CRAP to compare against. It reports each exemption count
on stderr (staged function(s) carry a ratchet mark and were not gated), and says the same
about a staged file no [[scope]] claims, so a new top-level directory cannot go ungated
in silence.
The Crapkit root can sit below the Git top. A config in packages/api gates
that package's staged files as project-relative paths such as app/m.py.
Path and root rules
also cover absolute arguments, literal filenames and Git diff settings.
Git runs hooks outside your shell's activated venv. Bare python must resolve to an
interpreter that has crapkit installed, or spell it out
(exec /path/to/venv/Scripts/python -m crapkit hook-precommit).
Route 1: .git/hooks/pre-commit (local, not committed)
cat > .git/hooks/pre-commit <<'EOF'
#!/bin/sh
exec python -m crapkit hook-precommit
EOF
chmod +x .git/hooks/pre-commitThe same file from PowerShell. Out-File and > write a byte-order mark (UTF-16 on
5.1) in front of the shebang, and git then answers every commit with cannot spawn .git/hooks/pre-commit and lets it through; Set-Content -Encoding ascii does not. Git
runs the hook with its own sh, so the interpreter is spelled with forward slashes and
quoted, and no chmod is needed on Windows:
$python = (Get-Command python).Source -replace '\\', '/'
Set-Content -Path .git/hooks/pre-commit -Encoding ascii -NoNewline -Value "#!/bin/sh`nexec '$python' -m crapkit hook-precommit`n"crapkit doctor warns when the hook file git would spawn starts with a byte-order mark.
Route 2: a committed hooks directory
The whole route, from a repo that has no githooks/ yet:
mkdir -p githooks
cat > githooks/pre-commit <<'EOF'
#!/bin/sh
exec python -m crapkit hook-precommit
EOF
chmod +x githooks/pre-commit
printf 'githooks/pre-commit text eol=lf\n' >> .gitattributes
git add .gitattributes githooks/pre-commit
git update-index --chmod=+x githooks/pre-commit
git commit -m "add crapkit gate hook"
git config core.hooksPath githooksThe --chmod goes between the add and the commit. It writes the executable bit to
the index, so a commit that already happened does not carry it: run it after and git ls-tree HEAD still says 100644, which is a hook Unix checkouts silently skip. The
.gitattributes line is the harder half of the same failure: under Windows' default
core.autocrlf the hook checks out CRLF and #!/bin/sh\r dies on Linux and macOS with a
bad-interpreter error. crapkit doctor warns when a file under core.hooksPath is not
100755 in the index and prints the update-index line for it.
Git will not read a hooks path out of a committed file, so that git config line belongs
in your CONTRIBUTING setup steps. Every clone arms the gate with it.
Route 3: the pre-commit framework
crapkit ships a .pre-commit-hooks.yaml declaring id: crapkit-gate. In your
.pre-commit-config.yaml:
repos:
- repo: https://github.com/JeanFrancoisGagne/crapkit
# crapkit's release step rewrites this line to the tag it just cut
rev: v0.7.6
hooks:
- id: crapkit-gateThat file arms nothing on its own. The framework writes .git/hooks/pre-commit when you
tell it to, and until then git commit runs no gate and says nothing:
pip install pre-commit
pre-commit installpre-commit install is the line every clone needs, the way Route 2 needs its
git config core.hooksPath line.
rev is a git ref pre-commit resolves against that remote. Pin a release tag, not a
branch: pre-commit autoupdate only moves between tags, and a moving main would change
your gate under you.
Route 4: CI
A CI job runs on a fresh clone, which has no .crapkit/ store, so bare crapkit verify
exits 1. Running coverage first would make the PR's own tree the baseline, a gate that
can never fail. The portable baseline is the mechanism:
# on the default branch, after a passing verify: commit this file
crapkit verify --emit-baseline crapkit-baseline.tsv
# in the PR job, against the committed baseline
crapkit verify --baseline-tsv crapkit-baseline.tsv --github--github emits ::error file=... annotations that land on the PR diff; --sarif PATH
writes SARIF 2.1.0 for code-scanning upload. Refresh the committed baseline whenever the
default branch's verify passes.
Two things the job has to do before those lines run. Install crapkit, pip install crapkit, and pin the version the way Route 3 pins rev: an unpinned install moves your
gate on whatever day a release lands. Fetch the whole history. actions/checkout
clones one commit by default, verify reads the diff against the baseline's commit out of
git, and a shallow clone does not have that commit:
$ crapkit verify --baseline-tsv crapkit-baseline.tsv
crapkit: baseline commit a74260f321f is not an ancestor of HEAD in this shallow clone, which does not hold it; set fetch-depth: 0 on the checkout or run git fetch --unshallowThat is exit 4 on a git clone --depth 1 of a repo whose baseline verifies at full depth.
On a full clone the same exit blames what it used to, a rebase or an amend that rewrote
history, and asks for a fresh baseline instead.
Set fetch-depth: 0 on the checkout step, which is what crapkit's own
.github/workflows/ci.yml does.
The whole PR job, on GitHub Actions:
on: pull_request
jobs:
crapkit:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0 # verify needs the baseline's commit
- uses: actions/setup-python@v5
with:
python-version: "3.12"
- run: pip install crapkit
- run: pip install -e ".[dev]" # your own test dependencies
- run: crapkit verify --baseline-tsv crapkit-baseline.tsv --githubThe second install is the one people leave out. verify reruns your lanes, so the job
needs whatever your test command needs: the coverage plugin, npm ci, a database, all of
it. Without them the lane writes no artifact and verify exits 5 quoting the runner's own
error, which is a broken job and not a verdict.
What a refusal looks like
$ git commit -m "add route"
crapkit gate: 1 staged function(s) exceed the complexity ceiling of 6:
ccn 7 app/m.py:9 route( a , b , c , d )
decompose before committing (coverage cannot save a function above the target).That commit exited 1, not 6. Git collapses any failed hook to 1, so 6 is a code you
only ever see by running the hook yourself: crapkit hook-precommit exits 6 on a
violation and 0 otherwise. The stderr block above is the same either way.
CRAPKIT_OVERRIDE_REASON is not a bypass. Setting it routes the commit through the full
three-record audit: an alert line through alert_command, a ratchet entry staged into the
commit, and a row in the override log. All three land or nothing does, and an unset
alert_command refuses the override outright. See
docs/ratchet.md.
The GitHub Action
action.yml at this repository's root is a composite action, so a reviewer sees crapkit's numbers on the pull request without installing anything. Four lines add it to a workflow, and every input has a default:
- uses: actions/checkout@v4
with:
fetch-depth: 0
- uses: JeanFrancoisGagne/crapkit@v0.7.6The whole job those four lines sit in:
on: pull_request
jobs:
crapkit:
runs-on: ubuntu-latest
permissions:
pull-requests: write # the comment, and nothing else
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0 # the diff, and verify's baseline commit
- uses: actions/setup-python@v5
with:
python-version: "3.12" # the interpreter the install below lands in
- run: pip install -e ".[dev]" # whatever your lanes need to run
- uses: JeanFrancoisGagne/crapkit@v0.7.6
with:
gate: "false"That pip install step is the one people leave out, and it is the same one Route 4 above
names: the action installs crapkit and nothing else, so your lanes still need whatever
your test command needs. Without it the lane writes no artifact and the comment says so.
fetch-depth: 0 is the other one. actions/checkout clones a single commit; the action
reads the pull request's changed files out of git and verify reads the diff against the
baseline's commit. With a shallow clone the file list comes back empty and the comment
ranks the whole repository instead of the diff.
The action installs crapkit from $GITHUB_ACTION_PATH, which is its own checkout of the
ref you pinned in uses:. So a pin left at last month's tag scores your tree with last
month's crapkit rather than with whatever released since, and pinning a tag is the whole
version policy; the snippets above name the current release.
What the comment looks like
One comment per pull request, edited in place on every push. A hidden
<!-- crapkit-action --> line is how the next run finds it, so a fifteen-push branch
carries one comment and not fifteen. On a push event there is no pull request to carry
it, and the same text goes to the job log instead.
Rendered from three saved payloads: a pull request that adds an untested route() (ccn 8)
beside a ratchet-marked legacy_router(), in a repository whose diff_uncovered_max is 3.
The payloads are under tests/fixtures/action_comment/, and the unit suite pins this block
to their render:
<!-- crapkit-action -->
## crapkit
4 functions in 2 files, 2 over ceiling 6, CRAP load 149.59, grade F.
**verify failed, exit 6: complexity gate.**
- gate: `app/calc.py:34` `route( a , b , c , d )` ccn 8, cov 0%, crap 72.0 -> decompose
- uncovered lines in `app/calc.py`: 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45
Run 3 against baseline 1, 1 changed file: 1 gate violation, 0 ratchet regressions, 0 new test failures, 11 uncovered changed lines.
### Worklist: 1 changed file
| File | Function | ccn | risk | remedy |
|---|---|---:|---:|---|
| `app/calc.py:34` | `route( a , b , c , d )` | 8 | 4.0 | decompose |
| `app/calc.py:19` | `legacy_router( a , b , c , d , e )` | 8 | 4.0 | decompose (accepted debt) |The first line is the run crapkit coverage wrote: functions and files, how many sit over
the ceiling (over ceiling 6, or over their ceilings (6; reports 12, util 4) when scopes
set their own), CRAP load and grade, with a failed lane's first error line appended as
; lane 'js' failed: .... When coverage --json died before a summary, the line quotes the
error object it printed instead: `crapkit coverage` exited 5: lane 'py' cannot import pytest-cov; pip install pytest-cov.
The verdict opens with the exit code and the rule it stands for (complexity gate,
ratchet regressions, new test failures, diff-coverage ceiling N), then one bullet per
finding: each gate violation with its function, ccn, coverage, CRAP and remedy; each
ratchet regression as recorded -> fresh; each new test failure by id; and the first twenty
uncovered changed lines, one bullet per file, with a count of the rest. The counts line
closes it. A verify that passed is one line: **verify passed.** Run 2 against baseline 1, 7 changed files.
The rows are the ranked worklist for the files the pull request changed, worst first,
top of them, with the rows a finding names listed first. risk is ccn times churn
weight, the number crapkit worklist ranks on, and remedy is the run's own verdict for
that function: decompose, split-lines, add-tests or ok. (accepted debt) marks a function the
committed ratchet carries a mark for, so an untouched legacy_router does not read like
the pull request's own new function. A pull request that touches no ranked function gets
the heading and no table.
The two file counts describe the same diff, counted twice. 39 changed files is
git diff --name-only base.sha...HEAD, the branch's own commits, and it is what the
table is filtered to. The count on the verdict line is what verify measured from the
same fork point. With delta: "false" the second one is 0, because there is nothing
behind the checkout to measure from.
The inputs
Input | Default | What it does |
|
|
|
|
| scores the pull request's base commit first, so the verdict covers the commits the pull request adds. Costs a second lane run; |
|
| worklist rows rendered in the table |
|
| the interpreter |
gate: "false" is the default on purpose. A team adopts the action before it has decided
which findings should stop a merge, and a check that fails on day one gets turned off on
day two.
What the verdict line covers
On a pull request, the commits the pull request adds. The action scores the fork point
first, then the checkout, then runs crapkit verify --base <fork>, which measures the
diff from there and takes the fork point's run as its baseline. So the gate judges the
functions in the diff a reviewer is reading, and a repository that was already over its
ceiling before the branch started does not fail every pull request that touches it.
The fork point is git merge-base of base.sha and HEAD, not base.sha itself.
base.sha is the base branch's tip when the event fired, so a base branch that moved
after the branch forked carries commits HEAD never saw, and a run there would be neither
the baseline verify wants nor a diff anyone is reviewing.
The base run happens in a detached worktree under RUNNER_TEMP, and its store is copied
over the checkout's so both runs sit in one place. The cost is two lane runs on a pull
request: your suite runs once at the fork point and once on the checkout. Set delta: "false" to skip the base run, and the verdict falls back to the checkout against its own
run, which reports the tree's own health and judges no changed function. The comment says
so in place of verify passed:
**verify judged no changed function:** the base run was not made (no base commit). Run 2 against baseline 2, 0 changed files.Three things leave the base run unmade: a shallow clone that does not hold the fork
point, a fork point older than your crapkit.toml, and a lane that will not run against
that tree. The step logs crapkit base scoring exited N and writes the reason to
crapkit-base.reason in the words the comment then quotes, shallow clone does not hold the fork point of <sha>; set fetch-depth: 0 on the checkout, no usable crapkit.toml at the fork point <sha>: ..., or lane failed at the fork point <sha>: ... with the lane's
first error line. The verdict falls back the same way delta: "false" does, and the
ratchet still runs, so exit 7 there is a finding. What differs is the job's status. With
gate: "true" on a pull request whose base run was attempted and not made, the exit step
exits 1 and prints the reason, because actions/checkout's default depth-1 clone would
otherwise turn every pull request into a green check that judged nothing. A push event
and delta: "false" never attempt the base run, so they keep verify's own code; a push
has no base commit and no pull request to comment on.
One requirement the base run adds: the lane has to measure the tree it runs in. A lane
that reaches an installed copy of your package instead of the checkout will measure the
pull request's code while standing on the base commit, and the two runs then describe the
same tree. crapkit verify refuses a run whose artifact names files outside the tree
(exit 5), which catches the loud version of this; a lane pinned to a path outside the
worktree is the quiet one. Point the lane at the tree, or set delta: "false".
--reuse-artifacts is what keeps each of those runs to one pass of your suite. coverage
ran the lanes moments earlier on that tree, and verify parses those artifacts rather than
running the whole suite a second time for the same numbers.
Which is why crapkit coverage has to exit 0 for there to be a verdict. When it exits
anything else (a lane that failed, an artifact it refused), the action does not run
verify: on a runner that keeps its workspace between jobs (clean: false), a
verify --reuse-artifacts over a failed measurement read the artifact the dead lane had
left from an earlier run, passed over it, and made that run the trusted baseline. The
comment then carries coverage's failure in place of the verdict:
**no verdict: `crapkit coverage` exited 5 (lane 'py' failed: lane 'py' wrote no artifact on its last attempt; the .crapkit/cov/py.json on disk predates it); verify did not run.**The parenthesis is the first line of the lane failure the summary carries. When every
lane failed, coverage prints no summary at all and the lane errors are only in the job
log, and the line says so. With gate: "true" the job exits with coverage's code.
The other gate that judges a delta is the portable baseline in Route 4:
commit crapkit-baseline.tsv on the default branch and run crapkit verify --baseline-tsv crapkit-baseline.tsv in a step of your own. It needs no second lane run, and it needs
someone to keep that file current.
The comment is posted with gh api and the job's own GITHUB_TOKEN, which needs
pull-requests: write. Two things it cannot do: a pull request from a fork gets a
read-only token, so the POST is a 403 there, and a self-hosted runner without the gh CLI
on PATH fails that step. Both leave the rendered text in the job log.
Subcommands
crapkit clean --dry-run --json previews policy-based cleanup. See
resource policies
for shared analysis workers, process lifetime, bounded logs and retained evidence.
Every subcommand takes --repo PATH, and the flag goes after the subcommand. Without
it the root is the nearest crapkit.toml at or above the current directory
(ADR 0002): from a
monorepo workspace crapkit worklist reads the root configuration that claims the
workspace, says crapkit: using crapkit.toml at /repo on stderr, and reads a relative path
argument from where you stand. claude-hook is the one exception: it has no --repo,
because it takes its root from the file named in the hook payload it reads.
$ crapkit worklist --repo /path/to/repo --scope util --top 1
worklist @ a7c5c85ac37 (run 1, floor ccn>=5, churn 12mo) - 1 of 3 active (--top 1), 0 dormant
risk 5.4 ccn 5 crap 30.0 cov 0% 5c/1a util/stats.py:1 bucket( value , low , high )Before it, argparse reads the path as the subcommand name and exits 2 without ever
mentioning --repo:
$ crapkit --repo /path/to/repo worklist --top 1
crapkit: error: argument command: invalid choice: '/path/to/repo' (choose from 'inventory', 'coverage', ...)--json prints one sorted-keys JSON object on stdout, always carrying a schema field.
Command | What it does |
| Applies configured retention to recognized idle default test evidence and recovers abandoned temporary mutation worktrees. Preserves active runs, caller-managed output and intentional mutation pools. |
| Sniffs tracked source into per-directory scopes, writes a self-validated starter |
| Checks the config still describes the repo: unknown keys (with the accepted spellings), zero-file scopes, tracked source no scope claims, scopes no lane covers, lane cwds and commands that no longer resolve, lizard importable, oversized files. It reads each lane command with the shell that will run it, so a quoted interpreter path is one word and a runner after |
| One lizard pass over every in-scope file into a SQLite snapshot run, cached by content hash. |
| Runs the lanes, joins branch coverage onto a fresh inventory, writes a scored run. A failed lane is recorded, not fatal: its scopes fall back to |
| The full verdict against the trusted baseline: gate on touched functions, ratchet, no new test failures, optional diff-coverage ceiling. The three baseline selectors are mutually exclusive; |
| The risk map: every admitted function ranked by |
| The actionable queue as JSON, with churn, budget estimates and uncovered lines. Same run and same admission floor as |
| The open claims, and the way to hand one back without waiting for a verify. |
| The start-editing packet for one function: its own |
| A function's score across runs plus its mark. |
| Fresh complexity for named files over the latest run's stale coverage, joined by name. Advisory: it writes no run. |
| The mark lifecycle: seed new debt, prune gone code (a mark whose file git renamed follows it), merge as a git driver, move re-paths marks, report reads burn-down from the file's own git history. See docs/ratchet.md. |
| Run history, and retention. |
| The override audit trail: who granted what, when, and why. |
| Totals per trusted run: functions, over-target count, CRAP load, average, per-scope rollup. It reads a per-run rollup table rather than rescanning every scored row, and fills that table for any run missing one, so it writes to the store (best effort: a read-only |
| The delta between the two newest runs with identical lane sets. Silent when nothing changed. |
| One self-contained HTML page written to |
| Near-duplicate functions by normalized line shingles with containment scoring. Defaults: |
| File pairs that keep landing in the same commits. Defaults: |
| Diff-scoped mutation testing: flips comparisons, boundary shifts, boolean connectives and boolean literals on changed lines, runs |
| Runs each owning scope's |
| The cc-only gate on staged blobs. No coverage, no snapshot, no repo-wide cache. Exit 6 on a violation. |
| Reads one Claude Code PostToolUse payload from stdin and judges the file it edited: ccn against the scope ceiling, on functions the edit changed, minus functions a ratchet mark already covers. Advisory only: the edit has landed, and |
| Rescores tracked files as they change (mtime polling, default 2s, subprocess-isolated so a half-saved syntax error never kills the watcher). |
| The help git, npm and docker answer to. With no TOPIC it prints the command list; with one it prints that subcommand's own help, the same page as |
| A stdio MCP server with no extra dependency, exposing twelve read-side tools named |
Reading the output
Flags: why a coverage number is missing
Flag | Meaning | Scored |
| A lane artifact spoke about this function. | Real |
| A lane covers the scope, but its artifact is silent on this function, which normally means no test imports the file. |
|
| No lane's |
|
| The scope sets |
|
The coverage summary counts all four as measured / untested / no_lane / cc_only.
Remedy: what to do about it
Remedy | Condition | Action |
|
| Split it. No amount of coverage clears this. |
|
| Put each definition on its own lines, then measure again. Coverage cannot tell functions on one line apart, so the score stays at uncovered whatever the tests do. |
|
| Cover the branches. |
|
| Nothing. |
Grade and CRAP load
The grade is the share of functions over their ceiling: A+ at exactly zero, A under
2%, B under 5%, C under 10%, D under 20%, F at 20% or more. crap_load beside it
is the plain sum of every function's CRAP score, so it moves when a function gets better
even if the letter does not.
Risk: what ranks the worklist
risk = ccn * churn weight. The weight is a time-weighted sum over the file's commits in
the churn window: each commit contributes a logistic weight rising to 0.5 for the newest
commit in the log and falling to near zero for the oldest, so five edits last month
outrank fifty from two years ago. The window anchors on the newest commit, never on the
wall clock, so a fixed tree ranks identically forever.
Age is not the input, position in the log is. A log whose commits all share one timestamp has no range to weight against, so each commit counts once: a one-commit repo weighs every file 1.0, ranks by ccn, and promotes nothing under the floor, because a top 10% of equal weights would be every file. Commits minutes apart already rank. This repo was eight commits old, all made the same day:
$ crapkit worklist --scope util
worklist @ a7c5c85ac37 (run 1, floor ccn>=5, churn 12mo) - 3 of 3 active (worklist_top 50), 0 dormant
risk 5.4 ccn 5 crap 30.0 cov 0% 5c/1a util/stats.py:1 bucket( value , low , high )
risk 4.5 ccn 9 crap 90.0 cov 0% 1c/1a util/curve.py:1 curve( scores , mode , floor , ceiling , skip_none )
risk 4.3 ccn 4 crap 4.2 cov 75% 5c/1a util/stats.py:13 spread( values , cap ) okbucket at ccn 5 outranks curve at ccn 9 because five commits touched it and one
touched curve. That is the whole point of weighting by churn. spread carries the ok
marker: already at or under its ceiling, listed anyway, and next-item would not hand it
out.
The list splits in two: active (files with commits in the window) and dormant
(zero churn, kept out of the queue but counted). Two rules reach under the
worklist_floor. A file whose churn weight sits in the top 10% is promoted down to ccn 3,
which is why spread appears above at ccn 4. And a function over its ceiling is admitted
whatever its ccn, so the floor can never hold back debt.
The trusted baseline
Every verify measures the working tree against one earlier run, the trusted
baseline. crapkit runs list marks which one that is today.
Which runs qualify. A coverage run, or a verify that passed. A failed verify
never qualifies, and neither does a partial run (a lane failed, so some scope fell back
to no-lane) nor a hook override record, which carries no scored rows at all. In runs list, verdict=- marks a run that produces no verdict rather than one that failed: only
verify renders a verdict. Four readers ask this one question and get this one answer: the
baseline pick here, ratchet seed, prune, and the tighten damping that compares a mark
against the same commit's previous run. A mark can no longer be signed off a run verify
refused.
What advances it. Any qualifying run. coverage writes one wherever HEAD is, so a
dashboard cron advances the baseline exactly as CI does. A passing verify advances it
and tightens the ratchet on the way.
The taint rule. A failed verify recorded findings against a tree. Until some
verify passes, runs made after that failure do not become the baseline: choosing one
would move the comparison point past the findings, the flagged function would stop
counting as touched, and nothing would look at it again. verify says which run it
refused and falls back to the newest run in front of the failure.
$ crapkit runs list
run 1 @ 88012a148f6 2026-08-23T09:27:46Z coverage verdict=- lanes=py baseline
run 2 @ 803bdde8556 2026-08-23T09:27:53Z verify verdict=FAILED lanes=py
run 3 @ 803bdde8556 2026-08-23T09:28:02Z coverage verdict=- lanes=py
$ crapkit verify
warning: run 3 is not the baseline: verify run 2 FAILED with 1 finding(s) and no passing verify has cleared it since — measuring against run 1 @ 88012a148f6 instead, so those findings stay visible. Fix them, or pass `--baseline 3` to accept the newer run deliberately.
verify FAILED @ d89068de7f3 vs baseline 88012a148f6 (2 changed files)
GATE crap 72.0 ccn 8 cov 0% calc/legacy.py:7 legacy_router( a , b , c , d , e ) -> decompose
findings: 1 committed / 0 dirty (uncommitted edits and untracked files)Run 3 is a coverage run somebody took on the tree run 2 refused, and it scores the same
ccn-8 function. Without the rule it would have become the baseline, legacy_router would
have stopped being a touched function, and that gate line would never print again.
The escape, twice. Fix the findings and let a verify pass, which clears the taint
for good. Or accept the newer run on purpose with verify --baseline 3: an explicit id
bypasses the rule, and the run history records which run the verdict used. Nothing here
touches a repo that has never run verify: with no failure to protect, coverage alone
always advances the baseline.
When the id you pass cannot serve. A --baseline ID naming a real run that is not a
candidate says which run it is, why, and which ones can:
$ crapkit verify --baseline 3
crapkit: run 3 is an inventory run (no coverage was measured) and cannot serve as a baseline; trusted runs: 1, 2; pass `--baseline 2` for the newestExit codes
Code | Meaning |
0 | OK. For |
1 | Overloaded. Three unrelated things, listed below the table. |
2 | Usage error from argparse: unknown flag, missing positional. Raised before crapkit's own error handling. |
3 | Config error: |
4 | Git error: not a repository, a baseline commit rewritten out of the history. |
5 | Tool error: lizard not importable, a lane that produced no artifact, one that measured a different tree, one that measured this tree and reported it in absolute paths (the join is root-relative, so those match nothing either; the refusal names the runner's own switch, |
6 | Gate violation. A function the diff touched is over its ceiling and past any ratchet mark it carries: an edit that leaves a marked function at or under its mark is the debt the repo signed for and is exempt. Also |
7 | Ratchet regression the diff never touched. A marked function scores worse than its recorded high-water mark; a touched one past its mark reports 6. |
8 | New test failures against the baseline run. Failures the baseline already had do not count. |
9 | Diff-coverage ceiling breached: |
Exit 1 means one of three things
CI cannot tell a crash from a clean policy verdict on the code alone. Which one you got depends on the command:
Command | What exit 1 means |
| A |
| The debt policy was breached. Also a verdict. |
anything else | An unexpected error: "no snapshot yet, run |
verify reports the first of 6, 7, 8, 9 that fires, in that order. A gate violation
and a ratchet regression together report 6. A run that takes any of them fails, so it
neither advances the baseline nor tightens the ratchet, exit 9 included.
Quickstart: Python
A repo with calc/grade.py, tests/test_grade.py, and a pyproject.toml. Commit first;
crapkit reads git ls-files. Install the coverage plugin first, because the lane init
writes runs pytest --cov and those flags come from pytest-cov:
pip install pytest-cov(pip install "crapkit[py]" pulls both at once when crapkit shares the suite's venv.)
If your suite drives its own CLI through subprocess.run, add [tool.coverage.run] patch = ["subprocess"] to pyproject.toml and keep coverage>=7.10.6: pytest-cov 7.0.0
dropped subprocess measurement, so without that key every entry point scores 0% and nothing
warns. docs/lanes.md has the whole rule.
1. Scaffold the config
$ crapkit init
wrote crapkit.toml with 1 scope(s): calc
detected 1 lane(s) from this repo's own files: py - next: run `crapkit coverage`
added to .gitignore: .crapkit/, .coverage, __pycache__/init sniffs tracked source into one scope per top-level source directory, and detects a
coverage lane from what the repo already has: a pytest marker file (pyproject.toml,
pytest.ini, setup.cfg) writes a live [[lane]], and so does a test script or
vitest/jest in package.json. A lockfile beside them names the environment: uv.lock,
poetry.lock, pdm.lock or Pipfile.lock makes the lane uv run python -m pytest … (and
the matching run for the rest), because a bare python binds to whichever venv the shell
has active rather than the one the repo pins — see
The interpreter a lane binds to. Whatever
it detects, it also leaves commented templates for the runners it did not find, and those
carry the same launcher, so uncommenting one cannot hand the bare python back. Every lane
it writes reports into .crapkit/cov/, which is why the .gitignore list is so short: see
Where artifacts live.
[crapkit]
target = 6
[[scope]]
name = "calc"
paths = ["calc"]
languages = ["python"]
[exclude]
# A leading **/ matches zero or more directories, so each glob below reaches the
# repo root and every nested copy. Test directories leave the corpus on their own.
globs = [
"**/node_modules/**",
"**/dist/**",
"**/build/**",
"**/vendor/**",
"**/generated/**",
"**/__generated__/**",
"**/*.generated.*",
"**/*.test.*",
"**/*.spec.*",
"**/test_*.py",
"**/*_test.py",
"**/conftest.py",
"**/*_test.go",
"**/*.config.ts",
"**/*.config.js",
"**/*.config.mts",
]
[[lane]]
name = "py"
command = "python -m pytest --cov --cov-branch --cov-report=json:.crapkit/cov/py.json --junitxml=.crapkit/cov/junit-py.xml --continue-on-collection-errors"
artifact = ".crapkit/cov/py.json"
results_artifact = ".crapkit/cov/junit-py.xml"
parser = "coveragepy"
scopes = ["calc"]
# Declare one [[lane]] per coverage command, then run `crapkit coverage`.
# [[lane]]
# name = "js"
# command = "npx vitest run --coverage --coverage.reportsDirectory=.crapkit/cov/js --coverage.reportOnFailure --reporter=default --reporter=junit --outputFile=.crapkit/cov/js/junit.xml"
# artifact = ".crapkit/cov/js/coverage-final.json"
# results_artifact = ".crapkit/cov/js/junit.xml"
# parser = "istanbul"
# scopes = ["<your-scope>"]
# `crapkit test-scoped FILES` runs one command per scope, with {files}
# replaced by that scope's files, each quoted; a template with no {files}
# runs as written, which is how a scope whose tests live elsewhere runs them.
[crapkit.scoped_tests]
# calc: no test file under calc/, so the whole suite runs, from tests/
calc = "python -m pytest tests -q -p no:cacheprovider"
The last block is the one an agent loop needs. crapkit test-scoped exits 3 for a file
whose scope declares no template, and AGENTS.md
makes it step 4 of the burn-down loop. Every key is in
docs/configuration.md.
2. Check the config against the repo
$ crapkit doctor
ok config keys all recognized
ok scope 'calc': 1 file
ok every tracked source file belongs to a scope
ok 1 lane(s) declared
ok lane 'py': python -> /home/you/ledger/.venv/bin/python (pytest 8.3.3, pytest-cov 7.1.0)
ok lizard 1.24.0
doctor: no problems founddoctor prints one line per check and exits 1 only on a FAIL. WARN and note report
and exit 0.
3. Score the repo, and read the queue
$ crapkit coverage
run 1 @ fae4db93108: 2 functions scored: 2 measured, 1 over ceiling 6, CRAP load 41.0, grade F
-> next: crapkit worklist
$ crapkit worklist
worklist @ fae4db93108 (run 1, floor ccn>=5, churn 12mo) - 1 of 1 active (worklist_top 50), 0 dormant
risk 14.0 ccn 14 crap 38.5 cov 50% 1c/1a calc/grade.py:7 classify( score , attempts , late , bonus )Columns: risk, ccn, the function's crap and cov off the ranked run (- on an
inventory-only run), <commits>c/<authors>a in the churn window, path:line, the
function's long name, then a marker on rows the burn-down queue will not hand out (ok,
no-lane). The header counts the active rows against their total, so 50 of 3980 active (worklist_top 50) says what the cap hid, and reads (--top N) when the flag set the cap.
--json also carries ccn_std, weight and ratchet_mark.
worklist is the risk map, not a to-do list. It ranks finished rows too, so it does
not empty when the burn-down does. next-item is the other view of that run: it drops the
no-lane rows, ranks by crap, and its empty: true is the stop condition.
4. Take the top item
$ crapkit next-item
{"commit": "fae4db93108b4841a00959f9117430679e7250ca", "empty": false, "item": {"authors": 1, "ccn": 14, "ccn_std": 14, "cognitive": 13, "commits": 1, "cov": 0.5, "crap": 38.5, "end": 28, "est_splits": 3, "est_uncovered_paths": 7, "flag": "measured", "function": "classify( score , attempts , late , bonus )", "handle": "classify", "nesting": 3, "nloc": 22, "path": "calc/grade.py", "remedy": "decompose", "scope": "calc", "start": 7, "target": 6, "uncovered_lines": [9, 11, 15, 17, 19, 24, 25, 26, 27, 28]}, "run_id": 1, "schema": 1, "skipped_no_lane": 0, "stale": false}remedy: "decompose", est_splits: 3 (this needs roughly three pieces to fit under 6),
and uncovered_lines naming the ten lines no test walks. handle is the name form to
pass back, and stale: false says the run still describes HEAD. Every field is in
docs/agent-json.md.
5. Seed the ratchet
Arm the debt gate before fixing anything. ratchet seed records every over-target
function at its current score, and from then on nothing may get worse.
$ crapkit ratchet seed
crapkit-ratchet.tsv: added 1, tightened 0 - 1 mark(s) vs run 1 (fae4db93108)
$ git add crapkit.toml crapkit-ratchet.tsv .gitignore && git commit -m "adopt crapkit"6. Fix it and verify
Extract until every piece sits at or under the ceiling. Here classify became
_validate, _adjusted, _band and a classify that only sequences them, with the
table of cases pushed into parametrized tests. Commit the fix, then:
$ crapkit verify
verify OK @ 8d10c13303d vs baseline fae4db93108 (5 changed files)
$ crapkit coverage
run 3 @ 8d10c13303d: 5 functions scored: 5 measured, 0 over ceiling 6, CRAP load 19.0, grade A+
-> next: crapkit worklistCRAP load 41.0 to 19.0, grade F to A+. verify reruns the lanes and checks three things
against the trusted baseline: every function the diff touched sits at or under its
ceiling, no marked function got worse, and no test that passed in the baseline fails now.
Exit 0 advances the baseline and tightens crapkit-ratchet.tsv in place, so the repaid
mark leaves the file: follow up with git commit -am "ratchet: classify repaid". The full
mark lifecycle is in docs/ratchet.md.
crapkit next-item now comes back empty: true with a reasons object saying which
ending you got. That is most of the stop condition, not all of it:
AGENTS.md states the whole rule and reads the rest of
reasons.
Quickstart: TypeScript
A vitest repo with src/grade.ts and test/grade.test.ts.
1. Scaffold the config
$ crapkit init
wrote crapkit.toml with 1 scope(s): src
detected 1 lane(s) from this repo's own files: js - next: run `crapkit coverage`
added to .gitignore: .crapkit/The lane init wrote is
npm run test -- --coverage --coverage.reportsDirectory=.crapkit/cov/js --coverage.reportOnFailure --reporter=default --reporter=junit --outputFile=.crapkit/cov/js/junit.xml.
It reads vitest's json reporter from .crapkit/cov/js/coverage-final.json; the
reportsDirectory flag is what keeps that report out of your root. The junit half is the
lane's results_artifact, which the crashed-worker and no-new-failures checks read; both
reporters are named because --reporter=junit alone would replace the console output you
watch the suite through. Anything that produces
an istanbul coverage-final.json works; see docs/lanes.md for the
jest and pytest recipes, a package
one directory down, and a
crapkit root below the repo top.
2. Install a coverage provider
This is the step that stops most TypeScript users. vitest ships no coverage provider
by default. Without one, init and doctor are both happy and coverage dies with
exit 5:
$ crapkit coverage
crapkit: lane 'js' FAILED: lane 'js' produced no artifact at .crapkit/cov/js/coverage-final.json (command exit 1); lane log: /repo/.crapkit/lane-js.log; last output: $ npm run test -- --coverage --coverage.reportsDirectory=.crapkit/cov/js --coverage.reportOnFailure --reporter=default --reporter=junit --outputFile=.crapkit/cov/js/junit.xml
MISSING DEPENDENCY Cannot find dependency '@vitest/coverage-v8'
(exit 1)
crapkit: every lane failed (1 of 1); the errors are aboveThat failure writes no run. Every lane failed, so coverage exits before it opens a
store: there is no .crapkit/crap.sqlite yet and the run ids below still start at 1.
Install the provider, and pin the major yourself. Unpinned, npm resolves the newest provider against your older vitest and refuses the tree:
npm i -D "@vitest/coverage-v8@<your vitest major>"Question | Answer |
Which provider? | Either works. |
Which crapkit parser? | Both feed |
Which version? | The provider's major has to match vitest's. On vitest 2 that is |
The artifact crapkit wants is coverage-final.json, written by vitest's json coverage
reporter, which is on by default. If your vitest config sets coverage.reporter
explicitly, keep "json" in the list.
vitest writes no coverage report at all when the run fails. The lane init wrote
already carries --coverage.reportOnFailure, so a red test still produces the artifact.
If you write the lane by hand, or you would rather keep the switch beside your other
coverage settings, coverage.reportOnFailure = true in the vitest config does the same
job; either one is enough. The full block is in
docs/lanes.md.
3. Score the repo
$ crapkit coverage
run 1 @ 8bfbe613fcd: 2 functions scored: 2 measured, 1 over ceiling 6, CRAP load 56.68, grade F
-> next: crapkit worklist
$ crapkit worklist
worklist @ 8bfbe613fcd (run 1, floor ccn>=5, churn 12mo) - 1 of 1 active (worklist_top 50), 0 dormant
risk 15.0 ccn 15 crap 52.4 cov 45% 1c/1a src/grade.ts:8 classify ( row Row )classify is ccn 15 against a ceiling of 6: one function holding the late-and-retry
penalty, the letter bands, the demotion rule and the null case.
4. Seed the ratchet and commit
ratchet seed records every over-target function at the score it has today, so nothing
can get worse while you burn this one down.
$ crapkit ratchet seed
crapkit-ratchet.tsv: added 1, tightened 0 - 1 mark(s) vs run 1 (8bfbe613fcd)
$ git add crapkit.toml crapkit-ratchet.tsv .gitignore && git commit -m "adopt crapkit"5. Fix it
Above the ceiling, coverage cannot help, so classify gets split rather than tested.
penalty, band and demote come out as their own exported functions, and classify
keeps the null case and the bonus:
export function classify(row: Row): string {
if (row.score === null) {
return "N/A";
}
let score = row.score - penalty(row.attempts, row.late);
if (row.bonus && score < 90) {
score += 3;
}
return demote(band(score), row);
}rescore --gate judges that edit on complexity alone, before the slow step:
$ crapkit rescore src/grade.ts --gate
rescore vs run 1 @ 8bfbe613fcd (coverage STALE, complexity fresh)
ccn cov crap remedy function
5 0% 30.0 add-tests src/grade.ts:22 band ( score )
5 0% 30.0 add-tests src/grade.ts:38 demote ( letter , row Row )
4 0% 20.0 add-tests src/grade.ts:8 penalty ( attempts , late )
4 45% 6.7 add-tests src/grade.ts:48 classify ( row Row )
4 75% 4.2 ok src/grade.ts:59 average ( scores Array )Exit 0: every piece is at or under 6. The crap column is loud because its coverage half
is still run 1's, from before three of those functions existed, and add-tests is the
literal instruction for step 6.
6. Cover the new pieces
rescore --gate passed on complexity, not on coverage. penalty, band and demote are
three functions no test has ever called, so each gets a table test:
describe("band", () => {
it.each([[95, "A"], [85, "B"], [75, "C"], [65, "D"], [10, "F"]])(
"scores %i as %s", (score, expected) => expect(band(score)).toBe(expected));
});Run the suite once before the slow step:
$ npx vitest run
Test Files 1 passed (1)
Tests 21 passed (21)Skip this step and step 7 fails rather than passes. Run on a copy of this repo with step 6
left out, verify reruns the lanes against the real tree and three functions the old
suite never called come back over the ceiling:
$ crapkit verify
verify FAILED @ 0296156ff21 vs baseline 0e646697946 (1 changed files)
GATE crap 17.8 ccn 5 cov 20% src/grade.ts:38 demote ( letter , row Row ) -> add-tests
GATE crap 12.4 ccn 5 cov 33% src/grade.ts:22 band ( score ) -> add-tests
GATE crap 10.8 ccn 4 cov 25% src/grade.ts:8 penalty ( attempts , late ) -> add-tests7. Verify
$ crapkit verify
verify OK @ 2af3433d979 vs baseline 8bfbe613fcd (3 changed files)
$ crapkit coverage
run 3 @ 2af3433d979: 5 functions scored: 5 measured, 0 over ceiling 6, CRAP load 22.0, grade A+
-> next: crapkit worklistCRAP load 56.68 to 22.0, grade F to A+, and the mark seeded in step 4 is gone: verify
dropped it once classify scored under the ceiling, rewriting the tracked
crapkit-ratchet.tsv in place. Commit it with your change. Marks only ever fall.
A verify may also print warning: N changed line(s) have no coverage above its verdict;
that block is advisory unless diff_uncovered_max is set
(docs/configuration.md).
It prints warning: N function(s) over the ceiling carry no ratchet mark when the tree
holds debt ratchet seed never signed: the gate judges touched functions only and the
ratchet check compares marks only, so coverage loss on such a function would pass unseen.
The count is unmarked_over_target in --json, fires no exit code, and is zero on a repo
with no debt.
Documentation
Page | Covers |
Start here for anything deeper. The illustrated handbook: what crapkit is, how every piece works, and where each command earns its keep. Also at docs/handbook.html, self-contained, so it opens straight from a clone. | |
The judgment layer over the quickstarts: scope granularity, exclude vs lane, scoped_tests wiring, the first-verify taint hazard. | |
Every | |
The lane model, vitest and jest and pytest recipes, artifact reuse, flake retest, containers. | |
Worker budgets, command cleanup, log rotation, test evidence retention and safe cleanup. | |
Seeding, pruning, the git merge driver, metric stamps, debt policy, overrides. | |
Existing installations: analysis and key versions, saved state, plugin alignment and Windows upgrades. | |
Lossless exports, portable baselines and ratchets, including filenames with delimiters. | |
The machine surface: | |
Where crapkit sits next to radon, xenon, wily, coverage.py and SonarQube, and how they run together. | |
The burn-down loop an agent runs, and the rules for changing crapkit itself. | |
Three skills and the MCP server for Claude Code and Codex, with advisory PostToolUse hook instructions for Claude Code. |
crapkit.schema.json is the authority on the config file shape.
Development
pip install -e ".[dev]"
git config core.hooksPath git-hooks
python tools/testing/run.pyThe dev extra includes pytest, pytest-cov, pytest-xdist and coverage.py. The shared
runner owns the unit and E2E schedule; use --unit-workers 1 for serial unit
reproduction or --coverage for combined branch coverage and JUnit. The git config
line arms the complexity gate. See
CONTRIBUTING.md
for development and the verified implementation report
for complete Windows source and Linux wheel results, focused benchmarks and their limits.
Maintainer and project background
crapkit is created and maintained by Jean-François Gagné. Read the project background for the problem it addresses and how it fits into his work on software and AI.
License
MIT. See LICENSE.
Available Tools
12 toolscheck_configConfig and repo health checkARead-onlyIdempotent
Checks that crapkit.toml agrees with the repo: typo keys, empty scopes, missing lane cwds, runners that fail to start. Run it first when any tool answers strangely or the ranking misses a file, and list_runs when only the history is in question. It needs no run, probes each runner once, runs no lane, and any problem arrives with isError true. repo can be any directory under the checkout, and one with no crapkit.toml above it answers a pointer, never a parent's config.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No | path to the scored repo's root (default: the repo the server was started in) |
Output Schema
| Name | Required | Description |
|---|---|---|
| lanes | No | per declared lane |
| store | No | the run store |
| schema | No | payload schema version, 1 |
| problems | No | FAIL findings as sentences naming the fix; non-empty is exit 1 and the result carries isError true |
| versions | No | the tools behind every number |
| warnings | No | WARN findings: unmeasured directories, scopes with a lane but no scoped_tests template, artifacts written outside .crapkit, lanes without results_artifact; exit stays 0 |
| resources | No | effective resource policy; limits and estimates, not sampled utilization |
| newest_run | No | the newest run, or null when nothing has run |
| analysis_version | No | the analysis semantics version; with the lizard version it stamps every ratchet mark, so a bump refuses old marks until the repo re-seeds |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint false, and the description adds behavior beyond those: it needs no run, probes each runner once, runs no lane, and reports any problem via isError true. It also clarifies the no-parent-config fallback behavior. Nothing contradicts the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and then moves through usage, behavior, and parameter details in a logical order. Every sentence adds unique information and there is no filler, making it dense but efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter, non-destructive check tool with an output schema, the description covers purpose, when to use it, what behavior to expect, error signaling, and parameter semantics. The output schema can handle the return structure, so nothing needed for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already documents the repo parameter and its default, and the description adds further meaning: repo can be any directory under the checkout, and absence of a crapkit.toml above it produces a pointer rather than inheriting a parent's config. This goes beyond the schema's description and meaningfully clarifies the parameter's semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Checks') and a specific resource ('crapkit.toml') and enumerates the concrete failure categories it catches: typo keys, empty scopes, missing lane cwds, and runners that fail to start. It also names the sibling list_runs and distinguishes when that tool is the right choice, so an agent can tell this tool apart from its siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit guidance is given: run this tool first when any tool answers strangely or when ranking misses a file, and use list_runs instead when only history is in question. It also clarifies that repo can be any directory under the checkout and that a directory with no crapkit.toml above it yields a pointer rather than falling back to a parent's config, which removes ambiguity about invocation context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_gateCommit gate verdict for an edited fileARead-onlyIdempotent
Checks whether an edited file clears the hook's commit gate: fresh ccn per changed function against its scope's ceiling less pardoned ratchet debt. Call it after an edit once get_function_brief states the rule. CLI verify gives the repo-wide verdict. It runs no tests, and a breach reads gate.ok false, not an error. path is repo-relative or absolute inside repo, outside or missing is a config error. A tracked file is judged on its diff from HEAD, an untracked one in full, an unchanged or unscoped one judges 0. repo may be any directory under the checkout.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | repo-relative source file to judge as edited | |
| repo | No | path to the scored repo's root (default: the repo the server was started in) |
Output Schema
| Name | Required | Description |
|---|---|---|
| gate | No | the verdict block |
| note | No | fixed reminder that coverage is the baseline run's and complexity the working tree's; crapkit verify gives the real verdict |
| schema | No | payload schema version, 1 |
| functions | No | every function in the file, rescored |
| baseline_run | No | id of the run whose coverage was reused |
| baseline_commit | No | that run's commit, full sha |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses significant behavior beyond annotations: it runs no tests, a breach returns gate.ok false rather than an error, path outside/missing causes a config error, and tracked vs untracked files are judged differently. These details are not implied by readOnlyHint, idempotentHint, or destructiveHint, adding valuable operational context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-structured, front-loading the core purpose and usage, then covering edge cases. Every sentence adds essential information—purpose, timing, scope, error handling, and file-state logic—with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description need not explain return values. It covers usage timing, input constraints, and behavioral edge cases (tracked/untracked, unchanged/unscoped) comprehensively. For a tool with this complexity, the description provides everything an agent needs to call it correctly without additional inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds meaning beyond the schema by clarifying that path can be repo-relative or absolute inside the repo, and that outside or missing is a config error. It also states repo may be any directory under the checkout, which extends the schema's default behavior. This goes beyond the schema's basic descriptions, though not exhaustively covering every parameter nuance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description precisely states the tool's purpose: checking whether an edited file clears the hook's commit gate, with a specific rule (fresh ccn against scope ceiling minus pardoned debt). It also differentiates from siblings by referencing get_function_brief as a prerequisite and CLI verify as a repo-wide alternative, making it clear when this tool is the right choice.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit usage guidance is provided: 'Call it after an edit once get_function_brief states the rule' and 'CLI verify gives the repo-wide verdict.' This tells the agent when to invoke this tool and when to use a different one, plus clarifies that it runs no tests, preventing expectations of test results.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_function_briefStart-editing packet for one functionARead-onlyIdempotent
Returns one function's start-editing packet from the newest trusted run: scored row, uncovered lines and the refresh, test, gate and verify command lines, none of them run. Use it once a function is chosen. Skip it for picking what to fix, that is get_next_item, and for a score across runs, get_function_history. Every call shingles the repo for twins, seconds on a large corpus. name must live in path. name takes the long name, a bare identifier, a start line or NAME#2, exact match first. A miss lists the file's functions instead of erroring.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | the bare identifier (classify, or route for a Rust `route cmd : & Cmd`) or the whole long_name get_next_item printed (classify( score , late )); both resolve, exact match first | |
| path | Yes | repo-relative source file, forward slashes | |
| repo | No | path to the scored repo's root (default: the repo the server was started in) |
Output Schema
| Name | Required | Description |
|---|---|---|
| lane | No | the lane whose artifact produced cov and uncovered_lines, verbatim from the config; null when no lane covers the scope |
| path | No | the resolved file, repo-relative |
| churn | No | the file's churn, or null when it had no commits in the window |
| notes | No | prose the config carries for whoever edits here |
| stale | No | true when the run's commit is not HEAD, so every number here describes an older tree; run commands.refresh first |
| commit | No | that run's commit, full sha |
| handle | No | short name form: the bare identifier or (anonymous)#N, the same value get_next_item prints |
| params | No | parameters in declaration order, so a test can call the function without opening the file |
| remedy | No | decompose, split-lines, add-tests or ok: the branch the session takes; scored.remedy carries the same value |
| run_id | No | id of the run these numbers come from: the newest trusted run (a coverage run, or a verify run whose verdict passed) |
| schema | No | payload schema version, 1 |
| scored | No | the whole scored row from the run |
| source | No | the function's own text, start to end inclusive, newlines intact: editable without reading the file |
| target | No | this scope's effective ccn ceiling, the same value as gate_rule.ceiling |
| attempts | No | every claim ever taken on this function, oldest first; [] on a first attempt, and a row with closed null is a claim still open |
| commands | No | the rest of the loop as whole command lines for this file and scope; run them as given |
| coupling | No | up to 5 change-coupling partners at support 5 and confidence 0.5, strongest first; empty when none qualify; list_coupled_files has the repo-wide list |
| function | No | the resolved lizard long name, whichever name form was asked with |
| regrowth | No | did this get fixed before |
| versions | No | what produced these numbers |
| gate_rule | No | what check_gate will judge this edit by |
| est_splits | No | 0 when ccn <= target, else ceil(ccn / target) |
| file_totals | No | the file rolled up |
| ratchet_mark | No | the committed ratchet mark, or null when the function carries none or the repo has no ratchet file |
| file_functions | No | every scored function in the same file: where an extracted helper lands, and which names are taken |
| uncovered_lines | No | line numbers no test ran; [] when the span is fully covered; null when no artifact could answer, then uncovered_lines_note says why |
| duplication_twins | No | up to 10 near-duplicate functions at similarity 0.8, best first; empty is normal; list_duplicate_functions has the repo-wide pairs |
| est_uncovered_paths | No | round((1 - cov) x ccn) |
| uncovered_lines_note | No | present only when uncovered_lines is null: the reason and the move (stale artifact, no test imports the file, coverage_optional scope) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds valuable behavioral context beyond annotations: the packet's commands are 'none of them run', every call 'shingles the repo for twins' with a cost warning ('seconds on a large corpus'), and a miss 'lists the file's functions instead of erroring'. This is useful non-obvious behavior that annotations cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: the core return value is stated first, followed by usage routing, cost warning, and parameter semantics. Every sentence earns its place, and the sibling differentiation is packed into one sentence without bloat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema, so return values need not be re-explained. The description covers purpose, usage timing, exclusions, cost, name resolution, and miss behavior. For a read-only, idempotent tool with 100% schema coverage and an output schema, nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters. The description adds meaning beyond the schema by explaining the name resolution order ('exact match first'), the accepted name forms ('long name, a bare identifier, a start line or NAME#2'), and the constraint that 'name must live in path'. It also clarifies the miss behavior, which helps an agent interpret a non-error response.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Returns'), a specific resource ('one function's start-editing packet from the newest trusted run'), and enumerates the packet's contents (scored row, uncovered lines, command lines). It also explicitly distinguishes itself from siblings get_next_item and get_function_history, so an agent can tell them apart without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance ('Use it once a function is chosen') and explicit when-not-to-use guidance ('Skip it for picking what to fix, that is get_next_item, and for a score across runs, get_function_history'). It also names the alternative tools, leaving no ambiguity about routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_function_historyOne function's score across runsARead-onlyIdempotent
Returns one function's ccn, coverage, crap and flag in every run that measured it, oldest first, plus its ratchet mark. Use it to tell improving from decaying or regrown, and get_function_brief instead to start an edit. history true spawns git log -L capped at 10 commits, and tests true is null unless the lane recorded contexts. name resolves off the newest run that scored path, so a substring such as "eval" fans out to one entry per long name matched, and repo may be any directory under the measured checkout.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | the bare identifier (classify, or route for a Rust `route cmd : & Cmd`) or the whole long_name get_next_item printed (classify( score , late )); both resolve, exact match first | |
| path | Yes | repo-relative source file, forward slashes | |
| repo | No | path to the scored repo's root (default: the repo the server was started in) | |
| tests | No | also list the tests covering this function (coverage.py contexts), as tests | |
| history | No | also list the commits that touched this function (git log -L), as commits |
Output Schema
| Name | Required | Description |
|---|---|---|
| name | No | the name argument as given |
| path | No | the file asked about |
| schema | No | payload schema version, 1 |
| functions | No | the functions name resolved to, in store order |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnly/idempotent annotations, the description discloses meaningful behaviors: results are ordered oldest first, history spawns git log -L capped at 10 commits, tests is null unless contexts were recorded, and name resolution uses the newest run scoring path with substring fan-out. This adds real behavioral context the annotations alone do not provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core return value and usage, and every sentence contributes useful information. It is somewhat dense, especially in the trailing clauses about history/tests/name/repo, but remains well-structured and not padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the return shape is already covered. The description completes the picture by specifying ordering, optional-parameter behavior, name resolution rules, repo scope, and the sibling alternative. An agent has enough context to invoke this tool correctly without further inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although schema coverage is 100%, the description adds important semantic details beyond the parameter descriptions: history's git log cap, tests' null behavior, name's newest-run resolution and substring fan-out, and repo's flexibility to any directory under the checkout. This materially helps an agent choose and format parameter values correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action and resource: it returns one function's ccn, coverage, crap, flag, and ratchet mark across runs, oldest first. This clearly distinguishes the tool from siblings like get_function_brief and get_trend by focusing on per-run history for a single function.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says to use this tool to tell improving from decaying or regrown functions, and to use get_function_brief instead when starting an edit. This gives the agent both a positive use case and a concrete alternative, satisfying the when/when-not guidance requirement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_next_itemNext function to fixARead-onlyIdempotent
Returns the next function to fix as a work packet, the newest trusted run's worst row by crap. Use it to start a fix, list_worklist to survey the same run by risk, and get_function_brief once a function is chosen. It runs no tests, and empty true means the queue is spent, not the work, so read reasons. Filters cut before top counts: top 3 with exclude ["tests/"] returns the three worst rows outside tests, and an unknown scope name is a config error.
| Name | Required | Description | Default |
|---|---|---|---|
| top | No | return the next N packets as items instead of one item (N >= 1, default 1) | |
| repo | No | path to the scored repo's root (default: the repo the server was started in) | |
| scope | No | restrict the ranking to these declared scopes (exact [[scope]] names from crapkit.toml, one --scope each) | |
| exclude | No | skip rows whose path or long function name contains any of these fragments (one --exclude each) |
Output Schema
| Name | Required | Description |
|---|---|---|
| item | No | the one packet, present when empty is false and top is absent or 1 |
| empty | No | true when the queue has nothing to hand out; then reasons is present and item and items are absent |
| items | No | up to top packets in crap-descending order, present when empty is false and top is above 1 |
| stale | No | true when the run's commit is not HEAD, so cov, crap and uncovered_lines describe an older tree; crapkit coverage --reuse-unchanged (get_function_brief's commands.refresh) clears it |
| commit | No | that run's commit, full sha |
| run_id | No | id of the run these numbers come from: the newest trusted run (a coverage run, or a verify run whose verdict passed) |
| schema | No | payload schema version, 1 |
| reasons | No | why the queue is empty, present only when empty is true; the stop condition is empty true with skipped_claimed and no_lane_over_target both 0 or absent |
| skipped_claimed | No | rows another session's claim hid; present only when non-zero, and list_claims names the holders |
| skipped_no_lane | No | rows above the floor that no lane measures, kept out of the ranking because their cov 0 is a tooling gap, not a testing gap |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark it readOnly, non-destructive, and idempotent. The description adds behavior beyond that: 'It runs no tests', 'empty true means the queue is spent, not the work, so read reasons', 'Filters cut before top counts', and 'an unknown scope name is a config error.' No contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, each serving a distinct purpose: purpose, usage routing, behavioral caveat, filter semantics. Front-loaded and dense without fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and annotations covering safety, the description addresses usage, filtering, error conditions, and queue semantics. It is complete for an agent to decide when and how to invoke it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema descriptions cover all four parameters, but the description enriches them: it explains filter ordering ('Filters cut before top counts'), gives a concrete example (top 3 with exclude ["tests/"]), and states scope errors. This is additive beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Returns the next function to fix as a work packet, the newest trusted run's worst row by crap.' This names the verb, resource, and ranking criterion, and later contrasts it with list_worklist and get_function_brief, making its purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states 'Use it to start a fix, list_worklist to survey the same run by risk, and get_function_brief once a function is chosen.' This gives clear when-to-use and alternatives, plus the note 'It runs no tests' adds a usage caveat.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_ratchet_reportRatchet debt burn-downARead-onlyIdempotent
Reports the ratchet debt burn-down: open marks with their age, repayments and policy findings. Use it to judge whether marked debt is repaid or piling up, and get_function_history for one function's mark. It reads the marks file and its git log only, ages count back from the newest commit touching that file, never the clock, and no marks file means zeros. repo can be any directory under a measured checkout, and an unmeasured one answers isError true with the setup pointer.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No | path to the scored repo's root (default: the repo the server was started in) |
Output Schema
| Name | Required | Description |
|---|---|---|
| open | No | marks open on disk, working tree included: an uncommitted seed counts |
| oldest | No | up to 20 open marks, oldest first |
| schema | No | payload schema version, 1 |
| anchor_ts | No | unix seconds of the newest commit that touched the ratchet file; every age and window counts back from here, never from the wall clock; 0 with no ratchet history |
| uncommitted | No | marks the working tree and the newest committed ratchet file disagree on: added, repaid or tightened but not committed |
| dropped_total | No | marks repaid over the whole committed history |
| dropped_last_30d | No | marks repaid in the 30 days before anchor_ts |
| dropped_last_90d | No | marks repaid in the 90 days before anchor_ts |
| policy_violations | No | null when no debt policy is configured, [] when the policy ran clean, else the findings as sentences |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnly/idempotent/non-destructive annotations, the description reveals material behavior: it reads only the marks file and its git log, ages are computed from the newest commit touching that file rather than the clock, a missing marks file yields zeros, and unmeasured checkouts return isError true with a setup pointer. This exceeds what annotations alone communicate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every clause earns its place: core output, intended use, alternative tool, data source, age semantics, edge case, and error behavior are all packed without redundancy. The most important information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given an output schema already exists, the description covers all the essential call-time knowledge: what the report contains, how to use it, what data it reads, how age is computed, and behavior for missing files and unmeasured repos. Nothing critical is left for the agent to infer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already fully documents the single optional repo parameter with 100% coverage, so the baseline is 3. The description adds meaningful semantics by clarifying that repo can be any directory under a measured checkout and that unmeasured ones produce an isError response, which is value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Reports the ratchet debt burn-down' and enumerates exactly what it contains (open marks, age, repayments, policy findings). It also names get_function_history as the sibling for single-function marks, so the tool is clearly distinguished from nearby alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states the intended decision use case: 'Use it to judge whether marked debt is repaid or piling up.' It also gives a direct alternative with a condition: 'get_function_history for one function's mark.' This is clear when-to-use and when-to-use-elsewhere guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_trendDebt trend per runARead-onlyIdempotent
Returns per-run totals for every trusted run, oldest first, with a grade per scope. Use it for the whole repo's trajectory. Use get_function_history for one function and get_ratchet_report for marked debt. It runs no lane and spawns no git. The first call on a large store sums every run into a rollup cache and takes seconds. Later calls read the cache back. repo may be any directory under a checkout where crapkit init and crapkit coverage have run. An unmeasured one answers isError true with the setup pointer.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No | path to the scored repo's root (default: the repo the server was started in) |
Output Schema
| Name | Required | Description |
|---|---|---|
| runs | No | one row per trusted run, oldest first |
| schema | No | payload schema version, 1 |
| target | No | the [crapkit] target: the default ccn ceiling every scope inherits unless it sets its own |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark this as read-only, idempotent, and non-destructive, and the description adds useful behavior beyond that: first call on a large store builds a rollup cache and takes seconds, later calls read the cache, and unmeasured repos return isError true with the setup pointer. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence carries distinct information: core behavior, usage selection, side-effect profile, performance characteristics, parameter prerequisites, and error behavior. The description is front-loaded with the main result and purpose, with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simple schema, rich annotations, and presence of an output schema, the description covers all critical context: when to use it, alternatives, behavioral quirks, performance expectations, and failure mode. An agent has everything needed to select and invoke this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaningful parameter context: repo may be any directory under a checkout where crapkit init and crapkit coverage have run. This clarifies allowed values beyond the schema's path-to-root wording.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: returns per-run totals for every trusted run with a grade per scope, ordered oldest first. It also names sibling tools for different cases, making the distinction explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says to use this tool for the whole repo's trajectory and directs to get_function_history for one function and get_ratchet_report for marked debt. It also states the prerequisite that crapkit init and crapkit coverage must have run, and notes it runs no lane and spawns no git.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_claimsOpen queue claimsARead-onlyIdempotent
Lists open claims on queue items, oldest first. Use it when get_next_item answers empty or skipped_claimed above 0, and use get_function_brief to see one function's own attempts. No tool here writes a claim: the CLI releases a stale one with crapkit claims release PATH NAME, and verify closes one at the ceiling. repo may be any directory under the checkout, since the server walks up to the nearest crapkit.toml. No crapkit.toml above it answers an init pointer, and a checkout never scored answers a coverage pointer, both as isError true.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No | path to the scored repo's root (default: the repo the server was started in) |
Output Schema
| Name | Required | Description |
|---|---|---|
| open | No | number of open claims |
| claims | No | open claims, oldest first |
| schema | No | payload schema version, 1 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, and non-destructive behavior, but the description adds rich context beyond them: ordering, repo path resolution via walking up to crapkit.toml, and two specific error states (init pointer, coverage pointer) that return isError true. It also explains the absence of write capabilities and how claims are released or closed elsewhere.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the purpose and usage triggers, and every sentence carries relevant information. It is slightly dense with CLI command syntax and edge-case error behavior, but none of that is filler given the tool's role in a larger workflow.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description need not explain return values. It covers usage triggers, path semantics, error scenarios, and the write-boundary of the tool set. Everything an agent needs to call this correctly is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema by clarifying that 'repo may be any directory under the checkout, since the server walks up to the nearest crapkit.toml.' This is useful semantic context for an otherwise simple parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb, resource, and ordering ('Lists open claims on queue items, oldest first'), and explicitly distinguishes itself from sibling get_next_item by naming when it is needed. An agent can clearly tell what this returns without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit triggering conditions ('when get_next_item answers empty or skipped_claimed above 0') and names an alternative purpose for get_function_brief. It also clarifies that no tool writes a claim and points to the CLI for stale-claim release, leaving no ambiguity about when to invoke it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_coupled_filesFiles that change togetherARead-onlyIdempotent
Lists file pairs that keep landing in the same commits over the churn window, strongest first, at most 50. Use it before editing a file to learn what an edit drags along. Use list_duplicate_functions for copied code rather than co-change. It reads git log once per call, not the scored run. An empty list means no pair cleared both thresholds, not a missing run. A pair must clear both thresholds: min_support 5 needs five shared commits, min_confidence 0.5 means the rarer file moved with its partner half the time. repo defaults to the server's own root.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No | path to the scored repo's root (default: the repo the server was started in) | |
| min_support | No | minimum shared commits before a pair counts (default 5) | |
| min_confidence | No | minimum P(pair changes together), 0 to 1 (default 0.5) |
Output Schema
| Name | Required | Description |
|---|---|---|
| pairs | No | pairs clearing both thresholds, ordered by support x confidence descending, at most 50 |
| schema | No | payload schema version, 1 |
| window_months | No | months of git history the pairs were counted over (the config's churn_window_months) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description adds valuable behavioral context beyond that: it reads git log once per call rather than the scored run, an empty list means no pair cleared thresholds, and both thresholds must be cleared. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence earns its place: purpose, use case, sibling distinction, performance behavior, empty-result semantics, and threshold semantics. The most important information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with an output schema, annotations, and only three optional parameters, the description fully covers when to call it, what it returns at a high level, how to interpret empty results, cost implications, and parameter meanings. Nothing material is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the description goes further by explaining what each parameter means operationally: min_support 5 requires five shared commits, min_confidence 0.5 means the rarer file moved with its partner half the time, and repo defaults to the server's own root. This helps the agent choose thresholds appropriately rather than treating them as opaque numbers.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Lists file pairs that keep landing in the same commits over the churn window, strongest first, at most 50.' It also differentiates from list_duplicate_functions, so an agent can distinguish this tool from its sibling immediately.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells when to use it: 'Use it before editing a file to learn what an edit drags along.' It also gives a clear alternative and exclusion: 'Use list_duplicate_functions for copied code rather than co-change.'
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_duplicate_functionsNear-duplicate function pairsARead-onlyIdempotent
Lists near-duplicate function pairs in the newest run, at most 50. Use it before a refactor so twins are folded together, and get_function_brief for one function's twins. It shingles source on every call, seconds on a large repo, skips functions under 8 lines and same-file pairs, and an empty list means no pair reached similarity. similarity is shared shingles over the smaller function: 1.0 admits only a function found whole inside another, 0.8 four lines in five, and repo may be any directory under the checkout.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No | path to the scored repo's root (default: the repo the server was started in) | |
| similarity | No | containment threshold, shared over smaller, 0 to 1 (default 0.8) |
Output Schema
| Name | Required | Description |
|---|---|---|
| pairs | No | pairs at or above similarity, best containment first, at most 50 |
| run_id | No | the newest run whose rows were compared |
| schema | No | payload schema version, 1 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, and non-destructive behavior. The description adds valuable runtime context: it shingles on every call, takes seconds on large repos, skips short and same-file functions, and explains that an empty list means no pair met the threshold. This goes well beyond the structured annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every clause earns its place: result scope, usage context, performance, filtering rules, empty-result meaning, threshold semantics, and repo scope. It front-loads the core purpose before diving into details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers what the tool returns, the limit, empty-result meaning, performance expectations, filtering behavior, parameter semantics, and when to use it. With a rich output schema and annotations present, no critical context is missing for an agent to invoke this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Even though schema coverage is 100%, the description adds meaningful interpretation: similarity is shared shingles over the smaller function, with 1.0 and 0.8 explained through concrete examples. It also clarifies that repo may be any directory under the checkout, going beyond the schema's generic path description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb and resource: 'Lists near-duplicate function pairs in the newest run, at most 50.' It clearly distinguishes this from siblings like get_function_brief by focusing on pair-level duplication in the newest run.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says to use it before a refactor and points to get_function_brief as the alternative when only one function's twins are needed. This gives the agent both a trigger condition and a routing decision.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_runsScored run historyARead-onlyIdempotent
Lists every run in the store, oldest first by id. Use it to date the store or to see which commit the other tools answer from. Use get_trend for per-run totals and get_function_history for one function's scores per run. It reads the store only and spawns no git. repo may be any directory under the checkout, because the server walks up to the nearest crapkit.toml, and a relative path resolves from the server's start directory. No crapkit.toml above it answers an init pointer, and a checkout never scored answers a coverage pointer, both as isError true.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No | path to the scored repo's root (default: the repo the server was started in) |
Output Schema
| Name | Required | Description |
|---|---|---|
| runs | No | every run in the store, oldest first |
| schema | No | payload schema version, 1 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover safety (readOnly, idempotent, non-destructive, closed world), yet the description adds substantial context beyond them: no git is spawned, repo path resolution walks up to the nearest crapkit.toml, relative paths resolve from the server's start directory, and two failure modes return isError true. This is exactly the extra behavioral detail annotations cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and primary use case are front-loaded in the first sentence, and every subsequent sentence carries distinct information (alternatives, git behavior, path resolution, error pointers). It is dense and runs long in a single block, but no sentence is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need no explanation; the description instead covers the gaps that matter — path resolution, absence of git side effects, and the two isError pointer outcomes. Complete for a read-only listing tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with one optional param, so baseline is 3. The description goes further by explaining that repo accepts any directory under the checkout because the server walks up to the nearest crapkit.toml, which is path-resolution behavior the schema's one-line description does not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource with ordering semantics ('Lists every run in the store, oldest first by id') and immediately names the siblings it is not (get_trend, get_function_history). An agent can distinguish it from all eleven siblings without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit 'use it to date the store or to see which commit the other tools answer from' plus named alternatives with the condition that selects each (per-run totals vs one function's per-run scores). Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_worklistRisk ranking of every functionARead-onlyIdempotent
Lists the newest trusted run's whole risk ranking, every admitted function ordered by ccn times recency-weighted churn. Use it to survey a repo or split work by file, and get_next_item for the one crap-ranked packet to fix now. It runs no tests, keeps finished rows so it never empties, and reads the churn cache, not git. scope narrows before top caps, so scope ["core"] with top 20 returns the 20 riskiest in core, an unknown scope name is a config error, and repo may be any directory under the measured checkout.
| Name | Required | Description | Default |
|---|---|---|---|
| top | No | cap the active list (default: the config's worklist_top, 50) | |
| repo | No | path to the scored repo's root (default: the repo the server was started in) | |
| scope | No | restrict the ranking to these declared scopes (exact [[scope]] names from crapkit.toml, one --scope each) |
Output Schema
| Name | Required | Description |
|---|---|---|
| floor | No | the effective worklist_floor: rows under this ccn are listed only when over their ceiling or in a hot file |
| stale | No | true when the run's commit is not HEAD, so cov, crap and uncovered_lines describe an older tree; crapkit coverage --reuse-unchanged (get_function_brief's commands.refresh) clears it |
| active | No | the ranking: functions in files with churn in the window, risk descending, cut at top or worklist_top |
| commit | No | that run's commit, full sha |
| run_id | No | id of the run these numbers come from: the newest trusted run (a coverage run, or a verify run whose verdict passed) |
| schema | No | payload schema version, 1 |
| dormant_top | No | the first 10 dormant functions, same shape as active: sleeping hazards kept out of the queue |
| active_total | No | active rows admitted before the cap: what top or worklist_top hid |
| dormant_count | No | ranked functions whose file had no commits in the window |
| churn_window_months | No | months of git history the churn weights cover |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only, idempotent, and non-destructive hints. The description adds crucial behavioral context: 'reads the churn cache, not git', 'keeps finished rows so it never empties', and the precedence rule 'scope narrows before top caps'. These details go beyond annotations and are essential for correct expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no waste. The purpose is front-loaded, followed by usage, then behavioral and parameter notes. Every clause earns its place; the structure is logical and scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema (so return format is covered), all parameters are optional with defaults explained, and the description covers usage, behavior, and parameter semantics comprehensively. An agent has everything needed to invoke it correctly without further inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Despite 100% schema coverage, the description adds meaningful value: it explains precedence ('scope narrows before top caps'), gives a concrete example ('scope ["core"] with top 20 returns the 20 riskiest in core'), and clarifies repo constraints ('any directory under the measured checkout'). These enrich the schema definitions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Lists the newest trusted run's whole risk ranking' and defines the ordering criterion ('ccn times recency-weighted churn'). It clearly distinguishes from siblings like get_next_item (single item) and get_trend (trend analysis).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states usage intent: 'Use it to survey a repo or split work by file' and directs to the sibling get_next_item for the single worst item. It also clarifies what the tool does not do ('runs no tests') and that it 'keeps finished rows so it never empties', aiding when-to-use decisions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.7.6- Changed
check_gate4 fields changed- changed
Output schema / properties / functions / items / properties / remedy / descriptionPrevious value: -"decompose (ccn over ceiling), add-tests (coverage short) or ok (nothing left to do)"New value: +"decompose (ccn over ceiling), split-lines (another function shares its source lines, so coverage cannot tell them apart and no test lowers the score until the definitions sit on separate lines), add-tests (coverage short) or ok (nothing left to do)" - changed
Output schema / properties / functions / items / properties / remedy / enumPrevious value: -[ - "decompose", - "add-tests", - "ok" -]New value: +[ + "decompose", + "split-lines", + "add-tests", + "ok" +] - changed
Output schema / properties / gate / properties / breaches / items / properties / remedy / descriptionPrevious value: -"decompose (ccn over ceiling), add-tests (coverage short) or ok (nothing left to do)"New value: +"decompose (ccn over ceiling), split-lines (another function shares its source lines, so coverage cannot tell them apart and no test lowers the score until the definitions sit on separate lines), add-tests (coverage short) or ok (nothing left to do)" - changed
Output schema / properties / gate / properties / breaches / items / properties / remedy / enumPrevious value: -[ - "decompose", - "add-tests", - "ok" -]New value: +[ + "decompose", + "split-lines", + "add-tests", + "ok" +]
- Changed
get_function_brief6 fields changed- changed
Output schema / properties / file_functions / items / properties / remedy / descriptionPrevious value: -"decompose (ccn over ceiling), add-tests (coverage short) or ok (nothing left to do)"New value: +"decompose (ccn over ceiling), split-lines (another function shares its source lines, so coverage cannot tell them apart and no test lowers the score until the definitions sit on separate lines), add-tests (coverage short) or ok (nothing left to do)" - changed
Output schema / properties / file_functions / items / properties / remedy / enumPrevious value: -[ - "decompose", - "add-tests", - "ok" -]New value: +[ + "decompose", + "split-lines", + "add-tests", + "ok" +] - changed
Output schema / properties / remedy / descriptionPrevious value: -"decompose, add-tests or ok: the branch the session takes; scored.remedy carries the same value"New value: +"decompose, split-lines, add-tests or ok: the branch the session takes; scored.remedy carries the same value" - changed
Output schema / properties / remedy / enumPrevious value: -[ - "decompose", - "add-tests", - "ok" -]New value: +[ + "decompose", + "split-lines", + "add-tests", + "ok" +] - changed
Output schema / properties / scored / properties / remedy / descriptionPrevious value: -"decompose (ccn over ceiling), add-tests (coverage short) or ok (nothing left to do)"New value: +"decompose (ccn over ceiling), split-lines (another function shares its source lines, so coverage cannot tell them apart and no test lowers the score until the definitions sit on separate lines), add-tests (coverage short) or ok (nothing left to do)" - changed
Output schema / properties / scored / properties / remedy / enumPrevious value: -[ - "decompose", - "add-tests", - "ok" -]New value: +[ + "decompose", + "split-lines", + "add-tests", + "ok" +]
- Changed
get_next_item4 fields changed- changed
Output schema / properties / item / properties / remedy / descriptionPrevious value: -"decompose (ccn over ceiling), add-tests (coverage short) or ok (nothing left to do)"New value: +"decompose (ccn over ceiling), split-lines (another function shares its source lines, so coverage cannot tell them apart and no test lowers the score until the definitions sit on separate lines), add-tests (coverage short) or ok (nothing left to do)" - changed
Output schema / properties / item / properties / remedy / enumPrevious value: -[ - "decompose", - "add-tests", - "ok" -]New value: +[ + "decompose", + "split-lines", + "add-tests", + "ok" +] - changed
Output schema / properties / items / items / properties / remedy / descriptionPrevious value: -"decompose (ccn over ceiling), add-tests (coverage short) or ok (nothing left to do)"New value: +"decompose (ccn over ceiling), split-lines (another function shares its source lines, so coverage cannot tell them apart and no test lowers the score until the definitions sit on separate lines), add-tests (coverage short) or ok (nothing left to do)" - changed
Output schema / properties / items / items / properties / remedy / enumPrevious value: -[ - "decompose", - "add-tests", - "ok" -]New value: +[ + "decompose", + "split-lines", + "add-tests", + "ok" +]
- Changed
list_worklist4 fields changed- changed
Output schema / properties / active / items / properties / remedy / descriptionPrevious value: -"decompose, add-tests or ok; only decompose and add-tests rows with a lane reach get_next_item"New value: +"decompose, split-lines, add-tests or ok; every row but ok reaches get_next_item when a lane measures it" - changed
Output schema / properties / active / items / properties / remedy / enumPrevious value: -[ - "decompose", - "add-tests", - "ok" -]New value: +[ + "decompose", + "split-lines", + "add-tests", + "ok" +] - changed
Output schema / properties / dormant_top / items / properties / remedy / descriptionPrevious value: -"decompose, add-tests or ok; only decompose and add-tests rows with a lane reach get_next_item"New value: +"decompose, split-lines, add-tests or ok; every row but ok reaches get_next_item when a lane measures it" - changed
Output schema / properties / dormant_top / items / properties / remedy / enumPrevious value: -[ - "decompose", - "add-tests", - "ok" -]New value: +[ + "decompose", + "split-lines", + "add-tests", + "ok" +]
1 tool update
v0.7.2- Changed
check_config1 field changed- added
Output schema / properties / resourcesAdded value: +{ + "description": "effective resource policy; limits and estimates, not sampled utilization", + "properties": { + "available_cpus": { + "description": "CPUs visible to this process", + "type": "integer" + }, + "budget_directory": { + "description": "coordination directory; status does not create it", + "type": "string" + }, + "coordination": { + "description": "worker slot ownership scope", + "type": "string" + }, + "cpu_probe": { + "description": "CPU affinity or fallback probe used", + "type": "string" + }, + "default_chunks_per_worker": { + "description": "chunk target used by automatic pool sizing", + "type": "integer" + }, + "default_source_bytes_per_worker": { + "description": "source-byte target for automatic spawn sizing; null for other start methods", + "type": [ + "integer", + "null" + ] + }, + "estimated_pool_memory_mb": { + "description": "estimated memory for the allowed worker count", + "type": "integer" + }, + "inherited_analysis_workers": { + "description": "valid inherited worker ceiling or null", + "type": [ + "integer", + "null" + ] + }, + "log_max_bytes": { + "description": "byte limit per current and backup lane log; zero is unlimited", + "type": "integer" + }, + "memory_budget_mb": { + "description": "inherited memory sizing hint or null", + "type": [ + "integer", + "null" + ] + }, + "memory_is_hard_limit": { + "description": "false: memory policy sizes workers without an OS allocation limit", + "type": "boolean" + }, + "pool_worker_limit": { + "description": "per-pool ceiling after CPU, memory and inherited limits", + "type": "integer" + }, + "requested_analysis_workers": { + "description": "configured per-pool ceiling; zero selects the default", + "type": "integer" + }, + "serial_fallback": { + "description": "busy slots allow serial work without waiting", + "type": "boolean" + }, + "shared_pool_limit": { + "description": "shared numbered pool slot ceiling", + "type": "integer" + }, + "test_retention_count": { + "description": "default test evidence count limit; zero disables it", + "type": "integer" + }, + "test_retention_days": { + "description": "default test evidence age limit; zero disables it", + "type": "integer" + }, + "worker_memory_estimate_mb": { + "description": "memory estimate per analysis worker", + "type": "integer" + } + }, + "type": "object" +}
1 tool update
v0.7.0- Changed
get_function_brief2 fields changed- changed
Output schema / properties / commands / properties / refresh_writes_run / descriptionPrevious value: -"always true: refresh is the one command here that writes a run to the store; the other three change nothing"New value: +"always true: refresh writes a coverage run to the store; other commands can also write caches or test artifacts" - changed
Output schema / properties / commands / properties / scoped_tests / descriptionPrevious value: -"this scope's own test command with the file filled in, or null when the scope declares no [crapkit.scoped_tests] template"New value: +"the crapkit test-scoped call for this literal file, or null when the scope declares no [crapkit.scoped_tests] template"
22 tool updates
v0.6.0- Removed
brief - Added
check_config - Added
check_gate - Removed
coupling - Removed
doctor - Removed
duplication - Removed
explain - Removed
gate - Added
get_function_brief - Added
get_function_history - Added
get_next_item - Added
get_ratchet_report - Added
get_trend - Added
list_claims - Added
list_coupled_files - Added
list_duplicate_functions - Added
list_runs - Added
list_worklist - Removed
next_item - Removed
ratchet_report - Removed
runs - Removed
worklist
5 tool updates
v0.5.0- Changed
brief1 field changed- added
Input schema / requiredAdded value: +[ + "path", + "name" +]
- Changed
explain3 fields changed- added
Input schema / properties / historyAdded value: +{ + "description": "also list the commits that touched this function (git log -L), as `commits`", + "type": "boolean" +} - added
Input schema / properties / testsAdded value: +{ + "description": "also list the tests covering this function (coverage.py contexts), as `tests`", + "type": "boolean" +} - added
Input schema / requiredAdded value: +[ + "path", + "name" +]
- Added
gate - Changed
next_item1 field changed- added
Input schema / properties / scopeAdded value: +{ + "description": "restrict the ranking to these declared scopes (exact names from crapkit.toml, one --scope each)", + "items": { + "type": "string" + }, + "type": "array" +}
- Changed
worklist1 field changed- added
Input schema / properties / scopeAdded value: +{ + "description": "restrict the ranking to these declared scopes (exact names from crapkit.toml, one --scope each)", + "items": { + "type": "string" + }, + "type": "array" +}
9 tool updates
v0.4.13- First observed
brief - First observed
coupling - First observed
doctor - First observed
duplication - First observed
explain - First observed
next_item - First observed
ratchet_report - First observed
runs - First observed
worklist
TDQS
Scored across 12 tools
Each tool targets a distinct operation: history, gate, next item, worklist, brief, config, ratchet, trend, claims, coupling, duplicates, and runs. Even similarly named get_function_history and get_function_brief are clearly separated by context (across-run history vs edit-start packet).
All tool names follow a consistent verb_noun pattern in snake_case: get_ for single-item retrieval, list_ for enumerations, and check_ for validations. There are no mixed conventions or vague verbs.
12 tools is well within the ideal 3-15 range. Each tool covers a distinct aspect of the code-quality analysis domain (metrics, queue, config, coupling, duplicates) with no redundancy or bloat.
The surface covers all major read-only workflows: function inspection, work selection, trend analysis, gate checks, config validation, and structural analysis (coupling/duplication). Write operations like releasing claims or running tests are intentionally CLI-side, so there are no obvious gaps for the server's role.
Maintenance
Related MCP Connectors
AI-native git hosting — repos, PRs, issues, CI gates, and AI code review over MCP (60 tools).
Repo intel for AI coding agents: overview, PRs, contributors, hot files, CI, deps. Remote MCP.
GitHub repo maintainability verdicts—maintained, slowing, at-risk, abandoned—via MCP.
Read-only MCP access to stepcode.dev's how-to corpus: analysis pages by tool.
Related MCP Servers
AlicenseNot gradedqualityBmaintenanceEnables access to Codacy's code quality platform through natural language, providing repository management, security analysis, pull request reviews, and local CLI-based code analysis. Supports comprehensive code quality monitoring including issues, coverage, security vulnerabilities, and technical debt assessment across organizations and repositories.391 npm63MIT- AlicenseAqualityAmaintenancePHP static analysis MCP server with 11 tools for querying 60+ code quality metrics, detecting problems (God Class, dependency cycles, SOLID violations), analyzing dependencies, identifying refactoring priorities, and mapping test coverage — all from live analysis data.112 npm90MIT
- AlicenseAqualityAmaintenanceDocGuard, an official, zero-dependency Node.js MCP server (npm: docguard-cli, docguard mcp) exposing 5 read-only doc-governance tools: guard, score, explain, verify-claims, diagnose.71,567 npm1,278 PyPI30MIT
- AlicenseAqualityAmaintenanceDiffgate MCP server acts as a code review engine, enabling AI coding agents to analyze and validate code diffs before application. It enhances AI workflows by providing self-checking capabilities to optimise and secure code changes.766 npm5Apache 2.0