Skip to main content
Glama


Repowise is an ambitious project. We want every engineer, and every agent working beside them, to understand a codebase the way the person who has maintained it for five years does: what calls what, what tends to break, what nobody uses anymore, and why it was built this way. Cutting tokens was never the goal. It happens anyway, because an agent that can ask the index stops searching, opening and re-reading files to find out. Measured against the other context tools on the same agent tasks, it is also the largest saving.

Set it up in one line

Open Claude Code, Codex, Cursor or any MCP-capable agent in your repository and paste:

Read https://docs.repowise.dev/setup.md and set up Repowise in this repository.

The agent installs Repowise, indexes the repo with no API key, wires itself to the index, and asks you before anything costs money. Prefer to do it yourself:

uv tool install repowise          # or: pipx install repowise / pip install repowise
cd /path/to/your/repo
repowise init --yes --no-prose    # graph, git, health, dead code, docs. No key, no spend.
repowise serve                    # local dashboard + MCP server

init wires Claude Code automatically. Then ask your agent "Use Repowise get_overview to summarize this repository" or "What breaks if I change src/auth.py?" Full setup, every agent, optional model-written docs →


Related MCP server: CodeGraph

What it does

One index, three ways to use it. Find the question you came with; each one links to the page that answers it.

Understand the code

You ask

Repowise gives you

How does checkout work in this repo?

A cited answer built from the call graph and the generated docs, in one call. Search and answers

What calls this function, and what does it call?

A call graph across 26 parsed languages, every edge stamped with how it was resolved and how far to trust it, plus traced execution flows from each entry point. The graph

Can I get docs for this codebase?

A wiki for every module and file, rendered from the code's structure with no key, or written by a model when you choose. It updates incrementally after each commit. Docs

Which of our docs are wrong?

Markdown checked against the tree: every reference to a file or symbol the code no longer has, with the line to edit. Doc drift

Why is it built this way?

Decisions mined from ADRs, # WHY: comments, commit and PR history and your agent sessions, each tied to the code it governs and flagged when it goes stale. Decisions

Who knows this code?

Owners, bus factor, knowledge-loss risk when the main author goes quiet, and suggested reviewers. Ownership

Can I see the architecture?

An explorable dependency map, C4 views, and a Structurizr export, no model involved. Dashboard

Change it safely

You ask

Repowise gives you

What breaks if I change this?

Symbol-level blast radius: the callers of what you changed, the files that historically change with it but are missing from your diff, and the tests that reach it. Change risk

How risky is this PR?

Where the change ranks against your repository's own recent commits, with the reasons, as a directive your agent can act on. Change risk

Which tests should run?

The tests a diff actually exercises, from a coverage report if you have one and from the call graph if you do not. Test intelligence

Did this PR add untested lines?

Patch coverage, branch coverage on changed lines and path-scoped gates in CI, on GitHub, GitLab or any runner. CI gates

Will this API change break another repo?

HTTP, gRPC, topic and OpenAPI contracts matched across repositories, with a breaking-change guard and the consumer files it affects. Workspaces

Is anyone else editing these files?

Other open branches touching the same files or their co-change partners, each with the reason it is listed. Branch overlap

Did we just commit a secret?

Keys, tokens and risky calls found in the working tree and in full git history, with a pre-commit check and a CI gate. Security signals

Improve it continuously

You ask

Repowise gives you

What should we fix first?

A ranked queue weighing impact against effort, using churn, fan-in, coverage and bug history. Fix first

Where is the debt?

A 1 to 10 score for every file from 53 deterministic detectors, split into defect risk, maintainability and performance, validated against real bug history. Code health

Why is this slow?

N+1 queries, I/O in loops, blocking calls inside async code and quadratic loops, traced across function and file boundaries. Performance

How do I break this up safely?

Concrete refactoring plans: Extract Method, Extract Class, Move Method, Split File, Break Cycle, with the exact symbols that move and what moves with them. Ready to hand to an agent. Refactoring

What can we delete?

Unreachable files, unused exports and unused packages, each with a confidence tier and the evidence behind it. Dead code

Where do bugs keep landing?

Bug-fix commits traced to files and symbols, and a warning when your agent edits a bug magnet. Bug history

Are our tests testing anything?

Tests with no assertions, tests that only check their own mocks, and untested hotspots. Test-quality smells

  • Test coverage without running coverage. Most repositories never produce a coverage report. Repowise answers "is this tested, and by what" from the import and call graph, and labels every answer measured or inferred.

  • It checks its own health score on your repository. After each index it reports how many of the lowest-scoring files actually had bug fixes in your recent history, so a bad score on your codebase is visible to you.

  • It learns from your agent sessions. With transcript capture on, the corrections you keep repeating ("use the shared HTTP client") become tracked decisions it feeds back to the agent later. Transcripts never leave your machine.

  • It writes your CLAUDE.md and AGENTS.md from the real index and keeps them current, so even an agent with no MCP support starts informed.

  • It shrinks command output before your agent reads it. repowise distill pytest keeps every failure and drops the noise, and nothing is lost: an expand command restores any cut.

  • New git worktrees start indexed. A linked worktree seeds its index from the base checkout instead of starting cold.

  • It knows which external systems you depend on. Package manifests across PyPI, npm, Cargo, Go, NuGet, Maven and CMake feed a map of the services and libraries the code reaches.

Pick your front door

If you care about...

Start here

A coding agent that knows the repository

Task-shaped context in fewer calls, with decisions and risk delivered before the agent asks. For agents ↓

Safer pull requests and faster CI

Change risk, symbol-level callers, missing co-changes and the tests a diff needs. Change intelligence ↓

Paying down the code most likely to hurt you

A defect-validated health score, then the concrete fix. Code health ↓

An estate of many repositories

Contracts matched across repos, breaking-change guards, architecture rules in CI, one MCP endpoint for everything. Workspaces ↓

Rolling it out across a company

Self-hosted with nothing leaving your network, per-language accuracy, sizing, compliance status and licensing in one place. Teams and enterprise ↓


Your agent stops guessing

Every question your agent asks about a repository has an answer that could have been computed ahead of time. Who calls this function? What breaks if I change it? Why is it written this way? Which files are actually dangerous? Without an index, the agent rediscovers that answer on every task: grep, read, re-read, forget.

Repowise gives Claude Code, Codex, Cursor, VS Code and any other MCP host ten task-shaped MCP tools backed by one index of graph, git, docs and decisions. Most code tools are built around data entities, one file or one symbol at a time, which pushes agents into long chains of sequential calls. These are built around tasks: pass several targets in one call and get the whole picture back. The tool list ↓

About tokens. Every tool in this category promises to cut your token bill, and there are a lot of tools in this category. We think tokens are a symptom. An agent burns them because it does not know the codebase, so it searches, opens files, opens more files, and searches again. Give it an index that already knows, and the savings show up on their own. They also happen to be the best we have measured: in a paired agent loop over 43 questions on django/django, Repowise cut the agent's own output by 31.6% (p<0.0001) and got there in 3.8 tool calls instead of 7.2, ahead of every other context tool in the same run. On 42 sealed retrieval tasks it found 0.876 of the files a fix needed, against 0.610 for the next tool. Method and every row we lose →

Context arrives before the agent asks. Optional hooks push it into the session when it matters: the governing decision when your agent edits a file that decision covers, a warning when it touches a file with a run of recent bug fixes, a short briefing at session start, and a correction when it reaches for a path that does not exist.

It learns from how you work. Switch on transcript capture (repowise decision source set session --on) and Repowise reads your own agent transcripts for the corrections you keep making, turning the durable ones into tracked decisions it delivers back later. Transcripts never leave your machine; one batched model call per update turns the candidates that clear the deterministic gates into records, and --no-llm keeps the gates and drops that call.

Five layers, one index:

Layer

What it contributes

1. Graph

File and symbol dependencies across 26 AST-parsed languages, confidence-stamped call resolution, communities, centrality, cycles and execution flows

2. Git history

Hotspots, ownership, co-change, bus factor and bug-fix history: behavioural signals static analysis cannot see

3. Docs

A wiki for every module and file, hybrid search, and your own markdown checked against the tree for claims the code no longer supports

4. Decisions

Architectural rationale from ADRs, inline markers, commits, PRs and agent sessions, each claim traced to evidence

5. Health and change

53 deterministic detectors across defect risk, maintainability and performance, change risk, test impact, dead code and concrete refactoring plans

The structural wiki needs no model. Model-written prose is an optional upgrade, one page or directory at a time.

How the layers fit together → · How the graph earns trust →

Also: stop paying for output nobody reads

Most of what an agent reads back from a shell command is noise: 300 lines of passing tests wrapped around 4 failures, full commit bodies when it asked what changed recently. repowise distill <cmd> compresses command output before the agent reads it, errors first, exit code preserved.

repowise distill pytest          # 61% fewer tokens, all 11 failure lines kept
repowise distill git log -50     # 89% fewer tokens
repowise saved                   # what distillation saved you, in tokens and dollars

Every omission leaves an inline [repowise#<ref>] marker that repowise expand <ref> reverses in full, so the agent can pull the detail back without re-running the command. Small outputs pass through untouched. An opt-in hook rewrites noisy commands for the agent automatically.

Full guide: docs/agent/DISTILL.md →


Know what's dangerous before you merge

Four deterministic signals, all computed from the graph and git history, no LLM:

  • Change risk. Score any commit or base..HEAD range 0-10 from the shape of the diff, ranked against your repository's own recent commits. PR mode returns directives an agent can act on: may_break, missing_cochanges, missing_tests, tests_to_run. One command: repowise risk main..HEAD. (reference →)

  • Bug history. Which files and symbols actually get bug-fixed, and how recently. Doc, test and config commits are filtered out so the count means what it says, and a file with a run of recent fixes is flagged as a bug magnet while you edit it. (reference →)

  • Test intelligence. Which tests reach a file and which ones a diff exercises, from a coverage report or from the call graph. (reference →)

  • Change coordination. Which other open branches edit the files you are editing, every row saying why it is listed, and whether the diff in front of you is one change or several unrelated ones. repowise overlap and repowise risk. (reference →)

Which tests cover this file, without a coverage report

Ingest LCOV, Cobertura, Clover, JaCoCo or a Go coverprofile and you get the measured answer. Most repositories never produce one, so the graph answers instead: a test file that imports a source file reaches it, which is a recorded edge, where most tools fall back to matching file names.

repowise impacted-tests main..HEAD   # only the tests this diff exercises
repowise health                      # untested hotspots, graph-aware

Checked against a real coverage run --contexts=test on this repository: 95.7% precision on what reaches a file and 97.5% on the run list. Every row is stamped basis: "measured" or "inferred", measured wins where both can answer, and an empty answer means unknown, never "no tests". Test intelligence →

In CI, and on every pull request

Patch coverage, doc drift, security and change risk run as gates in your own pipeline through the GitHub Action, a GitLab template or plain CLI commands anywhere else, with annotations, SARIF and GitLab Code Quality output. The gates need no API key. Repowise in CI →

On GitHub you can also install the free Repowise PR Bot, a hosted GitHub App that puts the same analysis on every pull request. One comment, edited in place on every push, and a green PR gets no comment at all. It shows symbol-level blast radius (the contracts the PR changed and every caller outside the PR), the tests and co-change partners missing from the change, change risk against the repository's own history, and a public analysis page per PR. Zero LLM calls, so the same diff always gets the same review.

A real comment on a real PR: repowise-dev/repowise#1204 · its analysis page → · install the PR bot →


★ Know exactly what to fix

A score that says "this file is risky" is where most tools stop. Repowise scores every file, finds where the risk concentrates, and names the specific fix.

Every file is scored 1-10 by 53 deterministic detectors (McCabe complexity, brain methods, LCOM4 cohesion, god classes, clone detection, untested hotspots, change entropy, prior-defect history and more), read through three lenses: defect risk, maintainability and performance. Performance findings such as N+1 queries and I/O in loops are traced across functions and files through the call graph, which is where file-local linters lose them. Only 25 of the 53 detectors move the defect score, because that is the number carrying published accuracy claims.

Zero LLM calls, zero cloud. Detector weights are calibrated against a real defect corpus, not hand-tuned: every file scored at a commit before the bug window so nothing leaks backward, with file size as an explicit control, so a detector only earns weight for defect lift beyond a file being big.

It checks itself on your repository. After every index, Repowise compares its own flags with your git history and tells you what it found: "16 of the 20 lowest-health files had a bug fix in the last 6 months, 3.3x the 24% baseline." If that number is bad on your codebase, you will see it.

Then it names the fix. Extract Class, Extract Helper, Move Method, Break Cycle, Split File or Extract Method, with the exact methods, edges and symbols that move, the callers and co-changing files that move with them, and a ranking that puts a fix on a central hub above the same fix on a leaf. Extract Method runs a dataflow pass over the function to lift the exact span and infer a behaviour-preserving signature.

repowise next                          # what to fix first, ranked by impact and effort
repowise health                        # KPIs and lowest-scoring files
repowise health --refactoring-targets  # ranked, concrete plans
repowise health --trend                # snapshots plus declining-health alerts
repowise dead-code                     # what can go, by confidence tier

The dashboard renders each plan as a card with a copy-to-agent button. An optional model step, never in the indexing path, expands a plan into generated code and a unified diff.

Validated on 21 open-source repositories across 9 languages (2,826 files scored at a fixed point and checked against the following 6 months of bug fixes): ROC AUC 0.737 [0.683, 0.787]. Against CodeScene on the 2,770 files both tools scored, ranking by Repowise surfaces 2.3x the defects under a fixed review budget (paired, p = 0.003). CodeScene keeps a shorter, slightly more precise list. Full head-to-head and its limits →

Guides: code health · refactoring · dead code


See all of it

repowise serve starts the full web dashboard next to the MCP server. No separate setup, all local.

Also in there: Architecture and C4 views, the Knowledge Graph and a zoomable map, Risk, Hotspots, Coupling and Blast radius, Contributors and Ownership, Decisions with an evidence drawer and timeline, Symbols, Security, Dead code, Costs and Workspace. Every view and what it answers: docs/start/DASHBOARD.md →


One intelligence layer across your software estate

Real systems are not one repository, and the expensive failures live in the gaps between them. Change a backend contract and Repowise can name the frontend calls that consume it, the services downstream, the companion files missing from the change, and the architecture rule the new dependency violates, before it ships.

Workspace intelligence

What it answers

Contract map

Which services provide and consume each HTTP, gRPC, event, socket and data contract? Links keep their exact or candidate confidence and the source evidence.

Cross-repo blast radius

If this provider changes, which downstream services are in structural reach, and which may drift through historical co-change?

Breaking-change guard

Was an endpoint removed or an OpenAPI, proto or signature shape changed incompatibly, and which consumer files are linked to that contract?

Test impact

Which tests in the consumer repos should run for this provider change, measured from coverage or inferred from the call graph?

Architecture as code

Does the live system graph violate declared dependency rules or contain cycles? repowise workspace check gates CI.

Architecture health

How coupled is the estate? Propagation cost, the cyclic core, service roles and a deterministic 1-10 architecture score.

Federated context

One dashboard and one MCP server answer across every repository while keeping repo-level evidence.

The system map models services, not just repository boxes, and never conflates a real contract with "these files often changed together". HTTP field-level comparison covers a bounded OpenAPI 3.x subset; matched consumers prove endpoint exposure, not field use or runtime failure. Workspace guide and exact support matrix →

Worktrees and updates stay light: a linked worktree seeds its index from the base checkout, and post-commit hooks, file watching, webhooks or polling keep each repository and the cross-repo graph current. Keeping the index fresh →


Supported languages

26 languages parsed to an AST, 40 on a five-rung ladder, framework-aware where an ecosystem handler exists. Every language ships in the open-source distribution.

Rung

Languages

What you get

Full (13)

Python · TypeScript · JavaScript · Svelte · Vue · Java · Kotlin · Go · Rust · C++ · C# · Scala · Ruby

The whole pipeline: AST symbols, import resolution, a resolved call graph, heritage, docstrings, framework edges and code-health markers

Good (11)

C · Swift · PHP · Dart · Object Pascal · COBOL · GDScript · VB.NET · Elixir · F# · Objective-C

All of the above except the full health suite, within the language-specific ceilings in the full matrix

Partial (2)

Luau / Roblox · Razor / Blazor

Luau: AST symbols and require() resolution, Rojo and .luaurc aware. Razor: component symbols, @code and component-tag call edges, C# health markers; no import resolution yet

Lightweight (6)

Clojure · Haskell · Lean 4 · Erlang · HTML · QML

A real file-to-file import graph, and no symbol-level claims

Structural (8)

R · Zig · Julia · Elm · OCaml · Crystal · Nim · D

Git history: blame, hotspots, co-change, ownership, bug history

SQL and dbt projects get ref() / source() lineage, shell scripts get function-level symbols, HTML pages contribute their <script src> and <link href> dependencies, and OpenAPI, Protobuf, GraphQL, Dockerfile, Terraform and similar formats get dedicated handlers. Anything else is still tracked through git history.

Every call edge is stamped with how it was resolved and how much to trust it, from same_file at 0.95 down to a repo-wide name match at 0.50, labelled as the guess it is. Accuracy per language, graded against each language's own compiler: docs/BENCHMARKS.md#accuracy-by-language.

Full matrix: docs/layers/LANGUAGE_SUPPORT.md → · adding a language takes five small steps and no changes to the parser core: docs/architecture/language-support.md → · languages moving up the ladder: roadmap →


Supported agents and editors

Six agents wired end to end · two at the Full tier · every other MCP host one paste away.

Full is every surface Repowise has: MCP tools, skills, slash commands, a managed instructions file, hook-level interception of tool calls, and transcript mining after the session. Good is MCP tools and the config to reach them, without hooks or transcript mining. Anything else that speaks MCP is one snippet away: repowise agents print-config claude-code prints a server entry for Cline, Windsurf, Zed, Gemini CLI or any host that reads mcpServers. Integration matrix →

In VS Code, the Repowise extension shows what your change breaks before you push (riskiest files, what is downstream, forgotten companion files, missing tests, suggested reviewers), health in the gutter and status bar, callers and ownership on hover, and refactoring plans as CodeLens. One install also registers the MCP server, so the same index serves you and your agent. Install from the Marketplace or Open VSX and run Repowise: Set Up This Repository. VS Code guide →

In Claude Code, Lens ships inside the Repowise plugin and shows you what the index knows while Claude works: the file it is on and how many files depend on it, a change review after a turn that edits files (code health, tests to run, other branches on the same files), and /lens, a pane with Flow (what to check before accepting each turn), a map of the repo lit by what Claude searched, opened and edited and what its edits reach, Ask, and a session recap. Nothing reaches Claude unless you press a button, and Lens makes no model calls of its own. It needs Claude Code 2.1.287 or later, and the map needs repowise serve --no-ui running. Lens guide →


The ten MCP tools

Every response carries a _meta envelope with the indexed commit, the index age and a stale warning when the index has fallen behind your checkout, so your agent always knows how much to trust what it just read.

Tool

What only this tool answers

get_overview()

Architecture summary, module map, entry points, git health. The first call on an unfamiliar codebase.

get_answer(question)

Hybrid retrieval (full-text plus vector), graph expansion and one cited answer with a calibrated retrieval_quality. Search, read and reason in a single round-trip.

get_context(targets, include?)

Triage card for files, modules or symbols: summary, signatures, hotspot flag, governing decisions, symbol ids. include opens callers, callees, ownership and metrics. Batch many targets in one call.

get_symbol("file.py::Name")

One indexed symbol's source with exact line bounds.

search_codebase(query)

Hybrid search over code and docs, by symbol, path or concept.

get_risk(targets?, changed_files?)

Hotspots, dependents, co-change partners, ownership, test gaps, bug history. Pass changed_files for PR mode and get a directive back.

get_change_risk(revspec)

What a commit, range or uncommitted change made worse across defect risk, maintainability and performance, the tests that touch it, and how the diff ranks against recent commits.

get_why(query?, targets?)

Architectural decisions with their verbatim evidence. Falls back to git archaeology when no decision exists.

get_dead_code(...)

Unreachable code by confidence tier, with cross-repo consumers in workspace mode.

get_health(targets?, include?)

Health scores and findings across all three lenses, Fix first, coverage, trends, doc drift and refactoring plans.

Ten is a deliberate ceiling: a small, task-shaped surface is easier for an agent to choose from than a large one. Seven more tools (dependency paths, execution flows, refactoring code generation, finding triage, and three workspace architecture tools) are opt-in. Parameters, examples and when to use which: docs/agent/MCP_TOOLS.md →


Measured against the field

Open-source agent-context tools, the same repositories, the same pinned commits, the same questions, each tool given its full advertised surface. The full page carries the rows we lose beside the rows we win.

  • Finds the right files. 0.876 file coverage against CodeGraph's 0.610 on a sealed 42-instance split held out from every improvement round. 19 wins, 1 loss. Deterministic grading, no LLM judge. n=42, sign test p=0.00004.

  • Less work in a real agent loop. -31.6% output tokens against a bare agent, leaner on 37 of 44 questions. n=43, p&lt;0.0001. CodeGraph is a genuine second at -24.4%.

  • Fewer steps. 3.8 tool calls where the bare agent needed 7.2, and 3.0 files opened instead of 7.2: the mechanism behind the saving, visible directly.

  • A call graph the compiler agrees with. Against the Go team's own call graph and the TypeScript checker's own resolution, Repowise main reads 0.976 to 0.995 precision, with recall up in all 7 cells since August. In all seven, no tool that recovers as much of the graph gets more of it right. The tool with the highest recall in most Go cells is still not us, and the page says so.

  • Accuracy by language. Graded against each language's own toolchain: C 0.95 to 0.98, Python 0.96 to 0.98 on three of four repositories, Rust 0.85 to 0.90, C# 0.73 to 0.92, C++ 0.75, Java 0.73 to 0.78, each with its recall and its weak spots.

  • Scale. dotnet/runtime, 58,924 files, indexed in 120 minutes at 11.7 GiB peak memory on one laptop. Building just the call graph uses 75 MB at the median, the lowest of five tools on 35 of 35 repositories.

The full results, the methodology, and the rows we lose → · The research it is built on →

No single product competes with all of this, so there is no single table. Rows marked measured are head-to-head numbers that link to docs/BENCHMARKS.md, where the sample sizes, the tests and the rows we lose live. Unmarked rows are capability presence, not measurements.

As an agent context layer

repowise

CodeGraph

Serena

DeepWiki

Self-hostable, open source

✅ AGPL-3.0

✅

✅

❌ cloud only

Private repo, no cloud

✅

✅

✅

❌ OSS forks only

MCP tools served

10 default + 7 opt-in

1

29

3

Finds the gold files (measured, n=42 sealed)

✅ 0.876

0.610

not in this run

not measured

Output tokens vs a bare agent (measured, n=43)

✅ -31.6%

-24.4%

-14.8%

not measured

Memory to build the graph (measured, 5 tools, 35 repos)

✅ 75 MB, lowest on 35 of 35

757 MB

not measured

n/a, cloud

Time to build the graph (measured, same run)

2.77s, fastest on 14 of 35

3.65s, fastest on 16

not measured

n/a, cloud

Time to build the full index, django (measured)

⚠️ 366.8s, slowest here

✅ 16.4s

not measured

n/a, cloud

Call-edge precision, hand-graded (measured, 560 rows, 9 languages)

✅ 85.7%

58.6%

not measured

not measured

Call-edge precision, judged by a compiler (measured, 5 tools, 7 cells)

✅ nothing that finds as much gets more of it right, 7 of 7

lower precision in 7

not measured

not measured

Generated documentation

✅

❌

❌

✅

Proactive agent hooks

✅ Claude + Codex

❌

❌

❌

Generated CLAUDE.md / AGENTS.md

✅

❌

❌

❌

Command-output distillation

✅ reversible

❌

❌

❌

Architectural decision records

✅

❌

❌

❌

Multi-repo workspace intelligence

✅ contracts, co-change, federated MCP

❌

❌

❌

The two cost rows answer different questions. Building the call graph, we are the lightest tool measured, about ten times lighter than the next, and roughly as fast as the fastest. Building the whole index, CodeGraph is 22x faster than we are, because by then we have also built the git-history layer, the wiki, the decisions and the health pass. If a call graph is all you need, that is the right trade and you should take it. With model-written prose on, it is 135x.

The hand-graded precision row cuts both ways. About fourteen percent of the call edges we draw are wrong, and on seastar CodeGraph grades better than we do. The compiler row exists because we graded the hand-read one ourselves: on Go and TypeScript the answer key is the Go team's own call graph and the tsc checker's own resolution, which we neither wrote nor can tune. Precision alone is easy to win by drawing almost nothing, and recall alone by drawing everything, so the claim is the pair.

Competitors measured at CodeGraph 1.5.0, Graphify 0.9.31, Serena 1.6.2.dev0 and code-review-graph 2.3.7 in August 2026. Repowise compiler-graded figures are from main on 2026-10-03; the other Repowise rows are from 081a59fa, August 2026.

As a code health tool

repowise

CodeScene

Self-hostable, open source

✅ AGPL-3.0

⚠️ on-prem Docker, proprietary

Code health score (1-10)

✅ 53 detectors, 25 scoring

✅ 25-30

Brain Method / LCOM4 / god class

✅

✅

Defects found at a 20% review budget (measured, 2,770 files)

✅ 0.173

0.074

Effort-aware ranking, Popt (measured, p=0.003)

✅ 0.607

0.462

Precision at that budget (measured)

0.580

✅ 0.636, a shorter list

Discrimination, ROC AUC (measured, paired)

0.731

0.705, p=0.054, not significant

Business impact (resolution time)

❌ we could not replicate this on open data

✅ Code Red study

Git intelligence (hotspots, ownership, co-change)

✅

✅

Pre-merge change-risk scoring

✅ 0-10 + directives

✅

Concrete cross-file refactoring plans

✅ graph-aware + blast radius

⚠️ within-function only

Test-coverage intelligence

✅ LCOV/Cobertura/Clover/JaCoCo/Go

❌

Dead code detection

✅

❌

Serves it to an AI agent over MCP

✅

✅

CodeScene is the only other vendor in this category with a published empirical defect study, which is why it is the one we ran head to head. It flags about 27 files where we flag 132, so if you want a short list to act on, its threshold is the better fit.

Documentation generators

DeepWiki, Google Code Wiki and Swimm generate documentation from a repository, which overlaps one of our layers. We have not measured against them, so there is no table here.

The PR bot, against LLM review bots

Repowise PR Bot

CodeRabbit

Greptile

LLM calls per PR

✅ zero

❌ every review

❌ every review

Same diff, same review

✅ deterministic

❌ sampled output

❌ sampled output

Your code sent to a model provider

✅ never

❌ yes

❌ yes

Symbol-level blast radius

✅ call graph

❌

⚠️ prose, from context

Co-change partners missing from the PR

✅ git history

❌

❌

Change risk vs the repo's own history

✅ 0-10 + percentile

❌

❌

Silent on a clean PR

✅ by default

⚠️ configurable

⚠️ configurable

An LLM reviewer is a different product: it can read intent, and it can be wrong in a new way on every run. This one does set arithmetic over a call graph and a git history, so pushing the same diff twice gives the same review twice. Full side-by-side comparisons: repowise.dev/compare →


For teams and enterprises

AI makes producing a change cheaper. It does not make understanding its consequences cheaper. In a large estate that answer crosses repositories, ownership boundaries, service contracts, test suites and years of architectural history. Repowise gives developers, agents, reviewers and platform teams the same evidence about what exists, what depends on it, what is risky and what will break, from one index, in place of a separate health tool, dead-code tool, code-search layer for agents, docs generator and test-impact service.

Where it runs, and what leaves your network

Deployment

Where source is read

Where the index lives

What leaves your network

Self-hosted, open source (pip install repowise)

your machine

.repowise/ next to the repo

anonymous CLI telemetry you can switch off, and your own LLM provider only if you turn prose on

Self-hosted, commercial

your VPC or an air-gapped network

Postgres plus LanceDB or pgvector, inside your network

the same, plus integrations you configure. Nothing at all in air-gapped mode

Hosted (repowise.dev)

the hosted indexer

infrastructure we operate

your code goes to the platform

Graph, git, health, change risk, tests, dead code and PR review make zero LLM calls. Raw source is parsed in memory and never persisted. Prose is optional and runs on your own provider contract or fully offline through Ollama, chosen per repository. The threat model and data flows are in the security review pack. The published research each layer is built on, and how we checked it, is in LINEAGE.md.

How accurate, and how big

Call-graph accuracy is graded against each language's own compiler toolchain, and we publish precision and recall per language, with the misses, in the accuracy table. On current main, precision runs 0.95 to 0.99 for Go, TypeScript, C and most Python, 0.85 to 0.90 for Rust, and 0.73 to 0.92 for Java and C#, where overloaded methods are the main gap.

The largest repository indexed so far is dotnet/runtime: 58,924 files in 120 minutes at 11.7 GiB peak memory, on one 31 GB laptop with no API key. Time and memory by repository size: scale.

Source control and CI

Any git host works, because indexing reads a local checkout. On top of that:

  • GitHub: the PR bot (GitHub App), a GitHub Action for the CI gates, and webhook auto-sync.

  • GitLab: a CI template for the same gates and webhook auto-sync.

  • Bitbucket Pipelines and others: the gates run as plain CLI commands with the pipeline's own branch variables. (Repowise in CI)

  • Managed integrations for GitHub Enterprise, Azure DevOps, GitLab self-managed and Bitbucket are rolling out.

  • Not on git? Point repowise init at a plain directory, an export, or a Perforce or SVN workspace. The graph, docs, decisions and health layers build normally; the history layer needs a commit log, and Perforce and SVN adapters are on the roadmap.

What is shipping, and what is not yet

Status

Capability

Shipping in open source

Every deterministic layer, ten MCP tools, multi-repo workspaces, contract extraction and cross-repo blast radius, test intelligence, architecture conformance, local dashboard, auto-sync, full-history secret scanning.

GA commercially

Hosted graph-aware security, CVE prioritization, CycloneDX SBOM and VEX, PCI-DSS and SOC 2 control-coverage reports, audit export and webhook stream, Jira and Confluence, reference HA topology on your infrastructure, custom language extensions, SLA support, IP indemnification.

Rolling out

Managed GitHub Enterprise, Azure DevOps, GitLab and Bitbucket integrations. SAML/OIDC SSO and SCIM. Engineering-leader dashboards.

Planned

RBAC and multi-tenancy, a packaged air-gap install bundle, the Helm chart.

Repowise holds no SOC 2, ISO 27001 or other audited certification today. The SOC 2 and PCI-DSS reports above are control-coverage signals from your own findings, and every export says so. Every item is tracked with its status in the capability matrix.

Pricing, licence and support

Commercial licences are priced per indexed repository, with unlimited seats inside the licensed set. More and more of the code in a repository is written and read by agents and CI, so seat counts stop tracking the value. Details →

Do we have to open-source our code because of the AGPL? No. Using Repowise inside your company, including on internal servers, creates no obligation to publish your code. The obligations apply if you modify Repowise and offer it to others over a network, or ship it inside your own product. A commercial licence removes them.

Commercial support comes with a named contact, a response-time SLA and a quarterly architecture review. Security fixes land on the latest minor release; reports go to security@repowise.dev under the disclosure policy.

Commercial detail · Security review pack · Roadmap · hello@repowise.dev

repowise.dev runs the same engine fully managed. We run it on our own codebase in the open: live snapshot · explore public repos.

Already indexed locally? repowise publish puts the same repo on repowise.dev in one command: it asks repowise.dev to index the repo's GitHub remote, so nothing is uploaded from your machine and only what you have pushed is published. A hosted index usually takes about 10 minutes. Public repos publish on a free account (up to 2 repos, no card); private repos and more repos need Pro, free for 10 days with a card. What hosted adds →


Privacy

  • Deterministic or offline mode: with --no-prose, code-derived content stays on your infrastructure. The CLI reports anonymous, opt-out usage telemetry (command names and coarse environment only); turn it off with repowise telemetry disable, DO_NOT_TRACK=1, or by running fully offline. What's collected →

  • Optional LLM features: generated prose, decision extraction and refactoring code generation send code-derived prompts directly to the provider configured with your own key. Repowise does not proxy those calls.

  • What's stored: the graph, embeddings, generated wiki pages and git metadata. Raw source is processed transiently and never persisted.

  • Fully offline: Ollama plus a local embedding model means zero external calls.

Doing a security review? docs/business/SECURITY_COMPLIANCE.md →


CLI

repowise init [PATH]      # index a codebase (asks; --yes --no-prose needs no key)
repowise generate [PATH]  # write wiki pages with a model, on demand
repowise serve [PATH]     # MCP server + local dashboard
repowise update [PATH]    # incremental update (--workspace for every repo)
repowise watch            # re-index on file change
repowise search "<q>"     # hybrid search (fulltext / semantic / symbol / path)
repowise ask "<q>"        # a synthesized answer with citations
repowise context <files>  # triage card: layer, hotspot, fix history, freshness
repowise symbol <id>      # one symbol's body, with verified line bounds
repowise why <q|path>     # decisions, rationale, git archaeology
repowise next             # what to fix first
repowise health           # code-health KPIs and lowest-scoring files
repowise risk main..HEAD  # score a branch or PR range
repowise overlap          # other branches editing the same files
repowise impacted-tests   # only the tests a diff exercises
repowise dead-code        # unreachable code by confidence tier
repowise doc-drift        # documentation the code no longer supports
repowise security         # secrets and risky patterns, tree or full history
repowise decision list    # architectural decisions
repowise export --format structurizr  # the architecture as Structurizr DSL
repowise distill pytest   # compact, errors-first, reversible command output
repowise saved            # tokens and dollars saved by distillation
repowise workspace add    # multi-repo workspace management
repowise doctor           # check setup, API keys, index drift
repowise publish          # put this repo on repowise.dev (indexed from GitHub; public repos free)
repowise uninstall        # remove what repowise wrote, and say what it left

Every command and flag: docs/reference/CLI_REFERENCE.md · config: docs/reference/CONFIG.md · something not working: docs/start/TROUBLESHOOTING.md · all docs: docs/README.md


Contributing

git clone https://github.com/repowise-dev/repowise
cd repowise
uv sync --all-packages
uv run repowise --version
uv run pytest tests/unit/

New here? You do not have to read 3,000 files to start. We keep a public index of this repository built by Repowise itself, re-indexed on every push: explore repowise with repowise → (architecture, hotspots, ownership, decisions, and a ranked refactoring backlog you are welcome to pick from).

Full guide, including how to add languages and LLM providers: CONTRIBUTING.md · architecture: docs/architecture/


License

AGPL-3.0. Free for individuals, teams and companies using Repowise internally.

For commercial licensing (the enterprise security and compliance layer, SSO/SCIM, RBAC, workflow integrations, priority support and SLA, or embedding Repowise in a product without AGPL obligations), see docs/business/COMMERCIAL.md or contact hello@repowise.dev.


Built for engineers who got tired of watching their AI agent cat the same file for the fourth time.

Available Tools

10 tools
get_answerA

Answer a how, where, or why question in one evidence-grounded call.

High confidence is content-grounded and may be used directly. Medium
confidence keeps the smallest verification evidence; low confidence leads
with an actionable local conclusion and ranked evidence. Provider keys and
network access are optional: local source, symbols, FTS, rationale, and
data-shape evidence remain usable when embeddings or synthesis fail.

Responses fit 24,000 serialized characters. Pass ``include=["evidence"]``
for the deduplicated expanded projection, capped at 32,000. Reductions carry
totals, emitted counts, reasons, and an exact one-call recovery.

Args:
    question: Developer question.
    scope: Optional repository-relative path prefix.
    repo: Usually omitted; a workspace alias when needed.
    include: Optional ``["evidence"]`` expanded projection.
ParametersJSON Schema
NameRequiredDescriptionDefault
repoNo
scopeNo
includeNo
questionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does so thoroughly. It discloses confidence tiers, evidence grounding, optional provider keys/network, fallback behavior, response size caps, reduction details, and recovery semantics. This is far beyond a basic 'answers questions' statement.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Although long, every sentence earns its place: purpose, confidence behavior, offline fallbacks, size limits, reduction recovery, and parameter semantics. The description is front-loaded with the core purpose and organized with clear sections; nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has 4 parameters, no annotations, and an output schema. The description covers purpose, usage scope, param meaning, behavioral guarantees, failure/recovery paths, and output constraints. There is no obvious missing information an agent needs to select and invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, yet the description's Args section adds meaningful semantics for every parameter: 'question: Developer question', 'scope: Optional repository-relative path prefix', 'repo: Usually omitted; a workspace alias when needed', and 'include: Optional ["evidence"] expanded projection'. This compensates fully for the absent schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Answer a how, where, or why question in one evidence-grounded call.' This clearly distinguishes the tool from siblings like get_health or get_risk, which target different query types. The phrase 'how, where, or why' anchors exactly when this tool applies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear context on when to use the tool for how/where/why questions, and it excludes general change-risk/health/overview concerns implicitly. It does not explicitly name sibling alternatives or state 'when not to use this vs. X', but the question-type framing is strong enough to guide selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_change_riskA

Review a commit, base..head range, or uncommitted work.

Leads with ``directive`` (what to do) and ``health_delta`` (what this
change newly made worse). A finding is reported only when the diff explains
it, and each names its ``attribution`` basis; findings the change wrote
sort above pre-existing ones it only touched.

Trust ``health_delta.status``: ``partial`` means files were skipped and the
change is not cleared.

``impacted_tests`` keeps measured coverage and inferred candidates distinct.
``patch_coverage`` is the share of changed executable lines stored coverage
ran (no revspec: from the merge-base); ``hints`` name tests to extend.
``fix_history`` is the changed files' bug-fix record, ``overlap`` the past
fixes on these exact lines. ``branch_overlap`` names other branches editing
them. ``diff_shape`` is one line on size, not a danger verdict. An empty
diff returns ``status: "nothing_to_score"`` and names the tree it read.

Args:
    revspec: Commit or ``base..head`` range. Omit to review uncommitted
        work, or ``HEAD`` when the tree is clean.
    repo: Repository alias in workspace mode; omit for the default.
    extensions: File suffixes to count, e.g. ``[".py", ".ts"]``.
    exclude_patterns: Gitignore-style paths to omit, e.g. ``["tests/"]``.
    baseline: Recent commits sampled for percentile ranking; 0 disables it.
    include: ``"findings"`` for every change finding, ``"diagnostics"`` for
        raw score mechanics, ``"scales"`` for units. All identical on
        repeat, so ask once.
    finding_id: Expand one ``health_delta`` finding by its id.
ParametersJSON Schema
NameRequiredDescriptionDefault
repoNo
includeNo
revspecNo
baselineNo
extensionsNo
finding_idNo
exclude_patternsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it explains the ``health_delta.status: partial`` skip condition, the empty-diff ``nothing_to_score`` case, attribution rules (findings only when the diff explains them), and determinism on repeat calls. It stops short of stating permissions/auth or cost/latency characteristics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is long but front-loaded, opening with the core capability before the dense output-field walkthrough. Nearly every sentence adds information (status semantics, attribution, coverage definition), though the middle field-by-field block is jargon-heavy and could be trimmed without losing agent-relevant signal.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the extensive return-field explanation is a bonus rather than a necessity. Combined with the full parameter coverage and edge-case handling (partial status, empty diff), the definition is complete for correct invocation; only cross-tool routing against get_risk/get_health is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and there are 7 parameters, yet the Args section documents every one: revspec (including the omit/HEAD fallback), repo, extensions with an example, exclude_patterns with an example, baseline with the 0-disables semantic, include's three modes plus idempotence, and finding_id's expansion role. This fully compensates for the empty schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence gives a specific verb (Review) and scopes the resource precisely: a commit, a ``base..head`` range, or uncommitted work. An agent can tell what it analyzes, though it never explicitly contrasts itself with the close sibling ``get_risk`` (whole-repo risk) or ``get_health``, so sibling disambiguation is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied through argument semantics rather than stated: revspec should be omitted for uncommitted work or set to HEAD when the tree is clean, and ``include`` should be asked once because results are deterministic. There is no explicit when-to-use-this-vs-alternative guidance and no exclusions relative to ``get_risk``/``get_health``.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_contextA

Triage card for files / modules / symbols — relationships, not source bytes.

Returns title, summary, signatures with line numbers, hotspot bit, and
decision_record titles. fix_history appears only on files with counted bug
fixes (count, age, bug_magnet); hotspot is churn. Either one is a cue to
call get_risk. episodes counts the dated records bound to a target — what
happened here and why — and appears only when there is at least one;
get_why serves the bodies. A symbol target is counted as its file, and a
module aggregates everything beneath it.
Batch targets in one call. No source bytes by default: pass
include=["skeleton"] for the whole file body-elided and line-verified in
ONE call, or Read it. Do not call get_symbol per signature.

Default responses fit 24,000 serialized chars; nonempty ``include`` uses
32,000. Reductions carry counts and ``_meta.omitted`` recovery refs;
``_meta.recovery_unavailable`` names a storage failure.
Include-gated blocks are projections, not omissions.

Args:
    targets: file paths, module paths, or "path::Symbol" ids.
    include: opt-in blocks: full_doc | ownership | last_change | callers
        | callees | metrics | community | decisions | skeleton | health
        | doc_drift (documents naming this file).
        An unrecognised key is named in ignored_arguments.
    compact: default True; False adds structure+imports+docstrings.
    repo: usually omitted.
ParametersJSON Schema
NameRequiredDescriptionDefault
repoNo
compactNo
includeNo
targetsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses 'No source bytes by default', response size limits ('Default responses fit 24,000 serialized chars; nonempty include uses 32,000'), conditional fields ('fix_history appears only...', 'episodes... appears only when there is at least one'), and error recovery ('_meta.omitted recovery refs', '_meta.recovery_unavailable names a storage failure').

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with a clear purpose, then structured into return-field semantics, usage guidance, size constraints, and an Args section. Each sentence adds a distinct operational fact; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (4 params, many include options) and that an output schema is present, the description covers input semantics, conditional outputs, size limits, and guidance to alternatives. An agent can invoke it correctly without further documentation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must add meaning. It explains each arg: targets are 'file paths, module paths, or path::Symbol ids', include lists allowed blocks, compact toggles extra structure, and repo is 'usually omitted'. It also notes unknown include keys are named in ignored_arguments.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Triage card for files / modules / symbols — relationships, not source bytes' and then enumerates returned fields. It explicitly names sibling tools: 'call get_risk', 'get_why serves the bodies', and warns 'Do not call get_symbol per signature', distinguishing it from related tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states 'Either one is a cue to call get_risk' and 'get_why serves the bodies' to route to alternatives. It tells the agent to 'Batch targets in one call' and not to call get_symbol per signature. It also advises passing include=['skeleton'] for body instead of separate read.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_dead_codeA

Unused exports, unreachable files, zombie packages — tiered by confidence.

Run before a cleanup sprint, not a targeted fix. Findings tier
high/medium/low with per-directory and per-owner rollups; workspace
mode lowers confidence on findings other repos import.

Args:
    repo: usually omitted.
    kind: unreachable_file | unused_export | unused_internal | zombie_package.
        An unrecognised value is dropped and named in ignored_arguments,
        never applied as a filter that matches nothing.
    min_confidence: floor, default 0.4 (0.7 = cleanup-ready only). Also
        accepts a tier name: "high" | "medium" | "low".
    safe_only: deletion-ready findings only (no runtime-load risk).
    limit: max findings per tier (clamped to 25).
    tier: "high" | "medium" | "low", banded as min_confidence.
    directory: path-prefix filter.
    owner: primary-owner filter.
    group_by: "directory" | "owner" rollup.
    include_internals: also scan private symbols (more false positives).
    include_zombie_packages: monorepo package findings (default true).
    no_unreachable: skip file-level reachability findings.
    no_unused_exports: skip public-export findings.
    finding_id: stable ``id`` emitted by a dead-code finding.
ParametersJSON Schema
NameRequiredDescriptionDefault
kindNo
repoNo
tierNo
limitNo
ownerNo
group_byNo
directoryNo
safe_onlyNo
finding_idNo
min_confidenceNo
no_unreachableNo
include_internalsNo
no_unused_exportsNo
include_zombie_packagesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and delivers: unrecognised kind values are dropped and surfaced in ignored_arguments rather than silently matching nothing, limit is clamped to 25, workspace mode degrades confidence on cross-repo imports, and safe_only restricts to deletion-ready findings with no runtime-load risk. These are exactly the non-obvious behaviors an agent needs before trusting results.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then usage guidance, then a dense args block where every line adds information. Slightly list-heavy with mild overlap between tier and min_confidence, but nothing is wasted and scanning is easy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 14-parameter tool with an output schema (so return shape needn't be explained) and no annotations, the description covers filters, defaults, edge cases, and rollups. The only soft spot is that the read-only nature of the operation is left implicit in the name rather than stated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, and the description documents all 14 parameters anyway, including accepted enum values for kind/tier/group_by, the default 0.4 floor and the 0.7 cleanup-ready threshold, tier-name acceptance for min_confidence, the false-positive cost of include_internals, and the finding_id handoff. This fully compensates for the schema's missing descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening line names the specific resources (unused exports, unreachable files, zombie packages) and the organizing concept (confidence tiering), which distinguishes it immediately from siblings like get_health, get_risk, and get_symbol. An agent can tell this is a dead-code analysis tool without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Run before a cleanup sprint, not a targeted fix" gives a clear when-to-use plus an explicit when-not, and the note about workspace mode lowering confidence adds situational context. It falls short of 5 only because the targeted-fix alternative (e.g. get_symbol or search_codebase) is implied rather than named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_healthA

Code-health scores and findings from stored analysis.

No ``targets`` returns a dashboard; targets rank files and findings.
Never recomputes health: commit, then run ``repowise update``.
Every block and accepted value: docs/agent/MCP_TOOLS.md.

Args:
    targets: file paths or ``module:<name>``; unmatched ones land in
        ``unresolved``.
    include: ``biomarkers``|``refactoring``|``trend``|``coverage``|
        ``accuracy``|``signals``|``churn_complexity``|``doc_drift``,
        or a dimension incl. ``advisory``; ``performance`` and
        ``refactoring`` add queues.
    only: keys to keep; identity, totals, recovery survive.
        ``biomarkers``/``accuracy``/``refactoring`` alias their block key;
        ``performance``/``defect``/``maintainability``/``advisory``
        do not: they filter rows into ``unknown_only_keys``.
    repo: usually omitted.
    limit: max rows per ranked list, ``0`` for none.
    cursor: zero-based offset into a ranked list.
    finding_id/plan_id: stable ``id`` from a finding or plan.
    opportunity_id: ``perf...``/``refop...``: the unit, its steps or
        plan, evidence paged by ``only=["*_evidence"]``.
    refactoring_view: ``diversified`` (default)|``canonical``|
        ``file_spread``; _type/_confidence/_effort filter.
    performance_view/_context/_boundary/_confidence/_actionability/_sort:
        queue filters; a rejected value lists the accepted.
    scope / counts: default ``all``/``everything``. ``production`` drops
        test files; ``code_shape`` drops the git-derived half of the
        score and its findings.
ParametersJSON Schema
NameRequiredDescriptionDefault
onlyNo
repoNo
limitNo
scopeNoall
countsNoeverything
cursorNo
includeNo
plan_idNo
targetsNo
finding_idNo
opportunity_idNo
performance_sortNo
performance_viewNo
refactoring_typeNo
refactoring_viewNodiversified
refactoring_effortNo
performance_contextNo
performance_boundaryNo
performance_confidenceNo
refactoring_confidenceNo
performance_actionabilityNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It discloses key behaviors: unmatched targets land in 'unresolved', the 'only' parameter has special alias handling that can filter rows into 'unknown_only_keys', defaults for views and scope are given, and it explicitly states the tool never recomputes health. These go well beyond basic read-only semantics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but well-organized: a summary line, a usage note, and a parameter block. It front-loads the core purpose. Some repetition (e.g., the only-parameter details) could be trimmed, but each sentence adds meaningful content. It is dense rather than bloated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 21 parameters, no enums, and an output schema present, the description covers all the important operational details: parameter semantics, defaults, edge cases (unmatched targets, unknown only_keys), and data freshness. It even points to documentation for full details. An agent can call this tool correctly without additional lookup.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must explain all parameters, and it does. The Args section details targets, include, only, repo, limit, cursor, finding_id/plan_id, opportunity_id, refactoring_view, performance_* filters, and scope/counts, including default values and side effects (e.g., 'performance'/'defect' alias behavior). This is far more informative than the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Code-health scores and findings from stored analysis,' which is a specific verb+resource. It immediately distinguishes the tool from sibling tools that handle search (search_codebase), answers (get_answer), or specific artifacts (get_symbol, get_dead_code). The two modes (dashboard vs. ranked files) add further precision.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives operational guidance, such as 'Never recomputes health: commit, then run repowise update' and explains the no-targets vs. targets behavior. However, it never explicitly contrasts this tool with siblings or states when to prefer it over them. The guidance is about how to use it, not when to select it versus alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_overviewA

Architecture map for an unfamiliar repo — first call when you don't know your way around.

Returns the synthesised overview summary, key modules, entry points,
architecture layers, code health, repo-wide git health, and
``next_actions`` (top work for the week and quarter).
Skip this on subsequent calls — once you have the map, jump straight to
``get_context`` / ``get_answer``.

Compact by default: ``content_md`` carries only the overview essay's summary
section, and the outline, onboarding, ownership and graph blocks ship only
on request. The response's ``more`` field names them.

Defaults fit 24,000 chars; nonempty ``include`` uses 32,000. Reductions
carry counts and recovery status in ``_meta``.
Include-gated blocks are projections, not omissions.

In workspace mode:
- Omit ``repo`` for the default repo's overview plus a workspace footer.
- ``repo="all"`` returns the cross-repo topology (co-changes, package deps,
  API contracts) — no single-repo detail.
- ``repo="<alias>"`` targets one specific repo.

Args:
    repo: Repository alias, path, or ID. Use ``"all"`` for workspace overview.
    include: Opt-in extras, any combination of:
        ``"content"`` — the full overview essay instead of its summary.
        ``"outline"`` — the stored wiki page tree, two rungs deep.
        ``"tour"`` — ``guided_tour`` + ``reading_order`` onboarding walks.
        ``"decisions"`` — ``key_decisions``; ``get_why`` is richer.
        ``"graph"`` — ``community_summary``, code-community clusters.
        ``"ownership"`` — ``knowledge_map``: top 3 owners by files owned.
ParametersJSON Schema
NameRequiredDescriptionDefault
repoNo
includeNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the full burden — and it delivers unusual detail: default vs. include-gated response size (24,000 / 32,000 chars), that reductions are reported with counts in `_meta`, and the important caveat that include-gated blocks are 'projections, not omissions.' It stops short of stating read-only/permission characteristics or any rate limits, which keeps it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the one-line purpose before any mechanics, and the prose is dense with no filler. It runs long and repeats the compactness/include trade-off in both the body and the Args block, which is minor redundancy rather than waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter, zero-required-parameter read tool with an output schema, the definition covers everything an agent needs: purpose, ordering relative to siblings, response-size behavior, workspace branching, and every include value. The output schema handles return shape, and the description still summarizes it helpfully.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate fully, and it does: `repo` is defined as alias/path/ID with three distinct workspace-mode behaviors, and `include` enumerates all six accepted values with a one-line explanation of what each returns and how it relates to siblings (e.g. 'decisions' vs. get_why).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a concrete metaphor-plus-verb framing ('Architecture map for an unfamiliar repo') that names the resource and its scope precisely, then explicitly positions itself as the first call before get_context/get_answer. An agent can distinguish it from all nine siblings without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit when ('first call when you don't know your way around'), an explicit when-not ('Skip this on subsequent calls'), and names the alternatives to jump to instead ('get_context' / 'get_answer'). Workspace-mode selection is also spelled out for each repo value.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_riskA

What history says about touching these files — bug fixes, churn, owners.

Fuses git temporal signals (``hotspot_score``/``owner_pct`` are 0-1; trend;
bus factor) with graph topology. ``dependents`` are directed structural
reach (source depends on target), ``consumers`` require typed contract links,
and ``co_change_partners`` are historical correlation only. Those counts
are a floor over the indexed graph. Structural reach is not proof of
runtime breakage. The response also includes security
findings. Pass changed_files for PR mode: the response leads with a
directive block (may_break, missing_cochanges, missing_tests,
tests_to_run) — read it first. Each test_recommendations row carries a
measured or inferred basis, and coverage availability is explicit. To
score a commit or ``base..head`` range instead, use ``get_change_risk``.

In PR mode ``structural_impact_score`` is an uncalibrated 0-10 structural
heuristic, never a runtime-breakage probability; ``overall_risk_score`` is
its deprecated exact alias.

Default responses fit 24,000 serialized chars; nonempty ``include`` uses
32,000. Reductions carry counts and ``_meta.omitted`` recovery refs;
``_meta.recovery_unavailable`` names a storage failure.
Include-gated blocks are projections, not omissions.

Args:
    targets: file paths to assess.
    repo: usually omitted.
    changed_files: PR-changed files for blast-radius mode.
    include: opt-in blocks - "graph", "churn", "scales" (units and
        calibration for every scalar; identical per call, so ask once).
ParametersJSON Schema
NameRequiredDescriptionDefault
repoNo
includeNo
targetsYes
changed_filesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so extensively: it discloses response size limits (24,000 default, 32,000 with include), truncation semantics ('Reductions carry counts and _meta.omitted recovery refs; _meta.recovery_unavailable names a storage failure'), that include-gated blocks are projections rather than omissions, and the epistemic caveats ('Structural reach is not proof of runtime breakage'; 'structural_impact_score is an uncalibrated 0-10 structural heuristic, never a runtime-breakage probability'; 'overall_risk_score is its deprecated exact alias'). This is unusually rich behavioral context for a large-payload read tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded well, but the body is dense and repetitive: PR mode is introduced twice, structural_impact_score is discussed in two separate places, and the exact character-budget figures plus recovery-ref mechanics occupy substantial space. Much of it earns its place, but the redundancy and length keep it below the top scores.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a four-parameter, output-schema-backed tool with zero annotation coverage, the description covers inputs, output shape, metadata/recovery semantics, and the meaning of the headline scores. An output schema exists, so it needn't restate return fields, and it appropriately focuses on the epistemic and truncation caveats an agent needs to interpret results correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% (the schema only supplies titles and types), so the description must compensate and does: it explains targets, why repo is 'usually omitted', the role of changed_files in PR mode, and enumerates include values ('graph', 'churn', 'scales' with units/calibration). It stops short of format-level detail (e.g., path conventions for targets), so a 4 rather than a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states the resource and the evidence it fuses ('What history says about touching these files — bug fixes, churn, owners'), which is a concrete, non-tautological purpose. It explicitly distinguishes itself from the sibling get_change_risk ('To score a commit or base..head range instead, use get_change_risk'), so an agent can route between them. The only mild weakness is that the 'verb' is diffuse — the description reads more like a data-fusion contract than a crisp action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It names a concrete trigger ('Pass changed_files for PR mode') and points to the divergent sibling for a different scope (commit or base..head → get_change_risk). It also tells the caller to 'read [the directive block] first' and that 'scale' includes are call-invariant ('ask once'). It does not spell out when NOT to call it or prerequisites, so it falls short of a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_symbolA

Follow-up read of one symbol whose id another response already gave you.

**Not an entry point.** ``get_answer`` already ships ``symbol_bodies``, and
for a whole file ``get_context(include=["skeleton"])`` or a plain Read is
one call instead of many. Reach here for a body that was elided, or for a
``continuation`` / omission ref. Never walk a file symbol by symbol.

Returns verified, line-numbered source for one indexed symbol, live range,
or omission ref. Ambiguity returns every candidate; an index miss returns
live fallback lines. A truncated result carries the exact continuation to
pass straight back.

Args:
    symbol_id: "path/to/file.py::Name", "path/to/file.py:140-180" for a
        live range, or an omission ref.
    context_lines: extra lines before/after (0-50).
    repo: usually omitted.
    query: omission refs only, regex/substring filter on lines.
    id: accepted alias for ``symbol_id``.
    depth: 1 (default) is this symbol alone; 2-3 also returns the bodies
        it calls, transitively, in ``callee_bodies``.
    reference: structured source reference emitted by this tool. Its id
        and repository are accepted together without caller translation.
ParametersJSON Schema
NameRequiredDescriptionDefault
idNo
repoNo
depthNo
queryNo
referenceNo
symbol_idNo
context_linesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It explains return contents, ambiguity resolution, fallback behavior for index misses, truncation/continuation handling, and depth semantics, all beyond what the bare schema shows.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but excellently structured: purpose and exclusions first, return behavior second, and a compact argument list. Every sentence adds value, and the not-an-entry-point warning earns its prominent placement.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 7 optional parameters, no annotations, and subtle modes like omission refs and depth expansion, the description covers use cases, routing, input formats, edge cases, and continuation handling. The output schema exists, so return-structure details need no extra explanation, but the description still provides them where relevant.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must fully compensate, and it does. Every parameter is explained: symbol_id formats, context_lines bounds, repo's usual omission, query's restriction to omission refs, id alias, depth behavior, and reference usage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource: a follow-up read of one symbol whose id was already provided by another response. It explicitly distinguishes itself from get_answer, get_context, and plain Read, so an agent can select it correctly without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description states clearly when to use this tool: for elided bodies, continuations, or omission refs, not as an entry point. It names alternative tools and even warns against walking a file symbol by symbol, giving strong routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_whyA

Why this code is shaped this way — decision records + evidence commits.

Call before refactors or pattern divergences. Query modes: a question
("why is auth using JWT?"), a file path (governing decisions + origin
story + alignment score), a question anchored to targets, or no query
(decision health dashboard). Falls back to git archaeology when no
decisions exist for a path — never empty. Evidence-bearing rows carry an
explicit ``provenance`` and self-contained ``evidence_refs``; matching ids
mean shared evidence, not independent corroboration. Every decision row
carries ``authority``: ``accepted`` means somebody signed it, ``candidate``
means nobody has yet. ``answer_basis`` names the strongest lane the response
rests on (decision, episode, rationale, archaeology, documentation,
candidate); only ``decision`` is a ruling, and ``candidate`` is the weakest
-- it means nothing cleared that bar.

Args:
    query: question, file/module path, or omit for the dashboard.
    targets: optional file paths to anchor the search, or to ask about on
        their own when there is no query.
    repo: usually omitted.
    id: decision or ``ev_...`` evidence id emitted by this or another tool.
    reference: structured evidence reference. Its id and repository are
        accepted together without caller translation.
ParametersJSON Schema
NameRequiredDescriptionDefault
idNo
repoNo
queryNo
targetsNo
referenceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It explicitly discloses fallback behavior ('Falls back to git archaeology when no decisions exist for a path — never empty'), evidence semantics ('matching ids mean shared evidence, not independent corroboration'), and the meaning of authority and answer_basis fields. This is rich, transparent behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every sentence earns its place. It is front-loaded with purpose and usage, followed by behavioral nuances, then a structured Args section. The information density is high without redundancy, making it well-suited for an AI agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (5 optional parameters, no annotations, 0% schema coverage), the description is remarkably complete. It covers all parameters, explains output semantics (authority, answer_basis, provenance), and describes fallback behavior. The presence of an output schema further relieves the need to detail return values, so nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate entirely. It does so admirably, explaining each parameter in the Args block: query (question, path, or omit), targets (anchor or standalone), repo (usually omitted), id (decision or ev_... format), and reference (id and repository accepted together). This goes far beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: it retrieves decision records and evidence commits explaining why code is shaped a certain way. It also gives a specific use case ('Call before refactors or pattern divergences'). However, it does not explicitly differentiate from sibling tools like get_context or get_answer, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear usage context with 'Call before refactors or pattern divergences' and explains the various query modes (question, file path, anchored targets, no query). It does not, however, mention when not to use the tool or explicitly compare to alternatives, so it earns a 4 rather than a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_codebaseA

Find code by concept, symbol, or path — hybrid codebase search.

For QUESTIONS ("how does X work", "where is Y handled", "why is Z like
this"), call get_answer instead: it runs this same hybrid retrieval
internally and synthesizes a cited answer, so searching first is a wasted
round-trip. Use this tool when you want the raw ranked hits themselves —
enumerating matches, resolving an identifier to a symbol_id, or scoping a
later get_context call.

mode="auto" (default) routes the query: identifier-shaped queries search
the indexed symbols (returns symbol_id/file/line bounds — pipe into
get_symbol), path-shaped queries resolve files (pipe into get_context),
and conceptual queries run wiki-semantic search. Mixed queries run hybrid,
symbol hits first. Decision records rank below file pages unless the query
is why-shaped.

`candidates` lists up to `limit` distinct openable file paths, best first.
Some results are pages, not files; this is what to Read.

Args:
    query: identifier, path, or natural-language query.
    limit: max results (default 5).
    page_type: restrict to one page type. Common: file_page (per-file
        docs, always present) or module_page (subsystem/concept pages).
        Any stored type filters (repo_overview, layer_page, scc_page,
        api_contract, infra_page, symbol_spotlight).
    kind: implementation | test | config | doc (concept/symbol modes).
    repo: alias, or "all" for workspace-wide.
    mode: auto | concept | symbol | path | hybrid.
    symbol_kind: filter symbol hits by kind (function|class|method|...).
ParametersJSON Schema
NameRequiredDescriptionDefault
kindNo
modeNoauto
repoNo
limitNo
queryYes
page_typeNo
symbol_kindNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does so thoroughly. It explains mode routing heuristics, decision-record ranking, candidates semantics (up to limit, best first), and that some results are pages rather than files. It also notes that this result set is what to Read, providing context for downstream actions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but highly organized: a one-line summary, a clear usage-orientation paragraph, a mode-behavior paragraph, a candidates paragraph, and a terse bullet-style Arg list. No redundant fluff; every sentence adds value. The structured flow aids scanability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (7 parameters, multiple modes, output schema), the description covers usage boundaries, mode routing, filtering, return semantics, and integration with sibling tools. It is complete enough for an agent to select and invoke the tool correctly without additional context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, but the description's Args section explains every parameter with meaningful details—e.g., page_type gives common valid values, mode enumerates auto|concept|symbol|path|hybrid. It compensates fully for the schema's lack of descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Find code by concept, symbol, or path — hybrid codebase search,' specifying a precise verb and resource. It clearly distinguishes the tool from the sibling get_answer by framing it as the raw-hit retrieval tool rather than a question-answering tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit guidance says to avoid this tool for factual questions ('call get_answer instead... searching first is a wasted round-trip') and directs use when 'you want the raw ranked hits themselves' or need to resolve symbol IDs or scope later get_context calls. This is concrete and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.53.0
    • Changedget_health1 field changed
      • addedInput schema / properties / performance_actionability
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Performance Actionability"
        +}
  2. 1 tool updatev0.50.0
    • Changedget_health2 fields changed
      • addedInput schema / properties / counts
        Added value: +{
        +  "default": "everything",
        +  "title": "Counts",
        +  "type": "string"
        +}
      • addedInput schema / properties / scope
        Added value: +{
        +  "default": "all",
        +  "title": "Scope",
        +  "type": "string"
        +}
  3. 8 tool updatesv0.48.0
    • Changedget_answer1 field changed
      • addedInput schema / properties / include
        Added value: +{
        +  "anyOf": [
        +    {
        +      "items": {
        +        "type": "string"
        +      },
        +      "type": "array"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Include"
        +}
    • Changedget_change_risk2 fields changed
      • addedInput schema / properties / finding_id
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Finding Id"
        +}
      • addedInput schema / properties / include
        Added value: +{
        +  "anyOf": [
        +    {
        +      "items": {
        +        "type": "string"
        +      },
        +      "type": "array"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Include"
        +}
    • Changedget_dead_code1 field changed
      • addedInput schema / properties / finding_id
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Finding Id"
        +}
    • Changedget_health13 fields changed
      • addedInput schema / properties / cursor
        Added value: +{
        +  "default": 0,
        +  "title": "Cursor",
        +  "type": "integer"
        +}
      • addedInput schema / properties / finding_id
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Finding Id"
        +}
      • addedInput schema / properties / opportunity_id
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Opportunity Id"
        +}
      • addedInput schema / properties / performance_boundary
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Performance Boundary"
        +}
      • addedInput schema / properties / performance_confidence
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Performance Confidence"
        +}
      • addedInput schema / properties / performance_context
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Performance Context"
        +}
      • addedInput schema / properties / performance_sort
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Performance Sort"
        +}
      • addedInput schema / properties / performance_view
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Performance View"
        +}
      • addedInput schema / properties / plan_id
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Plan Id"
        +}
      • addedInput schema / properties / refactoring_confidence
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Refactoring Confidence"
        +}
      • addedInput schema / properties / refactoring_effort
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Refactoring Effort"
        +}
      • addedInput schema / properties / refactoring_type
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Refactoring Type"
        +}
      • changedInput schema / properties / refactoring_view / default
        Previous value: -"canonical"New value: +"diversified"
    • Changedget_risk1 field changed
      • addedInput schema / properties / include
        Added value: +{
        +  "anyOf": [
        +    {
        +      "items": {
        +        "type": "string"
        +      },
        +      "type": "array"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Include"
        +}
    • Changedget_symbol1 field changed
      • addedInput schema / properties / reference
        Added value: +{
        +  "anyOf": [
        +    {
        +      "additionalProperties": true,
        +      "type": "object"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Reference"
        +}
    • Changedget_why2 fields changed
      • addedInput schema / properties / id
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Id"
        +}
      • addedInput schema / properties / reference
        Added value: +{
        +  "anyOf": [
        +    {
        +      "additionalProperties": true,
        +      "type": "object"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Reference"
        +}
    • Removedlist_repos
  4. 1 tool updatev0.45.0
    • Changedget_health1 field changed
      • addedInput schema / properties / refactoring_view
        Added value: +{
        +  "default": "canonical",
        +  "title": "Refactoring View",
        +  "type": "string"
        +}
  5. 2 tool updatesv0.44.0
    • Changedget_change_risk3 fields changed
      • addedInput schema / properties / revspec / anyOf
        Added value: +[
        +  {
        +    "type": "string"
        +  },
        +  {
        +    "type": "null"
        +  }
        +]
      • changedInput schema / properties / revspec / default
        Previous value: -"HEAD"New value: +null
      • removedInput schema / properties / revspec / type
        Removed value: -"string"
    • Changedget_dead_code2 fields changed
      • addedInput schema / properties / min_confidence / anyOf
        Added value: +[
        +  {
        +    "type": "number"
        +  },
        +  {
        +    "type": "string"
        +  }
        +]
      • removedInput schema / properties / min_confidence / type
        Removed value: -"number"
  6. 1 tool updatev0.41.0
    • Changedget_symbol1 field changed
      • addedInput schema / properties / depth
        Added value: +{
        +  "default": 1,
        +  "title": "Depth",
        +  "type": "integer"
        +}
  7. 5 tool updatesv0.39.0
    • Addedget_answer
    • Changedget_dead_code1 field changed
      • changedInput schema / properties / min_confidence / default
        Previous value: -0.5New value: +0.4
    • Changedget_health1 field changed
      • addedInput schema / properties / only
        Added value: +{
        +  "anyOf": [
        +    {
        +      "items": {
        +        "type": "string"
        +      },
        +      "type": "array"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Only"
        +}
    • Addedlist_repos
    • Addedsearch_codebase
  8. 5 tool updatesv0.33.0
    • Removedgenerate_refactoring_code
    • Removedget_answer
    • Addedget_change_risk
    • Removedlist_repos
    • Removedsearch_codebase
  9. 1 tool updatev0.31.0
    • Changedget_symbol5 fields changed
      • addedInput schema / properties / id
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Id"
        +}
      • addedInput schema / properties / symbol_id / anyOf
        Added value: +[
        +  {
        +    "type": "string"
        +  },
        +  {
        +    "type": "null"
        +  }
        +]
      • addedInput schema / properties / symbol_id / default
        Added value: +null
      • removedInput schema / properties / symbol_id / type
        Removed value: -"string"
      • removedInput schema / required
        Removed value: -[
        -  "symbol_id"
        -]
  10. 11 tool updatesv0.1.0
    • First observedgenerate_refactoring_code
    • First observedget_answer
    • First observedget_context
    • First observedget_dead_code
    • First observedget_health
    • First observedget_overview
    • First observedget_risk
    • First observedget_symbol
    • First observedget_why
    • First observedlist_repos
    • First observedsearch_codebase

TDQS

A4.4/5.0

Scored across 10 tools

Disambiguation4/5

Each tool has a distinct role, and descriptions go to unusual lengths to cross-reference boundaries (get_answer vs search_codebase vs get_context vs get_symbol, get_risk vs get_change_risk, get_why vs get_overview decisions). The overlap in the "understand this code" space is real but well-documented, so an agent can route correctly with careful reading.

Naming Consistency5/5

Nine tools follow get_<noun> (get_answer, get_health, get_risk, get_change_risk, get_dead_code, get_overview, get_context, get_symbol, get_why) and the tenth, search_codebase, is the same verb_noun pattern. The scheme is fully predictable throughout.

Tool Count5/5

Ten tools sit squarely in the ideal 3-15 range and each maps to a real analysis concern (overview, search, Q&A, context, symbols, health, risk, change risk, dead code, rationale). No tool feels redundant or bolted on.

Completeness4/5

The surface covers the code-intelligence lifecycle well: onboarding (overview), discovery (search), comprehension (answer/context/symbol), history/rationale (why), and quality/risk (health, risk, change risk, dead code). Minor gaps exist — no explicit repo/workspace listing or file-read primitive — but these are largely handled by other tools or external Read.

Maintenance

ActivityActive
ResponsivenessResponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Persistent codebase knowledge layer for AI agents. Pre-digests codebases into structured knowledge (symbols, dependency graphs, co-change patterns, architectural decisions) and serves via MCP. 28 languages, 14 tools, ~85% token reduction.
    7 npm
    8
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI agents to search code by meaning, explore codebase structure, store and query knowledge with temporal facts, and read source code through a set of MCP tools.
    310 npm
    7
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Provides code intelligence for AI coding agents by indexing repositories into a hybrid knowledge graph, enabling agents to query dependencies, impact, and context through 28 MCP tools.
    3
    Apache 2.0
  • A
    license
    Not graded
    quality
    A
    maintenance
    Provides AI agents with a function-level dependency graph of the codebase through 30 MCP tools, enabling structural queries about code dependencies, callers, and impact analysis.
    1,863 npm
    97
    Apache 2.0