Skip to main content
Glama

Cartograph

エージェントネイティブなコードインテリジェンス。 あらゆるリポジトリをクエリ可能なコードグラフに変換し、MCP を通じてコーディングエージェントに提供します。これにより、エージェントは*「これを変更したら何が壊れる?」*と尋ねることができ、grepして推測する代わりに済みます。

tree-sitter + SQLite。埋め込みなし、ベクターストアなし、APIキーなし、サーバーなし、コストなし。

→ ライブデモ — このリポジトリの実インデックスからプッシュのたびに生成されています。

CI Python 3.11+ License MIT


問題

コーディングエージェントに、見慣れない大きなリポジトリを与えて、その行動を観察してみてください:grepをし、ファイルを読み、またgrepをし、別のファイルを読む。パーサーなら一度の呼び出しで教えてくれる構造を再構築するためにコンテキストを消費し、それでも、変更が壊す3モジュール先の呼び出し元を見逃してしまいます。

通常の解決策はRAGです:コードベースを埋め込み、「類似」チャンクを取得します。しかし*「この関数を呼んでいるのは誰?」*は類似性の問題ではありません。それには正確な答えがあり、その答えはコールグラフにあります。

Cartographはグラフを構築し、エージェントが実際に動作する方法に合わせた10のツールを提供します。

$ cartograph blast src/cartograph/graph/store.py

## Blast radius — file `src/cartograph/graph/store.py`

17 dependent file(s), 31 affected symbol(s), 7 test file(s).

**Tests to run first**
- `tests/test_cli.py`
- `tests/test_docs.py`
- `tests/test_incremental.py`
- `tests/test_mcp.py`
- `tests/test_resolver.py`
- `tests/test_traversal.py`
- `tests/test_views.py`

**Dependent files** (by import distance)
- `src/cartograph/graph/resolver.py` · d1
- `src/cartograph/indexer/pipeline.py` · d1
- `src/cartograph/service.py` · d1
- `src/cartograph/cli.py` · d2
…

編集前の1回の呼び出し。テストスイートが赤くなった後の7回のgrepではありません。


Related MCP server: codeweave-mcp

クイックスタート

uv tool install cartograph-mcp     # or: pipx install cartograph-mcp

cartograph index ~/code/my-repo    # builds .cartograph/cartograph.db
cartograph arch                    # modules, layers, cycles, hotspots
cartograph blast src/auth/token.py # what a change here could break
cartograph callers validate_token  # reverse call tree

エージェントに組み込む

Claude Code:

claude mcp add cartograph -- cartograph serve /path/to/repo

または任意のMCPクライアントから、mcp.jsonを使用:

{
  "mcpServers": {
    "cartograph": {
      "command": "cartograph",
      "args": ["serve", "/path/to/repo"]
    }
  }
}

serveは初回実行時にインデックスが存在しなければインデックスを作成します。その後、エージェントに*「トークンバリデータを変更したら何が壊れる?」*と尋ねると、推測する代わりにblast_radiusを呼び出します。


10のツール

ツール

回答

find_symbol

Xはどこで定義されているか?(構造的重要度でランク付け)

search_code

名前、シグネチャ、docstringに対する全文検索(BM25)

get_symbol

1つのシンボル:シグネチャ、ドキュメント、メンバー、呼び出し元、呼び出し先、ソース

who_calls

逆コールツリー — シグネチャを変更する前に

what_it_calls

順方向コールツリー — すべてのファイルを読まずにコードを理解する

blast_radius

変更が壊す可能性があるもの、そして実行すべきテスト

related_symbols

「他に何を読むべきか?」をパーソナライズされたPageRankで

file_summary

ファイルが定義しているもの、インポートしているもの、そして誰がそれをインポートしているか

architecture_overview

モジュール、レイヤリング、インポートサイクル、ホットスポット、エントリポイント

index_stats

インデックスの健全性とルール別のエッジ解決の内訳

さらに、MCPリソース(cartograph://architecture、cartograph://stats)と、馴染みのないリポジトリに対するグラフ優先の最初のパスのためのorientプロンプトがあります。

対応言語: Python、TypeScript、TSX、JavaScript、Go。


議論に値する設計判断

1. 信頼度は第一級のカラム

型チェッカーがなければ、store.who_calls()がGraphStore.who_callsを意味することを知ることはできません。仮説をランク付けすることしかできません。だから、そのふりをする代わりに、すべてのエッジはそれを生成したルールと信頼度を記録します:

ルール

信頼度

直感

same-file

0.95

定義が同じスコープ内にある

import

0.90

ファイルがこの名前を明示的にインポートしている

receiver-type

0.85

Foo.bar() で Foo が既知のコンテナである

same-module

0.75

同じパッケージ内の兄弟ファイル

unique-global

0.60

リポジトリ内でこの名前を持つシンボルがただ1つあり、修飾なしの呼び出し

name-only

0.45

1件マッチするが、型付けされていないレシーバに対するもの

ambiguous

≤0.40

N個の候補があり、それぞれ1/Nの重みでN個のエッジとして保持

external

0.00

サードパーティまたは標準ライブラリのインポートに由来

unresolved

0.00

本当に不明(動的、または型付きメソッド)

呼び出し元(ツールの利用者)は、自分自身の運用ポイントを選びます。who_callsのデフォルトは≥0.5 — 精度優先。エージェントは答えに基づいて行動するからです。blast_radiusは0.3まで下げます — 再現率優先。影響を受けるテストを見逃すことが高くつくミスであり、誤検知はレビュアーが一目見るだけのコストだからです。

name-only層は実際のバグが理由で存在します。組み込みsetに対するseen.add(...)が、単に名前がたまたま一意だったというだけの理由で、リポジトリ内のクラスのaddメソッドに解決され、それが自信ありの呼び出し元として表示されました。型を付けられないレシーバ上のメソッド名は証拠にならないため、現在は精度ラインの下に置かれています。(テスト)

external層は、メトリクスに関する正直さのために存在します。ほとんどのリポジトリでは、「unresolved」バケットはtyper.Optionとsqlite3.executeが大半を占めます。これらをまとめると、カバレッジが実際よりもはるかに悪く見えるため、Cartographは内部解決率を報告します — リポジトリのシンボルに到達し得る呼び出しサイトのうち、実際に到達した割合です。

2. パースはインクリメンタル、解決は決してインクリメンタルではない

ファイルのsha256が変わったときだけ再パースされます。しかし、生の参照はrefsテーブルに事実として保存され、edgesは何かが変更されるたびに(refs × symbols)の純粋関数として再計算されます。

これこそが「編集のたびに再インデックス」を信頼できるものにしています。解決もインクリメンタルだったら、1つのファイルを編集したときに、別のファイルのエッジが移動したシンボルを指したままになる可能性があります。グローバルな再解決は、それを構造的に不可能にします。(テスト)

コストは現実のものなので、安全なショートカットは正確に1つだけあります。ファイルが追加・再パース・削除されていない場合、両方の入力テーブルは不変であり、解決は証明可能なほど同一であるため、スキップされます。これにより、Djangoのno-op再インデックスが、バイト単位で同一のグラフで7.5秒から0.67秒に短縮されました。

3. 埋め込みの代わりにPageRank

「どのgetのこと?」は構造的な質問です。40の呼び出しサイトが依存しているgetこそがエージェントが望むものであり、コールグラフはすでにそれを知っています。したがって、シンボルのランキングはコールグラフ上の重み付きPageRankです — 安定していて、説明可能で、無料です。モデルも、インデックス構築も、ベクターストアもありません。

related_symbolsは同じ考え方を拡張します。1つのシンボルをシードにしたパーソナライズドPageRankで、グラフを無向として扱います。関数を変更しようとしているとき、その呼び出し元と呼び出し先の両方が関連するコンテキストだからです。これはセマンティック検索の構造的類似物であり、埋め込みは不要です。

4. ツールはJSONではなくMarkdownを、トークン予算の下で返す

消費者はコンテキストウィンドウです。40シンボルのJSON配列は、ブラケットや繰り返されるキーに何千ものトークンを費やし、モデルはそれを結局再フォーマットします。ここにあるすべてのビューは、ハードなトークン予算を持つコンパクトなMarkdownです。

重要なのは、すべての切り詰めが明示されることです。87件の呼び出し元のうち20件だけをマーカーなしで渡されたエージェントは、残りの67件は存在しないと自信を持って結論づけ、何かを削除してしまいます。

5. トラバーサルはPythonではなくSQLiteで実行される

深さ4のwho_callsは再帰CTEなので、トラバーサル全体がSQLiteのCループ内に留まります。Djangoの252kエッジのグラフでは約5msです。エッジテーブルをPythonに読み込んで走査するのでは、そうはいきません。


ベンチマーク

実際のリポジトリ、Mシリーズのラップトップ、シングルプロセス。Cold = ゼロからの完全インデックス、Warm = no-op再インデックス。

リポジトリ

ファイル数

KLOC

シンボル数

エッジ数

Cold

Warm

DB

内部解決率

django

2,973

534

45,394

252,441

11.9s

0.67s

80 MB

83.2%

gin (Go)

98

24

1,610

9,179

0.32s

0.03s

2.5 MB

88.1%

flask

83

18

1,624

4,271

0.21s

0.03s

1.7 MB

87.4%

クエリレイテンシ(中央値、5回、ウォーム):

リポジトリ

find_symbol

who_calls d3

blast_radius

architecture_overview

django

12.3ms

5.1ms

5.6ms

68.5ms

gin

0.4ms

0.4ms

0.5ms

1.2ms

flask

0.5ms

1.1ms

1.3ms

1.8ms

scripts/bench.pyで再現できます。


アーキテクチャ

flowchart LR
  subgraph index["cartograph index"]
    W[walker<br/>git ls-files] --> P[tree-sitter<br/>+ .scm queries]
    P --> X[extract<br/>defs · refs · imports]
  end
  X --> DB[(SQLite<br/>symbols · refs<br/>edges · FTS5)]
  DB --> R[resolver<br/>rule cascade]
  R --> DB
  DB --> RK[PageRank<br/>Tarjan SCC]
  RK --> DB
  DB --> S[service facade]
  S --> V[views<br/>token-budgeted MD]
  V --> M[MCP server<br/>10 tools]
  V --> C[CLI]
  M --> A((coding agent))

モジュール

責務

indexer/walker.py

ファイル検出 — 正しい.gitignoreセマンティクスのためにgit ls-filesに委ねる

indexer/languages.py

言語ごとに1つのアダプタ:拡張子、クエリ、docstring、モジュールキー、インポート解決

indexer/extract.py

AST → シンボル/参照/インポート、言語非依存

queries/*.scm

tree-sitterのキャプチャパターン — 言語ごとの知識をデータとして

graph/schema.sql

グラフ:files、symbols、refs、edges、imports、FTS5

graph/resolver.py

信頼度カスケード

graph/algorithms.py

PageRank、パーソナライズドPageRank、反復的Tarjan SCC、レイヤリング

graph/store.py

再帰CTEトラバーサル、ランク付きルックアップ、集計

service.py

CLIとMCPサーバーが乖離しないための単一ファサード

views.py

トークン予算付きMarkdown

組み合わせ爆発するクエリを使わないスコープ解決

queries/*.scmを小さく保つ秘訣は、スコープをクエリに一切エンコードしないことです。キャプチャされたすべての定義は、そのtree-sitterノードIDでインデックス化され、参照の囲むシンボルはparentチェーンを辿ってシンボルに当たるまで探索することで見つかります。これは参照ごとにO(ツリー深さ)であり、クロージャ、メソッド、内部クラス、アロー関数を追加のパターンなしで自然に処理します。

言語の追加

LanguageAdapterをサブクラス化し(約40行)、.scmファイルを置きます。GoAdapterが最短の完全な例です。tests/test_queries.pyが、あなたのクエリを文法に対して自動的にコンパイルし、何かをキャプチャすることを検証します。


開発

git clone https://github.com/GokulRaj2210/cartograph-mcp && cd cartograph-mcp
uv sync
uv run pytest -q          # 209 tests
uv run ruff check .
uv run mypy               # strict

CIはPython 3.11/3.12/3.13(さらにmacOS)でスイートを実行し、その後ドッグフーディングします。つまり、このリポジトリをインデックス化し、インポートサイクルで失敗し、no-op再インデックスが何も再パースしないことを検証し、実際のstdio経由でMCPサーバーを駆動します。また、ビルドしたwheelをクリーンなvenvにインストールしてインデックス化します。パッケージ化された.scmファイルはwheelから漏れやすいが、ローカルでは気づきにくいからです。

サイクルゲートはすでにその価値を証明しています。このリポジトリで私が導入したstore → resolver → storeサイクルを検出し、ゲートを緩めるのではなく、問題のヘルパーを移動することで修正されました。

注目すべきテスト

  • tests/test_queries.py — すべての .scm が、それを読み込むすべての文法に対してコンパイルされ、何かをキャプチャすることを検証します。JavaScriptでは有効なパターン ((class_heritage (identifier))) は、TypeScriptではスーパータイプを extends_clause でラップするため、不可能なパターンです。その1行は、TypeScriptのシンボルを黙ってゼロ生成していました。

  • tests/test_incremental.py — 編集、削除、またはシンボルがファイル間を移動した後も、古いエッジが残らないことを検証します。

  • tests/test_resolver.py — すべてのルールが発動し、どのルールも信頼度を過大に主張しないことを検証します。

  • tests/test_cli.py — リーダーとインデクサーが同時にデータベースを保持できることを検証します。

  • tests/test_docs.py — 生成されたデモページがタグの釣り合いが取れた整形式HTMLであることを検証します。これにより、Markdownレンダラーの min_confidence に関するタグ交差バグが発見されました。


制限事項

率直に言えば、精度を過大に宣伝するコードインテリジェンスツールは、役に立たないどころか有害だからです:

  • 型推論なし。 self.conn.execute(...) は conn の型を知らなければリポジトリのシンボルに解決できません。これらは unresolved に入り、内部解決率~85%で残る大部分を占めます。

  • 動的ディスパッチは見えない。 getattr(obj, name)()、デコレータレジストリ、DIコンテナはエッジとして現れません。

  • 言語間エッジは追跡されない。 TypeScriptフロントエンドがPythonエンドポイントを呼び出す場合、それらは2つの独立したサブグラフになります。

  • 定義のみであり、すべての参照ではない。 値として使用されるシンボル(コールバックとして渡される)は、呼び出されるシンボルよりもグラフ内で弱い存在です。

ロードマップ:RustおよびJavaアダプター、言語サーバーが利用可能な場合に正確な解決を可能にするオプションのLSP拡張、およびPR規模の影響範囲を対象とした --changed-since <ref> モード。


なぜこれが存在するのか

大規模リポジトリにおけるコーディングエージェントの最大の弱点—コードの構造モデルが無いこと—が、より大きなモデルやベクターデータベースではなく、静的解析と適切に設計されたツールサーフェスで修正できるかどうかを知りたかったのです。ほとんど、それは可能です。

ライセンス

MIT

Available Tools

10 tools
architecture_overviewA

Orient yourself in an unfamiliar repo: modules, layers, cycles, hotspots.

Start here. One call replaces a dozen exploratory file reads: you get module sizes and layering, import cycles, the highest-PageRank symbols (the risky ones to change) and the repo's entry points.

ParametersJSON Schema
NameRequiredDescriptionDefault
include_diagramNoInclude a Mermaid diagram of the module graph

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the safety/behavior burden. It discloses what the call produces and signals efficiency by replacing 'a dozen exploratory file reads', making the operation's analytic, non-mutating nature clear through the 'you get...' framing. It stops short of stating any performance or read-only caveats explicitly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences with no filler; the purpose is front-loaded and the supporting details (what it returns) are listed compactly. Each clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-optional-parameter tool with an output schema, the description covers the key contextual information: when to use it, what to expect, and why it is valuable. Nothing critical is missing for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the single include_diagram parameter is fully documented in the schema. The description adds no parameter-specific guidance beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource ('Orient yourself in an unfamiliar repo') and enumerates concrete outputs (module sizes/layering, import cycles, PageRank hotspots, entry points). It clearly differentiates from symbol-level siblings like find_symbol and who_calls by positioning itself as the repo-level starting point.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Start here' and 'one call replaces a dozen exploratory file reads' provide explicit context for when to use it: early exploration of an unfamiliar codebase. It does not explicitly state when not to use it or name an alternative, so it misses the full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

blast_radiusA

Impact analysis: what a change here could break, and which tests to run.

Combines the reverse import graph with the reverse call graph, then highlights test files specifically. Recall-first by design (confidence >=0.3): the expensive mistake is a missed impacted test, not an extra one.

Call this before editing shared code and after finishing, to pick tests.

ParametersJSON Schema
NameRequiredDescriptionDefault
depthNoTransitive import/call depth
limitNoMax results
targetYesA file path or a symbol name/qualname

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It explains the internal approach (combining reverse import graph with reverse call graph), the recall-first bias with a specific confidence threshold of >=0.3, and the rationale that missed impacted tests are worse than extra ones. It does not explicitly state that the operation is read-only or safe, but the impact-analysis framing implies it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and front-loaded: the purpose is in the first sentence, methodology and behavior in the second, and usage guidance in the final sentence. Every sentence adds distinct value with no repetition or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema is present and the parameter schema fully describes the inputs, the description provides the necessary context: what the tool computes, how it prioritizes recall, what it highlights, and when to call it. An agent has enough to invoke it correctly and interpret its role relative to siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already documents target, depth, and limit with meaningful descriptions. The tool description adds no parameter-specific guidance beyond the schema, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a clear purpose: 'Impact analysis: what a change here could break, and which tests to run.' It also differentiates itself from siblings by explaining it combines the reverse import graph with the reverse call graph and specifically highlights test files, which sets it apart from who_calls, what_it_calls, and related_symbols.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit context for when to use the tool: 'Call this before editing shared code and after finishing, to pick tests.' It does not explicitly name alternatives or state when not to use it, but the workflow guidance is clear and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

file_summaryA

Outline of one file: what it defines, what it imports, who imports it.

Cheaper than reading the file when you only need to know whether it is relevant, and it adds the reverse-import view that reading cannot give you.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesFile path, or any distinctive part of one

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of explaining behavior. It discloses what the outline contains (definitions, imports, importers) and notes that it is cheaper than full file reading. It does not discuss edge cases like partial paths, errors, or cache behavior, but for a simple summary tool this is reasonably transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences put the core purpose first and the cost/use-case benefit second. Every sentence earns its place; there is no filler or redundant restating of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with an output schema, the description explains what the result contains and why one would choose this tool. It could be slightly stronger about how this compares to adjacent sibling tools, but nothing essential is missing for a basic invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the sole parameter is already well described. The description adds no new parameter-level detail, which is acceptable since the schema fully documents the path parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool's scope: an outline of one file covering definitions, imports, and reverse-imports. This distinguishes it from generic search or symbol tools by naming the specific resource and output aspects.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly frames the tool as a cheaper alternative to reading a file when only relevance matters, and highlights the reverse-import advantage. It does not name sibling tools or provide explicit when-not-to-use guidance, but the intended scenario is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_symbolA

Locate where a symbol is DEFINED, with its file:line, signature and doc.

This is the right first call for "where is X?" -- it is exact and ranked by structural importance, so if a repo has six functions called run, the one the codebase actually revolves around comes first.

Use search_code instead when you only know roughly what the thing does ("the retry logic") rather than what it is called.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNoFilter by kind: function, method, class, interface, struct, enum, type, const
langNoFilter by language: python, typescript, tsx, javascript, go
nameYesSymbol name or qualified name, exact or partial
limitNoMax results

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the burden. It discloses a key behavioral trait: results are 'ranked by structural importance', illustrated with the six-run-functions example. It also mentions exactness and the output shape, though it does not discuss limitations like auth or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no filler: the main purpose is front-loaded, the ranking behavior is immediately explained, and the alternative tool condition is given once. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists and the parameter schema is fully documented, so the description need not restate return types or parameter details. It supplies the missing context: when to use, how results are ranked, and when to switch to search_code, making it complete for an AI agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters. The description adds little beyond contextual emphasis on exactness and ranking; it does not deepen meaning for kind, lang, name, or limit beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Locate where a symbol is DEFINED', with concrete outputs (file:line, signature, doc). It also distinguishes from the sibling search_code by positioning itself as the exact lookup for known symbol names, so an agent can tell when to use it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says this is the right first call for 'where is X?' and names the alternative: use search_code when you only know roughly what the thing does. This gives clear selection criteria without the agent needing to infer.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_symbolA

Full detail for one symbol: signature, doc, members, callers and callees.

Prefer this over reading the whole file: you get the definition plus its immediate graph neighbourhood, which is usually all the context needed to make a safe edit. Set include_source=true when you intend to modify it.

ParametersJSON Schema
NameRequiredDescriptionDefault
depthNoCaller/callee depth to include
symbolYesSymbol id, qualified name (`module:Class.method`), `path:name`, or bare name
include_sourceNoInclude the full source text of the definition

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the burden of behavioral disclosure. It usefully states the returned artifacts (definition plus graph neighborhood) and the include_source toggle, but does not explain depth behavior, error cases, or cost of deep traversal. This is adequate but not deeply transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly packed paragraphs with no filler. The core purpose is in the first sentence, and the practical guidance follows immediately. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists and parameter docs are complete, the description covers the essential context: what the tool returns, why to prefer it, and when to enable source. It doesn't cover depth semantics or error behavior, but those are partially covered in the schema and are minor for a read-only lookup tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaningful usage semantics for include_source ('when you intend to modify it') that goes beyond the schema, and the symbol parameter's accepted forms are already well documented in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Full detail for one symbol' with concrete contents (signature, doc, members, callers, callees). This clearly differentiates get_symbol from siblings like search_code, who_calls, and what_it_calls by scoping it to a single symbol's combined context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit usage guidance: prefer this over reading the whole file, and set include_source=true when you intend to modify the symbol. It does not explicitly name all sibling alternatives or when those would be better, but the 'prefer this over...' framing gives clear decision context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

index_statsA

Index health: size, coverage, and the edge-resolution breakdown by rule.

Worth a call when graph answers look thin -- a low resolution rate or a stale indexed_at tells you the index needs rebuilding rather than the code being unusual.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of behavioral disclosure. It explains what the tool reports, including size, coverage, resolution breakdown, and indexed_at, and adds diagnostic meaning beyond a simple field list. It does not explicitly state that the tool is read-only, but for a stats tool this is strongly implied by the content described.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: the first sentence states exactly what the tool reports, and the second sentence gives actionable usage guidance. Every sentence earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter tool with an output schema available, the description provides everything needed to decide when and how to use it. It explains the tool's purpose, the data it returns, and the diagnostic scenario in which it is useful, leaving no meaningful gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there are no parameter meanings to clarify. The description still adds conceptual value by naming the key output dimensions (size, coverage, edge-resolution breakdown, indexed_at), which is appropriate for a parameterless tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the resource and content ('Index health: size, coverage, and the edge-resolution breakdown by rule'), which immediately distinguishes it from the symbol-focused sibling tools. However, it lacks an explicit verb like 'reports' or 'returns', so it falls just short of the strongest purpose clarity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives an explicit trigger: 'Worth a call when graph answers look thin'. It also explains how to interpret results ('low resolution rate or a stale indexed_at tells you the index needs rebuilding rather than the code being unusual'), which is excellent practical guidance for when this tool is the right choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_codeA

Full-text search across symbol names, signatures and docstrings (BM25).

Use when you know the intent but not the identifier. Results are re-ranked by call-graph importance, so central symbols outrank incidental mentions.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax results
queryYesFree-text query over names, signatures and docstrings

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are absent, so the description carries full behavioral disclosure. It discloses that search uses BM25 and that results are re-ranked by call-graph importance, which is valuable non-obvious behavior. It could mention pagination or query-syntax details, but the core operation and ordering semantics are transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no filler: the first defines scope, the second states when to use it, and the third explains ranking behavior. Every sentence earns its place, and the key use-case guidance is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema and only two straightforward parameters, the description is nearly complete. It covers the tool's purpose, use case, searchable content, and result ordering. It does not explicitly state exclusions or name the exact-identifier sibling, but the sibling context and 'not the identifier' phrasing make the intended boundary clear.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema describes both parameters completely, so the baseline is 3. The description reinforces that `query` is free-text and explains why certain matches outrank others, but it does not add per-parameter syntax or formatting detail beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb ('Full-text search') and a precise resource scope ('symbol names, signatures and docstrings'). The phrase 'Use when you know the intent but not the identifier' clearly distinguishes it from exact-identifier lookup tools such as find_symbol.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit usage condition: use it when the intent is known but the identifier is not. It does not name the alternative tool directly, but the contrast with exact-lookup siblings is strongly implied by the wording and the sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

what_it_callsA

Forward call tree: what this symbol depends on, transitively.

Use it to understand an unfamiliar function without reading every file it touches, and to spot the layer a piece of code really sits in.

ParametersJSON Schema
NameRequiredDescriptionDefault
depthNoTransitive callee depth
limitNo
symbolYesSource symbol (name, qualname or id)
min_confidenceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It discloses the transitive, graph-walking nature of the tool, but does not mention performance characteristics, result size limits, or other runtime behavior beyond what the schema hints at.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded, with the core definition in the first sentence and practical guidance in the second. No filler or redundant restatement of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description provides enough context for an agent to understand the tool's purpose and basic invocation. Some gaps remain around parameter semantics and explicit sibling differentiation, but the output schema and schema constraints partially fill those gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 50%: symbol and depth are documented, but limit and min_confidence lack descriptions. The tool description does not compensate by explaining these parameters or clarifying their units/purpose.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: build a forward call tree of what a symbol transitively depends on. This distinguishes it from reverse-call tools like who_calls, though it does not explicitly name siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides concrete use cases: understanding an unfamiliar function without reading every file, and identifying the layer a piece of code sits in. It gives clear context but does not state when to prefer an alternative tool or when not to use this one.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

who_callsA

Reverse call tree: everything that reaches this symbol, transitively.

The tool to use before changing a signature, tightening a validation, or deleting anything. Each edge reports the rule that produced it; treat sub-0.5 edges as leads rather than facts.

ParametersJSON Schema
NameRequiredDescriptionDefault
depthNoTransitive caller depth
limitNoMax results
symbolYesTarget symbol (name, qualname or id)
min_confidenceNoMinimum edge confidence (0.5 = precision-first)

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It discloses that each edge reports the rule that produced it and warns that sub-0.5 edges are leads rather than facts. It does not discuss cost or traversal size, but the output schema covers result shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three compact sentences deliver the definition, the trigger scenario, and the confidence caveat. The description is front-loaded with the core purpose and every sentence adds distinct value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only analysis tool with a full output schema and fully documented parameters, the description covers what the tool computes, when to use it, and how to interpret weak results. Nothing essential is missing for selecting and invoking it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

All four parameters already have schema descriptions, so the baseline is 3. The description adds meaningful semantics for min_confidence, explicitly saying sub-0.5 edges should be treated as leads, and implies that depth and limit control transitive expansion.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action and resource: 'Reverse call tree: everything that reaches this symbol, transitively.' This clearly distinguishes it from forward-call tools like what_it_calls without needing extra inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives concrete guidance on when to use the tool: 'The tool to use before changing a signature, tightening a validation, or deleting anything.' It does not explicitly list exclusions or alternatives, but the use-case framing is clear and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 10 tool updatesv0.1.0
    • First observedarchitecture_overview
    • First observedblast_radius
    • First observedfile_summary
    • First observedfind_symbol
    • First observedget_symbol
    • First observedindex_stats
    • First observedrelated_symbols
    • First observedsearch_code
    • First observedwhat_it_calls
    • First observedwho_calls

TDQS

A4/5.0

Scored across 10 tools

Disambiguation4/5

Tool purposes are largely distinct and descriptions explicitly route agents to the right one, but find_symbol/get_symbol and who_calls/blast_radius have adjacent responsibilities that could occasionally cause misselection. Overall, the overlap is minor and well-documented.

Naming Consistency3/5

All names are readable snake_case, but the set mixes verb-object names (find_symbol, search_code, get_symbol), question-style names (who_calls, what_it_calls), and noun-phrase names (blast_radius, file_summary, architecture_overview). This is not chaotic, but it lacks a single consistent naming pattern.

Tool Count5/5

Ten tools is a well-scoped surface for a code-graph analysis server. Each tool addresses a distinct job—search, symbol detail, call trees, impact analysis, overview, index health—without redundancy or bloat.

Completeness4/5

The toolchain covers symbol discovery, detailed lookup, dependency analysis, impact assessment, file outlining, architecture orientation, and index health, giving strong coverage of the code-understanding workflow. Minor gaps like direct raw-file access or listing all symbols in a file must be worked around via file_summary and get_symbol.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

  • Code intelligence platform for AI agents. 20 tools for architecture, security & impact analysis.

  • Codebase graphs, caller impact analysis, and recorded project context for AI coding agents.

  • The Cortex MCP server provides read-only access to real-time engineering context from the Cortex developer portal, allowing AI coding assistants to answer natural language questions about your organization's catalog (microservices, libraries, domains, teams, infrastructure), scorecards (engineering standards and best practices), initiatives (goals and deadlines), and Engineering Intelligence metrics. It includes tools for querying documentation, tracking personal entities, and accessing AI-assisted insights across the entire Cortex ecosystem.

  • Coding agents in multi-service codebases routinely rebuild existing helpers, trust stale type definitions, and modify API contracts without knowing who consumes them. Carrick solves this by indexing your entire TypeScript ecosystem across service and repository boundaries. By integrating deeply with the TypeScript compiler, Carrick traces every route, type, and cross-service call while recording function behaviour so agents search by intent rather than name. Delivered via MCP for AI agents and LSP for IDEs, Carrick ensures models see existing endpoints and utilities before generating new code. The scanner is source-available and runs from your CLI or CI pipeline.

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    RepoNova is an MCP server that builds a persistent knowledge graph of your codebase, enabling AI agents to query code structure, dependencies, and semantics through 11 specialized tools.
    174 npm
    7
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    An MCP server that gives AI agents structured code understanding and precise code intelligence via local indexing of AST, call graphs, and semantic search.
    147 npm
    4
    Apache 2.0
  • A
    license
    A
    quality
    B
    maintenance
    An MCP server that generates ranked, token-budgeted code structure maps using Tree-sitter AST analysis and PageRank, enabling AI agents to quickly understand unfamiliar codebases.
    2
    20 npm
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    MCP server for local-first code intelligence, providing structural code graph, semantic search, and impact analysis to AI agents.
    2
    MIT