Cartograph
Cartograph
面向智能体的代码智能。 将任意仓库变成可查询的代码图,并通过 MCP 提供给编码智能体——这样智能体就可以问*“如果我改了这个,会破坏什么?”*,而不是靠 grep 和猜测。
tree-sitter + SQLite。无需嵌入、无需向量存储、无需 API 密钥、无需服务器、零成本。
→ 在线演示 — 每次推送时都会基于此仓库的真实索引生成。
问题所在
给编码智能体一个陌生的大型仓库,看它会做什么:grep、读文件、再 grep、再读另一个文件。它消耗大量上下文来重建结构,而解析器本来只需一次调用就能告诉它——而且它仍然会漏掉在三个模块之外、因它的改动而被破坏的调用方。
通常的解决方案是 RAG:对代码库做嵌入,检索“相似”的代码块。但*“谁调用了这个函数?”*并不是一个相似性问题。它有精确的答案,而这个答案就存在于调用图中。
Cartograph 构建这个图,然后交给智能体十个贴合它们实际工作方式的工具。
$ cartograph blast src/cartograph/graph/store.py
## Blast radius — file `src/cartograph/graph/store.py`
17 dependent file(s), 31 affected symbol(s), 7 test file(s).
**Tests to run first**
- `tests/test_cli.py`
- `tests/test_docs.py`
- `tests/test_incremental.py`
- `tests/test_mcp.py`
- `tests/test_resolver.py`
- `tests/test_traversal.py`
- `tests/test_views.py`
**Dependent files** (by import distance)
- `src/cartograph/graph/resolver.py` · d1
- `src/cartograph/indexer/pipeline.py` · d1
- `src/cartograph/service.py` · d1
- `src/cartograph/cli.py` · d2
…一次调用,在编辑之前。而不是在测试套件变红之后再跑七次 grep。
Related MCP server: codeweave-mcp
快速开始
uv tool install cartograph-mcp # or: pipx install cartograph-mcp
cartograph index ~/code/my-repo # builds .cartograph/cartograph.db
cartograph arch # modules, layers, cycles, hotspots
cartograph blast src/auth/token.py # what a change here could break
cartograph callers validate_token # reverse call tree接入智能体
Claude Code:
claude mcp add cartograph -- cartograph serve /path/to/repo或者通过 mcp.json 接入任意 MCP 客户端:
{
"mcpServers": {
"cartograph": {
"command": "cartograph",
"args": ["serve", "/path/to/repo"]
}
}
}serve 在首次运行时(如果不存在索引)会进行索引。然后问你的智能体*“如果我改了 token 校验器,会破坏什么?”*,它会调用 blast_radius,而不是靠猜。
十个工具
工具 | 回答 |
| X 在哪里定义?(按结构重要性排序) |
| 对名称、签名、docstring 进行全文搜索(BM25) |
| 单个符号:签名、文档、成员、调用方、被调用方、源码 |
| 反向调用树——在修改签名之前 |
| 正向调用树——无需阅读每个文件就能理解代码 |
| 改动可能破坏什么,以及应该运行哪些测试 |
| “我还应该读什么?”通过个性化 PageRank |
| 文件定义了哪些内容、导入了什么,以及谁导入了它 |
| 模块、分层、导入循环、热点、入口点 |
| 索引健康度以及按规则划分的边解析明细 |
此外还有 MCP 资源(cartograph://architecture、cartograph://stats)和一个 orient 提示词,用于对陌生仓库进行以图为先的初步探索。
支持语言: Python、TypeScript、TSX、JavaScript、Go。
值得讨论的设计决策
1. 置信度是一等列
没有类型检查器,你无法确定 store.who_calls() 就是 GraphStore.who_calls。你只能对假设进行排序。因此,与其假装确定,每条边都记录了产生它的规则和置信度:
规则 | 置信度 | 直觉 |
| 0.95 | 定义就在作用域内 |
| 0.90 | 文件显式导入了该名称 |
| 0.85 |
|
| 0.75 | 同一包中的兄弟文件 |
| 0.60 | 仓库中恰好有一个符号叫这个名字,裸调用 |
| 0.45 | 只有一个匹配,但位于无类型接收者上 |
| ≤0.40 | N 个候选,保留为 N 条边,每条权重 1/N |
| 0.00 | 根植于第三方/标准库导入 |
| 0.00 | 真正未知(动态,或类型化方法) |
调用方随后选择自己的工作点。who_calls 默认采用 ≥0.5——精确率优先,因为智能体会基于答案行动。blast_radius 降至 0.3——召回率优先,因为漏掉受影响的测试是代价高昂的错误,而误报只会让审查者多看一眼。
name-only 这一层之所以存在,是因为一个真实的 bug。内置 set 上的 seen.add(...) 被解析到了仓库中某个类的 add 方法,仅仅因为这个名字恰好唯一——并且它显示为一个高置信度的调用方。无法确定类型的接收者上的方法名不能作为证据,所以它现在位于精确率线之下。(测试)
external 的存在是为了让指标保持诚实:在大多数仓库中,“unresolved”桶主要由 typer.Option 和 sqlite3.execute 占据。如果把它们混在一起,会让覆盖率看起来比实际差得多,因此 Cartograph 报告内部解析率——在可能命中仓库符号的调用点中,实际命中的比例。
2. 解析是增量的;符号解析则从不增量
只有当文件的 sha256 变化时,才会重新解析。但原始引用会作为事实存储在 refs 表中,而 edges 会在任何内容发生变化时,作为(refs × symbols)的纯函数重新计算。
这正是让“每次编辑后重新索引”变得可信的原因。如果解析也是增量的,编辑一个文件可能会让另一个文件中的边指向一个已经移动的符号。全局重新解析在结构上杜绝了这种可能。(测试)
这种代价是真实存在的,所以只有一个安全的捷径:如果没有文件被添加、重新解析或删除,那么两个输入表都不变,解析结果可证明完全相同——因此可以跳过。这让 Django 的无操作重新索引从 7.5 秒降至 0.67 秒,且生成的图逐字节一致。
3. 用 PageRank 代替嵌入
“你指的是哪个 get?”是一个结构性问题。四十个调用点所依赖的那个 get 才是智能体想要的,而调用图早就知道这一点。因此,符号排名是对调用图进行加权 PageRank——稳定、可解释、且零成本。无需模型、无需构建索引、无需向量存储。
related_symbols 扩展了同样的思路:以单个符号为种子的个性化 PageRank,将图视为无向图,因为当你即将修改一个函数时,它的调用方和被调用方都是相关的上下文。它是语义搜索的结构化对应物,并且不需要嵌入。
4. 工具返回 Markdown 而非 JSON,且在 token 预算内
消费者是上下文窗口。一个包含 40 个符号的 JSON 数组会在花括号和重复的键上浪费数千个 token,而且模型反正也会重新格式化。这里的每个视图都是带有硬性 token 预算的紧凑 Markdown。
关键是,每一次截断都会明确告知。如果一个智能体拿到 87 个调用方中的 20 个且没有任何标记,它会自信地断定其他 67 个不存在,然后删除某些东西。
5. 遍历在 SQLite 中运行,而非 Python
深度为 4 的 who_calls 是一个递归 CTE,因此整个遍历都停留在 SQLite 的 C 循环中。在 Django 拥有 252k 条边的图上,这大约需要 ~5ms。如果要把边表拉入 Python 再遍历,就不可能这么快。
基准测试
真实仓库,M 系列笔记本电脑,单进程。冷启动 = 从头完整索引;热启动 = 无操作重新索引。
仓库 | 文件数 | KLOC | 符号数 | 边数 | 冷启动 | 热启动 | 数据库 | 内部解析率 |
2,973 | 534 | 45,394 | 252,441 | 11.9s | 0.67s | 80 MB | 83.2% | |
gin (Go) | 98 | 24 | 1,610 | 9,179 | 0.32s | 0.03s | 2.5 MB | 88.1% |
83 | 18 | 1,624 | 4,271 | 0.21s | 0.03s | 1.7 MB | 87.4% |
查询延迟(5 次中位数,热启动):
仓库 |
|
|
|
|
django | 12.3ms | 5.1ms | 5.6ms | 68.5ms |
gin | 0.4ms | 0.4ms | 0.5ms | 1.2ms |
flask | 0.5ms | 1.1ms | 1.3ms | 1.8ms |
使用 scripts/bench.py 复现。
架构
flowchart LR
subgraph index["cartograph index"]
W[walker<br/>git ls-files] --> P[tree-sitter<br/>+ .scm queries]
P --> X[extract<br/>defs · refs · imports]
end
X --> DB[(SQLite<br/>symbols · refs<br/>edges · FTS5)]
DB --> R[resolver<br/>rule cascade]
R --> DB
DB --> RK[PageRank<br/>Tarjan SCC]
RK --> DB
DB --> S[service facade]
S --> V[views<br/>token-budgeted MD]
V --> M[MCP server<br/>10 tools]
V --> C[CLI]
M --> A((coding agent))模块 | 职责 |
| 文件发现——委托给 |
| 每种语言一个适配器:扩展名、查询、docstring、模块键、导入解析 |
| AST → 符号/引用/导入,与语言无关 |
| tree-sitter 捕获模式——每种语言的知识,以数据形式存在 |
| 图: |
| 置信度级联 |
| PageRank、个性化 PageRank、迭代 Tarjan 强连通分量、分层 |
| 递归 CTE 遍历、排序查找、聚合 |
| 统一门面,确保 CLI 和 MCP 服务器不会漂移 |
| 受 token 预算约束的 Markdown |
无需组合查询的作用域界定
让 queries/*.scm 保持小巧的诀窍:作用域从不编码在查询中。每个捕获的定义都按其 tree-sitter 节点 id 建立索引,而引用的所在符号则通过沿其 parent 链向上走直到命中一个符号来找到。这是每个引用的 O(树深度),并且天然支持闭包、方法、内部类和箭头函数——无需针对每种形状编写模式。
添加一种语言
继承 LanguageAdapter(约 40 行)并放入一个 .scm 文件。GoAdapter 是最短的完整示例。然后 tests/test_queries.py 会自动针对语法编译你的查询,并断言它们能捕获到内容。
开发
git clone https://github.com/GokulRaj2210/cartograph-mcp && cd cartograph-mcp
uv sync
uv run pytest -q # 209 tests
uv run ruff check .
uv run mypy # strictCI 在 Python 3.11/3.12/3.13(以及 macOS)上运行测试套件,然后吃自己的狗粮:它为这个仓库建立索引,在出现导入循环时失败,断言无操作重新索引不会重新解析任何内容,并通过真实的 stdio 驱动 MCP 服务器。它还会将构建的 wheel 安装到干净的 venv 中并用它建立索引,因为打包的 .scm 文件很容易被遗漏在 wheel 之外,而且在本地几乎不可能注意到。
这个循环检查已经证明了它的价值——它捕获了我在这个仓库中引入的一个 store → resolver → store 循环,后来通过移动有问题的辅助函数而不是放宽检查来修复。
值得注意的测试
tests/test_queries.py— 每个.scm都能针对每一个加载它的语法进行编译,并捕获到一些内容。在 JavaScript 中有效的模式((class_heritage (identifier)))在 TypeScript 中是一个 不可能的模式,因为 TypeScript 将父类型包装在extends_clause中。那一行在 TypeScript 中静默地产生了零个符号。tests/test_incremental.py— 在编辑、删除或符号在文件间移动后,不会留下陈旧的边。tests/test_resolver.py— 每一条规则都会触发,且没有一条会过度宣称其置信度。tests/test_cli.py— 读取器和索引器可以同时持有数据库。tests/test_docs.py— 生成的演示页面是格式良好的 HTML,标签平衡,这正是 Markdown 渲染器在min_confidence上的交叉标签 bug 被发现的方式。
局限性
说得直白些:因为一款过度吹嘘其精度的代码智能工具比毫无用处更糟糕。
不进行类型推断。 在不知道
conn类型的情况下,self.conn.execute(...)无法解析为仓库符号。这些会落入unresolved,它们是内部解析率约为 ~85% 时剩余的主要部分。动态分派不可见。
getattr(obj, name)()、装饰器注册表和 DI 容器不会以边的形式出现。不跟踪跨语言边。 TypeScript 前端调用 Python 端点时,两者是两张断开的子图。
仅定义,并非每个引用。 作为值使用的符号(作为回调传递)在图中比被调用的符号更弱。
路线图:Rust 和 Java 适配器、在语言服务器可用时可选的 LSP 增强以实现精确解析,以及一个 --changed-since <ref> 模式,用于评估 PR 的影响范围。
为什么存在
我想知道,编码代理在大型代码库上最大的弱点——缺乏代码的结构化模型——是否可以通过静态分析和设计良好的工具表面来修复,而不是依靠更大的模型或向量数据库。大多数情况下,可以。
许可证
MIT
Available Tools
10 toolsarchitecture_overviewA
Orient yourself in an unfamiliar repo: modules, layers, cycles, hotspots.
Start here. One call replaces a dozen exploratory file reads: you get module sizes and layering, import cycles, the highest-PageRank symbols (the risky ones to change) and the repo's entry points.
| Name | Required | Description | Default |
|---|---|---|---|
| include_diagram | No | Include a Mermaid diagram of the module graph |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the safety/behavior burden. It discloses what the call produces and signals efficiency by replacing 'a dozen exploratory file reads', making the operation's analytic, non-mutating nature clear through the 'you get...' framing. It stops short of stating any performance or read-only caveats explicitly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences with no filler; the purpose is front-loaded and the supporting details (what it returns) are listed compactly. Each clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-optional-parameter tool with an output schema, the description covers the key contextual information: when to use it, what to expect, and why it is valuable. Nothing critical is missing for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the single include_diagram parameter is fully documented in the schema. The description adds no parameter-specific guidance beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('Orient yourself in an unfamiliar repo') and enumerates concrete outputs (module sizes/layering, import cycles, PageRank hotspots, entry points). It clearly differentiates from symbol-level siblings like find_symbol and who_calls by positioning itself as the repo-level starting point.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Start here' and 'one call replaces a dozen exploratory file reads' provide explicit context for when to use it: early exploration of an unfamiliar codebase. It does not explicitly state when not to use it or name an alternative, so it misses the full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
blast_radiusA
Impact analysis: what a change here could break, and which tests to run.
Combines the reverse import graph with the reverse call graph, then highlights test files specifically. Recall-first by design (confidence >=0.3): the expensive mistake is a missed impacted test, not an extra one.
Call this before editing shared code and after finishing, to pick tests.
| Name | Required | Description | Default |
|---|---|---|---|
| depth | No | Transitive import/call depth | |
| limit | No | Max results | |
| target | Yes | A file path or a symbol name/qualname |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It explains the internal approach (combining reverse import graph with reverse call graph), the recall-first bias with a specific confidence threshold of >=0.3, and the rationale that missed impacted tests are worse than extra ones. It does not explicitly state that the operation is read-only or safe, but the impact-analysis framing implies it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded: the purpose is in the first sentence, methodology and behavior in the second, and usage guidance in the final sentence. Every sentence adds distinct value with no repetition or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema is present and the parameter schema fully describes the inputs, the description provides the necessary context: what the tool computes, how it prioritizes recall, what it highlights, and when to call it. An agent has enough to invoke it correctly and interpret its role relative to siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema already documents target, depth, and limit with meaningful descriptions. The tool description adds no parameter-specific guidance beyond the schema, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear purpose: 'Impact analysis: what a change here could break, and which tests to run.' It also differentiates itself from siblings by explaining it combines the reverse import graph with the reverse call graph and specifically highlights test files, which sets it apart from who_calls, what_it_calls, and related_symbols.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit context for when to use the tool: 'Call this before editing shared code and after finishing, to pick tests.' It does not explicitly name alternatives or state when not to use it, but the workflow guidance is clear and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
file_summaryA
Outline of one file: what it defines, what it imports, who imports it.
Cheaper than reading the file when you only need to know whether it is relevant, and it adds the reverse-import view that reading cannot give you.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | File path, or any distinctive part of one |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of explaining behavior. It discloses what the outline contains (definitions, imports, importers) and notes that it is cheaper than full file reading. It does not discuss edge cases like partial paths, errors, or cache behavior, but for a simple summary tool this is reasonably transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences put the core purpose first and the cost/use-case benefit second. Every sentence earns its place; there is no filler or redundant restating of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with an output schema, the description explains what the result contains and why one would choose this tool. It could be slightly stronger about how this compares to adjacent sibling tools, but nothing essential is missing for a basic invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the sole parameter is already well described. The description adds no new parameter-level detail, which is acceptable since the schema fully documents the path parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool's scope: an outline of one file covering definitions, imports, and reverse-imports. This distinguishes it from generic search or symbol tools by naming the specific resource and output aspects.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly frames the tool as a cheaper alternative to reading a file when only relevance matters, and highlights the reverse-import advantage. It does not name sibling tools or provide explicit when-not-to-use guidance, but the intended scenario is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_symbolA
Locate where a symbol is DEFINED, with its file:line, signature and doc.
This is the right first call for "where is X?" -- it is exact and ranked by
structural importance, so if a repo has six functions called run, the one
the codebase actually revolves around comes first.
Use search_code instead when you only know roughly what the thing does
("the retry logic") rather than what it is called.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | Filter by kind: function, method, class, interface, struct, enum, type, const | |
| lang | No | Filter by language: python, typescript, tsx, javascript, go | |
| name | Yes | Symbol name or qualified name, exact or partial | |
| limit | No | Max results |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the burden. It discloses a key behavioral trait: results are 'ranked by structural importance', illustrated with the six-run-functions example. It also mentions exactness and the output shape, though it does not discuss limitations like auth or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no filler: the main purpose is front-loaded, the ranking behavior is immediately explained, and the alternative tool condition is given once. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists and the parameter schema is fully documented, so the description need not restate return types or parameter details. It supplies the missing context: when to use, how results are ranked, and when to switch to search_code, making it complete for an AI agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters. The description adds little beyond contextual emphasis on exactness and ranking; it does not deepen meaning for kind, lang, name, or limit beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Locate where a symbol is DEFINED', with concrete outputs (file:line, signature, doc). It also distinguishes from the sibling search_code by positioning itself as the exact lookup for known symbol names, so an agent can tell when to use it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says this is the right first call for 'where is X?' and names the alternative: use search_code when you only know roughly what the thing does. This gives clear selection criteria without the agent needing to infer.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_symbolA
Full detail for one symbol: signature, doc, members, callers and callees.
Prefer this over reading the whole file: you get the definition plus its
immediate graph neighbourhood, which is usually all the context needed to
make a safe edit. Set include_source=true when you intend to modify it.
| Name | Required | Description | Default |
|---|---|---|---|
| depth | No | Caller/callee depth to include | |
| symbol | Yes | Symbol id, qualified name (`module:Class.method`), `path:name`, or bare name | |
| include_source | No | Include the full source text of the definition |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden of behavioral disclosure. It usefully states the returned artifacts (definition plus graph neighborhood) and the include_source toggle, but does not explain depth behavior, error cases, or cost of deep traversal. This is adequate but not deeply transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly packed paragraphs with no filler. The core purpose is in the first sentence, and the practical guidance follows immediately. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema exists and parameter docs are complete, the description covers the essential context: what the tool returns, why to prefer it, and when to enable source. It doesn't cover depth semantics or error behavior, but those are partially covered in the schema and are minor for a read-only lookup tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds meaningful usage semantics for include_source ('when you intend to modify it') that goes beyond the schema, and the symbol parameter's accepted forms are already well documented in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Full detail for one symbol' with concrete contents (signature, doc, members, callers, callees). This clearly differentiates get_symbol from siblings like search_code, who_calls, and what_it_calls by scoping it to a single symbol's combined context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit usage guidance: prefer this over reading the whole file, and set include_source=true when you intend to modify the symbol. It does not explicitly name all sibling alternatives or when those would be better, but the 'prefer this over...' framing gives clear decision context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
index_statsA
Index health: size, coverage, and the edge-resolution breakdown by rule.
Worth a call when graph answers look thin -- a low resolution rate or a stale
indexed_at tells you the index needs rebuilding rather than the code being
unusual.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of behavioral disclosure. It explains what the tool reports, including size, coverage, resolution breakdown, and indexed_at, and adds diagnostic meaning beyond a simple field list. It does not explicitly state that the tool is read-only, but for a stats tool this is strongly implied by the content described.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: the first sentence states exactly what the tool reports, and the second sentence gives actionable usage guidance. Every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with an output schema available, the description provides everything needed to decide when and how to use it. It explains the tool's purpose, the data it returns, and the diagnostic scenario in which it is useful, leaving no meaningful gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there are no parameter meanings to clarify. The description still adds conceptual value by naming the key output dimensions (size, coverage, edge-resolution breakdown, indexed_at), which is appropriate for a parameterless tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the resource and content ('Index health: size, coverage, and the edge-resolution breakdown by rule'), which immediately distinguishes it from the symbol-focused sibling tools. However, it lacks an explicit verb like 'reports' or 'returns', so it falls just short of the strongest purpose clarity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives an explicit trigger: 'Worth a call when graph answers look thin'. It also explains how to interpret results ('low resolution rate or a stale indexed_at tells you the index needs rebuilding rather than the code being unusual'), which is excellent practical guidance for when this tool is the right choice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_codeA
Full-text search across symbol names, signatures and docstrings (BM25).
Use when you know the intent but not the identifier. Results are re-ranked by call-graph importance, so central symbols outrank incidental mentions.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max results | |
| query | Yes | Free-text query over names, signatures and docstrings |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are absent, so the description carries full behavioral disclosure. It discloses that search uses BM25 and that results are re-ranked by call-graph importance, which is valuable non-obvious behavior. It could mention pagination or query-syntax details, but the core operation and ordering semantics are transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no filler: the first defines scope, the second states when to use it, and the third explains ranking behavior. Every sentence earns its place, and the key use-case guidance is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema and only two straightforward parameters, the description is nearly complete. It covers the tool's purpose, use case, searchable content, and result ordering. It does not explicitly state exclusions or name the exact-identifier sibling, but the sibling context and 'not the identifier' phrasing make the intended boundary clear.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema describes both parameters completely, so the baseline is 3. The description reinforces that `query` is free-text and explains why certain matches outrank others, but it does not add per-parameter syntax or formatting detail beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb ('Full-text search') and a precise resource scope ('symbol names, signatures and docstrings'). The phrase 'Use when you know the intent but not the identifier' clearly distinguishes it from exact-identifier lookup tools such as find_symbol.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit usage condition: use it when the intent is known but the identifier is not. It does not name the alternative tool directly, but the contrast with exact-lookup siblings is strongly implied by the wording and the sibling list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
what_it_callsA
Forward call tree: what this symbol depends on, transitively.
Use it to understand an unfamiliar function without reading every file it touches, and to spot the layer a piece of code really sits in.
| Name | Required | Description | Default |
|---|---|---|---|
| depth | No | Transitive callee depth | |
| limit | No | ||
| symbol | Yes | Source symbol (name, qualname or id) | |
| min_confidence | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses the transitive, graph-walking nature of the tool, but does not mention performance characteristics, result size limits, or other runtime behavior beyond what the schema hints at.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded, with the core definition in the first sentence and practical guidance in the second. No filler or redundant restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description provides enough context for an agent to understand the tool's purpose and basic invocation. Some gaps remain around parameter semantics and explicit sibling differentiation, but the output schema and schema constraints partially fill those gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 50%: symbol and depth are documented, but limit and min_confidence lack descriptions. The tool description does not compensate by explaining these parameters or clarifying their units/purpose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: build a forward call tree of what a symbol transitively depends on. This distinguishes it from reverse-call tools like who_calls, though it does not explicitly name siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides concrete use cases: understanding an unfamiliar function without reading every file, and identifying the layer a piece of code sits in. It gives clear context but does not state when to prefer an alternative tool or when not to use this one.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
who_callsA
Reverse call tree: everything that reaches this symbol, transitively.
The tool to use before changing a signature, tightening a validation, or deleting anything. Each edge reports the rule that produced it; treat sub-0.5 edges as leads rather than facts.
| Name | Required | Description | Default |
|---|---|---|---|
| depth | No | Transitive caller depth | |
| limit | No | Max results | |
| symbol | Yes | Target symbol (name, qualname or id) | |
| min_confidence | No | Minimum edge confidence (0.5 = precision-first) |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses that each edge reports the rule that produced it and warns that sub-0.5 edges are leads rather than facts. It does not discuss cost or traversal size, but the output schema covers result shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences deliver the definition, the trigger scenario, and the confidence caveat. The description is front-loaded with the core purpose and every sentence adds distinct value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only analysis tool with a full output schema and fully documented parameters, the description covers what the tool computes, when to use it, and how to interpret weak results. Nothing essential is missing for selecting and invoking it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
All four parameters already have schema descriptions, so the baseline is 3. The description adds meaningful semantics for min_confidence, explicitly saying sub-0.5 edges should be treated as leads, and implies that depth and limit control transitive expansion.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action and resource: 'Reverse call tree: everything that reaches this symbol, transitively.' This clearly distinguishes it from forward-call tools like what_it_calls without needing extra inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives concrete guidance on when to use the tool: 'The tool to use before changing a signature, tightening a validation, or deleting anything.' It does not explicitly list exclusions or alternatives, but the use-case framing is clear and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
10 tool updates
v0.1.0- First observed
architecture_overview - First observed
blast_radius - First observed
file_summary - First observed
find_symbol - First observed
get_symbol - First observed
index_stats - First observed
related_symbols - First observed
search_code - First observed
what_it_calls - First observed
who_calls
TDQS
Scored across 10 tools
Tool purposes are largely distinct and descriptions explicitly route agents to the right one, but find_symbol/get_symbol and who_calls/blast_radius have adjacent responsibilities that could occasionally cause misselection. Overall, the overlap is minor and well-documented.
All names are readable snake_case, but the set mixes verb-object names (find_symbol, search_code, get_symbol), question-style names (who_calls, what_it_calls), and noun-phrase names (blast_radius, file_summary, architecture_overview). This is not chaotic, but it lacks a single consistent naming pattern.
Ten tools is a well-scoped surface for a code-graph analysis server. Each tool addresses a distinct job—search, symbol detail, call trees, impact analysis, overview, index health—without redundancy or bloat.
The toolchain covers symbol discovery, detailed lookup, dependency analysis, impact assessment, file outlining, architecture orientation, and index health, giving strong coverage of the code-understanding workflow. Minor gaps like direct raw-file access or listing all symbols in a file must be worked around via file_summary and get_symbol.
Maintenance
Related MCP Connectors
Code intelligence platform for AI agents. 20 tools for architecture, security & impact analysis.
Codebase graphs, caller impact analysis, and recorded project context for AI coding agents.
The Cortex MCP server provides read-only access to real-time engineering context from the Cortex developer portal, allowing AI coding assistants to answer natural language questions about your organization's catalog (microservices, libraries, domains, teams, infrastructure), scorecards (engineering standards and best practices), initiatives (goals and deadlines), and Engineering Intelligence metrics. It includes tools for querying documentation, tracking personal entities, and accessing AI-assisted insights across the entire Cortex ecosystem.
Coding agents in multi-service codebases routinely rebuild existing helpers, trust stale type definitions, and modify API contracts without knowing who consumes them. Carrick solves this by indexing your entire TypeScript ecosystem across service and repository boundaries. By integrating deeply with the TypeScript compiler, Carrick traces every route, type, and cross-service call while recording function behaviour so agents search by intent rather than name. Delivered via MCP for AI agents and LSP for IDEs, Carrick ensures models see existing endpoints and utilities before generating new code. The scanner is source-available and runs from your CLI or CI pipeline.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceRepoNova is an MCP server that builds a persistent knowledge graph of your codebase, enabling AI agents to query code structure, dependencies, and semantics through 11 specialized tools.174 npm7MIT
- AlicenseNot gradedqualityDmaintenanceAn MCP server that gives AI agents structured code understanding and precise code intelligence via local indexing of AST, call graphs, and semantic search.147 npm4Apache 2.0
- AlicenseAqualityBmaintenanceAn MCP server that generates ranked, token-budgeted code structure maps using Tree-sitter AST analysis and PageRank, enabling AI agents to quickly understand unfamiliar codebases.220 npmMIT
- AlicenseNot gradedqualityAmaintenanceMCP server for local-first code intelligence, providing structural code graph, semantic search, and impact analysis to AI agents.2MIT