Skip to main content
Glama

Cartograph

에이전트 네이티브 코드 인텔리전스. 모든 저장소를 쿼리 가능한 코드 그래프로 바꾸고 MCP를 통해 코딩 에이전트에 제공하세요. 그러면 에이전트는 grep을 해대며 추측하는 대신 “이걸 바꾸면 뭐가 깨질까?” 를 물을 수 있습니다.

tree-sitter + SQLite. 임베딩도, 벡터 저장소도, API 키도, 서버도, 비용도 없습니다.

→ 라이브 데모 — 이 저장소의 실제 인덱스에서 푸시 때마다 생성됩니다.

CI Python 3.11+ License MIT


문제

코딩 에이전트에게 크고 익숙하지 않은 저장소를 주면 어떤 일을 하는지 지켜보세요: grep을 하고, 파일을 읽고, 또 grep을 하고, 다른 파일을 읽습니다. 파서가 한 번에 알려줄 수 있는 구조를 재구성하는 데 컨텍스트를 소진합니다. 그리고 변경으로 인해 망가뜨린 세 모듈 떨어진 호출자를 여전히 놓칩니다.

일반적인 해결책은 RAG입니다: 코드베이스를 임베딩하고 “비슷한” 청크를 검색합니다. 하지만 “이 함수를 누가 호출하나요?” 는 유사성 질문이 아닙니다. 정확한 답이 있고, 그 답은 호출 그래프에 있습니다.

Cartograph는 그래프를 구축한 다음 에이전트가 실제로 작업하는 방식에 맞춰진 열 가지 도구를 제공합니다.

$ cartograph blast src/cartograph/graph/store.py

## Blast radius — file `src/cartograph/graph/store.py`

17 dependent file(s), 31 affected symbol(s), 7 test file(s).

**Tests to run first**
- `tests/test_cli.py`
- `tests/test_docs.py`
- `tests/test_incremental.py`
- `tests/test_mcp.py`
- `tests/test_resolver.py`
- `tests/test_traversal.py`
- `tests/test_views.py`

**Dependent files** (by import distance)
- `src/cartograph/graph/resolver.py` · d1
- `src/cartograph/indexer/pipeline.py` · d1
- `src/cartograph/service.py` · d1
- `src/cartograph/cli.py` · d2
…

편집 전에 단 한 번의 호출. 테스트 스위트가 빨간불이 켜진 뒤의 일곱 번의 grep이 아닙니다.


Related MCP server: codeweave-mcp

빠른 시작

uv tool install cartograph-mcp     # or: pipx install cartograph-mcp

cartograph index ~/code/my-repo    # builds .cartograph/cartograph.db
cartograph arch                    # modules, layers, cycles, hotspots
cartograph blast src/auth/token.py # what a change here could break
cartograph callers validate_token  # reverse call tree

에이전트에 연결하기

Claude Code:

claude mcp add cartograph -- cartograph serve /path/to/repo

또는 mcp.json을 통해 모든 MCP 클라이언트:

{
  "mcpServers": {
    "cartograph": {
      "command": "cartograph",
      "args": ["serve", "/path/to/repo"]
    }
  }
}

serve는 인덱스가 없으면 첫 실행 시 인덱싱합니다. 그런 다음 에이전트에게 “토큰 검증기를 바꾸면 뭐가 깨질까?” 라고 물어보세요. 그러면 추측하는 대신 blast_radius를 호출합니다.


열 가지 도구

도구

답변

find_symbol

X가 정의된 위치는? (구조적 중요도순)

search_code

이름, 시그니처, 독스트링에 대한 전문 검색 (BM25)

get_symbol

하나의 심볼: 시그니처, 문서, 멤버, 호출자, 피호출자, 소스

who_calls

역방향 호출 트리 — 시그니처를 변경하기 전에

what_it_calls

정방향 호출 트리 — 모든 파일을 읽지 않고 코드 이해

blast_radius

변경으로 깨질 수 있는 것, 그리고 실행해야 할 테스트

related_symbols

개인화된 PageRank를 통한 “또 무엇을 읽어야 하나?”

file_summary

파일이 정의하고, 가져오고, 누가 가져오는지

architecture_overview

모듈, 계층화, 임포트 사이클, 핫스팟, 진입점

index_stats

인덱스 건강 상태와 규칙별 엣지 해석 내역

또한 MCP 리소스(cartograph://architecture, cartograph://stats)와 익숙하지 않은 저장소에 대한 그래프 우선 첫 탐색을 위한 orient 프롬프트가 있습니다.

지원 언어: Python, TypeScript, TSX, JavaScript, Go.


논쟁할 가치가 있는 설계 결정

1. 신뢰도는 일급 컬럼입니다

타입 체커 없이는 store.who_calls()가 GraphStore.who_calls를 의미한다고 알 수 없습니다. 가설의 순위를 매길 수만 있을 뿐입니다. 그래서인 척하지 않고, 모든 엣지는 그 엣지를 만들어낸 규칙과 신뢰도를 기록합니다:

규칙

신뢰도

직관

same-file

0.95

정의가 스코프 안에 바로 있다

import

0.90

파일이 이 이름을 명시적으로 임포트했다

receiver-type

0.85

Foo가 알려진 컨테이너인 Foo.bar()

same-module

0.75

같은 패키지의 형제 파일

unique-global

0.60

정확히 하나의 저장소 심볼만 이 이름을 가지며, 수식 없는 호출

name-only

0.45

하나의 일치하지만 타입이 없는 리시버에 대한 호출

ambiguous

≤0.40

N개의 후보, 각각 1/N 신뢰도의 N개 엣지로 유지

external

0.00

서드파티/표준 라이브러리 임포트에 뿌리를 둠

unresolved

0.00

진짜로 알 수 없음 (동적이거나 타입이 있는 메서드)

그러면 호출자(caller)는 자신의 운영 기준점을 선택합니다. who_calls는 기본적으로 ≥0.5입니다 — 정밀도 우선, 에이전트가 답을 바탕으로 행동하기 때문입니다. blast_radius는 0.3으로 낮춥니다 — 재현율 우선, 영향을 받는 테스트를 놓치는 것이 값비싼 실수이고 오탐(false positive)은 리뷰어가 한 번 훑어보는 비용만 들기 때문입니다.

name-only 계층은 실제 버그 때문에 존재합니다. 내장 set의 seen.add(...)가 이름이 우연히 유일하다는 이유만으로 저장소 클래스의 add 메서드로 해석되었고, 그것이 확신도 높은 호출자로 나타났습니다. 타입을 알 수 없는 리시버의 메서드 이름은 증거가 아니므로, 이제 정밀도 기준선 아래에 위치합니다. (test)

external은 메트릭에 대한 정직함을 위해 존재합니다. 대부분의 저장소에서 “unresolved” 버킷은 typer.Option과 sqlite3.execute가 지배합니다. 그것들을 한데 묶으면 적용 범위가 실제보다 훨씬 나빠 보이므로, Cartograph는 내부 해석(internal resolution) 을 보고합니다 — 저장소 심볼에 도달할 수 있는 호출 사이트 중 실제로 도달한 비율입니다.

2. 파싱은 증분적이고 해석은 결코 증분적이지 않다

파일은 sha256이 변경될 때만 다시 파싱됩니다. 그러나 원시 참조는 refs 테이블에 사실(facts) 로 저장되고, edges는 무언가 변경될 때마다 (refs × symbols)의 순수 함수로 다시 계산됩니다.

이것이 “편집할 때마다 다시 인덱싱”을 신뢰할 수 있게 만듭니다. 해석도 증분식이라면 한 파일을 편집할 때 다른 파일의 엣지가 이동한 심볼을 가리키는 채로 남을 수 있습니다. 전역 재해석은 이를 구조적으로 불가능하게 만듭니다. (test)

비용은 실재하므로 안전한 지름길은 정확히 하나뿐입니다. 파일이 추가되거나, 다시 파싱되거나, 제거되지 않았다면 두 입력 테이블은 변경되지 않았고 해석 결과는 증명 가능할 정도로 동일합니다 — 따라서 건너뜁니다. 그 결과 Django의 무작동(no-op) 재인덱스가 7.5초에서 0.67초로 줄었고 그래프는 바이트 단위로 동일했습니다.

3. 임베딩 대신 PageRank

“어떤 get을 의미했나요?”는 구조적 질문입니다. 40개의 호출 사이트가 의존하는 get이 에이전트가 원하는 것이며, 호출 그래프는 이미 그것을 알고 있습니다. 따라서 심볼 순위는 호출 그래프에 대한 가중 PageRank입니다 — 안정적이고 설명 가능하며 비용이 없습니다. 모델도, 인덱스 빌드도, 벡터 저장소도 없습니다.

related_symbols는 같은 아이디어를 확장합니다. 한 심볼에 시드된 개인화 PageRank로 그래프를 무향으로 취급합니다. 함수를 변경하려 할 때 그 함수의 호출자와 피호출자가 모두 관련 컨텍스트이기 때문입니다. 이는 의미론적 검색의 구조적 대응물이며 임베딩이 필요 없습니다.

4. 도구는 토큰 예산 하에서 JSON이 아닌 Markdown을 반환합니다

소비자는 컨텍스트 윈도우입니다. 40개 심볼의 JSON 배열은 중괄호와 반복되는 키에 수천 개의 토큰을 소비하며, 모델은 어차피 그것을 재구성합니다. 여기 있는 모든 뷰는 엄격한 토큰 예산을 가진 컴팩트한 Markdown입니다.

중요한 점은 모든 잘림(truncation)이 명시적으로 알려진다는 것입니다. 87명의 호출자 중 20명만 표시 없이 전달받은 에이전트는 나머지 67명이 존재하지 않는다고 확신하고 무언가를 삭제할 것입니다.

5. 순회는 Python이 아닌 SQLite에서 실행됩니다

깊이 4의 who_calls는 재귀 CTE이므로 전체 순회가 SQLite의 C 루프 안에 머무릅니다. Django의 252k 엣지 그래프에서는 약 ~5ms입니다. 엣지 테이블을 Python으로 가져와 순회한다면 그렇게 되지 않을 것입니다.


벤치마크

실제 저장소, M-시리즈 노트북, 단일 프로세스. Cold = 처음부터 전체 인덱스; warm = 무작동 재인덱스.

저장소

파일

KLOC

심볼

엣지

Cold

Warm

DB

내부 해석

django

2,973

534

45,394

252,441

11.9s

0.67s

80 MB

83.2%

gin (Go)

98

24

1,610

9,179

0.32s

0.03s

2.5 MB

88.1%

flask

83

18

1,624

4,271

0.21s

0.03s

1.7 MB

87.4%

쿼리 지연 시간 (5회 중앙값, warm):

저장소

find_symbol

who_calls d3

blast_radius

architecture_overview

django

12.3ms

5.1ms

5.6ms

68.5ms

gin

0.4ms

0.4ms

0.5ms

1.2ms

flask

0.5ms

1.1ms

1.3ms

1.8ms

scripts/bench.py로 재현할 수 있습니다.


아키텍처

flowchart LR
  subgraph index["cartograph index"]
    W[walker<br/>git ls-files] --> P[tree-sitter<br/>+ .scm queries]
    P --> X[extract<br/>defs · refs · imports]
  end
  X --> DB[(SQLite<br/>symbols · refs<br/>edges · FTS5)]
  DB --> R[resolver<br/>rule cascade]
  R --> DB
  DB --> RK[PageRank<br/>Tarjan SCC]
  RK --> DB
  DB --> S[service facade]
  S --> V[views<br/>token-budgeted MD]
  V --> M[MCP server<br/>10 tools]
  V --> C[CLI]
  M --> A((coding agent))

모듈

책임

indexer/walker.py

파일 탐색 — 올바른 .gitignore 의미를 위해 git ls-files에 위임

indexer/languages.py

언어별 어댑터 하나: 확장자, 쿼리, 독스트링, 모듈 키, 임포트 해석

indexer/extract.py

AST → 심볼/참조/임포트, 언어 비종속적

queries/*.scm

tree-sitter 캡처 패턴 — 언어별 지식을 데이터로

graph/schema.sql

그래프: files, symbols, refs, edges, imports, FTS5

graph/resolver.py

신뢰도 캐스케이드

graph/algorithms.py

PageRank, 개인화 PageRank, 반복 Tarjan SCC, 계층화

graph/store.py

재귀 CTE 순회, 순위 검색, 집계

service.py

CLI와 MCP 서버가 어긋나지 않게 하는 단일 파사드

views.py

토큰 예산 Markdown

조합 쿼리 없이 스코프 다루기

queries/*.scm을 작게 유지하는 비결: 스코프는 절대 쿼리에 인코딩되지 않습니다. 캡처된 모든 정의는 tree-sitter 노드 id로 인덱싱되고, 참조의 포함 심볼은 parent 체인을 따라가며 심볼을 만날 때까지 찾습니다. 참조당 O(트리 깊이)이며 클로저, 메서드, 내부 클래스, 화살표 함수를 추가 비용 없이 처리합니다 — 형태별 패턴이 필요 없습니다.

언어 추가하기

LanguageAdapter(~40줄)를 서브클래싱하고 .scm 파일을 넣으세요. GoAdapter는 가장 짧은 완전한 예제입니다. 그러면 tests/test_queries.py가 문법에 대해 쿼리를 자동으로 컴파일하고 무언가를 캡처하는지 확인합니다.


개발

git clone https://github.com/GokulRaj2210/cartograph-mcp && cd cartograph-mcp
uv sync
uv run pytest -q          # 209 tests
uv run ruff check .
uv run mypy               # strict

CI는 Python 3.11/3.12/3.13(macOS 포함)에서 스위트를 실행한 다음 dogfooding을 합니다. 이 저장소를 인덱싱하고, 임포트 사이클이 있으면 실패하고, 무작동 재인덱스가 아무것도 다시 파싱하지 않음을 단언하고, 실제 stdio로 MCP 서버를 구동합니다. 또한 빌드된 wheel을 깨끗한 venv에 설치하고 그걸로 인덱싱합니다. 패키징된 .scm 파일은 wheel에서 빠뜨리기 쉽고 로컬에서는 알아차리기 불가능하기 때문입니다.

사이클 게이트는 이미 제 역할을 톡톡히 했습니다. 이 저장소에서 제가 만든 store → resolver → store 사이클을 잡아냈고, 게이트를 완화하는 대신 문제의 헬퍼를 옮겨서 수정했습니다.

주목할 만한 테스트

  • tests/test_queries.py — 모든 .scm은 이를 로드하는 모든 그래머에 대해 컴파일되어 무언가를 캡처합니다. JavaScript에서 유효한 패턴((class_heritage (identifier)))은 TypeScript에서 불가능한 패턴인데, TypeScript는 슈퍼타입을 extends_clause로 감쌉니다. 그 한 줄은 조용히 TypeScript 심볼을 0개 생성했습니다.

  • tests/test_incremental.py — 편집, 삭제 또는 파일 간 심볼 이동 후에도 오래된 엣지가 없습니다.

  • tests/test_resolver.py — 모든 규칙이 발동되며, 어떤 규칙도 자신의 신뢰도를 과장하지 않습니다.

  • tests/test_cli.py — 리더와 인덱서가 동시에 데이터베이스를 보유할 수 있습니다.

  • tests/test_docs.py — 생성된 데모 페이지는 태그 균형이 맞는 올바른 형식의 HTML이며, 이를 통해 min_confidence에서 발생한 Markdown 렌더러의 교차 태그 버그가 발견되었습니다.


제한 사항

솔직히 말하면, 정밀도를 과장하는 코드 인텔리전스 도구는 쓸모없는 것보다 더 나쁩니다:

  • 타입 추론 없음. self.conn.execute(...)은 conn의 타입을 알지 못하면 저장소 심볼로 해석될 수 없습니다. 이러한 것들은 unresolved에 속하게 되며, 내부 해석률이 ~85%일 때 남은 것들의 대부분을 차지합니다.

  • 동적 디스패치는 보이지 않습니다. getattr(obj, name)(), 데코레이터 레지스트리, DI 컨테이너는 엣지로 나타나지 않습니다.

  • 크로스 언어 엣지는 추적되지 않습니다. TypeScript 프런트엔드가 Python 엔드포인트를 호출하는 것은 두 개의 분리된 서브그래프입니다.

  • 정의만 있고 모든 참조는 아닙니다. 값으로 사용되는 심볼(콜백으로 전달되는)은 호출되는 심볼보다 그래프에서 더 약합니다.

로드맵: Rust 및 Java 어댑터, 언어 서버를 사용할 수 있는 곳에서 정확한 해석을 위한 선택적 LSP 강화, PR 범위의 영향 반경을 위한 --changed-since <ref> 모드.


왜 이것이 존재하는가

저는 대규모 저장소에서 코딩 에이전트의 가장 큰 약점 — 코드에 대한 구조적 모델이 없다는 점 — 이 더 큰 모델이나 벡터 데이터베이스보다는 정적 분석과 잘 설계된 도구 표면으로 해결될 수 있는지 알고 싶었습니다. 대체로, 가능합니다.

라이선스

MIT

Available Tools

10 tools
architecture_overviewA

Orient yourself in an unfamiliar repo: modules, layers, cycles, hotspots.

Start here. One call replaces a dozen exploratory file reads: you get module sizes and layering, import cycles, the highest-PageRank symbols (the risky ones to change) and the repo's entry points.

ParametersJSON Schema
NameRequiredDescriptionDefault
include_diagramNoInclude a Mermaid diagram of the module graph

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the safety/behavior burden. It discloses what the call produces and signals efficiency by replacing 'a dozen exploratory file reads', making the operation's analytic, non-mutating nature clear through the 'you get...' framing. It stops short of stating any performance or read-only caveats explicitly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences with no filler; the purpose is front-loaded and the supporting details (what it returns) are listed compactly. Each clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-optional-parameter tool with an output schema, the description covers the key contextual information: when to use it, what to expect, and why it is valuable. Nothing critical is missing for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the single include_diagram parameter is fully documented in the schema. The description adds no parameter-specific guidance beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource ('Orient yourself in an unfamiliar repo') and enumerates concrete outputs (module sizes/layering, import cycles, PageRank hotspots, entry points). It clearly differentiates from symbol-level siblings like find_symbol and who_calls by positioning itself as the repo-level starting point.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Start here' and 'one call replaces a dozen exploratory file reads' provide explicit context for when to use it: early exploration of an unfamiliar codebase. It does not explicitly state when not to use it or name an alternative, so it misses the full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

blast_radiusA

Impact analysis: what a change here could break, and which tests to run.

Combines the reverse import graph with the reverse call graph, then highlights test files specifically. Recall-first by design (confidence >=0.3): the expensive mistake is a missed impacted test, not an extra one.

Call this before editing shared code and after finishing, to pick tests.

ParametersJSON Schema
NameRequiredDescriptionDefault
depthNoTransitive import/call depth
limitNoMax results
targetYesA file path or a symbol name/qualname

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It explains the internal approach (combining reverse import graph with reverse call graph), the recall-first bias with a specific confidence threshold of >=0.3, and the rationale that missed impacted tests are worse than extra ones. It does not explicitly state that the operation is read-only or safe, but the impact-analysis framing implies it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and front-loaded: the purpose is in the first sentence, methodology and behavior in the second, and usage guidance in the final sentence. Every sentence adds distinct value with no repetition or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema is present and the parameter schema fully describes the inputs, the description provides the necessary context: what the tool computes, how it prioritizes recall, what it highlights, and when to call it. An agent has enough to invoke it correctly and interpret its role relative to siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already documents target, depth, and limit with meaningful descriptions. The tool description adds no parameter-specific guidance beyond the schema, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a clear purpose: 'Impact analysis: what a change here could break, and which tests to run.' It also differentiates itself from siblings by explaining it combines the reverse import graph with the reverse call graph and specifically highlights test files, which sets it apart from who_calls, what_it_calls, and related_symbols.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit context for when to use the tool: 'Call this before editing shared code and after finishing, to pick tests.' It does not explicitly name alternatives or state when not to use it, but the workflow guidance is clear and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

file_summaryA

Outline of one file: what it defines, what it imports, who imports it.

Cheaper than reading the file when you only need to know whether it is relevant, and it adds the reverse-import view that reading cannot give you.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesFile path, or any distinctive part of one

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of explaining behavior. It discloses what the outline contains (definitions, imports, importers) and notes that it is cheaper than full file reading. It does not discuss edge cases like partial paths, errors, or cache behavior, but for a simple summary tool this is reasonably transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences put the core purpose first and the cost/use-case benefit second. Every sentence earns its place; there is no filler or redundant restating of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with an output schema, the description explains what the result contains and why one would choose this tool. It could be slightly stronger about how this compares to adjacent sibling tools, but nothing essential is missing for a basic invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the sole parameter is already well described. The description adds no new parameter-level detail, which is acceptable since the schema fully documents the path parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool's scope: an outline of one file covering definitions, imports, and reverse-imports. This distinguishes it from generic search or symbol tools by naming the specific resource and output aspects.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly frames the tool as a cheaper alternative to reading a file when only relevance matters, and highlights the reverse-import advantage. It does not name sibling tools or provide explicit when-not-to-use guidance, but the intended scenario is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_symbolA

Locate where a symbol is DEFINED, with its file:line, signature and doc.

This is the right first call for "where is X?" -- it is exact and ranked by structural importance, so if a repo has six functions called run, the one the codebase actually revolves around comes first.

Use search_code instead when you only know roughly what the thing does ("the retry logic") rather than what it is called.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNoFilter by kind: function, method, class, interface, struct, enum, type, const
langNoFilter by language: python, typescript, tsx, javascript, go
nameYesSymbol name or qualified name, exact or partial
limitNoMax results

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the burden. It discloses a key behavioral trait: results are 'ranked by structural importance', illustrated with the six-run-functions example. It also mentions exactness and the output shape, though it does not discuss limitations like auth or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no filler: the main purpose is front-loaded, the ranking behavior is immediately explained, and the alternative tool condition is given once. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists and the parameter schema is fully documented, so the description need not restate return types or parameter details. It supplies the missing context: when to use, how results are ranked, and when to switch to search_code, making it complete for an AI agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters. The description adds little beyond contextual emphasis on exactness and ranking; it does not deepen meaning for kind, lang, name, or limit beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Locate where a symbol is DEFINED', with concrete outputs (file:line, signature, doc). It also distinguishes from the sibling search_code by positioning itself as the exact lookup for known symbol names, so an agent can tell when to use it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says this is the right first call for 'where is X?' and names the alternative: use search_code when you only know roughly what the thing does. This gives clear selection criteria without the agent needing to infer.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_symbolA

Full detail for one symbol: signature, doc, members, callers and callees.

Prefer this over reading the whole file: you get the definition plus its immediate graph neighbourhood, which is usually all the context needed to make a safe edit. Set include_source=true when you intend to modify it.

ParametersJSON Schema
NameRequiredDescriptionDefault
depthNoCaller/callee depth to include
symbolYesSymbol id, qualified name (`module:Class.method`), `path:name`, or bare name
include_sourceNoInclude the full source text of the definition

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the burden of behavioral disclosure. It usefully states the returned artifacts (definition plus graph neighborhood) and the include_source toggle, but does not explain depth behavior, error cases, or cost of deep traversal. This is adequate but not deeply transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly packed paragraphs with no filler. The core purpose is in the first sentence, and the practical guidance follows immediately. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists and parameter docs are complete, the description covers the essential context: what the tool returns, why to prefer it, and when to enable source. It doesn't cover depth semantics or error behavior, but those are partially covered in the schema and are minor for a read-only lookup tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaningful usage semantics for include_source ('when you intend to modify it') that goes beyond the schema, and the symbol parameter's accepted forms are already well documented in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Full detail for one symbol' with concrete contents (signature, doc, members, callers, callees). This clearly differentiates get_symbol from siblings like search_code, who_calls, and what_it_calls by scoping it to a single symbol's combined context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit usage guidance: prefer this over reading the whole file, and set include_source=true when you intend to modify the symbol. It does not explicitly name all sibling alternatives or when those would be better, but the 'prefer this over...' framing gives clear decision context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

index_statsA

Index health: size, coverage, and the edge-resolution breakdown by rule.

Worth a call when graph answers look thin -- a low resolution rate or a stale indexed_at tells you the index needs rebuilding rather than the code being unusual.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of behavioral disclosure. It explains what the tool reports, including size, coverage, resolution breakdown, and indexed_at, and adds diagnostic meaning beyond a simple field list. It does not explicitly state that the tool is read-only, but for a stats tool this is strongly implied by the content described.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: the first sentence states exactly what the tool reports, and the second sentence gives actionable usage guidance. Every sentence earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter tool with an output schema available, the description provides everything needed to decide when and how to use it. It explains the tool's purpose, the data it returns, and the diagnostic scenario in which it is useful, leaving no meaningful gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there are no parameter meanings to clarify. The description still adds conceptual value by naming the key output dimensions (size, coverage, edge-resolution breakdown, indexed_at), which is appropriate for a parameterless tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the resource and content ('Index health: size, coverage, and the edge-resolution breakdown by rule'), which immediately distinguishes it from the symbol-focused sibling tools. However, it lacks an explicit verb like 'reports' or 'returns', so it falls just short of the strongest purpose clarity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives an explicit trigger: 'Worth a call when graph answers look thin'. It also explains how to interpret results ('low resolution rate or a stale indexed_at tells you the index needs rebuilding rather than the code being unusual'), which is excellent practical guidance for when this tool is the right choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_codeA

Full-text search across symbol names, signatures and docstrings (BM25).

Use when you know the intent but not the identifier. Results are re-ranked by call-graph importance, so central symbols outrank incidental mentions.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax results
queryYesFree-text query over names, signatures and docstrings

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are absent, so the description carries full behavioral disclosure. It discloses that search uses BM25 and that results are re-ranked by call-graph importance, which is valuable non-obvious behavior. It could mention pagination or query-syntax details, but the core operation and ordering semantics are transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no filler: the first defines scope, the second states when to use it, and the third explains ranking behavior. Every sentence earns its place, and the key use-case guidance is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema and only two straightforward parameters, the description is nearly complete. It covers the tool's purpose, use case, searchable content, and result ordering. It does not explicitly state exclusions or name the exact-identifier sibling, but the sibling context and 'not the identifier' phrasing make the intended boundary clear.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema describes both parameters completely, so the baseline is 3. The description reinforces that `query` is free-text and explains why certain matches outrank others, but it does not add per-parameter syntax or formatting detail beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb ('Full-text search') and a precise resource scope ('symbol names, signatures and docstrings'). The phrase 'Use when you know the intent but not the identifier' clearly distinguishes it from exact-identifier lookup tools such as find_symbol.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit usage condition: use it when the intent is known but the identifier is not. It does not name the alternative tool directly, but the contrast with exact-lookup siblings is strongly implied by the wording and the sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

what_it_callsA

Forward call tree: what this symbol depends on, transitively.

Use it to understand an unfamiliar function without reading every file it touches, and to spot the layer a piece of code really sits in.

ParametersJSON Schema
NameRequiredDescriptionDefault
depthNoTransitive callee depth
limitNo
symbolYesSource symbol (name, qualname or id)
min_confidenceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It discloses the transitive, graph-walking nature of the tool, but does not mention performance characteristics, result size limits, or other runtime behavior beyond what the schema hints at.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded, with the core definition in the first sentence and practical guidance in the second. No filler or redundant restatement of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description provides enough context for an agent to understand the tool's purpose and basic invocation. Some gaps remain around parameter semantics and explicit sibling differentiation, but the output schema and schema constraints partially fill those gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 50%: symbol and depth are documented, but limit and min_confidence lack descriptions. The tool description does not compensate by explaining these parameters or clarifying their units/purpose.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: build a forward call tree of what a symbol transitively depends on. This distinguishes it from reverse-call tools like who_calls, though it does not explicitly name siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides concrete use cases: understanding an unfamiliar function without reading every file, and identifying the layer a piece of code sits in. It gives clear context but does not state when to prefer an alternative tool or when not to use this one.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

who_callsA

Reverse call tree: everything that reaches this symbol, transitively.

The tool to use before changing a signature, tightening a validation, or deleting anything. Each edge reports the rule that produced it; treat sub-0.5 edges as leads rather than facts.

ParametersJSON Schema
NameRequiredDescriptionDefault
depthNoTransitive caller depth
limitNoMax results
symbolYesTarget symbol (name, qualname or id)
min_confidenceNoMinimum edge confidence (0.5 = precision-first)

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It discloses that each edge reports the rule that produced it and warns that sub-0.5 edges are leads rather than facts. It does not discuss cost or traversal size, but the output schema covers result shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three compact sentences deliver the definition, the trigger scenario, and the confidence caveat. The description is front-loaded with the core purpose and every sentence adds distinct value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only analysis tool with a full output schema and fully documented parameters, the description covers what the tool computes, when to use it, and how to interpret weak results. Nothing essential is missing for selecting and invoking it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

All four parameters already have schema descriptions, so the baseline is 3. The description adds meaningful semantics for min_confidence, explicitly saying sub-0.5 edges should be treated as leads, and implies that depth and limit control transitive expansion.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action and resource: 'Reverse call tree: everything that reaches this symbol, transitively.' This clearly distinguishes it from forward-call tools like what_it_calls without needing extra inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives concrete guidance on when to use the tool: 'The tool to use before changing a signature, tightening a validation, or deleting anything.' It does not explicitly list exclusions or alternatives, but the use-case framing is clear and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 10 tool updatesv0.1.0
    • First observedarchitecture_overview
    • First observedblast_radius
    • First observedfile_summary
    • First observedfind_symbol
    • First observedget_symbol
    • First observedindex_stats
    • First observedrelated_symbols
    • First observedsearch_code
    • First observedwhat_it_calls
    • First observedwho_calls

TDQS

A4/5.0

Scored across 10 tools

Disambiguation4/5

Tool purposes are largely distinct and descriptions explicitly route agents to the right one, but find_symbol/get_symbol and who_calls/blast_radius have adjacent responsibilities that could occasionally cause misselection. Overall, the overlap is minor and well-documented.

Naming Consistency3/5

All names are readable snake_case, but the set mixes verb-object names (find_symbol, search_code, get_symbol), question-style names (who_calls, what_it_calls), and noun-phrase names (blast_radius, file_summary, architecture_overview). This is not chaotic, but it lacks a single consistent naming pattern.

Tool Count5/5

Ten tools is a well-scoped surface for a code-graph analysis server. Each tool addresses a distinct job—search, symbol detail, call trees, impact analysis, overview, index health—without redundancy or bloat.

Completeness4/5

The toolchain covers symbol discovery, detailed lookup, dependency analysis, impact assessment, file outlining, architecture orientation, and index health, giving strong coverage of the code-understanding workflow. Minor gaps like direct raw-file access or listing all symbols in a file must be worked around via file_summary and get_symbol.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

  • Code intelligence platform for AI agents. 20 tools for architecture, security & impact analysis.

  • Codebase graphs, caller impact analysis, and recorded project context for AI coding agents.

  • The Cortex MCP server provides read-only access to real-time engineering context from the Cortex developer portal, allowing AI coding assistants to answer natural language questions about your organization's catalog (microservices, libraries, domains, teams, infrastructure), scorecards (engineering standards and best practices), initiatives (goals and deadlines), and Engineering Intelligence metrics. It includes tools for querying documentation, tracking personal entities, and accessing AI-assisted insights across the entire Cortex ecosystem.

  • Coding agents in multi-service codebases routinely rebuild existing helpers, trust stale type definitions, and modify API contracts without knowing who consumes them. Carrick solves this by indexing your entire TypeScript ecosystem across service and repository boundaries. By integrating deeply with the TypeScript compiler, Carrick traces every route, type, and cross-service call while recording function behaviour so agents search by intent rather than name. Delivered via MCP for AI agents and LSP for IDEs, Carrick ensures models see existing endpoints and utilities before generating new code. The scanner is source-available and runs from your CLI or CI pipeline.

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    RepoNova is an MCP server that builds a persistent knowledge graph of your codebase, enabling AI agents to query code structure, dependencies, and semantics through 11 specialized tools.
    174 npm
    7
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    An MCP server that gives AI agents structured code understanding and precise code intelligence via local indexing of AST, call graphs, and semantic search.
    147 npm
    4
    Apache 2.0
  • A
    license
    A
    quality
    B
    maintenance
    An MCP server that generates ranked, token-budgeted code structure maps using Tree-sitter AST analysis and PageRank, enabling AI agents to quickly understand unfamiliar codebases.
    2
    20 npm
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    MCP server for local-first code intelligence, providing structural code graph, semantic search, and impact analysis to AI agents.
    2
    MIT