Skip to main content
Glama

codetex-mcp

A commit-aware code context manager for LLMs. Indexes Git repositories into a multi-tier knowledge hierarchy — repo overviews, file summaries, and symbol details — stored in SQLite with vector search. Serves context to LLM clients via the Model Context Protocol (MCP) or a local CLI.

What It Does

codetex builds a structured, searchable index of your codebase that LLMs can query on demand:

  • Tier 1 — Repo Overview: Purpose, architecture, directory structure, key technologies, entry points

  • Tier 2 — File Summaries: Per-file purpose, public interfaces, dependencies, roles

  • Tier 3 — Symbol Details: Function/class signatures, parameters, return types, call relationships

Summaries are generated by an LLM (Anthropic Claude). Embeddings are computed locally with sentence-transformers for semantic search. Everything is stored in a single SQLite database with sqlite-vec for vector queries.

Incremental sync means only changed files are re-analyzed when you update your code.

Related MCP server: Serena MCP Server

Requirements

Installation

# With pip
pip install codetex-mcp

# With uv (recommended)
uv tool install codetex-mcp

Quick Start

1. Set your Anthropic API key

# Via environment variable
export ANTHROPIC_API_KEY=sk-ant-...

# Or via config
codetex config set llm.api_key sk-ant-...

2. Add a repository

# Local repo
codetex add /path/to/your/project

# Remote repo (clones to ~/.codetex/repos/)
codetex add https://github.com/user/repo.git

3. Index it

# Preview what indexing will cost (no API calls)
codetex index my-project --dry-run

# Build the full index
codetex index my-project

4. Query your codebase

# Repo overview (Tier 1)
codetex context my-project

# File summary (Tier 2)
codetex context my-project --file src/auth/login.py

# Symbol detail (Tier 3)
codetex context my-project --symbol authenticate_user

# Semantic search
codetex context my-project --query "how is authentication implemented?"

5. Keep it up to date

# Incremental sync — only re-analyzes changed files
codetex sync my-project

MCP Server Setup

The MCP server lets LLM clients (like Claude Code, Cursor, Windsurf, etc.) query your indexed codebases directly.

Claude Code

Add to your Claude Code MCP settings (~/.claude/claude_desktop_config.json):

{
  "mcpServers": {
    "codetex": {
      "command": "codetex",
      "args": ["serve"],
      "env": {
        "ANTHROPIC_API_KEY": "sk-ant-..."
      }
    }
  }
}

If you installed with uv tool, use the full path:

{
  "mcpServers": {
    "codetex": {
      "command": "/path/to/codetex",
      "args": ["serve"],
      "env": {
        "ANTHROPIC_API_KEY": "sk-ant-..."
      }
    }
  }
}

Find the path with which codetex or uv tool dir.

Other MCP Clients

Any client that supports MCP stdio transport can use codetex. The server command is:

codetex serve

Available MCP Tools

Once connected, the LLM has access to 7 tools:

Tool

Description

get_repo_overview

Tier 1 repo overview (architecture, technologies, entry points)

get_file_context

Tier 2 file summary with symbol list

get_symbol_detail

Tier 3 full symbol detail (signature, params, relationships)

search_context

Semantic search across all indexed context

get_repo_status

Index status (staleness, file/symbol counts, last indexed)

sync_repo

Trigger incremental sync from within the LLM session

list_repos

List all registered repositories

CLI Reference

codetex add <target>

Register a git repository. Accepts a local path or remote URL.

codetex add .                                    # Current directory
codetex add /path/to/repo                        # Local path
codetex add https://github.com/user/repo.git     # Remote (clones locally)
codetex add git@github.com:user/repo.git         # SSH remote

codetex index <repo-name>

Build a full index for a registered repository.

codetex index my-project                # Full index
codetex index my-project --dry-run      # Preview (files, symbols, estimated LLM calls/tokens)
codetex index my-project --path src/    # Index only files under src/

codetex sync <repo-name>

Incremental sync to the current HEAD. Only files changed since the last indexed commit are re-analyzed.

codetex sync my-project                 # Sync changes
codetex sync my-project --dry-run       # Preview what would change
codetex sync my-project --path src/     # Sync only changes under src/

codetex context <repo-name>

Query indexed context at any tier.

codetex context my-project                              # Tier 1: repo overview
codetex context my-project --file src/main.py           # Tier 2: file summary
codetex context my-project --symbol MyClass             # Tier 3: symbol detail
codetex context my-project --query "error handling"     # Semantic search

codetex status <repo-name>

Show index status: indexed commit, current HEAD, staleness, file/symbol counts, token usage.

codetex list

List all registered repositories with their index status.

codetex config show

Display the current configuration.

codetex config set <key> <value>

Update a configuration value.

codetex config set llm.api_key sk-ant-...
codetex config set llm.model claude-sonnet-4-5-20250929
codetex config set indexing.max_file_size_kb 1024
codetex config set indexing.max_concurrent_llm_calls 10

Configuration

Configuration is loaded in layers (last wins):

  1. Defaults — sensible out-of-the-box values

  2. TOML file — ~/.codetex/config.toml

  3. Environment variables — override everything

Config file

# ~/.codetex/config.toml

[storage]
data_dir = "~/.codetex"                  # Base directory for DB and cloned repos

[llm]
provider = "anthropic"                   # LLM provider (currently: anthropic)
model = "claude-sonnet-4-5-20250929"     # Model used for summarization
api_key = "sk-ant-..."                   # Anthropic API key

[indexing]
max_file_size_kb = 512                   # Skip files larger than this
max_concurrent_llm_calls = 5             # Parallel LLM requests during indexing
tier1_rebuild_threshold = 0.10           # Rebuild repo overview if >=10% of files changed on sync

[embedding]
model = "all-MiniLM-L6-v2"              # Sentence-transformers model for embeddings

Environment variables

Variable

Maps to

Example

ANTHROPIC_API_KEY

llm.api_key

sk-ant-...

CODETEX_DATA_DIR

storage.data_dir

/custom/path

CODETEX_LLM_PROVIDER

llm.provider

anthropic

CODETEX_LLM_MODEL

llm.model

claude-sonnet-4-5-20250929

CODETEX_MAX_FILE_SIZE_KB

indexing.max_file_size_kb

1024

CODETEX_MAX_CONCURRENT_LLM

indexing.max_concurrent_llm_calls

10

CODETEX_TIER1_THRESHOLD

indexing.tier1_rebuild_threshold

0.15

CODETEX_EMBEDDING_MODEL

embedding.model

all-MiniLM-L6-v2

File Exclusion

Files are filtered through multiple stages:

  1. Default excludes — node_modules/, __pycache__/, .git/, dist/, build/, .venv/, *.lock, *.min.js, *.pyc, *.so, etc.

  2. .gitignore — standard gitignore rules from your repo

  3. .codetexignore — same syntax as .gitignore, placed in your repo root. Use !pattern to un-ignore files

  4. File size — files exceeding max_file_size_kb are skipped

  5. Binary detection — files with null bytes in the first 8 KB are skipped

Language Support

Language

Tree-sitter (full AST)

Fallback (regex)

Python

Yes

Yes

JavaScript

Yes

Yes

TypeScript

Yes

Yes

Go

Yes

Yes

Rust

Yes

Yes

Java

Yes

Yes

Ruby

Yes

Yes

C/C++

Yes

Yes

All others

—

Yes

Tree-sitter grammars for all 8 languages are installed automatically. For other languages, the fallback parser uses regex patterns to extract functions, classes, and imports.

Architecture

CLI (Typer) ──┐
              ├──▶ Core Services (Indexer, Syncer, ContextStore, SearchEngine)
MCP (FastMCP)─┘         │              │              │
                    Analysis        LLM Provider    Embeddings
                 (tree-sitter +    (Anthropic)    (sentence-transformers)
                  regex fallback)       │              │
                         └──────────────┴──────────────┘
                                        │
                                   SQLite + sqlite-vec
  • Two entry points (CLI and MCP server) share the same core service layer

  • No DI framework — services are wired via a create_app() factory

  • All core services are async — CLI bridges with asyncio.run()

  • Embeddings are local — no external API calls for vector search (model auto-downloads on first run, ~90 MB)

  • Single SQLite database — 6 main tables + 2 vector tables (384-dimensional embeddings)

Development

git clone https://github.com/mrosata/codetex-mcp.git
cd codetex-mcp

# Install dependencies (including dev)
uv sync

# Run tests
uv run pytest

# Run tests with coverage
uv run pytest --cov=codetex_mcp

# Lint and format
uv run ruff check src/ tests/
uv run ruff format src/ tests/

# Type check
uv run mypy src/

Releasing

Releases are automated via GitHub Actions and python-semantic-release. Version bumps are driven by conventional commit messages on main.

Commit message format

Prefix

Effect

Example

fix: ...

Patch bump (0.1.0 → 0.1.1)

fix: handle missing gitignore

feat: ...

Minor bump (0.1.0 → 0.2.0)

feat: add Ruby tree-sitter support

feat!: ...

Major bump (0.1.0 → 1.0.0)

feat!: redesign context API

docs:, chore:, ci:, test:, refactor:

No release

docs: update README

A BREAKING CHANGE: line in the commit body also triggers a major bump.

How it works

  1. Push or merge a PR to main

  2. CI runs lint, type check, and tests

  3. The release workflow analyzes commits since the last tag

  4. If a version bump is needed, it:

    • Updates the version in pyproject.toml

    • Creates a git tag (e.g., v0.2.0)

    • Publishes a GitHub Release with a changelog

    • Builds and publishes the package to PyPI

If you need to release without the automation:

uv build
uv publish

License

MIT

Available Tools

7 tools
get_file_contextC

Return the Tier 2 file summary for a specific file.

Includes file purpose, public interfaces, dependencies, role classification, line count, and token count.

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathYes
repo_nameYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It implies a read-only operation but does not state it explicitly, nor does it describe behavior for missing files, repo state requirements, or side effects. The return content list is useful but does not cover these behavioral expectations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded with the core purpose, and the second sentence lists the included fields efficiently. There is no filler, though the unexplained 'Tier 2' term is a slight clarity cost.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple and has an output schema, but with no annotations and empty schema descriptions, the description must cover more ground. It omits usage context, parameter details, and any relationship to sibling tools like sync_repo or get_repo_status, leaving the agent to infer important context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description should compensate, but it only indirectly addresses file_path ('for a specific file') and says nothing about repo_name, path formats, or relationships between parameters. The parameter names are self-explanatory, but the description adds minimal semantic value beyond them.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Return') and names the resource ('Tier 2 file summary'), and the content list ('file purpose, public interfaces, dependencies, role classification, line count, token count') clarifies what is returned. It does not explicitly name sibling tools, but the resource type is clear enough to distinguish it from repo-level or symbol-level tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use or when-not-to-use guidance, no mention of alternatives like get_symbol_detail or search_context, and no prerequisites such as whether the repo needs to be synced first. The intended usage is only implied by the purpose statement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_repo_overviewA

Return the Tier 1 markdown overview for a repository.

Includes repository purpose, directory structure, key technologies, entry points, and architecture patterns.

ParametersJSON Schema
NameRequiredDescriptionDefault
repo_nameYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It describes the content returned (purpose, structure, etc.) but does not explicitly state that the operation is read-only, mention potential errors (e.g., repo not found), or clarify what 'Tier 1' means. For a simple read tool, this is adequate but not thorough; it does not contradict anything since no annotations exist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two lines, front-loads the core purpose ('Return the Tier 1 markdown overview'), and lists the included elements efficiently. It avoids fluff, though the term 'Tier 1' is unexplained and slightly opaque. Overall, it is concise and well-structured, earning a 4.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has an output schema (indicated by context), return-value details are likely covered there, so the description need not restate them. The description covers the essential content and is sufficient for a simple read operation with one parameter. However, it lacks any error-handling or edge-case information, and does not mention freshness or caching, which would be useful for a complete picture.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It adds no explicit explanation of the 'repo_name' parameter beyond its name, which is self-evident from the tool name and context. The description does not provide format, constraints, or examples, so it adds minimal value over the raw schema. A score of 3 reflects that the parameter is trivial and self-explanatory, but the description still fails to actively clarify it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb ('Return') and resource ('markdown overview for a repository'), and enumerates the content (purpose, directory structure, key technologies, entry points, architecture patterns). This distinguishes it from siblings like get_file_context (file-level) and search_context (search) without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus alternatives. It does not mention sibling tools, nor does it provide conditions like 'for high-level overview' or 'instead of get_file_context'. An agent is left to infer usage solely from the tool name and description, which is insufficient for proper selection among six siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_repo_statusB

Return index status for a repository.

Includes indexed commit, current HEAD, staleness indicator, file count, symbol count, total tokens, and last indexed time.

ParametersJSON Schema
NameRequiredDescriptionDefault
repo_nameYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, so the description carries the full burden of disclosing behavior. It makes clear this is a read-only status operation by saying 'Return index status' and listing output fields. However, it leaves the staleness indicator undefined and says nothing about error conditions, permissions, or whether the status could trigger indexing side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loads the main purpose, and uses a compact list of return fields. Every sentence earns its place. There is no filler or redundant restatement of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter status tool with an output schema, the description covers the essential return content. It is incomplete in two areas: it does not clarify the staleness indicator's meaning, and it gives no sibling differentiation, which matters because get_repo_overview appears to be a closely related read tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description never mentions repo_name or its expected format. The parameter is self-explanatory at a surface level, but there is no guidance on whether it should be a bare name, an owner/repo pair, or some other identifier. The description focuses entirely on the return payload rather than the input.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a clear verb and resource: 'Return index status for a repository.' It also enumerates the specific fields returned (indexed commit, HEAD, staleness, counts, tokens, time), which gives the agent a concrete sense of purpose. It does not explicitly contrast itself with siblings like get_repo_overview, but the status-focused scope is reasonably distinct.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives such as get_repo_overview or sync_repo. It does not mention exclusions, prerequisites, or situations where another sibling would be more appropriate. Usage context is only implied by the tool name and listed fields.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_symbol_detailC

Return the Tier 3 detail for a specific symbol.

Includes full signature, description, parameters with types, return type, call relationships, and file location.

ParametersJSON Schema
NameRequiredDescriptionDefault
repo_nameYes
symbol_nameYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the behavioral burden. It does disclose what the call returns (signature, parameters, relationships, file location) and implies a read-only operation, but it does not mention failure modes, exact-match requirements, permissions, or cost. This is adequate but not fully transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and front-loaded with the core purpose, followed by a useful content list. It is efficient, but the unexplained 'Tier 3' terminology introduces slight ambiguity and prevents a perfect structure score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although an output schema exists, the description lacks information about repo_name format, symbol_name disambiguation, whether the repo must be synced, and what 'Tier 3' means relative to the sibling tools. These gaps are significant for an agent deciding whether and how to invoke the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the parameters. It does not explain repo_name or symbol_name formats, ambiguities, or relationships, even though the parameter names are somewhat self-explanatory. The mention of 'parameters with types' refers to the returned symbol's parameters, not the tool's own inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Return') and resource ('Tier 3 detail for a specific symbol'), and it enumerates the retrieved content. It does not explicitly contrast itself with siblings like get_file_context or search_context, so it stops short of full differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as get_repo_overview, get_file_context, or search_context. No prerequisites, exclusions, or conditions are provided; usage is only implied by the tool's name and basic purpose.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_reposA

List all registered repositories with their status.

Returns a markdown table of repository names, remote URLs, indexed commits, and file counts.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are absent, so the description carries the behavioral disclosure burden. It does disclose the output format ('markdown table') and fields returned, which adds value. However, it does not explicitly state read-only behavior, absence of side effects, pagination, or any constraints, though 'list' implies a safe read.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with the core purpose front-loaded in the first sentence and output details in the second. There is no filler, redundancy, or unnecessary qualification.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter listing tool with an output schema, the description sufficiently covers the tool's purpose and return content. Missing usage guidance versus siblings and an explicit read-only statement are minor gaps given the tool's simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and the schema is empty, so the baseline is 4. No parameter explanation is needed; the description instead clarifies what the returned table contains, which is appropriate given the empty input schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('List') and resource ('all registered repositories') and enumerates the output columns ('repository names, remote URLs, indexed commits, and file counts'). Sibling differentiation is implicit through the word 'all' versus the more targeted sibling names, but no sibling is explicitly named, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance about when to choose this tool over siblings like get_repo_status or get_repo_overview. It only says what it does, not the context or exclusions that should drive selection. An agent must infer usage from tool names alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_contextC

Search for relevant code context using semantic similarity.

Returns a ranked list of matching files and symbols with relevance scores.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes
repo_nameYes
max_resultsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It mentions the return type and ranking but does not state whether the operation is read-only, any rate limits, required repository state (e.g., synced), or other side effects. For a search tool, the lack of explicit read-only or safety information is a notable gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two concise sentences that front-load the core purpose and return format. There is no wasted wording, and the structure is easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has three parameters, no annotations, and an output schema (not detailed), the description is minimal. It does not mention parameter usage, potential prerequisites (e.g., repository must be synced), or any nuances of the ranking. While the output schema covers return structure, the description still leaves critical context unaddressed for a tool with this complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate by explaining parameters. It does not. Although parameter names (repo_name, query, max_results) are self-explanatory, the description adds no detail on their semantics, expected formats, or constraints, leaving the agent to rely solely on the schema names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb ('Search'), resource ('code context'), and method ('semantic similarity'), and specifies the return type ('ranked list of matching files and symbols with relevance scores'). It distinguishes from siblings like get_file_context and get_symbol_detail by emphasizing semantic similarity rather than direct retrieval, though it does not explicitly name alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus the sibling tools. The description implies it is for semantic search but does not state explicit conditions, alternatives, or exclusions, leaving the agent to infer usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sync_repoB

Trigger an incremental sync for a repository.

Processes only files changed since the last indexed commit. Returns a summary of changes made.

ParametersJSON Schema
NameRequiredDescriptionDefault
repo_nameYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It explains the incremental nature and that a summary is returned, but it does not mention side effects, required permissions, whether the operation is destructive, or if it blocks or is async. This is a notable gap for a mutation-like action.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three concise sentences with no waste. It front-loads the action, then explains scope and return value. Every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter tool with an output schema, the description covers the core purpose and return value. However, it omits prerequisites (e.g., repo must already be indexed), potential side effects, and error conditions. It is minimally adequate but not rich enough for a tool that triggers a process.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the repo_name parameter. It only mentions 'repository' generically, without clarifying the format, source, or relationship to the parameter. The meaning is mostly derived from the parameter name itself, not enhanced by the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Trigger an incremental sync for a repository.' The 'incremental' qualifier and the distinction from sibling read-only tools (get_repo_status, get_repo_overview, search_context) make its purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use the tool ('Processes only files changed since the last indexed commit') but does not explicitly say when to prefer it over alternatives or provide when-not-to-use guidance. The context is clear but exclusions are absent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 7 tool updatesv0.5.0
    • First observedget_file_context
    • First observedget_repo_overview
    • First observedget_repo_status
    • First observedget_symbol_detail
    • First observedlist_repos
    • First observedsearch_context
    • First observedsync_repo

TDQS

A3.5/5.0

Scored across 7 tools

Disambiguation5/5

Each tool targets a distinct concern: repo-level overview, file context, symbol detail, semantic search, status, syncing, and repo listing. There is no meaningful overlap or ambiguity between tool boundaries.

Naming Consistency5/5

Tool names consistently follow a snake_case verb_noun pattern, mostly using get_ for retrieval actions. Minor variation like search_context and sync_repo still fit the same predictable style.

Tool Count5/5

Seven tools is well-scoped for a code context and indexing server. Each tool earns its place by covering a distinct retrieval or maintenance operation without unnecessary bloat.

Completeness4/5

The set covers repository overview, file context, symbol detail, semantic search, status, sync, and listing—covering the core workflows well. The main gap is the lack of explicit repo registration or full-sync tools, though these may be handled externally.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    B
    maintenance
    A semantic code context server that connects your local repository to AI assistants via the Model Context Protocol, enabling dynamic codebase exploration without copy-pasting.
    5
    90 npm
    3
    MIT
  • A
    license
    B
    quality
    D
    maintenance
    Provides IDE-like semantic code retrieval and editing tools for LLMs, enabling precise code understanding and manipulation in large codebases via the Model Context Protocol.
    25
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Turns a codebase into a queryable graph with semantic search, call graphs, and control/data flow analysis, served to AI coding agents via the Model Context Protocol.
    206 npm
    MIT
  • F
    license
    Not graded
    quality
    B
    maintenance
    Transforms local Git repositories into queryable, context-rich knowledge bases via AST-aware chunking and Git metadata, enabling AI assistants to search and understand codebases with semantic precision.
    -