codetex-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@codetex-mcpsearch for authentication implementation in my-project"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
codetex-mcp
A commit-aware code context manager for LLMs. Indexes Git repositories into a multi-tier knowledge hierarchy — repo overviews, file summaries, and symbol details — stored in SQLite with vector search. Serves context to LLM clients via the Model Context Protocol (MCP) or a local CLI.
What It Does
codetex builds a structured, searchable index of your codebase that LLMs can query on demand:
Tier 1 — Repo Overview: Purpose, architecture, directory structure, key technologies, entry points
Tier 2 — File Summaries: Per-file purpose, public interfaces, dependencies, roles
Tier 3 — Symbol Details: Function/class signatures, parameters, return types, call relationships
Summaries are generated by an LLM (Anthropic Claude). Embeddings are computed locally with sentence-transformers for semantic search. Everything is stored in a single SQLite database with sqlite-vec for vector queries.
Incremental sync means only changed files are re-analyzed when you update your code.
Related MCP server: Serena MCP Server
Requirements
Python 3.12+
Git
An Anthropic API key (for indexing)
Installation
# With pip
pip install codetex-mcp
# With uv (recommended)
uv tool install codetex-mcpQuick Start
1. Set your Anthropic API key
# Via environment variable
export ANTHROPIC_API_KEY=sk-ant-...
# Or via config
codetex config set llm.api_key sk-ant-...2. Add a repository
# Local repo
codetex add /path/to/your/project
# Remote repo (clones to ~/.codetex/repos/)
codetex add https://github.com/user/repo.git3. Index it
# Preview what indexing will cost (no API calls)
codetex index my-project --dry-run
# Build the full index
codetex index my-project4. Query your codebase
# Repo overview (Tier 1)
codetex context my-project
# File summary (Tier 2)
codetex context my-project --file src/auth/login.py
# Symbol detail (Tier 3)
codetex context my-project --symbol authenticate_user
# Semantic search
codetex context my-project --query "how is authentication implemented?"5. Keep it up to date
# Incremental sync — only re-analyzes changed files
codetex sync my-projectMCP Server Setup
The MCP server lets LLM clients (like Claude Code, Cursor, Windsurf, etc.) query your indexed codebases directly.
Claude Code
Add to your Claude Code MCP settings (~/.claude/claude_desktop_config.json):
{
"mcpServers": {
"codetex": {
"command": "codetex",
"args": ["serve"],
"env": {
"ANTHROPIC_API_KEY": "sk-ant-..."
}
}
}
}If you installed with uv tool, use the full path:
{
"mcpServers": {
"codetex": {
"command": "/path/to/codetex",
"args": ["serve"],
"env": {
"ANTHROPIC_API_KEY": "sk-ant-..."
}
}
}
}Find the path with which codetex or uv tool dir.
Other MCP Clients
Any client that supports MCP stdio transport can use codetex. The server command is:
codetex serveAvailable MCP Tools
Once connected, the LLM has access to 7 tools:
Tool | Description |
| Tier 1 repo overview (architecture, technologies, entry points) |
| Tier 2 file summary with symbol list |
| Tier 3 full symbol detail (signature, params, relationships) |
| Semantic search across all indexed context |
| Index status (staleness, file/symbol counts, last indexed) |
| Trigger incremental sync from within the LLM session |
| List all registered repositories |
CLI Reference
codetex add <target>
Register a git repository. Accepts a local path or remote URL.
codetex add . # Current directory
codetex add /path/to/repo # Local path
codetex add https://github.com/user/repo.git # Remote (clones locally)
codetex add git@github.com:user/repo.git # SSH remotecodetex index <repo-name>
Build a full index for a registered repository.
codetex index my-project # Full index
codetex index my-project --dry-run # Preview (files, symbols, estimated LLM calls/tokens)
codetex index my-project --path src/ # Index only files under src/codetex sync <repo-name>
Incremental sync to the current HEAD. Only files changed since the last indexed commit are re-analyzed.
codetex sync my-project # Sync changes
codetex sync my-project --dry-run # Preview what would change
codetex sync my-project --path src/ # Sync only changes under src/codetex context <repo-name>
Query indexed context at any tier.
codetex context my-project # Tier 1: repo overview
codetex context my-project --file src/main.py # Tier 2: file summary
codetex context my-project --symbol MyClass # Tier 3: symbol detail
codetex context my-project --query "error handling" # Semantic searchcodetex status <repo-name>
Show index status: indexed commit, current HEAD, staleness, file/symbol counts, token usage.
codetex list
List all registered repositories with their index status.
codetex config show
Display the current configuration.
codetex config set <key> <value>
Update a configuration value.
codetex config set llm.api_key sk-ant-...
codetex config set llm.model claude-sonnet-4-5-20250929
codetex config set indexing.max_file_size_kb 1024
codetex config set indexing.max_concurrent_llm_calls 10Configuration
Configuration is loaded in layers (last wins):
Defaults — sensible out-of-the-box values
TOML file —
~/.codetex/config.tomlEnvironment variables — override everything
Config file
# ~/.codetex/config.toml
[storage]
data_dir = "~/.codetex" # Base directory for DB and cloned repos
[llm]
provider = "anthropic" # LLM provider (currently: anthropic)
model = "claude-sonnet-4-5-20250929" # Model used for summarization
api_key = "sk-ant-..." # Anthropic API key
[indexing]
max_file_size_kb = 512 # Skip files larger than this
max_concurrent_llm_calls = 5 # Parallel LLM requests during indexing
tier1_rebuild_threshold = 0.10 # Rebuild repo overview if >=10% of files changed on sync
[embedding]
model = "all-MiniLM-L6-v2" # Sentence-transformers model for embeddingsEnvironment variables
Variable | Maps to | Example |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
File Exclusion
Files are filtered through multiple stages:
Default excludes —
node_modules/,__pycache__/,.git/,dist/,build/,.venv/,*.lock,*.min.js,*.pyc,*.so, etc..gitignore— standard gitignore rules from your repo.codetexignore— same syntax as.gitignore, placed in your repo root. Use!patternto un-ignore filesFile size — files exceeding
max_file_size_kbare skippedBinary detection — files with null bytes in the first 8 KB are skipped
Language Support
Language | Tree-sitter (full AST) | Fallback (regex) |
Python | Yes | Yes |
JavaScript | Yes | Yes |
TypeScript | Yes | Yes |
Go | Yes | Yes |
Rust | Yes | Yes |
Java | Yes | Yes |
Ruby | Yes | Yes |
C/C++ | Yes | Yes |
All others | — | Yes |
Tree-sitter grammars for all 8 languages are installed automatically. For other languages, the fallback parser uses regex patterns to extract functions, classes, and imports.
Architecture
CLI (Typer) ──┐
├──▶ Core Services (Indexer, Syncer, ContextStore, SearchEngine)
MCP (FastMCP)─┘ │ │ │
Analysis LLM Provider Embeddings
(tree-sitter + (Anthropic) (sentence-transformers)
regex fallback) │ │
└──────────────┴──────────────┘
│
SQLite + sqlite-vecTwo entry points (CLI and MCP server) share the same core service layer
No DI framework — services are wired via a
create_app()factoryAll core services are async — CLI bridges with
asyncio.run()Embeddings are local — no external API calls for vector search (model auto-downloads on first run, ~90 MB)
Single SQLite database — 6 main tables + 2 vector tables (384-dimensional embeddings)
Development
git clone https://github.com/mrosata/codetex-mcp.git
cd codetex-mcp
# Install dependencies (including dev)
uv sync
# Run tests
uv run pytest
# Run tests with coverage
uv run pytest --cov=codetex_mcp
# Lint and format
uv run ruff check src/ tests/
uv run ruff format src/ tests/
# Type check
uv run mypy src/Releasing
Releases are automated via GitHub Actions and python-semantic-release. Version bumps are driven by conventional commit messages on main.
Commit message format
Prefix | Effect | Example |
| Patch bump (0.1.0 → 0.1.1) |
|
| Minor bump (0.1.0 → 0.2.0) |
|
| Major bump (0.1.0 → 1.0.0) |
|
| No release |
|
A BREAKING CHANGE: line in the commit body also triggers a major bump.
How it works
Push or merge a PR to
mainCI runs lint, type check, and tests
The release workflow analyzes commits since the last tag
If a version bump is needed, it:
Updates the version in
pyproject.tomlCreates a git tag (e.g.,
v0.2.0)Publishes a GitHub Release with a changelog
Builds and publishes the package to PyPI
Manual release (not recommended)
If you need to release without the automation:
uv build
uv publishLicense
MIT
Available Tools
7 toolsget_file_contextC
Return the Tier 2 file summary for a specific file.
Includes file purpose, public interfaces, dependencies, role classification, line count, and token count.
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | ||
| repo_name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It implies a read-only operation but does not state it explicitly, nor does it describe behavior for missing files, repo state requirements, or side effects. The return content list is useful but does not cover these behavioral expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded with the core purpose, and the second sentence lists the included fields efficiently. There is no filler, though the unexplained 'Tier 2' term is a slight clarity cost.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple and has an output schema, but with no annotations and empty schema descriptions, the description must cover more ground. It omits usage context, parameter details, and any relationship to sibling tools like sync_repo or get_repo_status, leaving the agent to infer important context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description should compensate, but it only indirectly addresses file_path ('for a specific file') and says nothing about repo_name, path formats, or relationships between parameters. The parameter names are self-explanatory, but the description adds minimal semantic value beyond them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Return') and names the resource ('Tier 2 file summary'), and the content list ('file purpose, public interfaces, dependencies, role classification, line count, token count') clarifies what is returned. It does not explicitly name sibling tools, but the resource type is clear enough to distinguish it from repo-level or symbol-level tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use guidance, no mention of alternatives like get_symbol_detail or search_context, and no prerequisites such as whether the repo needs to be synced first. The intended usage is only implied by the purpose statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_repo_overviewA
Return the Tier 1 markdown overview for a repository.
Includes repository purpose, directory structure, key technologies, entry points, and architecture patterns.
| Name | Required | Description | Default |
|---|---|---|---|
| repo_name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It describes the content returned (purpose, structure, etc.) but does not explicitly state that the operation is read-only, mention potential errors (e.g., repo not found), or clarify what 'Tier 1' means. For a simple read tool, this is adequate but not thorough; it does not contradict anything since no annotations exist.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two lines, front-loads the core purpose ('Return the Tier 1 markdown overview'), and lists the included elements efficiently. It avoids fluff, though the term 'Tier 1' is unexplained and slightly opaque. Overall, it is concise and well-structured, earning a 4.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has an output schema (indicated by context), return-value details are likely covered there, so the description need not restate them. The description covers the essential content and is sufficient for a simple read operation with one parameter. However, it lacks any error-handling or edge-case information, and does not mention freshness or caching, which would be useful for a complete picture.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds no explicit explanation of the 'repo_name' parameter beyond its name, which is self-evident from the tool name and context. The description does not provide format, constraints, or examples, so it adds minimal value over the raw schema. A score of 3 reflects that the parameter is trivial and self-explanatory, but the description still fails to actively clarify it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('Return') and resource ('markdown overview for a repository'), and enumerates the content (purpose, directory structure, key technologies, entry points, architecture patterns). This distinguishes it from siblings like get_file_context (file-level) and search_context (search) without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus alternatives. It does not mention sibling tools, nor does it provide conditions like 'for high-level overview' or 'instead of get_file_context'. An agent is left to infer usage solely from the tool name and description, which is insufficient for proper selection among six siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_repo_statusB
Return index status for a repository.
Includes indexed commit, current HEAD, staleness indicator, file count, symbol count, total tokens, and last indexed time.
| Name | Required | Description | Default |
|---|---|---|---|
| repo_name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description carries the full burden of disclosing behavior. It makes clear this is a read-only status operation by saying 'Return index status' and listing output fields. However, it leaves the staleness indicator undefined and says nothing about error conditions, permissions, or whether the status could trigger indexing side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loads the main purpose, and uses a compact list of return fields. Every sentence earns its place. There is no filler or redundant restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter status tool with an output schema, the description covers the essential return content. It is incomplete in two areas: it does not clarify the staleness indicator's meaning, and it gives no sibling differentiation, which matters because get_repo_overview appears to be a closely related read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description never mentions repo_name or its expected format. The parameter is self-explanatory at a surface level, but there is no guidance on whether it should be a bare name, an owner/repo pair, or some other identifier. The description focuses entirely on the return payload rather than the input.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a clear verb and resource: 'Return index status for a repository.' It also enumerates the specific fields returned (indexed commit, HEAD, staleness, counts, tokens, time), which gives the agent a concrete sense of purpose. It does not explicitly contrast itself with siblings like get_repo_overview, but the status-focused scope is reasonably distinct.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives such as get_repo_overview or sync_repo. It does not mention exclusions, prerequisites, or situations where another sibling would be more appropriate. Usage context is only implied by the tool name and listed fields.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_symbol_detailC
Return the Tier 3 detail for a specific symbol.
Includes full signature, description, parameters with types, return type, call relationships, and file location.
| Name | Required | Description | Default |
|---|---|---|---|
| repo_name | Yes | ||
| symbol_name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the behavioral burden. It does disclose what the call returns (signature, parameters, relationships, file location) and implies a read-only operation, but it does not mention failure modes, exact-match requirements, permissions, or cost. This is adequate but not fully transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded with the core purpose, followed by a useful content list. It is efficient, but the unexplained 'Tier 3' terminology introduces slight ambiguity and prevents a perfect structure score.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although an output schema exists, the description lacks information about repo_name format, symbol_name disambiguation, whether the repo must be synced, and what 'Tier 3' means relative to the sibling tools. These gaps are significant for an agent deciding whether and how to invoke the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the parameters. It does not explain repo_name or symbol_name formats, ambiguities, or relationships, even though the parameter names are somewhat self-explanatory. The mention of 'parameters with types' refers to the returned symbol's parameters, not the tool's own inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Return') and resource ('Tier 3 detail for a specific symbol'), and it enumerates the retrieved content. It does not explicitly contrast itself with siblings like get_file_context or search_context, so it stops short of full differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives such as get_repo_overview, get_file_context, or search_context. No prerequisites, exclusions, or conditions are provided; usage is only implied by the tool's name and basic purpose.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_reposA
List all registered repositories with their status.
Returns a markdown table of repository names, remote URLs, indexed commits, and file counts.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are absent, so the description carries the behavioral disclosure burden. It does disclose the output format ('markdown table') and fields returned, which adds value. However, it does not explicitly state read-only behavior, absence of side effects, pagination, or any constraints, though 'list' implies a safe read.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with the core purpose front-loaded in the first sentence and output details in the second. There is no filler, redundancy, or unnecessary qualification.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter listing tool with an output schema, the description sufficiently covers the tool's purpose and return content. Missing usage guidance versus siblings and an explicit read-only statement are minor gaps given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and the schema is empty, so the baseline is 4. No parameter explanation is needed; the description instead clarifies what the returned table contains, which is appropriate given the empty input schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('List') and resource ('all registered repositories') and enumerates the output columns ('repository names, remote URLs, indexed commits, and file counts'). Sibling differentiation is implicit through the word 'all' versus the more targeted sibling names, but no sibling is explicitly named, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance about when to choose this tool over siblings like get_repo_status or get_repo_overview. It only says what it does, not the context or exclusions that should drive selection. An agent must infer usage from tool names alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_contextC
Search for relevant code context using semantic similarity.
Returns a ranked list of matching files and symbols with relevance scores.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ||
| repo_name | Yes | ||
| max_results | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It mentions the return type and ranking but does not state whether the operation is read-only, any rate limits, required repository state (e.g., synced), or other side effects. For a search tool, the lack of explicit read-only or safety information is a notable gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences that front-load the core purpose and return format. There is no wasted wording, and the structure is easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has three parameters, no annotations, and an output schema (not detailed), the description is minimal. It does not mention parameter usage, potential prerequisites (e.g., repository must be synced), or any nuances of the ranking. While the output schema covers return structure, the description still leaves critical context unaddressed for a tool with this complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate by explaining parameters. It does not. Although parameter names (repo_name, query, max_results) are self-explanatory, the description adds no detail on their semantics, expected formats, or constraints, leaving the agent to rely solely on the schema names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('Search'), resource ('code context'), and method ('semantic similarity'), and specifies the return type ('ranked list of matching files and symbols with relevance scores'). It distinguishes from siblings like get_file_context and get_symbol_detail by emphasizing semantic similarity rather than direct retrieval, though it does not explicitly name alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus the sibling tools. The description implies it is for semantic search but does not state explicit conditions, alternatives, or exclusions, leaving the agent to infer usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sync_repoB
Trigger an incremental sync for a repository.
Processes only files changed since the last indexed commit. Returns a summary of changes made.
| Name | Required | Description | Default |
|---|---|---|---|
| repo_name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It explains the incremental nature and that a summary is returned, but it does not mention side effects, required permissions, whether the operation is destructive, or if it blocks or is async. This is a notable gap for a mutation-like action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three concise sentences with no waste. It front-loads the action, then explains scope and return value. Every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool with an output schema, the description covers the core purpose and return value. However, it omits prerequisites (e.g., repo must already be indexed), potential side effects, and error conditions. It is minimally adequate but not rich enough for a tool that triggers a process.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the repo_name parameter. It only mentions 'repository' generically, without clarifying the format, source, or relationship to the parameter. The meaning is mostly derived from the parameter name itself, not enhanced by the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Trigger an incremental sync for a repository.' The 'incremental' qualifier and the distinction from sibling read-only tools (get_repo_status, get_repo_overview, search_context) make its purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool ('Processes only files changed since the last indexed commit') but does not explicitly say when to prefer it over alternatives or provide when-not-to-use guidance. The context is clear but exclusions are absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v0.5.0- First observed
get_file_context - First observed
get_repo_overview - First observed
get_repo_status - First observed
get_symbol_detail - First observed
list_repos - First observed
search_context - First observed
sync_repo
TDQS
Scored across 7 tools
Each tool targets a distinct concern: repo-level overview, file context, symbol detail, semantic search, status, syncing, and repo listing. There is no meaningful overlap or ambiguity between tool boundaries.
Tool names consistently follow a snake_case verb_noun pattern, mostly using get_ for retrieval actions. Minor variation like search_context and sync_repo still fit the same predictable style.
Seven tools is well-scoped for a code context and indexing server. Each tool earns its place by covering a distinct retrieval or maintenance operation without unnecessary bloat.
The set covers repository overview, file context, symbol detail, semantic search, status, sync, and listing—covering the core workflows well. The main gap is the lack of explicit repo registration or full-sync tools, though these may be handled externally.
Maintenance
Related MCP Connectors
Code intelligence for LLMs. Analyze, search, and retrieve code from any public git repository.
Codebase graphs, caller impact analysis, and recorded project context for AI coding agents.
Intelligent context infrastructure for AI teams: knowledge graph, sessions, tasks, documents.
Project memory, semantic code search, and grounded agent context.
Related MCP Servers
- AlicenseAqualityBmaintenanceA semantic code context server that connects your local repository to AI assistants via the Model Context Protocol, enabling dynamic codebase exploration without copy-pasting.590 npm3MIT
- AlicenseBqualityDmaintenanceProvides IDE-like semantic code retrieval and editing tools for LLMs, enabling precise code understanding and manipulation in large codebases via the Model Context Protocol.25MIT
- AlicenseNot gradedqualityBmaintenanceTurns a codebase into a queryable graph with semantic search, call graphs, and control/data flow analysis, served to AI coding agents via the Model Context Protocol.206 npmMIT
- FlicenseNot gradedqualityBmaintenanceTransforms local Git repositories into queryable, context-rich knowledge bases via AST-aware chunking and Git metadata, enabling AI assistants to search and understand codebases with semantic precision.-