knot
Summary: knot's MCP server gives AI agents read-only code intelligence over indexed repositories — semantic search, reverse-dependency lookup, file anatomy inspection, and cross-repo dependency discovery, backed by Qdrant (vectors) and Neo4j (graph).
search_hybrid_context— Semantic + structural code search by meaning ("user authentication"), returning files, signatures, docstrings, and dependencies. Best first step for exploring unfamiliar code; supports multi-repo scope and 1–20 results.find_callers— Reverse dependency lookup: who calls, extends, or implements an entity. Used for impact analysis, dead-code detection, refactoring safety, and (for JVM/C#) override discovery. Handles homonyms by grouping results per target.explore_file— File anatomy outline: all classes, interfaces, methods, and properties with signatures, docstrings, and line numbers, without reading the whole file.list_repositories— Inventory of indexed repos with entity/file counts, build system, and language; optional case-insensitive name filter.list_repo_dependencies— Cross-repo dependency graph (Maven, Gradle, Cargo, npm, NuGet) with forward/reverse lookup and transitive depth up to 10.Cross-cutting capabilities — Multi-language support (Java, Kotlin, Rust, C#, TypeScript/JS, Python, C/C++, Groovy, HTML/CSS, Markdown, build/config files), cross-repo linking, and repository scoping via single name, comma list, or
all/*.Constraints — Read-only with no side effects; requires a running knot-mcp with Qdrant and Neo4j initialized and target repos already indexed.
Supports indexing and analysis of Angular web components through HTML parsing, enabling cross-language linking between JavaScript, HTML, and CSS for full-stack SPA analysis.
Provides CSS/SCSS stylesheet indexing with class/ID selector extraction and variable tracking, enabling unified HTML/CSS discovery and cross-language search capabilities.
Supports CommonJS TypeScript file analysis as part of complete TypeScript/TSX/CTS language support for modern JavaScript/TypeScript codebases.
Provides Docker-based deployment options for universal compatibility across platforms, including containerized execution of the indexer, CLI tool, and MCP server.
Supports configuration via .env files for setting repository paths and database credentials during codebase indexing and analysis.
Provides comprehensive JavaScript/Node.js analysis including vanilla JS, Node.js, and module systems (.js, .mjs, .cjs, .jsx) with full cross-language linking capabilities.
Provides complete Kotlin codebase analysis with support for classes, interfaces, objects, companion objects, functions, methods, and properties using tree-sitter-kotlin-ng grammar.
Uses Markdown for documentation and skill files, including .knot-agent.md for teaching LLMs how to use the CLI tool for autonomous code analysis.
Integrates with Neo4j graph database for storing architectural relationships via call graphs, enabling structural navigation and reverse dependency analysis.
Supports Node.js module systems and JavaScript analysis as part of the hybrid web ecosystem with cross-language linking capabilities.
Extracts id and className attributes from JSX/TSX React components for unified HTML/CSS discovery and cross-language analysis in web applications.
Provides complete TypeScript/TSX analysis including modern JavaScript/TypeScript codebases with full cross-language linking and architectural relationship extraction.
knot
knot is a high-performance codebase indexer that extracts structural and semantic information from source code, enabling AI agents to understand, analyze, and navigate large code repositories. Currently supports Java, Kotlin, TypeScript, JavaScript/Node.js, Rust, Python, Groovy, C/C++, C#, HTML, and CSS/SCSS, plus Build Systems (Maven pom.xml, Gradle build.gradle, Jenkins pipeline, Cargo.toml, MSBuild .csproj + Directory.Packages.props), Configuration Files (YAML, JSON, .properties — optional), Kubernetes + Helm (optional), and Cross-Repo Dependency Linking with full cross-language linking.
For recent release notes see CHANGELOG.md.
The indexer automatically builds:
Vector Search Database (Qdrant) — semantic understanding via embeddings
Graph Database (Neo4j) — architectural relationships via call graphs
This dual-database approach powers both:
MCP (Model Context Protocol) Server — Exposes three tools to any LLM client (Claude, Gemini, ChatGPT, Cursor, etc.)
CLI Tool — Standalone
knotcommand for terminal and scripting environments
Knot in action
CLI — instant reverse dependency lookup
MCP — JSON-RPC protocol for AI agents
🧮 Token Efficiency — Measured, Not Claimed
An LLM agent exploring an unfamiliar codebase pays for every byte it reads. Without an index it greps and then reads whole files; with knot it receives a targeted answer. The difference was measured on three real indexed repositories across nine realistic exploration tasks:
Repo | Lang | Task | knot tokens | Read-the-code tokens | Reduction |
spring-ai | Java | discovery — how does the chat client run the advisor chain? | 1 092 | 10 168 | 89.3% |
spring-ai | Java | callers — who uses | 8 808 | 15 554 | 43.4% |
spring-ai | Java | explore — structure of | 4 865 | 7 838 | 37.9% |
puppeteer | TypeScript | discovery — how is a CDP session created? | 609 | 4 149 | 85.3% |
puppeteer | TypeScript | callers — who calls | 1 004 | 39 878 | 97.5% |
puppeteer | TypeScript | explore — structure of the | 7 287 | 25 300 | 71.2% |
knot | Rust | discovery — how are call intents resolved? | 594 | 14 824 | 96.0% |
knot | Rust | callers — who calls | 461 | 10 949 | 95.8% |
knot | Rust | explore — structure of the graph query module | 978 | 12 103 | 91.9% |
TOTAL | — | 9 tasks | 25 698 | 140 763 | 81.7% |
≈ 5.5× fewer tokens for the same nine questions — 115 000 tokens saved, enough to keep a long refactoring session inside a single context window.
Both sides are measured on the exact bytes an LLM would receive as tool
output, counted with OpenAI's cl100k_base tokenizer (tiktoken):
Task | knot side | Read-the-code side |
|
|
|
|
|
|
|
| full read of the file |
The baseline is deliberately generous, so the measured saving is a lower bound:
greps are restricted to the source files of the language (
-t java,-t ts,-t rust) — no changelogs, no generated docs, nonode_modules;for
discoverythe baseline is given oracle file selection: it reads only the files that answer the question, with zero wasted reads;for
callersit reads at most 5 files, while a rigorous impact analysis would need every file with a textual hit.
Honest caveats: knot's cost scales with the number of results, not with repo
size. The weakest row (spring-ai / ToolCallingManager, 43%) is a symbol with
156 references — knot enumerates all of them with exact call sites, while the
capped baseline reads only 5 files and still cannot tell a call from a comment.
The explore rows for large classes are also the least favourable, because
signatures plus docstrings are a large fraction of a well-documented file.
Repositories measured (as indexed): spring-ai 2 406 files / 25 733 entities,
puppeteer 1 832 files / 19 310 entities, knot 222 files / 4 000 entities.
Raw measurements are stored in
.perf_metrics/token_savings.json.
pip install tiktoken # optional: falls back to a chars/4 estimate
# edit the `root` paths in scripts/token_savings_tasks.json to match your checkouts
python3 scripts/token_savings_benchmark.py \
--config scripts/token_savings_tasks.json \
--save-json .perf_metrics/token_savings.jsonThe task definitions live in
scripts/token_savings_tasks.json and the
harness in
scripts/token_savings_benchmark.py;
point them at any repository you have indexed to measure your own codebase.
Related MCP server: MCP-RAG
✨ Key Features
🔍 Code Intelligence Tools
search_hybrid_context: Semantic + structural search. Find code by meaning, class name, method signature, docstrings, or comments. Returns full context including dependencies.find_callers: Reverse dependency lookup. Identify dead code, perform impact analysis, or understand the full call chain of any function/method. When multiple entities share the same name (e.g.,find_nearest_entity_by_linein different files), results are automatically grouped by target showing which specific entity each caller references. Supports cross-repository call resolution viaDEPENDS_ONgraph edges. For JVM languages (Java/Kotlin/Groovy) it also surfaces method-levelOVERRIDESedges bidirectionally — an Overridden by group listing subtype implementations/overrides and an Overrides group listing the supertype methods a method implements/overrides.explore_file: File anatomy inspection. Quickly see all classes, interfaces, methods, and functions in a file with signatures and documentation.list_repo_dependencies(MCP) /knot deps(CLI): Dependency graph visualization. Show which repositories depend on each other, forward and reverse, with transitive resolution.list_repositories/knot repos: Repository inventory. List every indexed repository along with its entity count, file count, build system, and primary language. Supports optional case-insensitive name filtering via--filter(CLI) orfilterparameter (MCP). Useful for orientation, sanity-checking indexing runs, and discovering which languages and build systems are present in the workspace.
🏗️ Multi-Language Support
Java: Full AST extraction with package-aware FQN resolution (e.g.,
com.example.app.UserService), class inheritance (EXTENDS), interface implementation (IMPLEMENTS), annotation tracking, and field-access method invocation resolutionKotlin: Complete support for Kotlin codebases with classes, interfaces, objects, companion objects, functions, methods, and properties. Fully compatible with tree-sitter-kotlin-ng grammar.
C#: Full C# support via
tree-sitter-c-sharp. Extracts classes, interfaces, structs, records (bothrecord classandrecord struct), enums, methods, constructors, properties, fields (withconstdetection), delegates, events, indexers, operators, local functions, and namespaces withCSharp*entity kinds. Namespace-qualified FQNs (MyApp.Services.UserService.GetUserAsync) work across both file-scoped (C# 10+) and block-form namespaces, including nested namespaces and nested types. Thebase_listheuristic splits: Base, IFaceintoEXTENDS/IMPLEMENTSusing theIPascalCaseconvention (structs and interfaces are deterministic), generic arguments are stripped (IRepository<User>→IRepository), XML doc comments (///) become docstrings, and attributes ([Obsolete]) are captured as decorators. Calls through field-typed receivers resolve to the exact implementation method, and C#virtual/overrideplus interface implementation produce method-levelOVERRIDESedges. MSBuild/NuGet:.csprojfiles are parsed for project identity and dependencies (Central Package Management viaDirectory.Packages.propsis supported); C# repos getbuild_system: "nuget"in the Repository node instead of the prior"none".TypeScript/TSX/CTS: Complete support for modern JavaScript/TypeScript codebases, including CommonJS TypeScript files
JavaScript/Node.js: Vanilla JS, Node.js, and module systems (
.js,.mjs,.cjs,.jsx)Hybrid Web Ecosystem: Cross-language linking between JavaScript, HTML, and CSS for full-stack SPA analysis
HTML: Custom elements (Web Components, Angular),
idandclassattribute indexing for cross-language CSS searchJSX/TSX Attributes: Extracts
idandclassNamefrom React components for unified HTML/CSS discoveryCSS/SCSS: Stylesheet indexing with class/ID selector extraction and variable tracking (CSS/SCSS variables, mixins, functions)
Rust: Struct, enum, union, trait, function, method, module extraction with trait implementation tracking (IMPLEMENTS relationships) and macro invocation references. Methods are indexed with the qualified FQN
Type::method(e.g.,KnotMcpHandler::new,WidgetA::new,Logger::new) and qualified calls from top-level functions resolve to the right target by receiver. Braced import/use capture —use foo::{Bar, Baz}anduse foo::Bar as Bazproduce explicit REFERENCES edges for all imported names, including traits imported solely to bring methods into scope. All Rust entity FQNs are now anchored at the owning crate and module path (e.g.knot::config::Config,knot::pipeline::parser::languages::rust::qualify_rust_fqns), so two crates that declare a type with the same bare name no longer collide. Files outsidesrc/(tests, benches, examples) receive a__fixture::<path>::<Entity>FQN prefix (e.g.__fixture::tests::testing_files::sample::Config), and files without aCargo.tomlancestor receive__loose::<path>::<Entity>, preventing name collisions with real source entities. CONTAINS relationships useenclosing_class_fqnfor exact disambiguation when multiple entities share the same class name. The on-disk index state file (.knot/index_state.json) carries aversionfield; opening a state file from an older version prints an error with instructions to runknot-indexer --clean.Python: Full Python extraction with class, function, method support, constants, module-level imports,
ValueReferencetracking for keyword arguments, class inheritance (EXTENDS), decorator extraction (@property,@staticmethod,@route(...),@dataclass), generic type hints (List[str],Optional[Dict],*args/**kwargs), Py2/Py3 exception syntax compatibility, andself.method()resolution with inherited method walking. Capturesclass_definition,function_definition(including async via optionalasyncmodifier), lambda assignments, and distinguishes methods from functions via parent context detection. Class instantiation (ClassName(...)) is automatically redirected toClassName.__init__sofind_callers ClassName.__init__lists every constructor call site (with fallback to inherited__init__via the extends chain); only class/struct kinds trigger the redirect — functions keep the legacy behavior.Groovy: Full Groovy language support via hybrid tree-sitter + ad-hoc lexical parser. Extracts classes, interfaces, traits, enums, typed/
def/quoted methods (incl. Spock specs), constructors, closures, script-level variables, fields/properties with visibility modifiers, nested classes, and decorators. Tracks package FQN and enclosing class relationships. Multi-line signatures (closure default params), assignment-vs-declaration disambiguation, innermost assignment for nested closures, UUID collision fix for duplicate method names,find_callersaccurately tracks private methods including those in anonymousnew AnActionclosures. Inheritance tracking: emitsEXTENDS/IMPLEMENTSreference intents forclass/interface/trait/enumheaders (single-line and multi-line) sofind_callerssurfaces real nextflow-style hierarchies — qualified parents (e.g.extends nextflow.plugin.BasePlugin) and generic-argument stripping (e.g.extends AbstractRepo<Order, Long> → extends AbstractRepo) are supported, and generic bounds (class Box<T extends Comparable>) are correctly not promoted to inheritance edges. Property accessors: bare property declarations (Path baseDir,boolean cacheable) are now indexed asGroovyProperty, and compiler-generatedgetX/setX/isXaccessors are synthesised as first-class method entities soOVERRIDESedges link Groovy properties to interface getter declarations. Comment-stripping prevents Javadoc continuation lines (* The pipeline script name) from producing phantom entities or corrupting scope tracking.Build Systems: Maven
pom.xml(dependencies + plugins via roxmltree), Gradlebuild.gradle(deps + plugins + tasks),Jenkinsfilepipeline (stages + steps), CargoCargo.toml(deps + workspace members + features), and MSBuild.csproj/Directory.Packages.propsextraction. MSBuild resolves project identity (<PackageId>→<AssemblyName>→ file stem), emits aBuildDependencyper<PackageReference>(attribute-form and version-less), and resolves Central Package Management versions from the nearestDirectory.Packages.propsancestor. UTF-8 BOMs are tolerated defensively. Identity markeridentity: package_idis carried in the signature when the project has an explicit<PackageId>so the cross-repo resolver prefers published packages over depth-tied unmarked candidates.Cargo.toml: Rust package manager support with package metadata, features, workspace members, and multi-format dependency parsing (simple, table, git, path).
Configuration Files: YAML (.yml/.yaml), JSON (.json), and Java Properties (.properties) with leaf-key granularity. Special handling for package.json (npm dependencies as BuildDependency, scripts as ConfigProperty).
Varnish Cache: Hand-written parsers for
.vcl(configuration),.vtc(test cases), and.vcc(VMOD C source). VCL extracts backends, probes, ACLs, subroutines (custom + built-in withvcl_*names, including aggregator entities for multi-part built-ins),importdirectives (withasaliases andfrompaths),includeedges,unuseddeclarations, VMOD instantiations, andreq.backend_hintassignments. VTC extractsvarnishtest/vtestcases, servers, clients, varnish instances, logexpect blocks, barriers, and-vcl+backendsynthesised backends (withis_test_context). VCC extracts$Module,$Function,$Object,$Method,$Event,$Restrict, ENUMs, and default parameters. References:Calls,Extends,Implements,References(with intentsVclSubCall,VclBackendRef,VclProbeRef,VclAclRef,VclInclude,VclVmodImport,VclUnusedRef,ValueReference); relationships:UsesBackend,UsesProbe,UsesAcl,Includes,ImportsVmod,DeclaredUnused. The Fastly VCL dialect is detected and skipped (returns empty entities).Kubernetes + Helm: K8s manifest parsing (Deployment, Service, ConfigMap, Secret, Ingress, Namespace) with label/annotation tracking and cross-resource references. Helm chart indexing (Chart.yaml metadata, values.yaml key-value pairs, template variable extraction via {{ .Values.X }}).
C/C++: Complete C/C++ support with namespace-aware FQN resolution (
Engine::MyClass::start), class/struct extraction, function/method tracking, macro definition and usage detection (uppercase identifier heuristic), type reference tracking (declarations,newexpressions), and full call graph analysis. Supports.c,.h,.cpp,.hpp,.cc,.cxx,.hh,.hxxextensions via tree-sitter-c and tree-sitter-cpp parsers. Includes intelligent auto-detection for.hheaders to parse them correctly as C or C++ based on their contents.Markdown: Documentation indexing with
MarkdownDocument(one per.md/.markdownfile) andMarkdownSection(one per ATX heading H1–H6). Section bodies — including paragraphs, fenced code blocks, lists, and tables — are captured intoembed_textfor full semantic search over documentation content, not just heading titles. FQNs are hierarchical and file-scoped (e.g.README.md::Setup > Installation > Linux), so same-named headings in different files or under different parents disambiguate cleanly. Section boundaries respect heading depth: a section's body extends until the next heading of equal or higher level, ensuring### Linuxunder## Installationdoes not bleed into a sibling## Configuration. Headings with inline markdown (backticks, em-dash, links, emoji) parse without losing their bodies, and realstart_line/end_linepositions are computed via tree-sitter for each section.
📚 Rich Comment Extraction
Captures docstrings (JavaDoc, JSDoc) preceding declarations
Extracts inline comments within method/function bodies
Respects nesting boundaries (class comments don't capture method comments)
Intelligently aggregates comment blocks
📊 Dual-Database Architecture
Qdrant: Vector search for semantic code understanding
Neo4j: Graph relationships for structural navigation
🚀 High Performance
Parallel Streaming Pipeline: Overlaps CPU-bound embedding with I/O-bound ingestion via MPSC channels
Incremental Indexing: Uses SHA-256 hashes to skip unchanged files
Real-time Watch Mode: Automatically re-indexes changed files in seconds via
--watchCPU Parallelism: AST extraction via Rayon
Scalable: Configurable batch processing and constant memory footprint (~2GB) regardless of repository size
Performance Benchmarking: Multi-level validation framework
Unit benchmarks: Criterion-based benchmarks for parse, embed, and graph write throughput (
benches/)E2E benchmarks: Full pipeline metrics capture with per-stage timing (
tests/benchmark_e2e.sh)CI regression tracking: Automated baseline comparison against tolerance thresholds (
scripts/compare_perf_metrics.sh)Token efficiency: LLM token cost of knot answers vs reading source files (
scripts/token_savings_benchmark.py) — see Token Efficiency
🛠️ Installation
Prerequisites
Component | Version | Notes |
Docker | 20.10+ | For running Qdrant and Neo4j |
qdrant | 1.x | Vector database (docker) |
neo4j | 5.x | Graph database (docker) |
Option A: Pre-compiled Binaries (macOS & Modern Linux)
Go to the Releases page and download the native executable for your platform.
Install knot binaries (CLI, MCP server, and indexer):
curl --proto '=https' --tlsv1.2 -LsSf https://github.com/raultov/knot/releases/latest/download/knot-installer.sh | shInstall agent-skills for your AI (Optional): Paste this into your LLM agent (Claude Code, OpenCode, Cursor, etc.):
Install the knot agent skills by following the instructions at: https://raw.githubusercontent.com/raultov/knot/master/README.md
The first command installs the knot binary to your PATH. The second (optional) allows your AI assistant to automatically download the agent skill index (.knot-agent.md) and run the installer to extract comprehensive guides for using knot CLI with AI agents and code analysis tools.
System Requirements:
Linux: glibc 2.38+ (Ubuntu 24.04+, Debian 13+, Fedora 39+, Arch)
macOS: Modern versions supported
Windows: Use Docker (Option B)
Option B: Docker (Universal Compatibility)
Docker images provide universal compatibility for any Linux distribution and Windows.
Docker Installation (All Binaries)
Build the image:
docker build -t knot:latest . --network=hostRun the indexer:
# Use --network host to connect to databases running on your host machine
docker run --rm \
-v /path/to/your/repo:/workspace \
-e KNOT_REPO_PATH=/workspace \
-e KNOT_NEO4J_PASSWORD=your-password \
--network host \
knot:latest \
knot-indexerRun the CLI tool:
docker run --rm \
-v /path/to/your/repo:/workspace \
-e KNOT_REPO_PATH=/workspace \
-e KNOT_NEO4J_PASSWORD=your-password \
--network host \
knot:latest \
knot search "user login flow"Run the MCP server:
docker run --rm \
-e KNOT_REPO_PATH=/workspace \
-e KNOT_NEO4J_PASSWORD=your-password \
--network host \
knot:latest \
knot-mcpNote: Uses Debian Trixie (glibc 2.38+) and includes ONNX Runtime for full functionality.
Option C: Install via Cargo
cargo install --git https://github.com/raultov/knotOption D: Build from Source
1. Start infrastructure with Docker:
docker compose up -d2. Clone and build:
git clone https://github.com/raultov/knot
cd knot
cargo build --release3. Configure:
mkdir -p ~/.config/knot
cp .env.example ~/.config/knot/.env
$EDITOR ~/.config/knot/.env # Set KNOT_REPO_PATH and Neo4j credentials4. Index a codebase:
./target/release/knot-indexer5. Query via CLI:
./target/release/knot search "your query"6. Start the MCP server:
./target/release/knot-mcp📖 Usage
🤖 Install Agent Skills (For AI Agents)
Option A: Let an LLM do it
Paste this into any LLM agent (Claude Code, OpenCode, Cursor, etc.):
Install the knot agent skills by following the instructions at: https://raw.githubusercontent.com/raultov/knot/master/README.md
Option B: Terminal (Manual)
curl -sO https://raw.githubusercontent.com/raultov/knot/master/.knot-agent.md && curl -fsSL https://raw.githubusercontent.com/raultov/knot/master/scripts/install-agent-skills.sh | bash📥 Quick Downloads (Binaries)
Download knot binaries (CLI + MCP server):
curl --proto '=https' --tlsv1.2 -LsSf https://github.com/raultov/knot/releases/latest/download/knot-installer.sh | sh📖 Agent-Skills Guides
Comprehensive documentation for using knot tools. The agent skills installer extracts:
search.md — Semantic code discovery guide with examples
callers.md — Reverse dependency lookup with critical usage rules
explore.md — File anatomy inspection guide
deps.md — Repository dependency graph guide
repos.md — Indexed repository inventory
workflows.md — Common patterns and best practices
For quick reference without downloading, see .knot-agent.md.
Using the CLI
The knot CLI provides the same capabilities as the MCP server via command-line commands, making it ideal for:
Terminal-only environments
Bash scripting and automation
CI/CD pipelines
Direct integration with other tools
Three main commands:
knot search — Semantic Code Search
knot search "user authentication" --max-results 10 --repo my-app
knot search "user authentication" --max-results 20 --repo "app-a,app-b" # Union across repos
knot search "user authentication" --max-results 20 --repo all # All indexed repos ('all' or '*')Find code entities by meaning, class names, docstrings, or comments.
knot callers — Reverse Dependency Lookup
knot callers "LoginService" --repo my-app
knot callers "LoginService" --repo "auth-service,billing-service"
knot callers "LoginService" --repo allFind all code that references a specific entity (dead code detection, impact analysis, call chains). When multiple entities share the same name in different files, results are automatically grouped by target with file locations and signatures.
Every caller entry is self-labeling: the owning repository is printed next to each row as (repo: <name>) — in the CLI table, the Markdown answer, and the resolution block — so rows stay attributable when the scope spans multiple repositories:
# References to `LoginService`
Resolved to 1 target by exact name match:
- `auth::service::LoginService` (class) at `src/service.rs:12` (repo: auth-service)
Found 1 reference(s) across all relationship types:
## Calls (1)
- **`signup`** (function) at `src/handlers.rs:88` (repo: auth-service)In the CLI table the Target column is labeled only for genuine cross-repo references (a caller in repo A referencing a target in repo B); the Caller column is always labeled when a repository is known.
knot explore — File Structure Inspection
knot explore "src/services/auth.ts" --repo my-appList all classes, methods, functions in a file with signatures and documentation.
knot deps — Repository Dependency Graph
knot deps my-app --depth 2 # Show forward dependencies (transitive)
knot deps my-app --reverse # Show who depends on this repoVisualize auto-discovered dependencies between indexed repositories with transitive resolution up to 3 levels deep.
knot repos — List Indexed Repositories
knot repos # Table with REPO / BUILD SYSTEM / LANGUAGE / FILES / ENTITIES
knot repos --filter app # Case-insensitive name filter (substring match)
knot repos --output json # Machine-readable list
knot repos --output markdown # GFM table for chat UIsShow the status of every repository currently indexed in the graph database — useful for orientation, sanity-checking that an indexing run completed, and discovering which languages and build systems are present across the workspace. Use --filter to quickly locate a specific repository when working with multiple indexed codebases.
Repository Scope Selection:
Both the CLI --repo/-r flag and MCP repo_name parameter support:
Single repository name:
--repo my-appComma-separated list:
--repo "repo-a,repo-b"(MCP also accepts["repo-a", "repo-b"])Sentinel:
--repo allor--repo "*"(searches every indexed repository)
Note: Multi-repo scope applies a global max_results limit across the union. Increase --max-results when searching across multiple repositories.
For detailed CLI usage guide, see .knot-agent.md — a machine-readable skill that teaches LLMs how to use knot CLI for autonomous code analysis.
Indexing a Codebase
Incremental Indexing (Default)
# First run: indexes all files
knot-indexer --repo-path /path/to/your/repo --neo4j-password secret
# Subsequent runs: only re-indexes changed files (fast!)
knot-indexer --repo-path /path/to/your/repo --neo4j-password secret
# NEW: Real-time Watch mode
knot-indexer --watch --repo-path /path/to/your/repo --neo4j-password secretHow it works:
Tracks file content via SHA-256 hashes in
.knot/index_state.jsonStores the downloaded
fastembedmodel in.knot/fastembed_cache/to keep the workspace cleanAutomatically detects: modified, added, and deleted files
Only re-parses and re-embeds changed files
Preserves graph relationships to unchanged files
Processes entities in memory-efficient 512-entity chunks
Performance:
Initial index (3800 files): ~60 minutes on standard hardware
Incremental update (3 files changed): ~5-10 seconds
Memory usage: Constant ~2GB regardless of repository size
Full Re-Index (Clean Mode)
# Force complete re-index (deletes all existing data)
knot-indexer --clean --repo-path /path/to/your/repo --neo4j-password secretUse --clean when:
You want to rebuild the entire index from scratch
You've changed Tree-sitter queries or embedding models
Troubleshooting indexing issues
Upgrade note (v1.5.1): File paths are now persisted as repo-relative paths with POSIX separators (e.g.
src/pipeline/embed.rs). Upgrading from v1.4.x triggers an automatic full re-index on first run — the on-disk.knot/index_state.jsoncarries a version field that the loader rejects when stale, andknot-indexerthen wipes the repo from both databases before rebuilding. No manual steps required. Entity UUIDs become machine-independent in the process: the same repo indexed from different checkout locations now produces identical UUIDs.
Indexing Progress
The indexer emits [Progress] log lines showing real-time completion across
the whole pipeline (parsing, embedding, ingestion, reference resolution).
The percentage is monotonically non-decreasing and reaches 100% only
once the run genuinely terminates.
Upgrade note (v1.6.2): The percentage now spans the entire pipeline via weighted bands. Previously it measured only file reading and saturated at
100%within seconds of starting, then froze for minutes while embedding and ingestion were still running. Seedocs/specs/indexing_progress_accuracy_plan.mdfor the full design.
Example with 5000 files where 1000 have been parsed and 5,000 entities are half-way through ingestion:
[Progress] [my-repo] 50.0% — files 5000/5000, entities 41600/83200, batch #325 (128 entities)Band table
Phase | Band | Driver |
|
| — |
Parsing |
|
|
Embedding + Ingestion |
|
|
|
| fixed (no sub-counters available) |
|
| forced |
| last computed value | frozen |
A final log line confirms completion:
[Progress] [my-repo] 100.0% — files 5000/5000, entities 83200/83200 — parsing and ingestion complete, resolving references...Library API (knot-server integration)
Callers that need to observe progress programmatically can use the ProgressTracker:
use std::sync::Arc;
use knot::pipeline::{ProgressTracker, run_indexing_pipeline_with_progress};
let progress = Arc::new(ProgressTracker::new());
let progress_clone = Arc::clone(&progress);
// Poll snapshot() from another task while the pipeline runs
tokio::spawn(async move {
loop {
let snap = progress_clone.snapshot();
println!(
"{:.1}% — files {}/{}, entities {}/{}",
snap.percent_complete,
snap.parsed_files,
snap.total_files,
snap.entities_ingested,
snap.total_entities
);
if snap.stage == IndexingStage::Completed || snap.stage == IndexingStage::Failed {
break;
}
tokio::time::sleep(std::time::Duration::from_millis(500)).await;
}
});
run_indexing_pipeline_with_progress(&cfg, &vdb, &gdb, &mut state, progress).await?;The snapshot() method is thread-safe (read-only locks + atomic loads) and returns a
IndexingProgress struct that serializes directly to JSON for REST endpoints.
Running E2E Integration Tests
To ensure indexer stability, run the E2E integration test suite:
# Run all language E2E tests (TypeScript, Java, JavaScript, Web, Kotlin, Rust, ...)
./tests/run_all_e2e_fast.sh
# Run only Kotlin E2E tests
./tests/run_kotlin_e2e.sh
# Run only Rust E2E tests
./tests/run_rust_e2e.sh
# Run only C# E2E tests
./tests/run_csharp_e2e.sh
# Run only Varnish E2E tests
./tests/run_varnish_e2e.shSee tests/KOTLIN_E2E_TESTS.md for detailed coverage and troubleshooting.
Using the MCP Server
The MCP server exposes five tools to any compatible AI client (built on rust-mcp-sdk 1.1 implementing MCP protocol 2025-11-25 with full tool annotations):
Embedding the tool surface: library consumers can serve the same five tools
from their own transport (e.g. an HTTP /mcp endpoint) without going through
the stdio server. KnotMcpHandler::tools() returns the canonical tool table
(no state or database connection required), and
KnotMcpHandler::dispatch(params) executes a tools/call without needing an
Arc<dyn McpServer> runtime handle. The stdio ServerHandler methods delegate
to these two entry points, so every surface stays identical by construction.
Tool 1: search_hybrid_context
Find code by meaning or keywords
Query: "How is user authentication implemented?"
Result: All auth-related code, signatures, docstrings, and dependenciesCapabilities:
Semantic search by functionality (vector embeddings)
Global multi-repository search by default (
repo_name: "all")Class/method/function name lookup
Docstring and inline comment search
Architectural pattern discovery
Full dependency context
Tool 2: find_callers
Find who calls a specific function
Query: "Find callers of getCurrentTimeInSeconds"
Result: All code that invokes this function + file locationsEach caller entry, target group header, and resolved target carries its repository as (repo: <name>), so results remain attributable under multi-repo scopes (repo_name: "all" or a comma list). The raw JSON (--output json) mirrors this with repo_name (referencing entity) and target_repo_name (referenced entity) fields on every row, plus repo_name on each resolution.targets[] entry.
Advanced: Search by Signature
# Find by full signature (Java)
echo '{"method":"tools/call","params":{"name":"find_callers","arguments":{"entity_name":"registerUser(String"}}}' | knot-mcp
# Find by parameter type (Kotlin)
echo '{"method":"tools/call","params":{"name":"find_callers","arguments":{"entity_name":"findById(Int"}}}' | knot-mcp
# Find by type annotation (TypeScript)
echo '{"method":"tools/call","params":{"name":"find_callers","arguments":{"entity_name":"(EventData"}}}' | knot-mcp
# Find by C# interface method (surfaces implementations + call sites)
echo '{"method":"tools/call","params":{"name":"find_callers","arguments":{"entity_name":"FindByIdAsync"}}}' | knot-mcpUse Cases:
Dead Code Detection: Zero callers = unused code
Impact Analysis: "What breaks if I modify this?"
Refactoring Safety: Find all references before removing
Override Discovery (JVM + C#): For Java/Kotlin/Groovy/C# methods, results include an Overridden by group (implementations/overrides in subtypes) and an Overrides group (the supertype methods a method implements/overrides). These are backed by real
OVERRIDESedges built at index time and resolved transitively at query time, so querying an interface/superclass method surfaces every implementation, and querying an implementation surfaces the declaration it overrides.
Tool 3: explore_file
Understand file structure
Query: "What's in BrowserService.ts?"
Result: All classes, methods, and functions with signatures and docsTool 4: list_repositories
Discover indexed codebases
Query: "What codebases are indexed?"
Result: Markdown table of all indexed repos with entity/file counts, language, and build systemTool 5: list_repo_dependencies
Traverse cross-repository dependency graphs
Query: "What repositories depend on auth-lib?"
Result: Repositories declaring build dependencies (pom.xml, build.gradle, Cargo.toml, package.json, NuGet)🔗 MCP Client Configuration
Supported Clients
knot works with any MCP-compatible AI client:
✅ Claude Desktop (Anthropic)
✅ Gemini CLI (Google)
✅ ChatGPT CLI / GPT (OpenAI)
✅ Cursor (AI IDE)
✅ Any standard MCP client
Configuration Examples
Claude Desktop
Add to claude_desktop_config.json:
{
"mcpServers": {
"knot": {
"command": "/absolute/path/to/knot/target/release/knot-mcp",
"env": {
"KNOT_REPO_PATH": "/path/to/indexed/repo",
"KNOT_QDRANT_URL": "http://localhost:6334",
"KNOT_NEO4J_URI": "bolt://localhost:7687",
"KNOT_NEO4J_USER": "neo4j",
"KNOT_NEO4J_PASSWORD": "your-password"
}
}
}
}Gemini CLI
{
"mcpServers": {
"knot": {
"command": "/absolute/path/to/knot/target/release/knot-mcp",
"env": {
"KNOT_REPO_PATH": "/path/to/indexed/repo",
"KNOT_QDRANT_URL": "http://localhost:6334",
"KNOT_NEO4J_URI": "bolt://localhost:7687",
"KNOT_NEO4J_USER": "neo4j",
"KNOT_NEO4J_PASSWORD": "your-password"
}
}
}
}ChatGPT / GPT CLI
Similar JSON configuration in your client's MCP configuration file.
⚙️ Configuration Reference
All options can be set via CLI flags, environment variables, or a ~/.config/knot/.env file.
Priority (highest to lowest): CLI flags > environment variables > .env file.
Env Variable | CLI Flag | Default | Description |
|
| (required) | Root directory of the repository to index |
|
| (auto-detected) | Repository name for multi-repo isolation (auto-detected from last path component) |
|
|
| Qdrant server URL |
|
|
| Qdrant collection name |
|
|
| Neo4j Bolt URI |
|
|
| Neo4j username |
|
| (required) | Neo4j password |
|
|
| Embedding vector dimension |
|
|
| Entities per batch |
|
|
| Force full re-index (delete all existing data) |
|
| (none) | Path to CA certificate bundle for corporate SSL proxies |
|
|
| Include YAML/JSON/properties/K8s/Helm files in the index |
| (env only) |
| Log level: |
🎨 Custom Tree-sitter Queries
The built-in extraction queries (queries/java.scm, queries/typescript.scm, queries/csharp.scm) can be overridden without recompiling:
KNOT_CUSTOM_QUERIES_PATH=/path/to/my/queries ./target/release/knot-indexerPlace java.scm, typescript.scm, and/or csharp.scm in your custom directory. Missing files fall back to built-in defaults.
🔐 Corporate SSL / CA Certificates
In restricted corporate environments with SSL-inspecting proxies, you may need to provide a custom CA certificate bundle so that knot can download the embedding model from HuggingFace.
Via environment variable:
export KNOT_CUSTOM_CA_CERTS=/etc/ssl/certs/corporate-bundle.pem
./target/release/knot-indexer --repo-path /path/to/repo --neo4j-password secretVia CLI flag:
./target/release/knot-indexer \
--custom-ca-certs /etc/ssl/certs/corporate-bundle.pem \
--repo-path /path/to/repo \
--neo4j-password secretVia .env file:
echo "KNOT_CUSTOM_CA_CERTS=/etc/ssl/certs/corporate-bundle.pem" >> ~/.config/knot/.env
./target/release/knot-indexerThis works for all three binaries: knot-indexer, knot-mcp, and knot.
🔄 Workflow Example
Step 1: Index a Java project
./target/release/knot-indexer --repo-path /home/user/my-java-app --neo4j-password secretStep 2: Query via CLI (Instant search)
./target/release/knot search "authentication logic"
./target/release/knot callers "UserService.login"Step 3: Start MCP server (For AI Agents)
./target/release/knot-mcpStep 4: Use with Claude Desktop
Claude will list the three tools in its Tools menu
Ask: "Search for all authentication logic"
Ask: "Find who calls the login method"
Ask: "Explore the structure of UserService.java"
🤖 Auto-Configuring AI Agents
knot includes a universal .prompt file in its root directory that automatically configures modern AI coding agents (Cursor, Cline, opencode, Claude, etc.) to use the knot-mcp tools correctly.
The directive explicitly instructs AI agents to prioritize:
search_hybrid_context— for semantic code discovery (instead ofgrep)find_callers— for reverse dependency analysis (instead of finding references manually)explore_file— for file structure inspection (instead of reading line-by-line)
This ensures that when you ask an AI agent to analyze, refactor, or understand your code, it leverages the full power of the vector and graph databases rather than falling back to context-blind regex searches. The .prompt file is universal and tool-agnostic, working with any LLM client that reads codebase directives.
🤝 Contributing
Contributions are welcome! Please ensure:
All code passes
cargo clippyandcargo fmtNo new
unsafecode (unsafe_code = "deny"at crate level; one audited exception insrc/utils/mod.rsfor corporate proxy CA bundle injection, documented via#[expect(unsafe_code, reason = "…")])Changes are compatible with Rust 2024 edition
All new functionality includes unit tests
Performance regressions are validated with the benchmark framework before submitting PRs
Development & Code Quality
make check # Run all local quality gates (fmt, clippy, test, dupes)
# Or run gates individually:
cargo clippy --all-targets -- -D warnings # Must pass
cargo fmt -- --check # Must pass
cargo test # Run all unit tests
cargo dupes check # Code duplication checkPerformance Benchmarking
The project includes a three-level benchmarking framework to validate optimizations and detect regressions:
Level 1 — Unit Benchmarks (Criterion):
cargo bench --bench pipeline_bench # Parse + prepare throughput per language
cargo bench --bench graph_upsert_bench # Neo4j UNWIND batching speedup (needs Neo4j)
cargo bench --bench channel_backpressure_bench # Bounded channel overheadLevel 2 — E2E Integration Benchmarks:
# Full pipeline metrics with memory and per-stage timing
./tests/benchmark_e2e.sh --focus rust_e2e --output-dir /tmp/perf_results
# Compare against baseline (fails CI if tolerance exceeded)
scripts/compare_perf_metrics.sh /tmp/perf_results .perf_metrics/baseline.jsonLevel 3 — Token Efficiency Benchmark:
# Measures knot tool output vs grep + file reads on indexed repositories
python3 scripts/token_savings_benchmark.py \
--config scripts/token_savings_tasks.json \
--save-json .perf_metrics/token_savings.jsonUnlike levels 1 and 2 (which measure indexing throughput), this one measures the
consumer side: how many LLM tokens an agent spends to answer a question with
knot versus by reading source files. Requires rg, a built knot binary, the
repositories in the config already indexed, and optionally tiktoken for exact
token counts. See Token Efficiency
for the published results.
Baseline files: .perf_metrics/baseline.json stores the last known good metrics (committed, updated on main/master merges). Tolerance thresholds in .perf_metrics/threshold_tolerances.json control regression gates (±5% time, ±10% memory by default).
CI Integration: The test-performance job in .github/workflows/ci.yml runs after all E2E correctness tests pass, comparing results against baseline and fails the build on regression.
📜 License
This project is licensed under the MIT License. See LICENSE for details.
🚀 Roadmap
For the full release history see CHANGELOG.md.
Upcoming
Long-Term Vision
Go support
IDE plugins (VS Code, IntelliJ, Vim)
Language Server Protocol (LSP) integration
Automated Code Review tool (MCP-based)
Ruby support
💬 Questions?
For issues, feature requests, or discussions, please open a GitHub issue.
Available Tools
5 toolsexplore_fileExplore file anatomyARead-onlyIdempotent
Read-only file anatomy inspection. Use this to list all classes, methods, and properties within a specific source file without reading its entire contents. Provides a structural bird's-eye view of a file, showing entity signatures and docstrings to quickly grasp a module's layout.
Usage: Use AFTER identifying an interesting file via 'search_hybrid_context' to understand its available methods, or before modifying a file. Do NOT use this for searching across multiple files.
Behaviour & Return: Read-only operation. Returns a Markdown-formatted outline of the file's entities, grouped by type (Classes, Methods, Interfaces), including line numbers for direct editor navigation. No side effects.
Path handling: file_path should be a repo-relative path (e.g. 'src/services/user.ts'). Absolute paths under your local checkout are also accepted; the tool strips the known local root automatically. The returned file_path is normalized to the same repo-relative form regardless of how it was queried. If the query is ambiguous across multiple repositories, the answer includes an 'ambiguous_path_candidates' list — retry with a longer path or pass repo_name.
Parameter guidance: 'file_path' must be a relative or absolute path to a valid source file. Include 'repo_name' if the file path might be ambiguous across multiple indexed repositories.
Supports Java, Kotlin, C#, and TypeScript codebases.
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | Path to the source file to explore. PREFERRED: a repo-relative path (e.g. 'src/services/user.ts'). ALSO ACCEPTED: an absolute path under the repository's local checkout (the tool strips KNOT_REPO_PATH / CWD automatically). | |
| repo_name | No | Optional but HIGHLY RECOMMENDED: repository scope. Accepts a single repository name (`'my-repo'`), a comma-separated list (`'repo-a,repo-b'`), or `'all'` (or `'*'`) to query every indexed repository. If you know the repository you are working on, include it in your FIRST query to avoid mixed results from other indexed projects. Omit to search across all repositories. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the 'No side effects' line is partly redundant. However, the description adds genuinely non-structured behavior: the Markdown outline return shape grouped by type with line numbers, path normalization to repo-relative form, and the 'ambiguous_path_candidates' retry behavior. That ambiguity handling is context the annotations cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then cleanly partitioned under Usage / Behaviour & Return / Path handling / Parameter guidance headers. Some redundancy: Path handling and Parameter guidance restate each other and repeat the schema's repo-relative preference, which costs a point against a strict conciseness bar.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by describing the return format (Markdown outline grouped by Classes/Methods/Interfaces with line numbers). It also discloses supported languages (Java, Kotlin, C#, TypeScript) and ambiguity fallback, so nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3 and the schema already documents both parameters. The description exceeds baseline by explaining ambiguity resolution for repo_name (retry with longer path or pass repo_name) and confirming that absolute paths are stripped against KNOWN_REPO_PATH/CWD, which the schema only gestures at.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('list all classes, methods, and properties within a specific source file') and explicitly contrasts with siblings ('Do NOT use this for searching across multiple files'), naming search_hybrid_context as the alternative. An agent can distinguish it from all five siblings without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit sequencing ('Use AFTER identifying an interesting file via search_hybrid_context... or before modifying a file') plus a clear exclusion ('Do NOT use this for searching across multiple files'). Both when-to-use and when-not-to-use are stated, satisfying the highest bar.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_callersFind callers (reverse dependencies)ARead-onlyIdempotent
Read-only reverse dependency lookup. Use this to find all code that references, calls, extends, or implements a specific entity. Answers 'who uses this code?' by querying the graph database. Differs from search tools by providing exact dependency tracking.
Usage: Use for impact analysis before refactoring or to detect dead code. Do NOT use this for semantic feature discovery—use 'search_hybrid_context' instead.
Matching is precedence-based: exact FQN (containing '.' or '::') → FQN suffix (Type.member) → exact name → signature prefix (accept(List) → fuzzy substring. The first tier that matches wins, so an exact name never returns fuzzy noise. Pass a qualified name (Namespace.Type.Member) to disambiguate homonyms. Responses state which tier matched and flag fuzzy results explicitly.
Behaviour & Return: Read-only graph traversal with no side effects. Returns Markdown grouped by relationship type (Calls, Extends, Implements, References, Overridden by, Overrides) with exact file paths and line numbers. Each caller entry and each resolved target states its repository as (repo: name), so rows are attributable when multiple repositories are in scope. For JVM code (Java/Kotlin/Groovy) and C#, 'Overridden by' lists method implementations/overrides in subtypes and 'Overrides' lists the supertype methods a method implements/overrides. When multiple entities with the same name exist (e.g., 'find_nearest_entity_by_line' in orphans.rs vs rust.rs), results are grouped by target entity showing which specific target each caller references. Each caller entry includes: name, kind, file_path:line_number, and signature. When multiple targets exist, each group shows the target's location and signature.
Parameter guidance: 'entity_name' supports exact names or signature fragments (e.g., 'handleRequest' or 'handle(Request'). Include 'repo_name' to filter results to the specific codebase being analyzed.
Supports Java, Kotlin, C#, Rust, and TypeScript codebases.
| Name | Required | Description | Default |
|---|---|---|---|
| repo_name | No | Optional but HIGHLY RECOMMENDED: repository scope. Accepts a single repository name (`'my-repo'`), a comma-separated list (`'repo-a,repo-b'`), or `'all'` (or `'*'`) to query every indexed repository. If you know the repository you are working on, include it in your FIRST query to avoid mixed results from other indexed projects. Omit to search across all repositories. | |
| entity_name | Yes | The name of the function, method, or class to find callers for |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint, and openWorldHint, yet the description adds substantial behavior beyond them: the five-tier precedence matching algorithm and the guarantee that an exact name never returns fuzzy noise, plus the return format (Markdown grouped by relationship type with file:line), per-row repo attribution, JVM/C# override semantics, and homonym grouping. This is genuinely additive context rather than restating annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and a clear Usage section before the dense Behavior & Return block. It is on the long side, with several illustrative examples that are helpful but slightly redundant, so it is well-organized rather than perfectly economical.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description carries the burden of explaining returns, and it does so thoroughly (grouping by relationship type, tier reporting, repo attribution, target grouping for homonyms). It also states the supported languages, leaving nothing essential for correct invocation unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 100%, so the baseline is 3, but the description adds meaning beyond the schema: it explains that entity_name accepts exact names or signature fragments (e.g., 'handle(Request') and that a qualified name disambiguates homonyms, and that repo_name scopes results. This enriches how the parameters should be populated rather than merely repeating their schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific operation ('reverse dependency lookup') on a specific resource ('all code that references, calls, extends, or implements a specific entity') and answers 'who uses this code?'. It explicitly distinguishes itself from search tools by providing exact dependency tracking, so an agent can differentiate it from search_hybrid_context without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit use cases ('impact analysis before refactoring or to detect dead code') and an explicit exclusion with the alternative named ('Do NOT use this for semantic feature discovery—use search_hybrid_context instead'). The when/when-not/alternative triangle is fully covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_repo_dependenciesList cross-repository dependenciesARead-onlyIdempotent
Read-only cross-repository dependency graph lookup. Shows which repositories depend on each other via build system declarations (Maven, Gradle, Cargo, npm, NuGet). Answers 'which repos does this repo depend on?' and 'which repos depend on this repo?'.
Usage: Use BEFORE cross-repo analysis to discover which other indexed repos are available for call tracing. Use reverse mode for impact analysis before making breaking changes in shared libraries.
Behaviour & Return: Read-only graph traversal with no side effects. Returns a JSON array of repository names. Empty results mean no DEPENDS_ON relationships exist for that repo.
Parameter guidance: 'repo_name' is required and must match the name used during indexing. 'max_depth' defaults to 3 (1 = direct only). 'reverse' toggles between forward and reverse dependency lookup.
Supports all build systems indexed by knot: Maven, Gradle, Cargo, npm, NuGet (.csproj + Central Package Management via Directory.Packages.props). C# repos that previously reported build_system: "none" now report "nuget" on re-index; knot-indexer --clean is recommended for immediate effect.
| Name | Required | Description | Default |
|---|---|---|---|
| reverse | No | If true, show repositories that depend ON this repo (reverse lookup). If false (default), show repositories this repo depends ON. Use reverse for impact analysis before breaking changes. | |
| max_depth | No | Maximum depth for transitive dependency traversal (default: 3). Use 1 for direct dependencies only. Higher values follow chains deeper. Must be between 1 and 10. | |
| repo_name | Yes | Repository name to show dependencies for. Must match the name used during indexing (e.g., 'my-java-repo', 'auth-service'). This is REQUIRED — there is no default. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and closed-world scope, so the safety profile is covered. The description still adds value beyond that: 'no side effects', the return shape (JSON array of repository names), and the meaning of empty results ('no DEPENDS_ON relationships exist'). It doesn't discuss depth limits on traversal or performance, which keeps it below a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Uses clear labeled sections (Usage, Behaviour & Return, Parameter guidance) with the core purpose front-loaded in the first sentence. The final paragraph on knot build-system support and re-indexing is genuinely bulky and somewhat tangential, costing it a point.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so ('JSON array of repository names', empty-result meaning), and annotations cover the safety profile. Build-system coverage (Maven, Gradle, Cargo, npm, NuGet) is spelled out. Minor gaps remain around traversal cost or ordering of results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema already documents repo_name, reverse, and max_depth with defaults and bounds. The description largely restates that ('max_depth defaults to 3 (1 = direct only)', 'reverse toggles between forward and reverse'), adding only light interpretation. Per the baseline rule, 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb and resource ('cross-repository dependency graph lookup') and defines its scope precisely: which repos depend on each other via build system declarations. It even frames the two exact questions it answers, which separates it from list_repositories (enumerate repos) and find_callers (trace calls) without needing to name them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use guidance ('Use BEFORE cross-repo analysis to discover which other indexed repos are available for call tracing') and conditions for reverse mode ('impact analysis before making breaking changes in shared libraries'). It lacks an explicit when-not-to-use statement or a direct named alternative from the sibling set, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_repositoriesList indexed repositoriesARead-onlyIdempotent
Read-only listing of all indexed repositories with optional name filtering. Shows repository metadata including entity count, file count, build system, and primary language. Answers 'what codebases have I indexed?' and 'which repositories match this name?'.
Usage: Use this tool FIRST to discover available codebases before searching or exploring. Once you know the repository name, switch to 'search_hybrid_context' for semantic search, 'find_callers' for reverse dependency lookup, 'explore_file' for file anatomy, or 'list_repo_dependencies' for cross-repo dependency graphs. Do NOT use this tool to search for code entities — use 'search_hybrid_context' instead.
Behaviour & Return: Read-only query with no side effects. Returns a Markdown table with columns: REPO, BUILD SYSTEM, LANGUAGE, FILES, ENTITIES. When no repositories match the filter, returns 'No repositories found.'
Parameter guidance: 'filter' is optional. When provided, only repositories whose name contains the filter string are returned (case-insensitive substring match). Omit to list all indexed repositories.
Supports all languages and build systems indexed by knot.
| Name | Required | Description | Default |
|---|---|---|---|
| filter | No | Optional filter to narrow down repositories by name (case-insensitive substring match). When provided, only repositories whose name contains this string are returned. Examples: 'auth' matches 'auth-service' and 'Auth-Lib', 'api' matches 'my-api'. Omit to list all indexed repositories. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds genuinely new context: the return is a Markdown table with named columns, and the no-match case returns 'No repositories found.' That extra disclosure is valuable since there is no output schema, though permission/indexing prerequisites are not addressed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then clearly sectioned into Usage, Behaviour & Return, and Parameter guidance. Slightly longer than necessary because the filter semantics and read-only nature repeat structured fields, but every section is scannable and earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description supplies the return shape and empty-result behavior; with four sibling tools, it routes the agent to the correct alternative; with one fully-documented param, nothing else is needed. Complete for a low-complexity read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single 'filter' param is fully documented in the schema (substring, case-insensitive, examples, optional). The description restates the same semantics rather than adding new meaning, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Read-only listing of all indexed repositories') plus the metadata surfaced and the questions it answers. An agent can distinguish it from search_hybrid_context, find_callers, explore_file, and list_repo_dependencies without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to use this tool FIRST to discover codebases, names the exact sibling to switch to for each follow-up intent, and includes an explicit 'Do NOT use this tool to search for code entities' exclusion. Full when/when-not/alternatives coverage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_hybrid_contextHybrid semantic + structural searchARead-onlyIdempotent
Read-only semantic and structural code search combining vector embeddings with graph analysis. Use this for initial codebase discovery to find features by their meaning (e.g., 'user authentication'). Locates code based on natural language descriptions instead of exact keywords, returning relevant files, signatures, and documentation.
⚠️ PREREQUISITE: This tool requires an active knot-mcp server with vector database (Qdrant) and graph database (Neo4j) initialized.
Behavior & Return: Performs a read-only dual query against vector DB (for semantic similarity) and graph DB (for architectural relationships). Returns Markdown-formatted results with file paths, line numbers, code snippets, and cross-repository dependencies. No side effects.
Usage: Use as your FIRST step when exploring unfamiliar code or discovering architectural patterns. Do NOT use this to find all usages of a specific function—use the 'find_callers' tool for that instead.
Parameter guidance: 'query' should be 2-5 words describing functionality. Increase 'max_results' to 10-20 for broad discovery, keep at 5 for focused search. Include 'repo_name' in your first query to avoid cross-repository pollution.
Supports Java, Kotlin, C#, and TypeScript codebases.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Search query describing what you're looking for (e.g., 'user authentication', 'API error handling') | |
| repo_name | No | Optional but HIGHLY RECOMMENDED: repository scope. Accepts a single repository name (`'my-repo'`), a comma-separated list (`'repo-a,repo-b'`), or `'all'` (or `'*'`) to query every indexed repository. If you know the repository you are working on, include it in your FIRST query to avoid mixed results from other indexed projects. Omit to search across all repositories. | |
| max_results | No | Maximum number of results to return (default: 5) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/idempotentHint/destructiveHint, but the description adds substantial context beyond them: an operational prerequisite (active knot-mcp server with Qdrant and Neo4j initialized), the dual vector+graph query mechanism, a concrete return shape (Markdown with file paths, line numbers, snippets, cross-repo dependencies), and supported languages (Java, Kotlin, C#, TypeScript).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Bold section labels (⚠️ PREREQUISITE, Behavior & Return, Usage, Parameter guidance) make it scannable and the critical routing rule is front-loaded. It is on the long side and lightly redundant, restating 'read-only'/'no side effects' that annotations already cover, but each section still carries load.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by describing the return payload (Markdown, file paths, line numbers, snippets, cross-repo deps), and it covers prerequisites, scoping, and language support. An agent has everything needed to call it correctly without further lookup.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real tuning guidance beyond the schema: 'query' should be 2-5 words, raise 'max_results' to 10-20 for broad discovery vs 5 for focused search, and include 'repo_name' on the first query to avoid cross-repository pollution. It stops short of explaining multi-repo list syntax (already in the schema).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('semantic and structural code search combining vector embeddings with graph analysis') and explicitly contrasts its niche against the sibling 'find_callers'. An agent can distinguish it from find_callers, explore_file, and list_repositories without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use ('FIRST step when exploring unfamiliar code or discovering architectural patterns') and explicit when-not-to-use with a named alternative ('Do NOT use this to find all usages of a specific function—use the find_callers tool for that instead'). Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
v1.9.4- Changed
explore_file1 field changed- changed
Input schema / properties / repo_name / typePrevious value: -"string"New value: +[ + "string", + "null" +]
- Changed
find_callers1 field changed- changed
Input schema / properties / repo_name / typePrevious value: -"string"New value: +[ + "string", + "null" +]
- Changed
list_repo_dependencies2 fields changed- changed
Input schema / properties / max_depth / typePrevious value: -"integer"New value: +[ + "integer", + "null" +] - changed
Input schema / properties / reverse / typePrevious value: -"boolean"New value: +[ + "boolean", + "null" +]
- Changed
list_repositories1 field changed- changed
Input schema / properties / filter / typePrevious value: -"string"New value: +[ + "string", + "null" +]
- Changed
search_hybrid_context2 fields changed- changed
Input schema / properties / max_results / typePrevious value: -"integer"New value: +[ + "integer", + "null" +] - changed
Input schema / properties / repo_name / typePrevious value: -"string"New value: +[ + "string", + "null" +]
3 tool updates
v1.8.1- Changed
explore_file1 field changed- changed
Input schema / properties / repo_name / descriptionPrevious value: -"Optional but HIGHLY RECOMMENDED: repository name to filter results to a specific codebase (e.g., 'my-java-repo'). If you know the repository you are working on, include this in your FIRST query to avoid mixed results from other indexed projects. Omit only to search across all repositories."New value: +"Optional but HIGHLY RECOMMENDED: repository scope. Accepts a single repository name (`'my-repo'`), a comma-separated list (`'repo-a,repo-b'`), or `'all'` (or `'*'`) to query every indexed repository. If you know the repository you are working on, include it in your FIRST query to avoid mixed results from other indexed projects. Omit to search across all repositories."
- Changed
find_callers1 field changed- changed
Input schema / properties / repo_name / descriptionPrevious value: -"Optional but HIGHLY RECOMMENDED: repository name to filter results to a specific codebase (e.g., 'my-java-repo'). If you know the repository you are working on, include this in your FIRST query to avoid mixed results from other indexed projects. Omit only to search across all repositories."New value: +"Optional but HIGHLY RECOMMENDED: repository scope. Accepts a single repository name (`'my-repo'`), a comma-separated list (`'repo-a,repo-b'`), or `'all'` (or `'*'`) to query every indexed repository. If you know the repository you are working on, include it in your FIRST query to avoid mixed results from other indexed projects. Omit to search across all repositories."
- Changed
search_hybrid_context1 field changed- changed
Input schema / properties / repo_name / descriptionPrevious value: -"Optional but HIGHLY RECOMMENDED: repository name to filter results to a specific codebase (e.g., 'my-java-repo'). If you know the repository you are working on, include this in your FIRST query to avoid mixed results from other indexed projects. Omit only to search across all repositories."New value: +"Optional but HIGHLY RECOMMENDED: repository scope. Accepts a single repository name (`'my-repo'`), a comma-separated list (`'repo-a,repo-b'`), or `'all'` (or `'*'`) to query every indexed repository. If you know the repository you are working on, include it in your FIRST query to avoid mixed results from other indexed projects. Omit to search across all repositories."
1 tool update
v1.5.1- Changed
explore_file1 field changed- changed
Input schema / properties / file_path / descriptionPrevious value: -"Absolute path to the source file to explore"New value: +"Path to the source file to explore. PREFERRED: a repo-relative path (e.g. 'src/services/user.ts'). ALSO ACCEPTED: an absolute path under the repository's local checkout (the tool strips KNOT_REPO_PATH / CWD automatically)."
1 tool update
v1.4.11- Added
list_repositories
4 tool updates
v1.4.0- Added
explore_file - Added
find_callers - Added
list_repo_dependencies - Added
search_hybrid_context
4 tool updates
v1.3.8- Removed
explore_file - Removed
find_callers - Removed
list_repo_dependencies - Removed
search_hybrid_context
4 tool updates
- Added
explore_file - Added
find_callers - Added
list_repo_dependencies - Added
search_hybrid_context
4 tool updates
v1.3.2- Removed
explore_file - Removed
find_callers - Removed
list_repo_dependencies - Removed
search_hybrid_context
4 tool updates
v1.2.8- Added
explore_file - Added
find_callers - Added
list_repo_dependencies - Added
search_hybrid_context
4 tool updates
v1.2.7- Removed
explore_file - Removed
find_callers - Removed
list_repo_dependencies - Removed
search_hybrid_context
1 tool update
v1.2.5- Added
list_repo_dependencies
TDQS
Scored across 5 tools
Each tool targets a distinct read-only operation: semantic search, reverse call lookup, file structure inspection, cross-repo dependency lookup, and repo listing. Descriptions explicitly cross-reference and warn against misuse, leaving no ambiguity.
All tool names follow a consistent snake_case verb_noun pattern (search_hybrid_context, find_callers, explore_file, list_repo_dependencies, list_repositories). Though search_hybrid_context adds a modifier, it remains predictable and readable.
Five tools form a well-scoped set covering core codebase discovery and dependency analysis without redundancy. Each earns its place and matches the expected range for a specialized code intelligence server.
The surface covers discovery, reverse dependencies, file anatomy, and repo-level dependency graphs, but lacks a forward-dependency (callee) lookup and full source retrieval. These are minor gaps that agents can partially work around via search snippets and caller data.
Maintenance
Related MCP Connectors
Codebase graphs, caller impact analysis, and recorded project context for AI coding agents.
Codebase intelligence for AI agents — dead code, blast radius, ownership.
Code intelligence platform for AI agents. 20 tools for architecture, security & impact analysis.
Code intelligence for coding agents: semantic, AST, graph, and full-text search. 279+ languages.
Related MCP Servers
- AlicenseBqualityDmaintenanceA complete MCP server for Retrieval-Augmented Generation with file management and vector memory for agents. Supports multiple document formats (PDF, DOCX, TXT, MD, CSV, JSON) with semantic search using Hugging Face embeddings and ChromaDB for efficient vector storage.1191MIT
- AlicenseNot gradedqualityNot gradedmaintenanceTurns Claude Desktop into a personal document question-answering system using local vector search. Index PDF, TXT, and Markdown documents into collections and get answers based strictly on your documents with zero hallucination.9-
- AlicenseNot gradedqualityCmaintenanceA graph-powered code intelligence engine that indexes codebases into a structural knowledge graph to provide AI agents with deep context on function calls, types, and execution flows. It offers local, zero-dependency tools for hybrid search, impact analysis, and dead code detection across Python, JavaScript, and TypeScript projects.811MIT
- FlicenseNot gradedqualityDmaintenanceA minimalist indexing tool that provides AI agents with semantic search and structural AST parsing for deep codebase understanding. It enables autonomous agents to navigate large codebases predictably using vector embeddings and native language server capabilities like definition and reference tracking.-