Skip to main content
Glama

knot

License: MIT Rust knot MCP server

knot is a high-performance codebase indexer that extracts structural and semantic information from source code, enabling AI agents to understand, analyze, and navigate large code repositories. Currently supports Java, Kotlin, TypeScript, JavaScript/Node.js, Rust, Python, Groovy, C/C++, C#, HTML, and CSS/SCSS, plus Build Systems (Maven pom.xml, Gradle build.gradle, Jenkins pipeline, Cargo.toml, MSBuild .csproj + Directory.Packages.props), Configuration Files (YAML, JSON, .properties — optional), Kubernetes + Helm (optional), and Cross-Repo Dependency Linking with full cross-language linking.

For recent release notes see CHANGELOG.md.

The indexer automatically builds:

  • Vector Search Database (Qdrant) — semantic understanding via embeddings

  • Graph Database (Neo4j) — architectural relationships via call graphs

This dual-database approach powers both:

  • MCP (Model Context Protocol) Server — Exposes four read-only tools (search, callers, explore, list_files/repos) to any LLM client (Claude, Gemini, ChatGPT, Cursor, etc.)

  • CLI Tool — Standalone knot command for terminal and scripting environments

Knot in action

CLI — instant reverse dependency lookup

MCP — JSON-RPC protocol for AI agents


🧮 Token Efficiency — Measured, Not Claimed

An LLM agent exploring an unfamiliar codebase pays for every byte it reads. Without an index it greps and then reads whole files; with knot it receives a targeted answer. The difference was measured on three real indexed repositories across nine realistic exploration tasks:

Repo

Lang

Task

knot tokens

Read-the-code tokens

Reduction

spring-ai

Java

discovery — how does the chat client run the advisor chain?

1 092

10 168

89.3%

spring-ai

Java

callers — who uses ToolCallingManager?

8 808

15 554

43.4%

spring-ai

Java

explore — structure of DefaultChatClient.java

4 865

7 838

37.9%

puppeteer

TypeScript

discovery — how is a CDP session created?

609

4 149

85.3%

puppeteer

TypeScript

callers — who calls createCDPSession?

1 004

39 878

97.5%

puppeteer

TypeScript

explore — structure of the Page API

7 287

25 300

71.2%

knot

Rust

discovery — how are call intents resolved?

594

14 824

96.0%

knot

Rust

callers — who calls format_references_result?

461

10 949

95.8%

knot

Rust

explore — structure of the graph query module

978

12 103

91.9%

TOTAL

—

9 tasks

25 698

140 763

81.7%

≈ 5.5× fewer tokens for the same nine questions — 115 000 tokens saved, enough to keep a long refactoring session inside a single context window.

Both sides are measured on the exact bytes an LLM would receive as tool output, counted with OpenAI's cl100k_base tokenizer (tiktoken):

Task

knot side

Read-the-code side

discovery

knot search "<question>" --repo <r> --output markdown

rg -l <keyword> (candidate list) + full read of the files that actually answer the question

callers

knot callers "<symbol>" --repo <r> --output markdown

rg -n "\b<symbol>\b" + full read of the first 5 distinct files with hits

explore

knot explore "<file>" --repo <r> --output markdown

full read of the file

The baseline is deliberately generous, so the measured saving is a lower bound:

  • greps are restricted to the source files of the language (-t java, -t ts, -t rust) — no changelogs, no generated docs, no node_modules;

  • for discovery the baseline is given oracle file selection: it reads only the files that answer the question, with zero wasted reads;

  • for callers it reads at most 5 files, while a rigorous impact analysis would need every file with a textual hit.

Honest caveats: knot's cost scales with the number of results, not with repo size. The weakest row (spring-ai / ToolCallingManager, 43%) is a symbol with 156 references — knot enumerates all of them with exact call sites, while the capped baseline reads only 5 files and still cannot tell a call from a comment. The explore rows for large classes are also the least favourable, because signatures plus docstrings are a large fraction of a well-documented file.

Repositories measured (as indexed): spring-ai 2 406 files / 25 733 entities, puppeteer 1 832 files / 19 310 entities, knot 222 files / 4 000 entities. Raw measurements are stored in .perf_metrics/token_savings.json.

pip install tiktoken            # optional: falls back to a chars/4 estimate
# edit the `root` paths in scripts/token_savings_tasks.json to match your checkouts
python3 scripts/token_savings_benchmark.py \
  --config scripts/token_savings_tasks.json \
  --save-json .perf_metrics/token_savings.json

The task definitions live in scripts/token_savings_tasks.json and the harness in scripts/token_savings_benchmark.py; point them at any repository you have indexed to measure your own codebase.


Related MCP server: MCP-RAG

✨ Key Features

🔍 Code Intelligence Tools

  • search_hybrid_context: Semantic + structural search. Find code by meaning, class name, method signature, docstrings, or comments. Returns full context including dependencies.

  • find_callers: Reverse dependency lookup. Identify dead code, perform impact analysis, or understand the full call chain of any function/method. Whenever the query resolves to more than one entity sharing the same name (e.g., find_nearest_entity_by_line in different files), results are automatically grouped by target (### Target: <fqn> at <file>:<line>) showing which specific entity each caller references — including when only one of the homonyms has callers. A genuinely single-target resolution keeps the concise ungrouped form. Supports cross-repository call resolution via DEPENDS_ON graph edges. For JVM languages (Java/Kotlin/Groovy) it also surfaces method-level OVERRIDES edges bidirectionally — an Overridden by group listing subtype implementations/overrides and an Overrides group listing the supertype methods a method implements/overrides.

  • explore_file: File anatomy inspection. Quickly see all classes, interfaces, methods, functions, constants, and other entities in a file with signatures and documentation. Attribute-usage markup tokens (html_class, html_id) are summarized in a separate section by default to avoid cluttering code lists; use --include-markup (CLI) or include_markup: true (MCP) to list them in full. Structured content returns markup_count and markup_references alongside ambiguous_path_candidates.

  • list_repo_dependencies (MCP) / knot deps (CLI): Dependency graph visualization. Show which repositories depend on each other, forward and reverse, with transitive resolution.

  • list_files (MCP) / knot files (CLI): Read-only file enumeration with entity counts, deterministically ordered. Accepts an optional path — a repo-relative directory prefix matched on a path boundary (src/api never matches src/api-notes.md) or a glob (src/**/*_test.rs). Prefixes and globs alike page the repository file list in pages of 2000 entries with per-page boundary filtering, ensuring no valid candidates are dropped. When output exceeds 2000 matches, a truncation notice is displayed below the table reporting the exact pre-truncation match count when known (or an explicit lower bound when the scan stopped early). JSON output carries total_matches and total_is_lower_bound. Pair with the new --path/path scope on search_hybrid_context to restrict a semantic search to a subtree.

  • list_repositories / knot repos: Repository inventory. List every indexed repository along with its entity count, file count, build system, and primary language. Supports optional case-insensitive name filtering via --filter (CLI) or filter parameter (MCP). Useful for orientation, sanity-checking indexing runs, and discovering which languages and build systems are present in the workspace.

🏗️ Multi-Language Support

  • Java: Full AST extraction with package-aware FQN resolution (e.g., com.example.app.UserService), class inheritance (EXTENDS), interface implementation (IMPLEMENTS), annotation tracking, and field-access method invocation resolution

  • Kotlin: Complete support for Kotlin codebases with classes, interfaces, objects, companion objects, functions, methods, and properties. Fully compatible with tree-sitter-kotlin-ng grammar.

  • C#: Full C# support via tree-sitter-c-sharp. Extracts classes, interfaces, structs, records (both record class and record struct), enums, methods, constructors, properties, fields (with const detection), delegates, events, indexers, operators, local functions, and namespaces with CSharp* entity kinds. Namespace-qualified FQNs (MyApp.Services.UserService.GetUserAsync) work across both file-scoped (C# 10+) and block-form namespaces, including nested namespaces and nested types. The base_list heuristic splits : Base, IFace into EXTENDS/IMPLEMENTS using the IPascalCase convention (structs and interfaces are deterministic), generic arguments are stripped (IRepository<User> → IRepository), XML doc comments (///) become docstrings, and attributes ([Obsolete]) are captured as decorators. Calls through field-typed receivers resolve to the exact implementation method, and C# virtual/override plus interface implementation produce method-level OVERRIDES edges. MSBuild/NuGet: .csproj files are parsed for project identity and dependencies (Central Package Management via Directory.Packages.props is supported); C# repos get build_system: "nuget" in the Repository node instead of the prior "none".

  • TypeScript/TSX/CTS: Complete support for modern JavaScript/TypeScript codebases, including CommonJS TypeScript files

  • JavaScript/Node.js: Vanilla JS, Node.js, and module systems (.js, .mjs, .cjs, .jsx)

  • Hybrid Web Ecosystem: Cross-language linking between JavaScript, HTML, and CSS for full-stack SPA analysis

  • HTML: Custom elements (Web Components, Angular), id and class attribute indexing for cross-language CSS search

  • JSX/TSX Attributes: Extracts id and className from React components for unified HTML/CSS discovery

  • CSS/SCSS: Stylesheet indexing with class/ID selector extraction and variable tracking (CSS/SCSS variables, mixins, functions)

  • Rust: Struct, enum, union, trait, function, method, module extraction with trait implementation tracking (IMPLEMENTS relationships) and macro invocation references. Methods are indexed with the qualified FQN Type::method (e.g., KnotMcpHandler::new, WidgetA::new, Logger::new) and qualified calls from top-level functions resolve to the right target by receiver. Braced import/use capture — use foo::{Bar, Baz} and use foo::Bar as Baz produce explicit REFERENCES edges for all imported names, including traits imported solely to bring methods into scope. All Rust entity FQNs are now anchored at the owning crate and module path (e.g. knot::config::Config, knot::pipeline::parser::languages::rust::qualify_rust_fqns), so two crates that declare a type with the same bare name no longer collide. Files outside src/ (tests, benches, examples) receive a __fixture::<path>::<Entity> FQN prefix (e.g. __fixture::tests::testing_files::sample::Config), and files without a Cargo.toml ancestor receive __loose::<path>::<Entity>, preventing name collisions with real source entities. CONTAINS relationships use enclosing_class_fqn for exact disambiguation when multiple entities share the same class name. The on-disk index state file (.knot/index_state.json) carries a version field; opening a state file from an older version prints an error with instructions to run knot-indexer --clean.

  • Python: Full Python extraction with class, function, method support, constants, module-level imports, ValueReference tracking for keyword arguments, class inheritance (EXTENDS), decorator extraction (@property, @staticmethod, @route(...), @dataclass), generic type hints (List[str], Optional[Dict], *args/**kwargs), Py2/Py3 exception syntax compatibility, and self.method() resolution with inherited method walking. Captures class_definition, function_definition (including async via optional async modifier), lambda assignments, and distinguishes methods from functions via parent context detection. Class instantiation (ClassName(...)) is automatically redirected to ClassName.__init__ so find_callers ClassName.__init__ lists every constructor call site (with fallback to inherited __init__ via the extends chain); only class/struct kinds trigger the redirect — functions keep the legacy behavior.

  • Groovy: Full Groovy language support via hybrid tree-sitter + ad-hoc lexical parser. Extracts classes, interfaces, traits, enums, typed/def/quoted methods (incl. Spock specs), constructors, closures, script-level variables, fields/properties with visibility modifiers, nested classes, and decorators. Tracks package FQN and enclosing class relationships. Multi-line signatures (closure default params), assignment-vs-declaration disambiguation, innermost assignment for nested closures, UUID collision fix for duplicate method names, find_callers accurately tracks private methods including those in anonymous new AnAction closures. Inheritance tracking: emits EXTENDS/IMPLEMENTS reference intents for class/interface/trait/enum headers (single-line and multi-line) so find_callers surfaces real nextflow-style hierarchies — qualified parents (e.g. extends nextflow.plugin.BasePlugin) and generic-argument stripping (e.g. extends AbstractRepo<Order, Long> → extends AbstractRepo) are supported, and generic bounds (class Box<T extends Comparable>) are correctly not promoted to inheritance edges. Property accessors: bare property declarations (Path baseDir, boolean cacheable) are now indexed as GroovyProperty, and compiler-generated getX/setX/isX accessors are synthesised as first-class method entities so OVERRIDES edges link Groovy properties to interface getter declarations. Comment-stripping prevents Javadoc continuation lines (* The pipeline script name) from producing phantom entities or corrupting scope tracking.

  • Build Systems: Maven pom.xml (dependencies + plugins via roxmltree), Gradle build.gradle (deps + plugins + tasks), Jenkinsfile pipeline (stages + steps), Cargo Cargo.toml (deps + workspace members + features), and MSBuild .csproj / Directory.Packages.props extraction. MSBuild resolves project identity (<PackageId> → <AssemblyName> → file stem), emits a BuildDependency per <PackageReference> (attribute-form and version-less), and resolves Central Package Management versions from the nearest Directory.Packages.props ancestor. UTF-8 BOMs are tolerated defensively. Identity marker identity: package_id is carried in the signature when the project has an explicit <PackageId> so the cross-repo resolver prefers published packages over depth-tied unmarked candidates.

  • Cargo.toml: Rust package manager support with package metadata, features, workspace members, and multi-format dependency parsing (simple, table, git, path).

  • Configuration Files: YAML (.yml/.yaml), JSON (.json), and Java Properties (.properties) with leaf-key granularity. Special handling for package.json (detected by filename: npm dependencies as BuildDependency, scripts as ConfigProperty, ProjectIdentity even for dependency-free library manifests).

  • Varnish Cache: Hand-written parsers for .vcl (configuration), .vtc (test cases), and .vcc (VMOD C source). VCL extracts backends, probes, ACLs, subroutines (custom + built-in with vcl_* names, including aggregator entities for multi-part built-ins), import directives (with as aliases and from paths), include edges, unused declarations, VMOD instantiations, and req.backend_hint assignments. VTC extracts varnishtest/vtest cases, servers, clients, varnish instances, logexpect blocks, barriers, and -vcl+backend synthesised backends (with is_test_context). VCC extracts $Module, $Function, $Object, $Method, $Event, $Restrict, ENUMs, and default parameters. References: Calls, Extends, Implements, References (with intents VclSubCall, VclBackendRef, VclProbeRef, VclAclRef, VclInclude, VclVmodImport, VclUnusedRef, ValueReference); relationships: UsesBackend, UsesProbe, UsesAcl, Includes, ImportsVmod, DeclaredUnused. The Fastly VCL dialect is detected and skipped (returns empty entities).

  • Kubernetes + Helm: K8s manifest parsing (Deployment, Service, ConfigMap, Secret, Ingress, Namespace) with label/annotation tracking and cross-resource references. Helm chart indexing (Chart.yaml metadata, values.yaml key-value pairs, template variable extraction via {{ .Values.X }}).

  • C/C++: Complete C/C++ support with namespace-aware FQN resolution (Engine::MyClass::start), class/struct extraction, function/method tracking, macro definition and usage detection (uppercase identifier heuristic), type reference tracking (declarations, new expressions), and full call graph analysis. Supports .c, .h, .cpp, .hpp, .cc, .cxx, .hh, .hxx extensions via tree-sitter-c and tree-sitter-cpp parsers. Includes intelligent auto-detection for .h headers to parse them correctly as C or C++ based on their contents.

  • Markdown: Documentation indexing with MarkdownDocument (one per .md/.markdown file) and MarkdownSection (one per ATX heading H1–H6). Section bodies — including paragraphs, fenced code blocks, lists, and tables — are captured into embed_text for full semantic search over documentation content, not just heading titles. FQNs are hierarchical and file-scoped (e.g. README.md::Setup > Installation > Linux), so same-named headings in different files or under different parents disambiguate cleanly. Section boundaries respect heading depth: a section's body extends until the next heading of equal or higher level, ensuring ### Linux under ## Installation does not bleed into a sibling ## Configuration. Headings with inline markdown (backticks, em-dash, links, emoji) parse without losing their bodies, and real start_line/end_line positions are computed via tree-sitter for each section.

📚 Rich Comment Extraction

  • Captures docstrings (JavaDoc, JSDoc) preceding declarations

  • Extracts inline comments within method/function bodies

  • Respects nesting boundaries (class comments don't capture method comments)

  • Intelligently aggregates comment blocks

📊 Dual-Database Architecture

  • Qdrant: Vector search for semantic code understanding

  • Neo4j: Graph relationships for structural navigation

🚀 High Performance

  • Parallel Streaming Pipeline: Overlaps CPU-bound embedding with I/O-bound ingestion via MPSC channels

  • Incremental Indexing: Uses SHA-256 hashes to skip unchanged files

  • Real-time Watch Mode: Automatically re-indexes changed files in seconds via --watch

  • Embed-Text Recall: Every entity's embed text carries its identifier surface — name, FQN and the camel/snake-case-tokenized form of both, plus the tokenized names of what it calls in its body. A function with no doc comment is still findable by the behaviour described in natural language ("authenticate user with email and password" surfaces a doc-less login through the normalize_email / verify_credentials vocabulary in its body). Rust signatures ((email: &str, password: &str)) survive extraction too. Requires one re-index when upgrading from an older index state version (automatic: old states are rejected with instructions).

  • CPU Parallelism: AST extraction via Rayon

  • Scalable: Configurable batch processing and constant memory footprint (~2GB) regardless of repository size

  • Performance Benchmarking: Multi-level validation framework

    • Unit benchmarks: Criterion-based benchmarks for parse, embed, and graph write throughput (benches/)

    • E2E benchmarks: Full pipeline metrics capture with per-stage timing (tests/benchmark_e2e.sh)

    • CI regression tracking: Automated baseline comparison against tolerance thresholds (scripts/compare_perf_metrics.sh)

    • Token efficiency: LLM token cost of knot answers vs reading source files (scripts/token_savings_benchmark.py) — see Token Efficiency


🛠️ Installation

Prerequisites

Component

Version

Notes

Docker

20.10+

For running Qdrant and Neo4j

qdrant

1.x

Vector database (docker)

neo4j

5.x

Graph database (docker)

Option A: Pre-compiled Binaries (macOS & Modern Linux)

Go to the Releases page and download the native executable for your platform.

Install knot binaries (CLI, MCP server, and indexer):

curl --proto '=https' --tlsv1.2 -LsSf https://github.com/raultov/knot/releases/latest/download/knot-installer.sh | sh

Install agent-skills for your AI (Optional): Paste this into your LLM agent (Claude Code, OpenCode, Cursor, etc.):

Install the knot agent skills by following the instructions at: https://raw.githubusercontent.com/raultov/knot/master/README.md

The first command installs the knot binary to your PATH. The second (optional) allows your AI assistant to automatically download the agent skill index (.knot-agent.md) and run the installer to extract comprehensive guides for using knot CLI with AI agents and code analysis tools.

System Requirements:

  • Linux: glibc 2.38+ (Ubuntu 24.04+, Debian 13+, Fedora 39+, Arch)

  • macOS: Modern versions supported

  • Windows: Use Docker (Option B)

Option B: Docker (Universal Compatibility)

Docker images provide universal compatibility for any Linux distribution and Windows.

Docker Installation (All Binaries)

Build the image:

docker build -t knot:latest . --network=host

Run the indexer:

# Use --network host to connect to databases running on your host machine
docker run --rm \
  -v /path/to/your/repo:/workspace \
  -e KNOT_REPO_PATH=/workspace \
  -e KNOT_NEO4J_PASSWORD=your-password \
  --network host \
  knot:latest \
  knot-indexer

Run the CLI tool:

docker run --rm \
  -v /path/to/your/repo:/workspace \
  -e KNOT_REPO_PATH=/workspace \
  -e KNOT_NEO4J_PASSWORD=your-password \
  --network host \
  knot:latest \
  knot search "user login flow"

Run the MCP server:

docker run --rm \
  -e KNOT_REPO_PATH=/workspace \
  -e KNOT_NEO4J_PASSWORD=your-password \
  --network host \
  knot:latest \
  knot-mcp

Note: Uses Debian Trixie (glibc 2.38+) and includes ONNX Runtime for full functionality.


Option C: Install via Cargo

cargo install --git https://github.com/raultov/knot

Option D: Build from Source

1. Start infrastructure with Docker:

docker compose up -d

2. Clone and build:

git clone https://github.com/raultov/knot
cd knot
cargo build --release

3. Configure:

mkdir -p ~/.config/knot
cp .env.example ~/.config/knot/.env
$EDITOR ~/.config/knot/.env  # Set KNOT_REPO_PATH and Neo4j credentials

4. Index a codebase:

./target/release/knot-indexer

5. Query via CLI:

./target/release/knot search "your query"

6. Start the MCP server:

./target/release/knot-mcp

📖 Usage

🤖 Install Agent Skills (For AI Agents)

Option A: Let an LLM do it

Paste this into any LLM agent (Claude Code, OpenCode, Cursor, etc.):

Install the knot agent skills by following the instructions at: https://raw.githubusercontent.com/raultov/knot/master/README.md

Option B: Terminal (Manual)

curl -sO https://raw.githubusercontent.com/raultov/knot/master/.knot-agent.md && curl -fsSL https://raw.githubusercontent.com/raultov/knot/master/scripts/install-agent-skills.sh | bash

📥 Quick Downloads (Binaries)

Download knot binaries (CLI + MCP server):

curl --proto '=https' --tlsv1.2 -LsSf https://github.com/raultov/knot/releases/latest/download/knot-installer.sh | sh

📖 Agent-Skills Guides

Comprehensive documentation for using knot tools. The agent skills installer extracts:

  • search.md — Semantic code discovery guide with examples

  • callers.md — Reverse dependency lookup with critical usage rules

  • explore.md — File anatomy inspection guide

  • deps.md — Repository dependency graph guide

  • repos.md — Indexed repository inventory

  • workflows.md — Common patterns and best practices

For quick reference without downloading, see .knot-agent.md.


Using the CLI

The knot CLI provides the same capabilities as the MCP server via command-line commands, making it ideal for:

  • Terminal-only environments

  • Bash scripting and automation

  • CI/CD pipelines

  • Direct integration with other tools

Three main commands:

knot search "user authentication" --max-results 10 --repo my-app
knot search "user authentication" --max-results 20 --repo "app-a,app-b"  # Union across repos
knot search "user authentication" --max-results 20 --repo all              # All indexed repos ('all' or '*')
knot search "user authentication" --kinds definition                        # Only functions/methods/types

Find code entities by meaning, class names, docstrings, or comments.

Ranking is kind-aware: function/method/class/struct definitions outrank markdown docs, test files, config properties and build-dependency entities for natural-language queries, and callers/helpers are shown as context attached to a definition — never as substitutes. The shared entry point of the highest-ranked helpers outranks those helpers, and an entity merely named after a generic verb (find, get, create, build, acquire, …) does not win on the bare verb unless its container (FQN) corroborates the query. Use --kinds to narrow the result types (aliases: definition, callable, class/type/struct, or exact kinds like rust_function).

knot callers — Reverse Dependency Lookup

knot callers "LoginService" --repo my-app
knot callers "LoginService" --repo "auth-service,billing-service"
knot callers "LoginService" --repo all

Find all code that references a specific entity (dead code detection, impact analysis, call chains). Whenever the query resolves to more than one target sharing that name, results are automatically grouped by target (### Target: <fqn> at <file>:<line>) with file locations and signatures — including when only one of the homonyms has callers.

Target resolution is code-only by default: documentation, configuration, build-system and Kubernetes/Helm entities (markdown_section, config_property, build_dependency, cargo_package, project_identity, k8s_*, helm_*, …) can never be presented as resolved targets. When the filter removes matches the response discloses them (Non-code matches hidden — N entities …) and names the fix (kinds=all to include them). A fuzzy query like cargo no longer fills the target list with Cargo.toml metadata. Fuzzy matching is also case-insensitive, so hikari finds Hikari-named artifacts. Queries shorter than 4 characters resolve by anchored name prefix rather than substring: knot callers use returns useSearch, useSources, useSaveCv and friends without dragging in every *STATUSES constant that merely contains use. Queries of 4+ characters keep the case-insensitive fuzzy substring fallback. Override the scope with --kinds all / MCP kinds:"all", or pass an explicit allow-list (--kinds callable, --kinds build_dependency).

The buckets cover every edge type the pipeline produces: Calls, Extends, Implements, References, Macro calls (Rust MACRO_CALLS), DOM references (JS → HTML element id), CSS class usage (JS → CSS class), script/stylesheet imports, the VCL edges (uses backend/probe/acl, includes, VMOD imports, declared-unused) plus Overridden by / Overrides — so init_vec's macro call sites and app-container's JS manipulators now show up instead of a false "may be unused".

Every caller entry is self-labeling: the owning repository is printed next to each row as (repo: <name>) — in the CLI table, the Markdown answer, and the resolution block — so rows stay attributable when the scope spans multiple repositories:

# References to `LoginService`

Resolved to 1 target by exact name match:
- `auth::service::LoginService` (class) at `src/service.rs:12`  (repo: auth-service)

Found 1 reference(s) across all relationship types:

## Calls (1)

- **`signup`** (function) at `src/handlers.rs:88`  (repo: auth-service)

In the CLI table the Target column is labeled only for genuine cross-repo references (a caller in repo A referencing a target in repo B); the Caller column is always labeled when a repository is known.

knot explore — File Structure Inspection

knot explore "src/services/auth.ts" --repo my-app

List all classes, methods, functions in a file with signatures and documentation.

knot deps — Repository Dependency Graph

knot deps my-app --depth 2           # Show forward dependencies (transitive)
knot deps my-app --reverse           # Show who depends on this repo

Visualize auto-discovered dependencies between indexed repositories with transitive resolution up to 3 levels deep.

knot repos — List Indexed Repositories

knot repos                          # Table with REPO / BUILD SYSTEM / LANGUAGE / FILES / ENTITIES
knot repos --filter app             # Case-insensitive name filter (substring match)
knot repos --output json            # Machine-readable list
knot repos --output markdown        # GFM table for chat UIs

Show the status of every repository currently indexed in the graph database — useful for orientation, sanity-checking that an indexing run completed, and discovering which languages and build systems are present across the workspace. Use --filter to quickly locate a specific repository when working with multiple indexed codebases.

Repository Scope Selection: Both the CLI --repo/-r flag and MCP repo_name parameter support:

  • Single repository name: --repo my-app

  • Comma-separated list: --repo "repo-a,repo-b" (MCP also accepts ["repo-a", "repo-b"])

  • Sentinel: --repo all or --repo "*" (searches every indexed repository)

Note: Multi-repo scope applies a global max_results limit across the union. Increase --max-results (range 1-100, enforced; larger values are clamped and a note is printed — no pagination, narrow with --kinds/--path/--repo or refine the query instead) when searching across multiple repositories.

For detailed CLI usage guide, see .knot-agent.md — a machine-readable skill that teaches LLMs how to use knot CLI for autonomous code analysis.

Indexing a Codebase

Incremental Indexing (Default)

# First run: indexes all files
knot-indexer --repo-path /path/to/your/repo --neo4j-password secret

# Subsequent runs: only re-indexes changed files (fast!)
knot-indexer --repo-path /path/to/your/repo --neo4j-password secret

# NEW: Real-time Watch mode
knot-indexer --watch --repo-path /path/to/your/repo --neo4j-password secret

How it works:

  • Tracks file content via SHA-256 hashes in .knot/index_state.json

  • Stores the downloaded fastembed model in .knot/fastembed_cache/ to keep the workspace clean

  • Automatically detects: modified, added, and deleted files

  • Only re-parses and re-embeds changed files

  • Preserves graph relationships to unchanged files

  • Processes entities in memory-efficient 512-entity chunks

Performance:

  • Initial index (3800 files): ~60 minutes on standard hardware

  • Incremental update (3 files changed): ~5-10 seconds

  • Memory usage: Constant ~2GB regardless of repository size

Full Re-Index (Clean Mode)

# Force complete re-index (deletes all existing data)
knot-indexer --clean --repo-path /path/to/your/repo --neo4j-password secret

Use --clean when:

  • You want to rebuild the entire index from scratch

  • You've changed Tree-sitter queries or embedding models

  • Troubleshooting indexing issues

Upgrade note (v1.5.1): File paths are now persisted as repo-relative paths with POSIX separators (e.g. src/pipeline/embed.rs). Upgrading from v1.4.x triggers an automatic full re-index on first run — the on-disk .knot/index_state.json carries a version field that the loader rejects when stale, and knot-indexer then wipes the repo from both databases before rebuilding. No manual steps required. Entity UUIDs become machine-independent in the process: the same repo indexed from different checkout locations now produces identical UUIDs.

Indexing Progress

The indexer emits [Progress] log lines showing real-time completion across the whole pipeline (parsing, embedding, ingestion, reference resolution). The percentage is monotonically non-decreasing and reaches 100% only once the run genuinely terminates.

Upgrade note (v1.6.2): The percentage now spans the entire pipeline via weighted bands. Previously it measured only file reading and saturated at 100% within seconds of starting, then froze for minutes while embedding and ingestion were still running. See docs/specs/indexing_progress_accuracy_plan.md for the full design.

Example with 5000 files where 1000 have been parsed and 5,000 entities are half-way through ingestion:

[Progress] [my-repo] 50.0% — files 5000/5000, entities 41600/83200, batch #325 (128 entities)

Band table

Phase

Band

Driver

Idle / Discovering / Classifying / CleaningStaleData

0%

—

Parsing

0% → 10%

parsed_files / total_files

Embedding + Ingestion

10% → 90%

entities_ingested / total_entities

ResolvingReferences

95%

fixed (no sub-counters available)

Completed

100%

forced

Failed

last computed value

frozen

A final log line confirms completion:

[Progress] [my-repo] 100.0% — files 5000/5000, entities 83200/83200 — parsing and ingestion complete, resolving references...

Library API (knot-server integration)

Callers that need to observe progress programmatically can use the ProgressTracker:

use std::sync::Arc;
use knot::pipeline::{ProgressTracker, run_indexing_pipeline_with_progress};

let progress = Arc::new(ProgressTracker::new());
let progress_clone = Arc::clone(&progress);

// Poll snapshot() from another task while the pipeline runs
tokio::spawn(async move {
    loop {
        let snap = progress_clone.snapshot();
        println!(
            "{:.1}% — files {}/{}, entities {}/{}",
            snap.percent_complete,
            snap.parsed_files,
            snap.total_files,
            snap.entities_ingested,
            snap.total_entities
        );
        if snap.stage == IndexingStage::Completed || snap.stage == IndexingStage::Failed {
            break;
        }
        tokio::time::sleep(std::time::Duration::from_millis(500)).await;
    }
});

run_indexing_pipeline_with_progress(&cfg, &vdb, &gdb, &mut state, progress).await?;

The snapshot() method is thread-safe (read-only locks + atomic loads) and returns a IndexingProgress struct that serializes directly to JSON for REST endpoints.

Running E2E Integration Tests

To ensure indexer stability, run the E2E integration test suite:

# Run all language E2E tests (TypeScript, Java, JavaScript, Web, Kotlin, Rust, ...)
./tests/run_all_e2e_fast.sh

# Run only Kotlin E2E tests
./tests/run_kotlin_e2e.sh

# Run only Rust E2E tests
./tests/run_rust_e2e.sh

# Run only C# E2E tests
./tests/run_csharp_e2e.sh

# Run only Varnish E2E tests
./tests/run_varnish_e2e.sh

See tests/KOTLIN_E2E_TESTS.md for detailed coverage and troubleshooting.

Using the MCP Server

The MCP server exposes five tools to any compatible AI client (built on rust-mcp-sdk 1.1 implementing MCP protocol 2025-11-25 with full tool annotations):

Embedding the tool surface: library consumers can serve the same five tools from their own transport (e.g. an HTTP /mcp endpoint) without going through the stdio server. KnotMcpHandler::tools() returns the canonical tool table (no state or database connection required), and KnotMcpHandler::dispatch(params) executes a tools/call without needing an Arc<dyn McpServer> runtime handle. The stdio ServerHandler methods delegate to these two entry points, so every surface stays identical by construction.

Tool 1: search_hybrid_context

Find code by meaning or keywords

Query: "How is user authentication implemented?"
Result: All auth-related code, signatures, docstrings, and dependencies

Capabilities:

  • Semantic search by functionality (vector embeddings)

  • Global multi-repository search by default (repo_name: "all")

  • Class/method/function name lookup

  • Docstring and inline comment search

  • Architectural pattern discovery

  • Full dependency context

Search ranking contract (kind-aware re-rank, query-time only):

  1. Definition channel — alongside the plain cosine scan, a second bounded Qdrant pass excluding all non-code kinds guarantees code definitions enter the candidate pool even on documentation-heavy repositories (a prose-saturated cosine window once left the entry-point signal dead). An explicit documentation-scoped search (kinds=markdown_section, …) never sees the code channel.

  2. Recall channels in one pool — cosine hits, the definition channel, the name/token probe (identifiers the query literally names) and the caller-recall bridge (callers of the top semantic roots — seeded from the union of channels, depth 1 + depth 2 in the call graph) all merge into one deduplicated pool before ranking.

  3. Kinds — callables outrank type declarations, which outrank prose and config/build/infra; test paths carry an additional penalty. A neutral kind (constant, …) whose graph node orchestrates ≥ 2 outgoing CALLS edges earns a boost: a TypeScript MCP tool is export const x = defineTool({...}), and behavior is not the wiring's kind. Prose/config/test never take structural boosts.

  4. Entry-point signal — how many of the top semantic roots a candidate calls, directly or through one helper, attributed to the caller covering a strictly greater root set with the full entry-point signature (≥ 2 distinct roots).

  5. Name-prefix contract is definition-only — a query matching an entity's name keeps leading slots only for definition kinds; prose, config/build and k8s/helm prefix hits are demoted into the pool and ranked on their own cosine (documentation-only topics with no competing definition still surface their best section). Test paths keep their slot but stay penalized inside the re-rank.

  6. Determinism — ties break on (file_path, start_line, uuid); the CLI (knot search) and the MCP tool share the same core (run_search_hybrid_context).

Live-index verification of the reported recall regression set runs via the opt-in harness tests/run_rank_recall_live.sh (requires indexed repositories; skipped otherwise).

Tool 2: find_callers

Find who calls a specific function

Query: "Find callers of getCurrentTimeInSeconds"
Result: All code that invokes this function + file locations

Each caller entry, target group header, and resolved target carries its repository as (repo: <name>), so results remain attributable under multi-repo scopes (repo_name: "all" or a comma list). The raw JSON (--output json) mirrors this with repo_name (referencing entity) and target_repo_name (referenced entity) fields on every row, plus repo_name on each resolution.targets[] entry.

Advanced: Search by Signature

# Find by full signature (Java)
echo '{"method":"tools/call","params":{"name":"find_callers","arguments":{"entity_name":"registerUser(String"}}}' | knot-mcp

# Find by parameter type (Kotlin)
echo '{"method":"tools/call","params":{"name":"find_callers","arguments":{"entity_name":"findById(Int"}}}' | knot-mcp

# Find by type annotation (TypeScript)
echo '{"method":"tools/call","params":{"name":"find_callers","arguments":{"entity_name":"(EventData"}}}' | knot-mcp

# Find by C# interface method (surfaces implementations + call sites)
echo '{"method":"tools/call","params":{"name":"find_callers","arguments":{"entity_name":"FindByIdAsync"}}}' | knot-mcp

Use Cases:

  • Dead Code Detection: Zero callers = unused code

  • Impact Analysis: "What breaks if I modify this?"

  • Refactoring Safety: Find all references before removing

  • Override Discovery (JVM + C#): For Java/Kotlin/Groovy/C# methods, results include an Overridden by group (implementations/overrides in subtypes) and an Overrides group (the supertype methods a method implements/overrides). These are backed by real OVERRIDES edges built at index time and resolved transitively at query time, so querying an interface/superclass method surfaces every implementation, and querying an implementation surfaces the declaration it overrides.

Truncation & completeness: the queried name is first resolved to concrete targets, capped at 25 by default (hard ceiling 500). When more targets match, the response states it explicitly and quantified:

> **Truncated** — 112 targets matched; showing the first 25 by FQN.

> **Counts below are partial** — they cover only the 25 of 112 targets shown.
> Re-run with a fully qualified name, or raise `max_targets`, for the complete set.

The relationship buckets (Calls, Extends, Implements, References) are complete for the resolved targets — there is no per-bucket cap — so the partial counts caveat tells you exactly what to do next: pass a fully qualified name to disambiguate homonyms, or raise max_targets (MCP parameter, default 25, maximum 500; CLI --max-targets) to opt into the full impact set.

The same shown-vs-total contract applies to search_hybrid_context: when the Sample callers: / Sample usages: blocks under an entity list fewer entries than the reported count, the header reads Sample callers — showing 3 of 21 (truncated): instead of implying the sample is the complete set.

Tool 3: explore_file

Understand file structure

Query: "What's in BrowserService.ts?"
Result: All classes, methods, constants, and functions with signatures and docs, plus compact markup summaries

Tool 4: list_repositories

Discover indexed codebases

Query: "What codebases are indexed?"
Result: Markdown table of all indexed repos with entity/file counts, language, and build system

Tool 4b: list_files

Enumerate a repository's files

Query: path = "src/hooks", repo = "my-app"
Result: Ordered table of the files under src/hooks with their entity counts

search_hybrid_context shares the same path parameter to scope results to that subtree.

Tool 5: list_repo_dependencies

Traverse cross-repository dependency graphs

Query: "What repositories depend on auth-lib?"
Result: Repositories declaring build dependencies (pom.xml, build.gradle, Cargo.toml, package.json, NuGet)

Indexing either side of a relationship creates the DEPENDS_ON edge: index a library after its consumers and they are linked retroactively (reverse sweep); an empty lookup is explained explicitly with a three-way classification — declared-but-unindexed dependencies are listed by name (uncapped), but a dependency that resolves to an indexed repository without an edge yet is reported as a stale graph with a re-index hint instead of being falsely called "not indexed"; the reverse direction names consumers that declare the repo without an edge — never a bare "No dependencies found."


🔗 MCP Client Configuration

Supported Clients

knot works with any MCP-compatible AI client:

  • ✅ Claude Desktop (Anthropic)

  • ✅ Gemini CLI (Google)

  • ✅ ChatGPT CLI / GPT (OpenAI)

  • ✅ Cursor (AI IDE)

  • ✅ Any standard MCP client

Configuration Examples

Claude Desktop

Add to claude_desktop_config.json:

{
  "mcpServers": {
    "knot": {
      "command": "/absolute/path/to/knot/target/release/knot-mcp",
      "env": {
        "KNOT_REPO_PATH": "/path/to/indexed/repo",
        "KNOT_QDRANT_URL": "http://localhost:6334",
        "KNOT_NEO4J_URI": "bolt://localhost:7687",
        "KNOT_NEO4J_USER": "neo4j",
        "KNOT_NEO4J_PASSWORD": "your-password"
      }
    }
  }
}

Gemini CLI

{
  "mcpServers": {
    "knot": {
      "command": "/absolute/path/to/knot/target/release/knot-mcp",
      "env": {
        "KNOT_REPO_PATH": "/path/to/indexed/repo",
        "KNOT_QDRANT_URL": "http://localhost:6334",
        "KNOT_NEO4J_URI": "bolt://localhost:7687",
        "KNOT_NEO4J_USER": "neo4j",
        "KNOT_NEO4J_PASSWORD": "your-password"
      }
    }
  }
}

ChatGPT / GPT CLI

Similar JSON configuration in your client's MCP configuration file.


⚙️ Configuration Reference

All options can be set via CLI flags, environment variables, or a ~/.config/knot/.env file. Priority (highest to lowest): CLI flags > environment variables > .env file.

Env Variable

CLI Flag

Default

Description

KNOT_REPO_PATH

--repo-path

(required)

Root directory of the repository to index

KNOT_REPO_NAME

--repo-name

(auto-detected)

Repository name for multi-repo isolation (auto-detected from last path component)

KNOT_QDRANT_URL

--qdrant-url

http://localhost:6334

Qdrant server URL

KNOT_QDRANT_COLLECTION

--qdrant-collection

knot_entities

Qdrant collection name

KNOT_NEO4J_URI

--neo4j-uri

bolt://localhost:7687

Neo4j Bolt URI

KNOT_NEO4J_USER

--neo4j-user

neo4j

Neo4j username

KNOT_NEO4J_PASSWORD

--neo4j-password

(required)

Neo4j password

KNOT_EMBED_MODEL

--embed-model

AllMiniLML6V2

Embedding model (AllMiniLML6V2 (default) or the opt-in BGEBaseENV15)

KNOT_EMBED_DIM

--embed-dim

(derived)

Deprecated (hidden): the dimension is derived from the model. A value that agrees warns; one that contradicts aborts.

KNOT_BATCH_SIZE

--batch-size

128

Entities per batch

KNOT_CLEAN

--clean

false

Force full re-index (delete all existing data)

KNOT_CUSTOM_CA_CERTS

--custom-ca-certs

(none)

Path to CA certificate bundle for corporate SSL proxies

KNOT_INCLUDE_CONFIG_FILES

--include-config-files

false

Include YAML/JSON/properties/K8s/Helm files in the index

RUST_LOG

(env only)

info

Log level: trace, debug, info, warn, error


🤖 Embedding Model Selection

knot supports exactly two embedding models, selected via KNOT_EMBED_MODEL or --embed-model. The default is chosen for backward compatibility: a v1.10.0 user upgrading finds zero re-index and zero configuration change — same model, same collection, same dimension.

Model

dim

Default?

Qdrant collection (derived)

AllMiniLML6V2

384

yes

knot_entities (unchanged since 1.0)

BGEBaseENV15 (opt-in)

768

no

knot_entities_bge768 (derived automatically)

The supported set is closed: the whole model universe for knot is this table (src/pipeline/embed/model.rs), and adding a future model must be a one-row change there. The four models available in v1.10.0 only (BGESmallENV15, MultilingualE5Small, JinaEmbeddingsV2BaseCode, NomicEmbedTextV15) were removed on purpose — setting one aborts startup with an error listing the two accepted names. The vector dimension is derived from the model: KNOT_EMBED_DIM / --embed-dim are hidden and deprecated (a value that agrees emits a deprecation warning; one that contradicts is a hard error because it means you believe a different model is active).

An explicit collection (KNOT_QDRANT_COLLECTION set by CLI, env or .env) always wins; when it is unset, the model's derived collection applies — BGE users could otherwise collide with a fixed-size knot_entities created at 384.

search_hybrid_context re-ranks independently of the model's cosine scale: each candidate pool is normalized before the boosts are applied, so switching model does not require re-tuning the ranker.

Startup guards (model markers)

Every index run records the producing model on the repository's :Repository node. On startup, knot and knot-mcp verify that the configured model matches the collection's real vector dimension and every marked repository:

  • Dimension mismatch with the collection — aborts with an actionable message (a collection's vector size is fixed at creation).

  • Every marked repository on another model — aborts, naming both models.

  • Some repositories on another model — warns and names the repositories that will not appear in semantic search. Scope-limited searches (search_hybrid_context with repo) specifically return a note instead of a silent empty result.

  • Legacy index without markers — the dimension infers the model (384 ⇒ AllMiniLML6V2, 768 ⇒ BGEBaseENV15); a consistent legacy index is never fail-closed and self-heals (writes the marker) on the next index run.

Runbook: opting into BGE-base

# 1. Choose the model (consumers read this too):
export KNOT_EMBED_MODEL=BGEBaseENV15

# 2. Every repository needs a clean re-index; the alternative collection
#    knot_entities_bge768 is derived automatically for every run afterwards:
KNOT_REPO_PATH=/path/to/repo KNOT_REPO_NAME=my-repo knot-indexer --clean

# 3. Restart every consumer (knot-mcp servers, knot CLI users).

Cost and recall measurements for both models live in docs/measurements/ (model_cost_1_11.md, model_matrix_1_11.md).


🎨 Custom Tree-sitter Queries

The built-in extraction queries (queries/java.scm, queries/typescript.scm, queries/csharp.scm) can be overridden without recompiling:

KNOT_CUSTOM_QUERIES_PATH=/path/to/my/queries ./target/release/knot-indexer

Place java.scm, typescript.scm, and/or csharp.scm in your custom directory. Missing files fall back to built-in defaults.


🔐 Corporate SSL / CA Certificates

In restricted corporate environments with SSL-inspecting proxies, you may need to provide a custom CA certificate bundle so that knot can download the embedding model from HuggingFace.

Via environment variable:

export KNOT_CUSTOM_CA_CERTS=/etc/ssl/certs/corporate-bundle.pem
./target/release/knot-indexer --repo-path /path/to/repo --neo4j-password secret

Via CLI flag:

./target/release/knot-indexer \
  --custom-ca-certs /etc/ssl/certs/corporate-bundle.pem \
  --repo-path /path/to/repo \
  --neo4j-password secret

Via .env file:

echo "KNOT_CUSTOM_CA_CERTS=/etc/ssl/certs/corporate-bundle.pem" >> ~/.config/knot/.env
./target/release/knot-indexer

This works for all three binaries: knot-indexer, knot-mcp, and knot.


🔄 Workflow Example

Step 1: Index a Java project

./target/release/knot-indexer --repo-path /home/user/my-java-app --neo4j-password secret

Step 2: Query via CLI (Instant search)

./target/release/knot search "authentication logic"
./target/release/knot callers "UserService.login"

Step 3: Start MCP server (For AI Agents)

./target/release/knot-mcp

Step 4: Use with Claude Desktop

  • Claude will list the read-only tool surface in its Tools menu

  • Ask: "Search for all authentication logic"

  • Ask: "Find who calls the login method"

  • Ask: "Explore the structure of UserService.java"

🤖 Auto-Configuring AI Agents

knot includes a universal .prompt file in its root directory that automatically configures modern AI coding agents (Cursor, Cline, opencode, Claude, etc.) to use the knot-mcp tools correctly.

The directive explicitly instructs AI agents to prioritize:

  • search_hybrid_context — for semantic code discovery (instead of grep)

  • find_callers — for reverse dependency analysis (instead of finding references manually)

  • explore_file — for file structure inspection (instead of reading line-by-line)

This ensures that when you ask an AI agent to analyze, refactor, or understand your code, it leverages the full power of the vector and graph databases rather than falling back to context-blind regex searches. The .prompt file is universal and tool-agnostic, working with any LLM client that reads codebase directives.


🤝 Contributing

Contributions are welcome! Please ensure:

  • All code passes cargo clippy and cargo fmt

  • No new unsafe code (unsafe_code = "deny" at crate level; one audited exception in src/utils/mod.rs for corporate proxy CA bundle injection, documented via #[expect(unsafe_code, reason = "…")])

  • Changes are compatible with Rust 2024 edition

  • All new functionality includes unit tests

  • Performance regressions are validated with the benchmark framework before submitting PRs

Development & Code Quality

make check                                  # Run all local quality gates (fmt, clippy, test, dupes)

# Or run gates individually:
cargo clippy --all-targets -- -D warnings  # Must pass
cargo fmt -- --check                        # Must pass
cargo test                                  # Run all unit tests
cargo dupes check                           # Code duplication check

Performance Benchmarking

The project includes a three-level benchmarking framework to validate optimizations and detect regressions:

Level 1 — Unit Benchmarks (Criterion):

cargo bench --bench pipeline_bench          # Parse + prepare throughput per language
cargo bench --bench graph_upsert_bench     # Neo4j UNWIND batching speedup (needs Neo4j)
cargo bench --bench channel_backpressure_bench  # Bounded channel overhead

Level 2 — E2E Integration Benchmarks:

# Full pipeline metrics with memory and per-stage timing
./tests/benchmark_e2e.sh --focus rust_e2e --output-dir /tmp/perf_results

# Compare against baseline (fails CI if tolerance exceeded)
scripts/compare_perf_metrics.sh /tmp/perf_results .perf_metrics/baseline.json

Level 3 — Token Efficiency Benchmark:

# Measures knot tool output vs grep + file reads on indexed repositories
python3 scripts/token_savings_benchmark.py \
  --config scripts/token_savings_tasks.json \
  --save-json .perf_metrics/token_savings.json

Unlike levels 1 and 2 (which measure indexing throughput), this one measures the consumer side: how many LLM tokens an agent spends to answer a question with knot versus by reading source files. Requires rg, a built knot binary, the repositories in the config already indexed, and optionally tiktoken for exact token counts. See Token Efficiency for the published results.

Baseline files: .perf_metrics/baseline.json stores the last known good metrics (committed, updated on main/master merges). Tolerance thresholds in .perf_metrics/threshold_tolerances.json control regression gates (±5% time, ±10% memory by default).

CI Integration: The test-performance job in .github/workflows/ci.yml runs after all E2E correctness tests pass, comparing results against baseline and fails the build on regression.


📜 License

This project is licensed under the MIT License. See LICENSE for details.


🚀 Roadmap

For the full release history see CHANGELOG.md.

Upcoming

Long-Term Vision

  • Go support

  • IDE plugins (VS Code, IntelliJ, Vim)

  • Language Server Protocol (LSP) integration

  • Automated Code Review tool (MCP-based)

  • Ruby support


💬 Questions?

For issues, feature requests, or discussions, please open a GitHub issue.

Available Tools

6 tools
explore_fileExplore file anatomyA
Read-onlyIdempotent

Read-only file anatomy inspection. Use this to list all classes, methods, and properties within a specific source file without reading its entire contents. Provides a structural bird's-eye view of a file, showing entity signatures and docstrings to quickly grasp a module's layout.

Usage: Use AFTER identifying an interesting file via 'search_hybrid_context' to understand its available methods, or before modifying a file. Do NOT use this for searching across multiple files.

Behaviour & Return: Read-only operation. Returns a Markdown-formatted outline of the file's entities, grouped by type (Classes, Methods, Interfaces, Constants, etc.), including line numbers for direct editor navigation. Attribute-usage markup references (html_class, html_id) are summarized in a separate section by default to avoid inflating code lists; set include_markup to true to list them in full. No side effects.

Path handling: file_path should be a repo-relative path (e.g. 'src/services/user.ts'). Absolute paths under your local checkout are also accepted; the tool strips the known local root automatically. The returned file_path is normalized to the same repo-relative form regardless of how it was queried. If the query is ambiguous across multiple repositories, the answer includes an 'ambiguous_path_candidates' list — retry with a longer path or pass repo_name.

Parameter guidance: 'file_path' must be a relative or absolute path to a valid source file. Include 'repo_name' if the file path might be ambiguous across multiple indexed repositories. Set 'include_markup' to true to expand attribute-usage markup tokens.

Supports Java, Kotlin, C#, and TypeScript codebases.

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathYesPath to the source file to explore. PREFERRED: a repo-relative path (e.g. 'src/services/user.ts'). ALSO ACCEPTED: an absolute path under the repository's local checkout (the tool strips KNOT_REPO_PATH / CWD automatically).
repo_nameNoOptional but HIGHLY RECOMMENDED: repository scope. Accepts a single repository name (`'my-repo'`), a comma-separated list (`'repo-a,repo-b'`), or `'all'` (or `'*'`) to query every indexed repository. If you know the repository you are working on, include it in your FIRST query to avoid mixed results from other indexed projects. Omit to search across all repositories.
include_markupNoOptional: set to true to list all attribute-usage markup tokens (html_class, html_id) in full rather than in a compact summary. Defaults to false.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false; the description adds that it returns a Markdown-formatted outline grouped by type with line numbers, summarizes markup tokens by default, and normalizes paths. It also discloses ambiguity handling via 'ambiguous_path_candidates'. These details go well beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is structured with clear paragraphs and front-loads the core purpose in the first sentence. Each section addresses a distinct concern (usage, return, path handling, params, languages), so the length is justified despite minor redundancy with the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is complex (path handling, ambiguity, markup control) and the description covers all of it: when to use, what it returns, how paths are normalized, what the markup flag does, and which languages are supported. The absence of an output schema is compensated by the explicit return-format description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already documents all three parameters thoroughly (100% coverage). The description's parameter guidance largely reiterates the schema, but it adds extra value by explaining path normalization, the 'ambiguous_path_candidates' retry hint, and the supported languages (Java, Kotlin, C#, TypeScript), which helps interpret valid file_path values.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Read-only file anatomy inspection' and explicitly says it lists all classes, methods, and properties within a specific source file without reading its entire contents. This states a specific verb and resource and differentiates it from file-listing and search siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly directs use AFTER search_hybrid_context and before modifying a file, and warns 'Do NOT use this for searching across multiple files.' It names the alternative and the condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_callersFind callers (reverse dependencies)A
Read-onlyIdempotent

Read-only reverse dependency lookup. Use this to find all code that references, calls, extends, or implements a specific entity. Answers 'who uses this code?' by querying the graph database. Differs from search tools by providing exact dependency tracking.

Usage: Use for impact analysis before refactoring or to detect dead code. Do NOT use this for semantic feature discovery—use 'search_hybrid_context' instead.

Matching is precedence-based: exact FQN (containing '.' or '::') → FQN suffix (Type.member) → exact name → signature prefix (accept(List) → name prefix (queries under 4 characters) → fuzzy substring (queries of 4+ characters). The first tier that matches wins, so an exact name never returns fuzzy noise. Queries shorter than 4 characters resolve by anchored name prefix instead of substring, avoiding substring noise on short strings. Pass a qualified name (Namespace.Type.Member) to disambiguate homonyms. Responses state which tier matched and flag fuzzy results explicitly.

Behaviour & Return: Read-only graph traversal with no side effects. Returns Markdown grouped by relationship type (Calls, Extends, Implements, References, Overridden by, Overrides) with exact file paths and line numbers. Each caller entry and each resolved target states its repository as (repo: name), so rows are attributable when multiple repositories are in scope. For JVM code (Java/Kotlin/Groovy) and C#, 'Overridden by' lists method implementations/overrides in subtypes and 'Overrides' lists the supertype methods a method implements/overrides. When the query resolves to more than one entity with that name (homonyms, e.g., 'find_nearest_entity_by_line' in orphans.rs vs rust.rs), results are grouped by target entity showing which specific target each caller references — even when only one of the homonyms has callers. Each caller entry includes: name, kind, file_path:line_number, and signature. When multiple targets exist, each group shows the target's location and signature.

Entity-kind scope: target resolution is code-only by default — documentation, configuration, build-system and Kubernetes/Helm metadata (markdown_section, config_property, build_dependency, cargo_package, project_identity, k8s_*, helm_*, …) can never be presented as resolved targets. When the filter removed matches, the response says so ('Non-code matches hidden — N entities …'), never silently. Pass kinds='all' (or '*') to disable the filter, or a comma-separated allow-list of exact kinds/aliases ('callable', 'config', 'docs', 'rust_function', 'build_dependency', …) to scope resolution explicitly. The response's resolution.kind_filter field states which scope applied ('code_default', 'any', 'explicit').

Relationship coverage: the buckets cover every edge type the pipeline produces — Calls, Extends, Implements, References, Macro calls (MACRO_CALLS), DOM references (JS → HTML id), CSS class usage (JS → CSS class), script/stylesheet imports, and the VCL edges (uses backend/probe/acl, includes, imports vmod, declared-unused) — plus Overridden by / Overrides.

Truncation & completeness: the queried name is first resolved to concrete targets (capped at 25 by default). When more targets match than fit the cap, the response states 'Truncated — N targets matched; showing the first M by FQN' and 'Counts below are partial — they cover only the M of N targets shown', so bucket counts are never mistaken for the complete impact set. Raise 'max_targets' (up to 500) to retrieve more targets when the notice reports truncation.

Parameter guidance: 'entity_name' supports exact names or signature fragments (e.g., 'handleRequest' or 'handle(Request'). Include 'repo_name' to filter results to the specific codebase being analyzed.

Supports Java, Kotlin, C#, Rust, and TypeScript codebases.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindsNoOptional entity-kind scope for target resolution. Omit for the default code-only scope (docs/config/build metadata are hidden from the target list and the response discloses them). Use 'all' (or '*') to disable filtering, or a comma-separated allow-list of exact kinds or aliases ('callable', 'class', 'config', 'docs', 'rust_function', 'build_dependency', ...).
repo_nameNoOptional but HIGHLY RECOMMENDED: repository scope. Accepts a single repository name (`'my-repo'`), a comma-separated list (`'repo-a,repo-b'`), or `'all'` (or `'*'`) to query every indexed repository. If you know the repository you are working on, include it in your FIRST query to avoid mixed results from other indexed projects. Omit to search across all repositories.
entity_nameYesThe name of the function, method, or class to find callers for
max_targetsNoMaximum number of resolved targets to include (default: 25, max: 500). Raise this when the response reports a truncated target list and you need the complete impact set.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and idempotentHint=true, and the description aligns: 'Read-only graph traversal with no side effects.' It adds substantial behavioral detail beyond annotations: precedence-based matching tiers, truncation behavior, hidden non-code match disclosure, grouping by relationship type, and per-repository attribution.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but well-structured with labeled sections (Usage, Behaviour & Return, Entity-kind scope, Relationship coverage, Truncation, Parameter guidance) and is front-loaded with purpose and usage. Each section adds meaningful detail for a complex tool, though some content could be tightened without losing value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex tool with no output schema, the description is remarkably complete: it covers matching semantics, return format, kind filtering, truncation and completeness caveats, relationship types, language support, and parameter guidance. An agent has sufficient information to invoke the tool correctly and interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the description enriches all parameters: entity_name supports exact names or signature fragments, repo_name is recommended for scoping, kinds has a detailed allow-list explanation, and max_targets is tied to truncation behavior ('Raise max_targets... when the notice reports truncation'). This goes well beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Read-only reverse dependency lookup' that finds 'all code that references, calls, extends, or implements a specific entity.' It explicitly differentiates itself from search tools by providing 'exact dependency tracking,' so an agent can distinguish it from siblings like search_hybrid_context without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage guidance is explicit and actionable: use it for impact analysis before refactoring or dead-code detection, and do NOT use it for semantic feature discovery, naming search_hybrid_context as the alternative. This gives the agent clear selection criteria relative to sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_filesList indexed repository filesA
Read-onlyIdempotent

Read-only listing of the files an indexed repository carries, with entity counts and deterministic order. Answers 'list every file under src/hooks' or 'which files live in src/api/**' without prior path knowledge.

Usage: Use this BEFORE 'search_hybrid_context' when you do not know the codebase layout; then pass the same prefix to the search's optional 'path' parameter to scope results to those files. Do NOT use this to enumerate entities — use 'explore_file' on a file from this listing for its anatomy.

Behaviour & Return: Read-only query with no side effects. Returns a Markdown table with columns: REPOSITORY, FILE, ENTITIES, ordered by (repository, path). When output exceeds 2000 files, the table displays the first 2000 matches alongside a truncation notice stating the exact total when known, or an explicit lower bound when the scan stopped early. When nothing matches the prefix, returns 'No indexed files matched the given path'.

Parameter guidance: 'path' is optional. PREFERRED: a repo-relative directory prefix (e.g. 'src/api') matched on a path boundary, or a glob ('src/**/*_test.rs'). Prefixes and globs page all indexed files in pages with per-page boundary filtering without dropping candidates. Absolute paths under the local checkout are accepted and normalized like 'explore_file'. Omit to list every indexed file (capped; the reply notes truncation).

Parameter guidance: 'repo_name' scopes the listing. Accepts a single repository name or a comma-separated list; include it when several indexed repositories may share path shapes.

Supports all languages indexed by knot.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoOptional repo-relative directory prefix (e.g. 'src/api', matched on a path boundary) or glob pattern (segments: `*` segment wildcard, `**` any depth, `?` one char). Absolute paths under the local checkout are accepted and normalized like 'explore_file'. Omit to list every indexed file.
repo_nameNoOptional but HIGHLY RECOMMENDED: repository scope. Accepts a single repository name (`'my-repo'`), a comma-separated list (`'repo-a,repo-b'`), or `'all'` (or `'*'`) to query every indexed repository. Include it when several indexed repositories may share path shapes.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark readOnly/idempotent/non-destructive, but the description adds behavioral details beyond hints: deterministic ordering by (repository, path), a 2000-file cap with a truncation notice showing exact or lower-bound totals, and the exact empty-result string. There is no contradiction with the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but deliberately sectioned (Usage, Behaviour & Return, Parameter guidance) and front-loaded with its core purpose. Some prose repeats annotation facts and schema descriptions, so it is not maximally tight, but the structure keeps the extra length navigable and each section earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter read-only listing tool, the description covers intended use, chaining with siblings, exact return columns and order, truncation behavior, empty-match response, and both parameter semantics. Since there is no output schema, explicitly documenting the return format and edge cases is valuable and fully handled.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so the baseline is 3, but the description goes well beyond schema text: it explains path-boundary prefix matching, glob semantics, paging behavior ('per-page boundary filtering without dropping candidates'), absolute-path normalization, and when repo_name is needed (shared path shapes across repositories). This materially improves call construction.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence states a specific verb ('listing'), a specific resource ('files an indexed repository carries'), and two distinguishing outputs (entity counts, deterministic order). It explicitly distinguishes its contribution from search_hybrid_context and explore_file, so it cannot be confused with siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'Usage' section states concretely when to use it ('BEFORE search_hybrid_context when you do not know the codebase layout'), how to chain it with that sibling, and an explicit negative instruction ('Do NOT use this to enumerate entities — use explore_file'). This is explicit, actionable routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_repo_dependenciesList cross-repository dependenciesA
Read-onlyIdempotent

Read-only cross-repository dependency graph lookup. Shows which repositories depend on each other via build system declarations (Maven, Gradle, Cargo, npm, NuGet). Answers 'which repos does this repo depend on?' and 'which repos depend on this repo?'.

Usage: Use BEFORE cross-repo analysis to discover which other indexed repos are available for call tracing. Use reverse mode for impact analysis before making breaking changes in shared libraries.

Behaviour & Return: Read-only graph traversal with no side effects. Returns a JSON array of repository names. Empty results mean no DEPENDS_ON relationships exist for that repo. Empty lookups are explained in the response text with a three-way classification: declares-but-resolves-without-edge (stale graph, re-index hint), declares-but-nothing-resolves (not indexed), or nothing declared; the reverse direction names consumers that declare the repo without an edge yet.

Parameter guidance: 'repo_name' is required and must match the name used during indexing. 'max_depth' defaults to 3 (1 = direct only) and applies to both directions — in reverse mode it follows dependents transitively. 'reverse' toggles between forward and reverse dependency lookup.

Supports all build systems indexed by knot: Maven, Gradle, Cargo, npm, NuGet (.csproj + Central Package Management via Directory.Packages.props). C# repos that previously reported build_system: "none" now report "nuget" on re-index; knot-indexer --clean is recommended for immediate effect.

ParametersJSON Schema
NameRequiredDescriptionDefault
reverseNoIf true, show repositories that depend ON this repo (reverse lookup). If false (default), show repositories this repo depends ON. Use reverse for impact analysis before breaking changes.
max_depthNoMaximum depth for transitive dependency traversal (default: 3, max: 10). Use 1 for direct dependencies only. Requests above 10 are clamped to 10 (and below 1 to 1). Applies to both directions: with reverse=true it follows dependents transitively.
repo_nameYesRepository name to show dependencies for. Must match the name used during indexing (e.g., 'my-java-repo', 'auth-service'). This is REQUIRED — there is no default.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the readOnly/idempotent annotations by disclosing that it is a side-effect-free graph traversal, specifying the JSON array return type, and explaining empty result semantics with a three-way classification including stale-graph and not-indexed cases. This is rich, actionable behavioral detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but exceptionally well structured with labeled paragraphs for usage, behavior, return values, parameter guidance, and supported build systems. Every sentence carries actionable information and is front-loaded with the core purpose before caveats.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is no output schema, the description fully compensates by explaining return shape, empty-result semantics, direction behavior, depth semantics, and supported build systems. For a 3-parameter tool with rich annotations, nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description mostly restates the schema's parameter guidance (repo_name required, max_depth default/applies to both directions, reverse toggling) rather than adding entirely new meaning, though it does usefully summarize the key constraints in one place.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('lookup') and precise resource ('cross-repository dependency graph'), and explicitly frames the two questions it answers: which repos does this repo depend on and which depend on it. This clearly distinguishes it from sibling tools like list_repositories and find_callers.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit usage guidance is provided: use before cross-repo analysis for call tracing, and use reverse mode for impact analysis before breaking changes. It does not explicitly name alternatives or state when not to use it, so it stops short of a 5, but the context is clear enough for an agent to select it appropriately.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_repositoriesList indexed repositoriesA
Read-onlyIdempotent

Read-only listing of all indexed repositories with optional name filtering. Shows repository metadata including entity count, file count, build system, and primary language. Answers 'what codebases have I indexed?' and 'which repositories match this name?'.

Usage: Use this tool FIRST to discover available codebases before searching or exploring. Once you know the repository name, switch to 'search_hybrid_context' for semantic search, 'find_callers' for reverse dependency lookup, 'explore_file' for file anatomy, or 'list_repo_dependencies' for cross-repo dependency graphs. Do NOT use this tool to search for code entities — use 'search_hybrid_context' instead.

Behaviour & Return: Read-only query with no side effects. Returns a Markdown table with columns: REPO, BUILD SYSTEM, LANGUAGE, FILES, ENTITIES. When no repositories match the filter, returns 'No repositories found.'

Parameter guidance: 'filter' is optional. When provided, only repositories whose name contains the filter string are returned (case-insensitive substring match). Omit to list all indexed repositories.

Supports all languages and build systems indexed by knot.

ParametersJSON Schema
NameRequiredDescriptionDefault
filterNoOptional filter to narrow down repositories by name (case-insensitive substring match). When provided, only repositories whose name contains this string are returned. Examples: 'auth' matches 'auth-service' and 'Auth-Lib', 'api' matches 'my-api'. Omit to list all indexed repositories.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds genuinely new context: the return is a Markdown table with named columns, and the no-match case returns 'No repositories found.' That extra disclosure is valuable since there is no output schema, though permission/indexing prerequisites are not addressed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then clearly sectioned into Usage, Behaviour & Return, and Parameter guidance. Slightly longer than necessary because the filter semantics and read-only nature repeat structured fields, but every section is scannable and earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description supplies the return shape and empty-result behavior; with four sibling tools, it routes the agent to the correct alternative; with one fully-documented param, nothing else is needed. Complete for a low-complexity read tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single 'filter' param is fully documented in the schema (substring, case-insensitive, examples, optional). The description restates the same semantics rather than adding new meaning, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Read-only listing of all indexed repositories') plus the metadata surfaced and the questions it answers. An agent can distinguish it from search_hybrid_context, find_callers, explore_file, and list_repo_dependencies without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says to use this tool FIRST to discover codebases, names the exact sibling to switch to for each follow-up intent, and includes an explicit 'Do NOT use this tool to search for code entities' exclusion. Full when/when-not/alternatives coverage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_hybrid_contextHybrid semantic + structural searchA
Read-onlyIdempotent

Read-only semantic and structural code search combining vector embeddings with graph analysis. Use this for initial codebase discovery to find features by their meaning (e.g., 'user authentication'). Locates code based on natural language descriptions instead of exact keywords, returning relevant files, signatures, and documentation.

⚠️ PREREQUISITE: This tool requires an active knot-mcp server with vector database (Qdrant) and graph database (Neo4j) initialized.

Behavior & Return: Performs a read-only dual query against vector DB (for semantic similarity) and graph DB (for architectural relationships). Returns Markdown-formatted results with file paths, line numbers, code snippets, and cross-repository dependencies. No side effects.

Usage: Use as your FIRST step when exploring unfamiliar code or discovering architectural patterns. Do NOT use this to find all usages of a specific function—use the 'find_callers' tool for that instead.

Ranking contract: results are kind-aware — function/method/class/struct definitions outrank markdown docs, test files, config properties and build-dependency entities for natural-language queries; callers and helpers appear as context attached to a definition, never as substitutes. The shared entry point of the highest-ranked helpers outranks those helpers a loose paraphrase surfaces.

Generic-verb guard: an entity merely named after a generic verb or noun (find/get/create/build/acquire/borrow/current/…) does not win on that name alone; the full name boost is paid only when the entity's container context (FQN) corroborates a second query token.

Recall contract: entity embeds carry identifier tokens and the tokenized call names of the entity's body, so a paraphrase of what a definition does (even one with no doc comment) still surfaces it. Entities whose identifier shares a word with the query enter the candidate pool by token match alone.

Result bound: 'max_results' is 1-100 (default 5) and is enforced — a larger request is clamped to 100 and the reply says so. There is no pagination: when the bound is not enough, narrow the scope with 'kinds' / 'path' / 'repo_name' or refine the query rather than raising the limit.

Parameter guidance: 'query' should be 2-5 words describing functionality. Increase 'max_results' to 10-20 for broad discovery, keep at 5 for focused search. Include 'repo_name' in your first query to avoid cross-repository pollution.

Supports Java, Kotlin, C#, and TypeScript codebases.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoOptional path filter. A repo-relative directory prefix ('src/api', matched on a path boundary so 'src/api-notes.md' never matches) or a glob ('src/**/*_test.rs'). Use 'list_files' first when you do not know the layout. Omit to search every file.
kindsNoOptional entity-kind filter. Accepts exact wire-format kinds (`'rust_function'`, `'markdown_section'`, `'kotlin_class'`, …) or aliases: `'definition'` (all functions/methods/types), `'callable'`/`'function'`/`'method'` (all callable kinds), `'class'`/`'type'`/`'struct'` (all type kinds). Comma-separate for multiple values. Omit to search all kinds.
queryYesSearch query describing what you're looking for (e.g., 'user authentication', 'API error handling')
repo_nameNoOptional but HIGHLY RECOMMENDED: repository scope. Accepts a single repository name (`'my-repo'`), a comma-separated list (`'repo-a,repo-b'`), or `'all'` (or `'*'`) to query every indexed repository. If you know the repository you are working on, include it in your FIRST query to avoid mixed results from other indexed projects. Omit to search across all repositories.
max_resultsNoMaximum number of results to return (default: 5, max: 100). Requests above 100 are clamped to 100 and the reply says so — there is no cursor or pagination; to look past the bound, narrow the search with 'kinds' / 'path' / 'repo_name' or refine the query.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint, but the description adds substantial context: no side effects, clamping behavior on max_results, absence of pagination, ranking kind-awareness, recall contract, and the prerequisite server/database setup. This goes far beyond what annotations convey and is consistent with them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is comprehensive but quite lengthy, with many distinct sections (prerequisite, behavior, usage, ranking contract, recall contract, result bound, parameter guidance, supported languages). While each section earns its place for a complex tool, the overall length dilutes focus. It is well-structured with headers and emoji but could be tightened without losing critical information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a hybrid search tool with no output schema, the description covers prerequisites, return format, ranking, recall, result bounds, parameter tuning, and language support. The agent has everything needed to invoke it correctly and interpret results, including explicit exclusions (no pagination, clamp behavior).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so baseline is 3, but the description adds practical guidance beyond the schema: recommends 2-5 word queries, suggests max_results ranges for broad vs focused discovery, and explains path boundary matching and glob usage. This meaningfully supplements the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it is a read-only semantic and structural code search combining vector embeddings with graph analysis, and explicitly contrasts itself with find_callers for usage-of-a-function. The verb+resource+method is specific and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells the agent to use this as the first step for unfamiliar code or architectural discovery, and gives a concrete 'do NOT use' with the alternative tool (find_callers). Also includes parameter guidance (query length, max_results values, repo_name) and a ranking contract that governs result interpretation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev1.11.1
    • Changedexplore_file1 field changed
      • addedInput schema / properties / include_markup
        Added value: +{
        +  "description": "Optional: set to true to list all attribute-usage markup tokens (html_class, html_id) in full rather than in a compact summary. Defaults to false.",
        +  "type": [
        +    "boolean",
        +    "null"
        +  ]
        +}
  2. 4 tool updatesv1.9.6
    • Changedfind_callers2 fields changed
      • addedInput schema / properties / kinds
        Added value: +{
        +  "description": "Optional entity-kind scope for target resolution. Omit for the default code-only scope (docs/config/build metadata are hidden from the target list and the response discloses them). Use 'all' (or '*') to disable filtering, or a comma-separated allow-list of exact kinds or aliases ('callable', 'class', 'config', 'docs', 'rust_function', 'build_dependency', ...).",
        +  "maxLength": 1024,
        +  "minLength": 1,
        +  "type": [
        +    "string",
        +    "null"
        +  ]
        +}
      • addedInput schema / properties / max_targets
        Added value: +{
        +  "default": 25,
        +  "description": "Maximum number of resolved targets to include (default: 25, max: 500). Raise this when the response reports a truncated target list and you need the complete impact set.",
        +  "maximum": 500,
        +  "minimum": 1,
        +  "type": [
        +    "integer",
        +    "null"
        +  ]
        +}
    • Addedlist_files
    • Changedlist_repo_dependencies1 field changed
      • changedInput schema / properties / max_depth / description
        Previous value: -"Maximum depth for transitive dependency traversal (default: 3). Use 1 for direct dependencies only. Higher values follow chains deeper. Must be between 1 and 10."New value: +"Maximum depth for transitive dependency traversal (default: 3, max: 10). Use 1 for direct dependencies only. Requests above 10 are clamped to 10 (and below 1 to 1). Applies to both directions: with reverse=true it follows dependents transitively."
    • Changedsearch_hybrid_context4 fields changed
      • addedInput schema / properties / kinds
        Added value: +{
        +  "description": "Optional entity-kind filter. Accepts exact wire-format kinds (`'rust_function'`, `'markdown_section'`, `'kotlin_class'`, …) or aliases: `'definition'` (all functions/methods/types), `'callable'`/`'function'`/`'method'` (all callable kinds), `'class'`/`'type'`/`'struct'` (all type kinds). Comma-separate for multiple values. Omit to search all kinds.",
        +  "maxLength": 255,
        +  "minLength": 1,
        +  "type": [
        +    "string",
        +    "null"
        +  ]
        +}
      • changedInput schema / properties / max_results / description
        Previous value: -"Maximum number of results to return (default: 5)"New value: +"Maximum number of results to return (default: 5, max: 100). Requests above 100 are clamped to 100 and the reply says so — there is no cursor or pagination; to look past the bound, narrow the search with 'kinds' / 'path' / 'repo_name' or refine the query."
      • changedInput schema / properties / max_results / maximum
        Previous value: -20New value: +100
      • addedInput schema / properties / path
        Added value: +{
        +  "description": "Optional path filter. A repo-relative directory prefix ('src/api', matched on a path boundary so 'src/api-notes.md' never matches) or a glob ('src/**/*_test.rs'). Use 'list_files' first when you do not know the layout. Omit to search every file.",
        +  "maxLength": 500,
        +  "minLength": 1,
        +  "type": [
        +    "string",
        +    "null"
        +  ]
        +}
  3. 5 tool updatesv1.9.4
    • Changedexplore_file1 field changed
      • changedInput schema / properties / repo_name / type
        Previous value: -"string"New value: +[
        +  "string",
        +  "null"
        +]
    • Changedfind_callers1 field changed
      • changedInput schema / properties / repo_name / type
        Previous value: -"string"New value: +[
        +  "string",
        +  "null"
        +]
    • Changedlist_repo_dependencies2 fields changed
      • changedInput schema / properties / max_depth / type
        Previous value: -"integer"New value: +[
        +  "integer",
        +  "null"
        +]
      • changedInput schema / properties / reverse / type
        Previous value: -"boolean"New value: +[
        +  "boolean",
        +  "null"
        +]
    • Changedlist_repositories1 field changed
      • changedInput schema / properties / filter / type
        Previous value: -"string"New value: +[
        +  "string",
        +  "null"
        +]
    • Changedsearch_hybrid_context2 fields changed
      • changedInput schema / properties / max_results / type
        Previous value: -"integer"New value: +[
        +  "integer",
        +  "null"
        +]
      • changedInput schema / properties / repo_name / type
        Previous value: -"string"New value: +[
        +  "string",
        +  "null"
        +]
  4. 3 tool updatesv1.8.1
    • Changedexplore_file1 field changed
      • changedInput schema / properties / repo_name / description
        Previous value: -"Optional but HIGHLY RECOMMENDED: repository name to filter results to a specific codebase (e.g., 'my-java-repo'). If you know the repository you are working on, include this in your FIRST query to avoid mixed results from other indexed projects. Omit only to search across all repositories."New value: +"Optional but HIGHLY RECOMMENDED: repository scope. Accepts a single repository name (`'my-repo'`), a comma-separated list (`'repo-a,repo-b'`), or `'all'` (or `'*'`) to query every indexed repository. If you know the repository you are working on, include it in your FIRST query to avoid mixed results from other indexed projects. Omit to search across all repositories."
    • Changedfind_callers1 field changed
      • changedInput schema / properties / repo_name / description
        Previous value: -"Optional but HIGHLY RECOMMENDED: repository name to filter results to a specific codebase (e.g., 'my-java-repo'). If you know the repository you are working on, include this in your FIRST query to avoid mixed results from other indexed projects. Omit only to search across all repositories."New value: +"Optional but HIGHLY RECOMMENDED: repository scope. Accepts a single repository name (`'my-repo'`), a comma-separated list (`'repo-a,repo-b'`), or `'all'` (or `'*'`) to query every indexed repository. If you know the repository you are working on, include it in your FIRST query to avoid mixed results from other indexed projects. Omit to search across all repositories."
    • Changedsearch_hybrid_context1 field changed
      • changedInput schema / properties / repo_name / description
        Previous value: -"Optional but HIGHLY RECOMMENDED: repository name to filter results to a specific codebase (e.g., 'my-java-repo'). If you know the repository you are working on, include this in your FIRST query to avoid mixed results from other indexed projects. Omit only to search across all repositories."New value: +"Optional but HIGHLY RECOMMENDED: repository scope. Accepts a single repository name (`'my-repo'`), a comma-separated list (`'repo-a,repo-b'`), or `'all'` (or `'*'`) to query every indexed repository. If you know the repository you are working on, include it in your FIRST query to avoid mixed results from other indexed projects. Omit to search across all repositories."
  5. 1 tool updatev1.5.1
    • Changedexplore_file1 field changed
      • changedInput schema / properties / file_path / description
        Previous value: -"Absolute path to the source file to explore"New value: +"Path to the source file to explore. PREFERRED: a repo-relative path (e.g. 'src/services/user.ts'). ALSO ACCEPTED: an absolute path under the repository's local checkout (the tool strips KNOT_REPO_PATH / CWD automatically)."
  6. 1 tool updatev1.4.11
    • Addedlist_repositories
  7. 4 tool updatesv1.4.0
    • Addedexplore_file
    • Addedfind_callers
    • Addedlist_repo_dependencies
    • Addedsearch_hybrid_context
  8. 4 tool updatesv1.3.8
    • Removedexplore_file
    • Removedfind_callers
    • Removedlist_repo_dependencies
    • Removedsearch_hybrid_context
  9. 4 tool updates
    • Addedexplore_file
    • Addedfind_callers
    • Addedlist_repo_dependencies
    • Addedsearch_hybrid_context

TDQS

A4.7/5.0

Scored across 6 tools

Disambiguation5/5

Each tool targets a distinct concern: repository listing, file listing, file anatomy, semantic search, exact caller lookup, and cross-repo dependency graphs. Descriptions explicitly cross-reference each other to steer usage, leaving no realistic overlap between tools.

Naming Consistency5/5

All tool names follow a consistent snake_case verb_noun pattern: list_files, list_repositories, explore_file, find_callers, search_hybrid_context, list_repo_dependencies. The naming clearly communicates the action and target for every tool.

Tool Count5/5

Six tools is well-scoped for a code-intelligence server covering repository discovery, file exploration, semantic search, and dependency analysis. Each tool has a clear place in the workflow and none feel redundant.

Completeness4/5

The surface covers the core read-only codebase exploration lifecycle: discover indexed repos, list files, inspect file structure, search semantically, find callers, and map repo dependencies. A minor gap is the absence of a raw source-content retrieval tool, since exploration is limited to structural outlines and snippets.

Maintenance

ActivityActive
ResponsivenessWithin a week

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    D
    maintenance
    A complete MCP server for Retrieval-Augmented Generation with file management and vector memory for agents. Supports multiple document formats (PDF, DOCX, TXT, MD, CSV, JSON) with semantic search using Hugging Face embeddings and ChromaDB for efficient vector storage.
    11
    9 npm
    1
    MIT
  • A
    license
    Not graded
    quality
    Not graded
    maintenance
    Turns Claude Desktop into a personal document question-answering system using local vector search. Index PDF, TXT, and Markdown documents into collections and get answers based strictly on your documents with zero hallucination.
    9 npm
    -
  • A
    license
    Not graded
    quality
    C
    maintenance
    A graph-powered code intelligence engine that indexes codebases into a structural knowledge graph to provide AI agents with deep context on function calls, types, and execution flows. It offers local, zero-dependency tools for hybrid search, impact analysis, and dead code detection across Python, JavaScript, and TypeScript projects.
    956 PyPI
    814
    MIT
  • F
    license
    Not graded
    quality
    D
    maintenance
    A minimalist indexing tool that provides AI agents with semantic search and structural AST parsing for deep codebase understanding. It enables autonomous agents to navigate large codebases predictably using vector embeddings and native language server capabilities like definition and reference tracking.
    -