Skip to main content
Glama
dfine-io

dfine-semantic

Official
by dfine-io

@dfine-io-gmbh/semantic-mcp

Semantic code search as an MCP server. It embeds a project's code into a local SQLite vector store and answers natural-language questions that a plain grep cannot match. Everything runs locally once the model is downloaded.

How it works

  • Code is split into chunks and embedded with jinaai/jina-embeddings-v2-base-code (768 dimensions) through @huggingface/transformers.

  • Vectors live in one local SQLite database per project (sqlite-vec).

  • The server speaks MCP over stdio, revision 2026-07-28 and the 2025 revisions, and works with any MCP client, for example Claude Code or Cursor.

  • Each search first re-indexes the files that changed since the last one. If that would take longer than about a minute, clients on MCP 2026-07-28 that support forms ask you first. Other clients re-index without asking.

  • Long runs report progress to clients that ask for it. Claude Code and other clients that reset their timeout on progress keep such a call open until it finishes.

Related MCP server: Smart Coding MCP

Requirements

  • Node.js 22 or newer

  • Git on the PATH. Every indexed root must be inside a git work tree, because files are listed with git ls-files.

  • Network access to download the model (about 640 MB) into ~/.dfine-semantic/models when it is not there yet

  • better-sqlite3 and sqlite-vec ship prebuilt binaries for macOS (arm64, x64), Linux (x64, arm64) and Windows (x64). Other targets build from source and need a C/C++ toolchain.

Use it with an MCP client

No install step is needed: npx fetches and runs the server. Add it to your MCP client config (.mcp.json for Claude Code project scope, or ~/.claude.json for user scope):

{
  "mcpServers": {
    "dfine-semantic": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "@dfine-io-gmbh/semantic-mcp"]
    }
  }
}

The server sends short usage instructions to the client when it connects.

Tools

Name

Purpose

semantic_search

Natural-language query, returns ranked file:line references

index_project

Index or refresh a project root, embedding only what changed

index_status

List indexed projects with file, chunk and size counts

find_duplicates

List near-identical code for files or line ranges (opt-in, see below)

semantic_search and find_duplicates also return their results as structured data, described by each tool's output schema. Its status field tells an agent that duplicate search is off for a project, which is not the same as finding no duplicates.

semantic_search covers .ts and .tsx by default. Pass include (for example [".md", ".vue"]) to search other indexed file types. index_project indexes the allow-listed extensions you pass and keeps that list, so a later call without extensions indexes the same file types.

Duplicate search (opt-in)

find_duplicates is off until you turn it on for a project. It needs a second index of short code windows, and building that index takes time and disk space, so projects that never use it never pay for it.

  1. Ask your agent to run index_project with duplicates: true once for the project. On an indexed repository with about 1,400 TypeScript files, this takes about 45 minutes and 70 MB. A project without a search index yet needs about 35 minutes more.

  2. Ask for duplicates of files or line ranges, for example src/a.ts:10-40, and your agent calls find_duplicates. Searches and index runs keep the windows of edited files current. When a result says windows are missing, run index_project again.

  3. Run index_project with duplicates: false to switch it off and delete the duplicate index.

Each result pairs a range of your file with a similar range in another file, plus a similarity score. Read both ranges before you merge anything. At the default threshold of 0.88, the tool found about 40% of the real duplicates in a measured TypeScript project; a higher threshold returns fewer false pairs. It covers .ts, .tsx, .js, .jsx and .mjs files and skips tests, specs and .d.ts files. Pass exclude with folders such as src/generated to leave generated or vendored code out.

Configuration

Variable

Default

Purpose

SEMANTIC_ALLOWED_ROOTS

cwd, ~/.claude

Extra absolute roots the server may index (comma-separated)

SEMANTIC_DATA_DIR

~/.dfine-semantic

Where the indexes and the model are stored

The working directory counts as a root unless it is / or your home folder. Indexes are keyed by project path and survive upgrades. Run index_project with force: true for a clean rebuild.

Upgrading from 0.1.3 or older

  • Restart every session that still runs 0.1.3 or older. An old server cannot read an upgraded index.

  • The model downloads once more into its new folder. Copies inside old npx caches can be deleted.

  • Files are now split into chunks differently. Searches keep using your existing index, and clients that support forms offer a rebuild. You can also run index_project.

  • find_duplicates is new and stays off until your agent runs index_project with duplicates: true.

Security

See SECURITY.md for the security model and how to report a vulnerability.

Local development

pnpm install
pnpm build        # tsc
pnpm lint:dlint   # dlint, the dfine linter on the TypeScript compiler
pnpm check        # tsc, dlint and prettier
pnpm start

License and support

MIT, see LICENSE. Questions or issues: support@dfine.io or the issue tracker.

Available Tools

4 tools
find_duplicatesFind duplicate codeA
Read-onlyIdempotent

List code in other files that nearly matches the given files or line ranges. Use it in reviews and refactors - pass changed line ranges ("src/a.ts:10-40") for sharper pairs. Treat every pair as a candidate - read both ranges before you call it a duplicate. Expect it to be off per project - ask the user before you turn it on. Turn it on with index_project duplicates: true - the first build can take over an hour on large projects. Covers .ts, .tsx, .js, .jsx, .mjs files; skips tests, specs and .d.ts files.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoAbsolute project root - omit to use the working directory
filesYesRoot-relative or absolute paths, optionally with lines: "src/a.ts:10-40"
limitNoMax pairs per file
excludeNoRoot-relative or absolute folders or files to leave out, e.g. ["src/generated"]
thresholdNoMin similarity - raise to 0.92 for fewer false pairs, lower for more recall

Output Schema

ParametersJSON Schema
NameRequiredDescription
notesYesSync notes, why a file or exclude entry has no pairs, and how to read the pairs
pairsYesCandidates, best first per file
statusYes"off": duplicate search is not enabled for this project, which is not the same as no duplicates

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and openWorldHint=false, so the safety profile is covered structurally. The description adds real operational context beyond that: first build latency ('can take over an hour on large projects'), the per-project off-by-default behavior, coverage of .ts/.tsx/.js/.jsx/.mjs and skipping of tests/specs/.d.ts, and the 'treat every pair as a candidate' caveat. It does not disclose the result/pair shape, but with an output schema present that is acceptable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core behavior in the first sentence and then layers usage, prerequisites, and caveats compactly. The dash-separated clauses are dense but each carries actionable information; minor tightening is possible but nothing is padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given five parameters, full schema coverage, an output schema, and annotations, the description supplies the missing operational layer: enabling prerequisite, cost warning, file-type scope, and false-positive caveat. It omits only return-shape details, which the output schema already covers.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description goes further by giving a concrete line-range syntax example ('src/a.ts:10-40') and explaining the intent behind line-range inputs (sharper pairs), which the schema does not convey, plus threshold-like tuning guidance ('raise to fewer false pairs').

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List code in other files that nearly matches the given files or line ranges'), and narrows scope to a near-match similarity search rather than general search. It also names the sibling index_project as the enabling mechanism, so an agent can distinguish it from semantic_search.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells the agent when to use it ('in reviews and refactors'), how to sharpen results (pass changed line ranges), when to ask the user before enabling ('ask the user before you turn it on'), and names the prerequisite (index_project duplicates: true). Both when-to-use and gating conditions are covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

index_projectIndex a projectA
DestructiveIdempotent

Build or refresh the semantic index of a project root. Embeds only new, edited or outdated files and drops deleted ones. Pass duplicates: true once to enable find_duplicates for this project.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesAbsolute path to project root
forceNoDiscard the stored index and rebuild it from scratch
duplicatesNotrue builds the duplicate index, false deletes it - omit to keep it as is
extensionsNoFile extensions to index - omit to keep the project's list

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false, so the safety profile is covered structurally. The description adds real behavioral value beyond that: it explains the incremental strategy (only new/edited/outdated files are embedded, deleted ones dropped) and that duplicates indexing is opt-in and persists across calls. It stops short of noting cost/duration or concurrency concerns for a long-running index build.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences: purpose first, then the incremental/destructive behavior, then the one opt-in flag that has side effects on another tool. No filler, and the most decision-relevant fact is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter mutation tool with no output schema and annotations that already carry the safety hints, the description covers purpose, mutation semantics, and the one non-obvious flag. It is complete enough to invoke correctly; only return/status reporting and expected duration are unaddressed, which is minor given index_status exists.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so path, force, duplicates, and extensions are all self-documenting in the schema. The description reinforces the duplicates parameter's cross-tool effect on find_duplicates, which is mild added meaning, but does not clarify force vs. default or extension-list persistence beyond what the schema says. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Build or refresh the semantic index of a project root') and immediately scopes what it touches. It is distinguishable from siblings like index_status (which presumably reports) and semantic_search (which queries) without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives one concrete usage instruction — pass duplicates: true once to enable find_duplicates — which ties this tool to a sibling. However, it never says when to call this versus index_status (e.g., check status before reindexing) or when a full rebuild (force) is warranted over the default incremental behavior.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

index_statusIndex statusA
Read-onlyIdempotent

List indexed projects with file, chunk and size counts - pass path for one project.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoOptional project path to check

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint, idempotentHint, openWorldHint=false), so the description is not obligated to explain that. It does add useful content beyond the annotations by disclosing what the call returns — file, chunk, and size counts per project — which matters given there is no output schema. It stops short of describing freshness, pagination, or behavior for large project sets.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, front-loaded with the listing behavior and followed immediately by the parameter rule. Nothing is wasted and no sentence is padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter read-only tool with no output schema, the description covers the essential elements: what is listed, what is returned, and how the optional path narrows scope. Minor gaps remain around result size/pagination and whether counts reflect a live or stale index.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single parameter is already documented as 'Optional project path to check'. The description's 'pass path for one project' reinforces the scoping semantics (omitting it broadens the result set) but adds no format or syntax detail beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('List indexed projects') and even names the returned fields ('file, chunk and size counts'), which is more than a restatement of the title. It does not explicitly distinguish itself from siblings like index_project or find_duplicates, but the read-only listing intent is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'pass path for one project' implies the two modes (omit path to list all, supply path for a single project), so usage is inferable rather than spelled out. There is no explicit when-to-use guidance relative to index_project or the other siblings, so it sits at the implied-usage tier.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.4
    • First observedfind_duplicates
    • First observedindex_project
    • First observedindex_status
    • First observedsemantic_search

TDQS

A4.1/5.0

Scored across 4 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: semantic_search retrieves code, index_project builds/refreshes the index, index_status reports state, and find_duplicates detects near-matches. There is no overlap between searching, indexing, status, and duplicate detection.

Naming Consistency4/5

All names use snake_case consistently, but the internal pattern varies slightly (adjective_noun, noun_verb, noun_noun, verb_noun). Still highly readable and predictable as a set.

Tool Count4/5

Four tools is lean but well-matched to a focused semantic-search-and-indexing server. Each tool earns its place without redundancy, though the surface is on the thin side.

Completeness4/5

The core lifecycle (index build, status, search, duplicate detection) is covered. Minor gaps exist, such as removing or clearing an index and per-file reindexing, but agents can work around these via index_project.

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    F
    maintenance
    Enables semantic code search across codebases with automatic incremental indexing. Searches return relevant code snippets with file paths and line numbers based on natural language queries.
    1
    802
    Apache 2.0
  • A
    license
    A
    quality
    F
    maintenance
    Provides intelligent semantic code search using local AI embeddings, enabling natural language queries to find relevant code by meaning rather than exact keywords. Indexes codebases in the background with smart project detection and privacy-first local processing.
    6
    31 npm
    200
    MIT
  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables local semantic code search across repositories using natural language, with AST-aware chunking and hybrid vector/FTS5 retrieval.
    -