dfine-semantic
OfficialClick on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@dfine-semanticwhere do we handle API errors?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
@dfine-io-gmbh/semantic-mcp
Semantic code search as an MCP server. It embeds a project's code into a
local SQLite vector store and answers natural-language questions that a plain grep cannot match.
Everything runs locally once the model is downloaded.
How it works
Code is split into chunks and embedded with
jinaai/jina-embeddings-v2-base-code(768 dimensions) through@huggingface/transformers.Vectors live in one local SQLite database per project (
sqlite-vec).The server speaks MCP over stdio, revision 2026-07-28 and the 2025 revisions, and works with any MCP client, for example Claude Code or Cursor.
Each search first re-indexes the files that changed since the last one. If that would take longer than about a minute, clients on MCP 2026-07-28 that support forms ask you first. Other clients re-index without asking.
Long runs report progress to clients that ask for it. Claude Code and other clients that reset their timeout on progress keep such a call open until it finishes.
Related MCP server: Smart Coding MCP
Requirements
Node.js 22 or newer
Git on the
PATH. Every indexed root must be inside a git work tree, because files are listed withgit ls-files.Network access to download the model (about 640 MB) into
~/.dfine-semantic/modelswhen it is not there yetbetter-sqlite3andsqlite-vecship prebuilt binaries for macOS (arm64, x64), Linux (x64, arm64) and Windows (x64). Other targets build from source and need a C/C++ toolchain.
Use it with an MCP client
No install step is needed: npx fetches and runs the server. Add it to your MCP client config
(.mcp.json for Claude Code project scope, or ~/.claude.json for user scope):
{
"mcpServers": {
"dfine-semantic": {
"type": "stdio",
"command": "npx",
"args": ["-y", "@dfine-io-gmbh/semantic-mcp"]
}
}
}The server sends short usage instructions to the client when it connects.
Tools
Name | Purpose |
| Natural-language query, returns ranked |
| Index or refresh a project root, embedding only what changed |
| List indexed projects with file, chunk and size counts |
| List near-identical code for files or line ranges (opt-in, see below) |
semantic_search and find_duplicates also return their results as structured data, described by
each tool's output schema. Its status field tells an agent that duplicate search is off for a
project, which is not the same as finding no duplicates.
semantic_search covers .ts and .tsx by default. Pass include (for example [".md", ".vue"])
to search other indexed file types. index_project indexes the allow-listed extensions you pass
and keeps that list, so a later call without extensions indexes the same file types.
Duplicate search (opt-in)
find_duplicates is off until you turn it on for a project. It needs a second index of short code
windows, and building that index takes time and disk space, so projects that never use it never pay
for it.
Ask your agent to run
index_projectwithduplicates: trueonce for the project. On an indexed repository with about 1,400 TypeScript files, this takes about 45 minutes and 70 MB. A project without a search index yet needs about 35 minutes more.Ask for duplicates of files or line ranges, for example
src/a.ts:10-40, and your agent callsfind_duplicates. Searches and index runs keep the windows of edited files current. When a result says windows are missing, runindex_projectagain.Run
index_projectwithduplicates: falseto switch it off and delete the duplicate index.
Each result pairs a range of your file with a similar range in another file, plus a similarity
score. Read both ranges before you merge anything. At the default threshold of 0.88, the tool found
about 40% of the real duplicates in a measured TypeScript project; a higher threshold returns fewer
false pairs. It covers .ts, .tsx, .js, .jsx and .mjs files and skips tests, specs and
.d.ts files. Pass exclude with folders such as src/generated to leave generated or vendored
code out.
Configuration
Variable | Default | Purpose |
|
| Extra absolute roots the server may index (comma-separated) |
|
| Where the indexes and the model are stored |
The working directory counts as a root unless it is / or your home folder. Indexes are keyed by
project path and survive upgrades. Run index_project with force: true for a clean rebuild.
Upgrading from 0.1.3 or older
Restart every session that still runs 0.1.3 or older. An old server cannot read an upgraded index.
The model downloads once more into its new folder. Copies inside old
npxcaches can be deleted.Files are now split into chunks differently. Searches keep using your existing index, and clients that support forms offer a rebuild. You can also run
index_project.find_duplicatesis new and stays off until your agent runsindex_projectwithduplicates: true.
Security
See SECURITY.md for the security model and how to report a vulnerability.
Local development
pnpm install
pnpm build # tsc
pnpm lint:dlint # dlint, the dfine linter on the TypeScript compiler
pnpm check # tsc, dlint and prettier
pnpm startLicense and support
MIT, see LICENSE. Questions or issues: support@dfine.io or the issue tracker.
Available Tools
4 toolsfind_duplicatesFind duplicate codeARead-onlyIdempotent
List code in other files that nearly matches the given files or line ranges. Use it in reviews and refactors - pass changed line ranges ("src/a.ts:10-40") for sharper pairs. Treat every pair as a candidate - read both ranges before you call it a duplicate. Expect it to be off per project - ask the user before you turn it on. Turn it on with index_project duplicates: true - the first build can take over an hour on large projects. Covers .ts, .tsx, .js, .jsx, .mjs files; skips tests, specs and .d.ts files.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Absolute project root - omit to use the working directory | |
| files | Yes | Root-relative or absolute paths, optionally with lines: "src/a.ts:10-40" | |
| limit | No | Max pairs per file | |
| exclude | No | Root-relative or absolute folders or files to leave out, e.g. ["src/generated"] | |
| threshold | No | Min similarity - raise to 0.92 for fewer false pairs, lower for more recall |
Output Schema
| Name | Required | Description |
|---|---|---|
| notes | Yes | Sync notes, why a file or exclude entry has no pairs, and how to read the pairs |
| pairs | Yes | Candidates, best first per file |
| status | Yes | "off": duplicate search is not enabled for this project, which is not the same as no duplicates |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and openWorldHint=false, so the safety profile is covered structurally. The description adds real operational context beyond that: first build latency ('can take over an hour on large projects'), the per-project off-by-default behavior, coverage of .ts/.tsx/.js/.jsx/.mjs and skipping of tests/specs/.d.ts, and the 'treat every pair as a candidate' caveat. It does not disclose the result/pair shape, but with an output schema present that is acceptable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core behavior in the first sentence and then layers usage, prerequisites, and caveats compactly. The dash-separated clauses are dense but each carries actionable information; minor tightening is possible but nothing is padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given five parameters, full schema coverage, an output schema, and annotations, the description supplies the missing operational layer: enabling prerequisite, cost warning, file-type scope, and false-positive caveat. It omits only return-shape details, which the output schema already covers.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description goes further by giving a concrete line-range syntax example ('src/a.ts:10-40') and explaining the intent behind line-range inputs (sharper pairs), which the schema does not convey, plus threshold-like tuning guidance ('raise to fewer false pairs').
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List code in other files that nearly matches the given files or line ranges'), and narrows scope to a near-match similarity search rather than general search. It also names the sibling index_project as the enabling mechanism, so an agent can distinguish it from semantic_search.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent when to use it ('in reviews and refactors'), how to sharpen results (pass changed line ranges), when to ask the user before enabling ('ask the user before you turn it on'), and names the prerequisite (index_project duplicates: true). Both when-to-use and gating conditions are covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
index_projectIndex a projectADestructiveIdempotent
Build or refresh the semantic index of a project root. Embeds only new, edited or outdated files and drops deleted ones. Pass duplicates: true once to enable find_duplicates for this project.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Absolute path to project root | |
| force | No | Discard the stored index and rebuild it from scratch | |
| duplicates | No | true builds the duplicate index, false deletes it - omit to keep it as is | |
| extensions | No | File extensions to index - omit to keep the project's list |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false, so the safety profile is covered structurally. The description adds real behavioral value beyond that: it explains the incremental strategy (only new/edited/outdated files are embedded, deleted ones dropped) and that duplicates indexing is opt-in and persists across calls. It stops short of noting cost/duration or concurrency concerns for a long-running index build.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences: purpose first, then the incremental/destructive behavior, then the one opt-in flag that has side effects on another tool. No filler, and the most decision-relevant fact is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter mutation tool with no output schema and annotations that already carry the safety hints, the description covers purpose, mutation semantics, and the one non-obvious flag. It is complete enough to invoke correctly; only return/status reporting and expected duration are unaddressed, which is minor given index_status exists.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so path, force, duplicates, and extensions are all self-documenting in the schema. The description reinforces the duplicates parameter's cross-tool effect on find_duplicates, which is mild added meaning, but does not clarify force vs. default or extension-list persistence beyond what the schema says. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Build or refresh the semantic index of a project root') and immediately scopes what it touches. It is distinguishable from siblings like index_status (which presumably reports) and semantic_search (which queries) without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives one concrete usage instruction — pass duplicates: true once to enable find_duplicates — which ties this tool to a sibling. However, it never says when to call this versus index_status (e.g., check status before reindexing) or when a full rebuild (force) is warranted over the default incremental behavior.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
index_statusIndex statusARead-onlyIdempotent
List indexed projects with file, chunk and size counts - pass path for one project.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Optional project path to check |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint, idempotentHint, openWorldHint=false), so the description is not obligated to explain that. It does add useful content beyond the annotations by disclosing what the call returns — file, chunk, and size counts per project — which matters given there is no output schema. It stops short of describing freshness, pagination, or behavior for large project sets.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the listing behavior and followed immediately by the parameter rule. Nothing is wasted and no sentence is padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read-only tool with no output schema, the description covers the essential elements: what is listed, what is returned, and how the optional path narrows scope. Minor gaps remain around result size/pagination and whether counts reflect a live or stale index.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single parameter is already documented as 'Optional project path to check'. The description's 'pass path for one project' reinforces the scoping semantics (omitting it broadens the result set) but adds no format or syntax detail beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('List indexed projects') and even names the returned fields ('file, chunk and size counts'), which is more than a restatement of the title. It does not explicitly distinguish itself from siblings like index_project or find_duplicates, but the read-only listing intent is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'pass path for one project' implies the two modes (omit path to list all, supply path for a single project), so usage is inferable rather than spelled out. There is no explicit when-to-use guidance relative to index_project or the other siblings, so it sits at the implied-usage tier.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
semantic_searchSemantic code searchARead-onlyIdempotent
Find code by meaning in an indexed project. Ask one full natural-language sentence, not keywords. Good: "How does the app validate share token permissions?" Bad: "shareToken auth validate". Returns file:line references by default - read the files you need. Searches .ts and .tsx unless include adds other indexed extensions.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Absolute project root - omit to use the working directory | |
| limit | No | Max results | |
| query | Yes | One natural-language sentence | |
| include | No | Extra indexed extensions to search, e.g. [".md", ".css"] | |
| threshold | No | Min similarity - raise to 0.5 for precision, lower for recall | |
| returnFullContent | No | Return chunk code instead of file:line references - keep limit under 20 |
Output Schema
| Name | Required | Description |
|---|---|---|
| notes | Yes | Sync, index and limit notes |
| status | Yes | "not_indexed": run index_project for this path first |
| hasMore | Yes | More matches pass the threshold |
| results | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, and closed-world, so the safety profile is covered. The description adds real behavioral context beyond that: results default to file:line references, the caller should then read the files, and only .ts/.tsx are searched unless include adds extensions. Minor gap: nothing about result ordering or the effect of threshold on output volume.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Five short lines, front-loaded with the core purpose, then query style, then return-format and scope defaults. The good/bad examples earn their place by making the query convention unambiguous; there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be described, and the description still usefully notes the default reference format. It covers query style, default extensions, and the read-the-files follow-up. It is slightly thin on how limit and threshold interact with the returned set.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter is already documented, and the baseline is 3. The description largely restates schema content (natural-language sentence, include for extra extensions, file:line default) and adds only the query-phrasing examples; threshold and path receive no additional meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — 'Find code by meaning in an indexed project' — and scopes it to an indexed project, which cleanly separates it from the indexing siblings (index_project, index_status) and from find_duplicates. An agent can tell what it retrieves without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete invocation guidance: 'Ask one full natural-language sentence, not keywords,' with a good and a bad example, plus the default extension scope and how include extends it. It lacks an explicit when-not-to-use or a direct pointer to a sibling alternative, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.4- First observed
find_duplicates - First observed
index_project - First observed
index_status - First observed
semantic_search
TDQS
Scored across 4 tools
Each tool has a clearly distinct purpose: semantic_search retrieves code, index_project builds/refreshes the index, index_status reports state, and find_duplicates detects near-matches. There is no overlap between searching, indexing, status, and duplicate detection.
All names use snake_case consistently, but the internal pattern varies slightly (adjective_noun, noun_verb, noun_noun, verb_noun). Still highly readable and predictable as a set.
Four tools is lean but well-matched to a focused semantic-search-and-indexing server. Each tool earns its place without redundancy, though the surface is on the thin side.
The core lifecycle (index build, status, search, duplicate detection) is covered. Minor gaps exist, such as removing or clearing an index and per-file reindexing, but agents can work around these via index_project.
Related MCP Connectors
Search indexed code, trace dependencies, assess change impact, and recall repository memory.
Ask a codebase what calls what: search, blast radius, paths between symbols, and diffs.
Code intelligence for coding agents: semantic, AST, graph, and full-text search. 279+ languages.
Project memory, semantic code search, and grounded agent context.
Related MCP Servers
- AlicenseAqualityFmaintenanceEnables semantic code search across codebases with automatic incremental indexing. Searches return relevant code snippets with file paths and line numbers based on natural language queries.1802Apache 2.0
- AlicenseAqualityFmaintenanceProvides intelligent semantic code search using local AI embeddings, enabling natural language queries to find relevant code by meaning rather than exact keywords. Indexes codebases in the background with smart project detection and privacy-first local processing.631 npm200MIT
- AlicenseAqualityDmaintenanceProvides semantic code search over codebases using local embeddings with natural language queries. Supports hybrid search, file watching, and respects .gitignore.1142 PyPI5MIT
- FlicenseNot gradedqualityDmaintenanceEnables local semantic code search across repositories using natural language, with AST-aware chunking and hybrid vector/FTS5 retrieval.-