nl-file-search
This server lets you search indexed local files by natural language and retrieve stored metadata for indexed items.
search_filesruns a natural-language semantic search over the local index (notes, images, videos, PDFs, etc.).Search can be narrowed by optional
modality(text,image,video, orpdf) and result countk(default 8).get_filereturns metadata for a specific indexed path already in the index, such as text or media page/time range info.It does not read arbitrary disk paths; only paths previously ingested are accessible.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@nl-file-searchsearch my notes for anything about the vintage red truck in the rain"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
nl-file-search
Natural-language search over local files. Phase 1 indexes markdown, plain text, images, videos, and PDFs with Gemini Embedding 2, stores vectors in SQLite (sqlite-vec), and exposes search through a CLI and a Cursor MCP server.
Work in progress. Designed by Gary Lucero. Coded by Cursor.
A later Python app can import the same nl_file_search.search module. Office documents are Phase 2; music is Phase 3. See background/ROADMAP.md.
Phase 1 file types
Type | Extensions | How it is indexed |
Text |
| Heading/paragraph chunks |
Images |
| Sent to Gemini as image bytes |
Video |
| Split into 120s clips with ffmpeg |
| One page per embedding |
Unknown extensions are skipped. HEIC is skipped. Secret-like names (.env, *.pem, credentials.json, SSH keys) are never read or sent to Gemini.
Related MCP server: Mimir
Requirements
Windows, macOS, or Linux
Python 3.12+
A Gemini API key (
GEMINI_API_KEY)ffmpeg on
PATHwhen you index videos (Windows:winget install Gyan.FFmpeg; macOS:brew install ffmpeg; Linux: your package manager)Network access for every ingest and every search (queries are embedded with the same model)
SQLite is not Windows-only. The same code uses Path.home() / "nl-file-search" on every OS.
Setup
Windows (PowerShell):
cd C:\source\repos\nl-file-search
python -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install -U pip
python -m pip install -e .macOS / Linux:
cd ~/src/nl-file-search
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -U pip
python -m pip install -e .Copy the example config and env files into ~/nl-file-search (your home directory on every OS). Run this from the repo, with the venv active. nl-search does not create these files.
python scripts/copy_profile.pyThen:
Edit
~/nl-file-search/config.yamland set the folders to index.Edit
~/nl-file-search/.envand setGEMINI_API_KEY=...(never commit this file).
If a key was ever pasted into a chat or ticket, revoke it in Google AI Studio and issue a new one.
Data lives outside the repo:
Path | |
Data folder |
|
SQLite index |
|
~/nl-file-search/
config.yaml # copied from config.example.yaml; you edit this
.env # copied from .env.example; you put the API key here
index.sqlite # created on the first successful nl-search ingest
logs/ # created when a command runs after config existsPaths in config.yaml
List each folder under sources with a path value. Use forward slashes in double quotes; that form works on Windows, macOS, and Linux:
sources:
- path: "~/Notes"
- path: "~/Pictures"
- path: "/absolute/path/to/folder"~ is expanded to your home directory. Paths are resolved to absolute locations when ingest runs.
On Windows, a backslash in a double-quoted string is a YAML escape (\U in C:\Users is not a folder separator). Use one of these instead:
Form | Example |
Forward slashes |
|
Single quotes |
|
Escaped backslashes |
|
Example config.yaml (also in config.example.yaml):
sources:
- path: "~/Notes"
- path: "~/Pictures"
exclude:
- "**/.git/**"
- "**/node_modules/**"
embed:
model: gemini-embedding-2
dimensions: 768
video:
max_seconds: 120CLI
nl-search ingest
nl-search ingest --path /path/to/folder
nl-search search "vintage red truck in the rain"
nl-search statusCursor MCP
Add a server in Cursor’s MCP settings (user or project). Point command at this repo’s venv Python:
{
"mcpServers": {
"nl-file-search": {
"command": "/absolute/path/to/nl-file-search/.venv/bin/python",
"args": ["-m", "nl_file_search.mcp_server"]
}
}
}On Windows, use .venv\\Scripts\\python.exe instead of .venv/bin/python. Restart Cursor MCP after saving. Do not put GEMINI_API_KEY in that config; the server reads ~/nl-file-search/.env.
Tools:
search_files— natural-language search; use this first when asking about indexed local filesget_file— indexed text or media metadata for a path already in the index (not an arbitrary disk read)
Security
The API key lives only in
~/nl-file-search/.env.The SQLite database lives only in
~/nl-file-search/index.sqlite.Ingest skips credential-like files and default junk directories (
.git,node_modules,.venv,__pycache__).get_fileonly returns rows already in the index. It will not open..\..\.envor other paths that were never ingested.Retrieved snippets go to Cursor the same way an open file would. Do not index folders that must never leave the machine.
Phase 2 and 3 (not built yet)
Phase 2: modern Office (
.docx,.xlsx,.pptx)Phase 3: music (
.mp3,.m4a,.flac)
Out of scope: Google Docs, Drive export, LibreOffice, legacy .doc / .xls / .ppt.
Available Tools
2 toolsget_fileGet FileA
Return indexed text or media metadata for a path already in the index.
Does not read arbitrary disk paths. Pass a path from search_files. Media results include path and page/time range, not binary bytes.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose two real traits: the index-only constraint (no arbitrary disk access) and the shape of media results ('path and page/time range, not binary bytes'). It omits error behavior for unindexed paths and any permission or rate-limit context, so it is good but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short, front-loaded sentences with zero filler; the primary purpose leads and constraints follow. Every sentence adds a distinct, actionable fact.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read tool with an output schema present, the description covers what the tool does, what input is valid, and the nature of the return (text vs. media metadata). Return-value detail is legitimately delegated to the output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the single 'path' parameter, and it does: the path must be one already in the index and obtained from search_files, not an arbitrary filesystem path. It still says nothing about accepted path syntax or format (absolute vs. relative), leaving a small gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Return indexed text or media metadata') plus a precise scope constraint ('for a path already in the index'). It implicitly distinguishes itself from the sibling left_files by specifying that the path must originate from search_files rather than an arbitrary disk location.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states the negative case ('Does not read arbitrary disk paths') and names the source of valid input ('Pass a path from search_files'), which routes the agent to the sibling for discovery and back here for retrieval. Nothing about when this tool applies is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_filesSearch FilesA
Search the local file index with natural language.
Use this first when the user asks to find notes, images, videos, PDFs, or other files that may have been ingested by nl-search. modality may be text, image, video, or pdf.
| Name | Required | Description | Default |
|---|---|---|---|
| k | No | ||
| query | Yes | ||
| modality | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully notes the index contains files 'may have been ingested by nl-search,' clarifying the searchable scope. It says nothing about ranking, the effect of k, or that results are scored/limited, leaving meaningful behavior undisclosed for a tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the core purpose followed by the trigger and the modality enumeration. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values needn't be explained. The description covers purpose, trigger, and modality values adequately. The main remaining gap is k's meaning, which is relevant for a tool that returns a bounded result set.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description should compensate. It does explain modality by listing valid values (text, image, video, pdf), which helps, but k and query are left entirely to the schema — and k (a non-obvious default-8 result count) is undocumented anywhere.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — 'Search the local file index with natural language' — which is clear and actionable. It also enumerates the content types covered. It doesn't explicitly contrast with the sibling get_file by name, so an agent must infer the search-vs-retrieve distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use this first when the user asks to find notes, images, videos, PDFs' gives a concrete trigger condition and the word 'first' implies a follow-up tool exists. However, it never names get_file or states when retrieval (rather than search) is the right choice, so the alternative is only suggested, not specified.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.1.0- First observed
get_file - First observed
search_files
TDQS
Scored across 2 tools
search_files and get_file have clearly distinct roles: search finds indexed files by natural language, while get_file retrieves content or metadata for a path returned by search. The descriptions reinforce this boundary by explicitly telling users to pass paths from search_files and not to use get_file for arbitrary disk paths.
Both tools follow a consistent verb_noun snake_case pattern: search_files and get_file. The singular/plural difference mirrors the operation, and the naming is predictable and readable.
Two tools are slightly below the typical 3-15 range, but they form a focused search-then-retrieve workflow. The scope is narrow enough that each tool earns its place without bloat.
The server covers the core lifecycle of finding indexed files and retrieving their indexed content or media metadata. Minor gaps remain around index management or broader listing, but these may be intentionally handled by the companion nl-search ingestion system.
Related MCP Connectors
Search and reason over your Obsidian-style Markdown vault, right from ChatGPT.
Securely search and manage workspace context files for AI agents and teams.
Ingest, manage, and retrieve documents for RAG-powered AI applications
Personal context for every AI: search, read, and write back to your private Markdown library.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceProvides AI assistants with semantic search and read access to local files and directories, enabling knowledge retrieval from indexed content.38 npm17MIT
- AlicenseNot gradedqualityAmaintenanceLocal-first memory and retrieval for private project knowledge. Enables indexing files, searching, and asking questions about project documents using local embeddings and LLM.6AGPL 3.0
- AlicenseNot gradedqualityDmaintenanceProvides local semantic search over files using embeddings, enabling directory indexing and natural language queries without external services.MIT
- AlicenseNot gradedqualityCmaintenanceEnables fully local, cross-lingual retrieval over documents and source code by indexing files and providing search and ingest tools, with all data staying on the machine.GPL 2.0