Reusable Privacy MCP Server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Reusable Privacy MCP Serversearch my local notes for password reset steps"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Reusable Privacy MCP Server
English | Português (Brasil)
A zero-data-leakage, privacy-first Model Context Protocol (MCP) server designed for any local document workspace.
Features
🔒 Zero External Data Exposure: Inverted keyword indexing and full-text searching run 100% offline.
🛡️ Automated PII & Secret Redaction: Automatically strips CPFs, CNPJs, Emails, Phone numbers, Passwords, and Secrets before sending snippets to LLMs.
🚫 Path Access Control: Blacklists sensitive subfolders (e.g.
Financeiro,Recursos Humanos,.git).⚡ Just-In-Time (JIT) Snippets: Line-bounded snippet retrieval prevents context stuffing.
⚙️ Config-Driven: Customize security rules per workspace using a local
.mcp-privacy.ymlfile.
Related MCP server: MemoryMesh
Getting Started
Install dependencies:
npm installBuild the Node.js project:
npm run build(Optional) Build as a standalone
.exe(No Node.js required on target machine):npm run build:exe # or: bun run build:exeThis compiles
privacy-mcp.exe. When launched withoutMCP_PROJECT_ROOT, it automatically targets the folder it is placed in or executed from (process.cwd()).
How to Register in Client Hosts (Claude Desktop / Cursor / Antigravity)
Option A: Using Node.js
Add this entry to your claude_desktop_config.json or MCP configuration:
{
"mcpServers": {
"workspace-docs": {
"command": "node",
"args": ["/path/to/privacy-mcp-server/dist/index.js"],
"env": {
"MCP_PROJECT_ROOT": "/path/to/your/documents"
}
}
}
}Option B: Using Standalone Executable (privacy-mcp.exe)
{
"mcpServers": {
"workspace-docs": {
"command": "/path/to/privacy-mcp-server/privacy-mcp.exe",
"env": {
"MCP_PROJECT_ROOT": "/path/to/your/documents"
}
}
}
}Tip: If MCP_PROJECT_ROOT is omitted in Option B, privacy-mcp.exe automatically protects and indexes whichever directory it is launched from.
Multi-MCP & Multi-Workspace Architecture
To scale this tool across multiple folders or combine it with other MCP tools:
MCP Gateway / Meta-MCP Router: Run a proxy server that aggregates
privacy-mcpwith other MCP servers (Git, SQLite, Web Search) into a single client entry point with tool namespacing (privacy_search,git_diff).Multi-Workspace Support: Configured via
mcp-workspaces.ymlto index multiple project roots under a single server instance usinglist_workspacesandsearch_all_workspaces.Modular Plugin Architecture: Embed multiple tools (privacy engine, git inspector, tabular data parser) directly into the single
.exebinary, enabled per workspace.Auto-Discovery Launcher: A lightweight system tray utility (
mcp-tray.exe) that watches folders for.mcp-privacy.ymland auto-registers them into Claude/Cursor configs.
For detailed design blueprints on these patterns, see docs/MULTI_MCP_ARCHITECTURE.md.
Workspace Configuration (.mcp-privacy.yml)
Place a .mcp-privacy.yml file in any workspace root directory:
name: "My Project Knowledge Base"
security:
restricted_directories:
- "Financeiro"
- "Recursos Humanos"
- "Certificado Digital"
- ".git"
- "node_modules"
allowed_extensions:
- ".md"
- ".json"
- ".txt"
- ".csv"
max_snippet_lines: 50
redaction:
enable_default_pii: trueAvailable Tools
3 toolslist_structureC
Safely lists allowed subdirectories and document files in the workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| subpath | No | Subpath relative to workspace root (e.g. '' or 'Comercial') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. 'Safely lists allowed ...' hints at a read-only, permission-filtered operation, but it never explains what 'allowed' means (auth/access scope), whether the listing is recursive or one-level, whether it paginates, or what the response contains.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single, front-loaded sentence with no filler. The only mildly wasteful word is the vague 'Safely', which asserts a property without explaining it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter read tool with no output schema and no annotations, the description is minimally adequate but leaves meaningful gaps: the meaning of 'allowed', the depth/scope of the listing, and the returned structure are all unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single optional parameter with 100% schema description coverage ('Subpath relative to workspace root (e.g. '' or 'Comercial')'), so the schema already fully documents it. The description adds no additional meaning about subpath syntax or default behavior, which is the expected baseline when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description pairs a specific verb ('lists') with a specific resource ('allowed subdirectories and document files in the workspace'), so an agent knows exactly what it returns. It does not name or distinguish itself from siblings search_docs and read_snippet, but the listing-versus-searching distinction is reasonably inferable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use this tool versus search_docs or read_snippet, nor any prerequisite or exclusion statements. The intended usage (browse the workspace hierarchy) is only implied by the description wording.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_snippetB
Retrieves a line-bounded, sanitized snippet from an allowed file.
| Name | Required | Description | Default |
|---|---|---|---|
| lineCount | No | Number of lines to read (capped at max_snippet_lines) | |
| startLine | No | Start line number (1-indexed) | |
| relativePath | Yes | Relative file path |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose two meaningful traits beyond the schema: output is 'sanitized' and reads are restricted to 'allowed' files. However, it says nothing about what happens for a disallowed path, error behavior, or what gets stripped during sanitization.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; every phrase (line-bounded, sanitized, allowed file) carries information. It is arguably under-sized for the behavioral quirks it hints at, keeping it just short of a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple 3-parameter read tool with no output schema and no annotations, the description covers the basics but leaves key questions open: what 'allowed' means, whether sanitization alters content, and how errors surface. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters (relativePath, startLine, lineCount) are already documented. The description's 'line-bounded' wording reinforces the line-range parameters but adds no format or default details beyond the schema; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Retrieves ... snippet') plus three qualifiers: line-bounded, sanitized, and from an allowed file. An agent can tell it reads file content, but nothing in the text distinguishes it from siblings list_structure or search_docs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use or when-not-to-use guidance, and no mention of alternatives such as search_docs. The only implicit cue is that the tool returns file content, leaving routing decisions to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_docsB
Searches documents locally using full-text keywords. Returns sanitized matching lines.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Keyword query to search for | |
| maxResults | No | Max results to return (default 10) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does disclose two useful traits: the search is local (no network/remote dependency) and results are "sanitized matching lines." It says nothing about permissions, indexing prerequisites, or result limits beyond the schema default.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, and the core action is front-loaded. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description's hint about the return shape (sanitized matching lines) is valuable. However, for a tool with no annotations and no output schema, it omits result format details, pagination/limit behavior, and failure modes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters (query, maxResults) are already documented. The description adds no format or syntax detail beyond what the schema provides, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource ("Searches documents") plus a mechanism ("full-text keywords"), so an agent knows exactly what the tool does. It does not, however, explicitly distinguish itself from siblings like read_snippet or list_structure.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus the alternatives (list_structure, read_snippet) or when it is inappropriate. "Locally" hints at scope but leaves routing entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v1.0.0- First observed
list_structure - First observed
read_snippet - First observed
search_docs
TDQS
Scored across 3 tools
Each tool targets a distinct operation: listing structure, searching content, and reading a bounded snippet. There is no overlap in purpose, making selection unambiguous.
All tools follow a consistent snake_case verb_noun pattern (list_structure, search_docs, read_snippet). The naming is predictable and easy to parse.
Three tools is within the typical well-scoped range (3–15) and each earns its place for a focused read-only document access server. No tool feels redundant or missing.
The surface covers listing, searching, and reading snippets, which is a complete read-only lifecycle for a privacy-focused document server. Minor gaps exist (e.g., no metadata or pagination tool), but agents can work around them.
Maintenance
Related MCP Connectors
Securely search and manage workspace context files for AI agents and teams.
Multi-engine search for AI agents. Trust scoring, local corpus, MCP-native. Self-hostable, BYOK.
Search and reason over your Obsidian-style Markdown vault, right from ChatGPT.
Private-by-default, local-first memory/context/task orchestrator for MCP apps and agents.
Related MCP Servers
- AlicenseAqualityAmaintenancePrivacy-first local document search using semantic search. Runs entirely on your machine with no cloud services, supporting PDF, DOCX, TXT, and Markdown files.96,063 npm407MIT
- AlicenseAqualityDmaintenanceA universal, local-first MCP hub that indexes personal files (documents, code, etc.) and provides private semantic search via hybrid dense+BM25 retrieval, enabling agents like Claude Desktop to query your data without sending it to the cloud.176MIT
- AlicenseNot gradedqualityBmaintenanceEnables private, offline semantic search across local files (documents, images, videos) using OCR and vector search, and optionally performs web research with cited sources.AGPL 3.0
- AlicenseAqualityCmaintenanceEnables local-first hybrid knowledge retrieval from authorized Markdown and plain-text files, combining full-text and vector search with reranking and traceable source references via a single search tool.1MIT