privacyscrubber-mcp
Based on the provided schema, this MCP server locally redacts sensitive data before LLM use and restores masked values afterward.
Sanitize raw text, code, or logs with
sanitize_text, replacing PII/secrets/credentials with placeholders like[EMAIL_1]and[API_KEY_1].Choose a detection profile:
General(free) or PRO profiles such asDev,Medical,Legal,Finance,Security,HR, etc.Sanitize a local file by absolute path with
sanitize_file, returning safe redacted content for AI analysis.Restore original private values in LLM responses with
reveal_textusing the local volatile RAM-only session map.Keep original sensitive data in memory only, avoiding sending real PII/secrets to remote LLMs.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@privacyscrubber-mcpscrub this text: 'Name: Alice, SSN: 123-45-6789'"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
@privacyscrubber/mcp-server
CISO-Approved Zero-Trust PII & Secrets Redaction MCP Server for Cursor, Windsurf, and Claude Desktop. Locally scrubs PII, secrets, credentials, and custom regex rules from files and text contexts before they reach remote LLM providers to prevent API leaks and ensure HIPAA/SOC 2 compliance at the developer endpoint.
โญ Support Zero-Trust Open Source: If PrivacyScrubber protects your API keys and code from leaks, please Star this repository or run
gh repo star moxno/privacyscrubber-mcpin your terminal!
๐ Zero-Trust Data Flow
All sensitive parameters, identifiers, and variables are intercepted locally inside your machine's RAM. They are replaced by tokens (e.g. [EMAIL_1]) before being sent to the AI. Once the AI responds, the tokens are safely swapped back to original values in your local context.
[Raw Input / Files] โโ> [MCP sanitize_text] โโ> [Masked Tokens] โโ> [LLM API]
โ โ
(In-Memory Map) (Result)
โ โ
[Original Output] <โโโ [MCP reveal_text] <โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโRelated MCP server: trustboost-pii-sanitizer
๐ Installation
1. Install via Smithery
To automatically configure and run with your preferred client, install using Smithery:
npx -y @smithery/cli install @privacyscrubber/mcp-server --write-to-clients2. Instant Run with NPX
Run the server directly without local installation:
npx -y @privacyscrubber/mcp-server3. Programmatic Node.js / TypeScript SDK (Lightweight Presidio Alternative)
Need direct, in-memory zero-trust PII sanitization in your backend microservice, Next.js app, or RAG vector pipeline rather than an MCP server? Use our official zero-dependency SDK:
npm install @privacyscrubber/sdkimport OpenAI from 'openai';
import { wrapOpenAI } from '@privacyscrubber/sdk';
// Transparently masks PII before sending to LLM and rehydrates responses:
const openai = wrapOpenAI(new OpenAI({ apiKey: process.env.OPENAI_API_KEY }));
// Outbound prompt is sanitized in local RAM before leaving your machine:
// "Schedule a call with [NAME_1] at [EMAIL_1] regarding API key [AWS_KEY_1]."
const completion = await openai.chat.completions.create({
model: 'gpt-4o',
messages: [{ role: 'user', content: 'Schedule a call with Alice Smith at alice@acme.com with AKIAIOSFODNN7EXAMPLE.' }]
});
// Incoming LLM answer is automatically rehydrated with "Alice Smith (alice@acme.com)":
console.log(completion.choices[0].message.content);โก Why @privacyscrubber/sdk vs Microsoft Presidio?
Microsoft Presidio is the Python standard, but deploying it in a Node.js / TypeScript stack requires running heavy Python microservices, Docker containers, and 500MB+ spaCy NLP models with 35โ120ms latency. @privacyscrubber/sdk runs 100% in-process with zero dependencies:
Feature / Metric |
| Microsoft Presidio | AWS Comprehend / Google Cloud DLP |
Runtime & Dependencies | 0 dependencies (~150KB) | Python + Docker + spaCy (~500MB) | Heavy Cloud SDKs |
Execution Latency | Sub-millisecond in RAM (<1ms) | 35โ120ms HTTP/gRPC roundtrip | 180โ400ms external cloud roundtrip |
Docker / Sidecar Needed | None (Pure in-process Node/WASM) | Mandatory Docker container | VPC endpoints & IAM configurations |
Network Egress | 0 Bytes (100% Air-gapped) | Local internal network hop | Full unencrypted payload to cloud |
Data Loss & Reversibility | 0 Loss (RAM-reversible via | Manual token vault configuration | Irreversible masking / hashing |
OpenAI / LangChain 1-Liner | Built-in ( | Complex custom pipeline glue | Custom proxy architecture |
DevOps Secrets Interception | Built-in (AWS, JWT, DB URIs, GitHub PAT) | Custom regex recognizers needed | Cloud-specific classifiers |
Streaming Rehydration | Built-in ( | Buffering / chunk split failures | Not supported in real-time streams |
๐ View @privacyscrubber/sdk on NPM | Full documentation, Express middleware, LangChain transforms, and 25 compliance profiles.
๐ค AI Orchestrator Recipes: LangChain, LlamaIndex, CrewAI & AutoGPT
If you are building autonomous agents, RAG vector pipelines, or backend services rather than single-user IDE prompts, use @privacyscrubber/sdk to sanitize data in-memory:
1. LangChain.js & LCEL Pipelines
import { ChatOpenAI } from '@langchain/openai';
import { createLangChainTransform } from '@privacyscrubber/sdk';
const transform = createLangChainTransform({ defaultProfile: 'Dev' });
// 1. Sanitize in local RAM before LLM call:
const { scrubbedText, tokenMap } = transform.preprocess(
"Deploying database with user admin and pwd postgresql://user:SecretPass123@db.internal:5432/prod"
);
// 2. Outbound LLM receives only [DB_URI_1]
const model = new ChatOpenAI({ model: 'gpt-4o' });
const response = await model.invoke(scrubbedText);
// 3. Inbound response is automatically restored with original secrets in local RAM:
const finalOutput = transform.postprocess(response.content, tokenMap);2. LlamaIndex.TS & RAG Vector Ingestion
import { sanitize } from '@privacyscrubber/sdk';
import { Document, VectorStoreIndex } from 'llamaindex';
// Strip customer PII and credentials BEFORE vector embedding into Pinecone / Chroma:
const rawText = "Customer John Doe (SSN: 042-88-9124, email: john@example.com) requested refund.";
const { scrubbedText, tokenMap, telemetry } = sanitize(rawText, { profile: 'Finance' });
// Vector store indexes only syntax-preserving tokens [NAME_1], [SSN_1], [EMAIL_1]:
const document = new Document({ text: scrubbedText, metadata: { psRiskLevel: telemetry.riskLevel } });
const index = await VectorStoreIndex.fromDocuments([document]);3. CrewAI & Multi-Agent Swarms (Sidecar Mode / Python Interop)
# Run lightweight zero-dependency local daemon (<1ms in-memory, 0 external network calls):
npx @privacyscrubber/sdk sidecar --port 3100# In your Python CrewAI / AutoGen Agent tools:
import requests
def sanitize_agent_input(prompt: str) -> tuple[str, dict]:
res = requests.post("http://127.0.0.1:3100/sanitize", json={"text": prompt, "profile": "Dev"}).json()
return res["scrubbedText"], res["tokenMap"]
def restore_agent_output(ai_response: str, token_map: dict) -> str:
res = requests.post("http://127.0.0.1:3100/restore", json={"text": ai_response, "tokenMap": token_map}).json()
return res["restoredText"]4. AutoGPT / Multi-Turn Autonomous Agents
import { PrivacyScrubberEngine, createGuardedTools } from '@privacyscrubber/sdk';
const engine = new PrivacyScrubberEngine({ defaultProfile: 'Dev' });
// Equips agent with tools that intercept commands, file reads, and git diffs before model exposure:
const guardedTools = createGuardedTools(engine);๐ข Enterprise & Production Licensing:
Self-Serve Developer SDK ($299/mo flat or $2,990/yr): Unlimited internal backend nodes, microservices, and RAG pipelines. Get SDK License
PrivacyScrubber TEAMS ($99/mo flat): Unlimited team seats, centralized policy enforcement, encrypted session handoff. Deploy TEAMS
Enterprise Air-Gapped License: On-premise source code distribution, zero-network custom models. Contact Enterprise
๐ก๏ธ Architecture & Security Deep-Dive (Zero-Trust vs Cloud DLP)
When AI IDEs (Cursor, Claude Desktop, Windsurf) connect to model providers, developer credentials, database connection strings, and internal customer PII are at continuous risk of prompt exfiltration. The @privacyscrubber/mcp-server enforces four immutable architectural guarantees:
Stdio Air-Gapped Transport: The MCP server communicates exclusively over local standard input/output (
stdio) child processes spawned by your IDE. It opens zero external listening ports and initiates zero remote network requests.Volatile RAM-Only Token Map: The mapping table between synthetic tokens (
[AWS_KEY_1],[EMAIL_1]) and raw cleartext is maintained exclusively in ephemeral node memory and is wiped the moment your IDE session closes.Deterministic AST Lookarounds vs Cloud Proxy Overhead: Unlike cloud DLP gateways (Nightfall, Skyflow) that add 200โ400ms latency and transmit unencrypted code to third parties,
@privacyscrubber/mcp-serverruns locally in <2ms with zero data egress.Permanent Standards & Academic Validation:
IETF Specification: draft-sibiryakov-ztds-protocol-01 โ Revision 01 ยท 2026-09-23 ยท 13 pages
CERN / Zenodo Foundation: DOI 10.5281/zenodo.22058770
Center for Open Science (OSF): DOI 10.17605/OSF.IO/5BYJF
Patent Pending: Israel Patent Office Application
IL 331905(WIPO DAS Code:B17B)
โ๏ธ Client Integrations
Claude Desktop
Add this to your Claude Desktop config file:
macOS:
~/Library/Application Support/Claude/claude_desktop_config.jsonWindows:
%APPDATA%\Claude\claude_desktop_config.json
{
"mcpServers": {
"privacyscrubber": {
"command": "npx",
"args": ["-y", "@privacyscrubber/mcp-server"],
"env": {
"PRIVACYSCRUBBER_KEY": "YOUR_OPTIONAL_PRO_LICENSE_KEY"
}
}
}
}Cursor / Windsurf
Navigate to Settings -> Features -> MCP.
Add new MCP server:
Name:
privacyscrubberType:
commandCommand:
npx -y @privacyscrubber/mcp-server
Optional: Set
PRIVACYSCRUBBER_KEYas an environment variable in your system shell.
Cline / Roo Code
Add to cline_mcp_settings.json:
{
"mcpServers": {
"privacyscrubber": {
"command": "npx",
"args": ["-y", "@privacyscrubber/mcp-server"],
"env": {
"PRIVACYSCRUBBER_KEY": "YOUR_OPTIONAL_PRO_LICENSE_KEY"
}
}
}
}Claude Code CLI
Add directly from your terminal:
claude mcp add privacyscrubber -- npx -y @privacyscrubber/mcp-server๐ ๏ธ Provided Tools & JSON-RPC Specifications
1. sanitize_text
Redacts PII, secrets, API keys, and credentials from a text block and populates the volatile local replacement mapping.
Arguments:
text(string, required): The raw content or logs to sanitize.profile(string, optional): Gated industry detection profile (e.g., 'General', 'Dev', 'Medical', 'Legal', 'Compliance'). Defaults to 'General'.
JSON-RPC Call Example:
{ "method": "tools/call", "params": { "name": "sanitize_text", "arguments": { "text": "Contact me at dev-key-1234 or jane.doe@company.com", "profile": "General" } } }Response Example:
{ "content": [ { "type": "text", "text": "Contact me at [SECRET_1] or [EMAIL_1]" } ] }
2. reveal_text
Detokenizes the AI response back to the original values locally.
Arguments:
text(string, required): The response from the LLM containing tokenized placeholders.
JSON-RPC Call Example:
{ "method": "tools/call", "params": { "name": "reveal_text", "arguments": { "text": "Please reach out to [EMAIL_1] regarding the update." } } }Response Example:
{ "content": [ { "type": "text", "text": "Please reach out to jane.doe@company.com regarding the update." } ] }
3. sanitize_file
Reads a local file, extracts text, sanitizes it, and returns the redacted template for LLM analysis.
Supported Formats: Plain text (source code, logs, CSV, JSON, markdown) and Microsoft Word (
.docx) documents.Arguments:
filePath(string, required): Absolute file path to read and sanitize.profile(string, optional): The industry detection profile.
4. guard_exec (Command Execution Firewall)
Safely executes terminal commands in an isolated child process, masking stdout/stderr PII, database credentials, and API keys in local RAM before passing them to the AI agent. Includes a CISO audit receipt in stderr.
Arguments:
command(string, required): The shell command to execute (e.g.cat .env,docker logs web,git diff).cwd(string, optional): Working directory.profile(string, optional): Detection profile (defaults toDev).timeout_ms(number, optional): Timeout in ms (defaults to15000).
5. guard_read_file (Credential-Masking File Reader)
Reads files (.env, configs, source code, database dumps) and tokenizes all passwords, JWTs, and PII in volatile memory, returning safe redacted content for AI reasoning.
Arguments:
file_path(string, required): Path to file.profile(string, optional): Detection profile (defaults toDev).max_lines(number, optional): Line cap for large files (defaults to500).
6. guard_git_diff (Pre-Commit & Diff Sanitizer)
Inspects staged (--cached) or unstaged repository diffs, redacting any newly introduced secrets or PII in local RAM before AI code review or commit message generation.
Arguments:
staged(boolean, optional): Iftrue, inspects staged changes (git diff --cached). Defaults tofalse.cwd(string, optional): Working directory.profile(string, optional): Detection profile (defaults toDev).
7. guard_apply_patch (Safe Patch Applicator)
Reverses token placeholders ([API_KEY_1], [SECRET_1]) in AI-generated code or text by looking up the local RAM session map, creating a .bak backup, and writing authentic cleartext directly to disk. The remote LLM never sees real secrets.
Arguments:
file_path(string, required): Path to target file.content(string, required): Content containing tokens to restore on disk.create_backup(boolean, optional): Backup existing file before write (defaults totrue).
8. create_agent_rules (1-Click Agent Rule Injection)
Automatically scaffolds CISO-grade Zero-Trust directives into .cursorrules, .windsurfrules, CLAUDE.md, .github/copilot-instructions.md, or .clinerules.
Arguments:
agent_types(array, optional):["all"],["cursor"],["windsurf"],["claude_code"],["copilot"],["cline"]. Defaults to["all"].workspace_dir(string, optional): Workspace directory.
9. check_status
Returns a visual dashboard showing your current tier, session request count, active profiles, and upgrade instructions. Use it at any time to check your license status or get setup help.
Arguments: (none required)
JSON-RPC Call Example:
{ "method": "tools/call", "params": { "name": "check_status", "arguments": {} } }Response Example (Free Tier):
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ PrivacyScrubber MCP Server v2.2.4 โ โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฃ โ ๐ Tier: FREE โ โ ๐ Session requests: 5 โ โ ๐ Input size limit: 15,000 characters/request โ โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฃ โ ๐ท๏ธ Profiles: General only โ PRO unlocks 25 more โ โ ๐ Custom rules: ๐ Locked โ requires PRO โ โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฃ โ ๐ณ Upgrade to PRO โ $110 Lifetime โ โ https://privacyscrubber.com/pricing?utm_source=npm&utm_medium=readme&utm_campaign=mcp_server โ โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฃ โ After purchase, add your key to MCP config: โ โ "PRIVACYSCRUBBER_KEY": "<your-key-here>" โ โ Full setup guide: โ โ https://privacyscrubber.com/features/mcp/?utm_source=npm&utm_medium=readme&utm_campaign=mcp_server โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
๐ก๏ธ Standalone CLI: ps-guard
PrivacyScrubber bundles ps-guard for Unix pipe, pre-commit, and agentic workflows:
# Pipe any output through RAM redaction
cat .env | npx ps-guard --profile dev
# Execute commands through the ZTDS safety wrapper
npx ps-guard -- npm test
# Review git diff with secrets redacted
npx ps-guard --diff --staged
# Generate rules for all AI IDEs (.cursorrules, .windsurfrules, CLAUDE.md, copilot, cline)
npx ps-guard --rules๐ Browser Extension & Web Client
Looking for real-time protection directly inside your web browser?
Chrome Extension: Get the PrivacyScrubber Chrome Extension to sanitize prompts directly inside ChatGPT, Claude, and Gemini in real-time.
Web Sandbox: Use the zero-server browser sanitization tools at PrivacyScrubber Homepage.
๐ License & Commercial Upgrade
By default, the server runs under the Free Tier (restricted to 15,000 characters per request and the basic General PII profile). To unlock 30 specialized engineering, medical, legal, and financial PII profiles, as well as team-wide custom rules, you can purchase a commercial license.
Feature Comparison
Feature | Free Tier | PRO Tier | TEAMS Tier |
Volatile Tokenization | โ Yes | โ Yes | โ Yes |
Standard PII Masking | โ Yes | โ Yes | โ Yes |
Max Character Length | 15,000 chars | โพ๏ธ Unlimited | โพ๏ธ Unlimited |
Industry Profiles | General Only | 30 Profiles | 30 Profiles |
Custom Regex Rules | โ Locked | โพ๏ธ Unlimited | โพ๏ธ Unlimited |
Team Rules Sync (GPO) | โ No | โ No | โ Yes (Shared Link) |
Licensing Cost | $0 | $110 Lifetime | $99/mo Flat Rate |
๐ Acquire a PRO / TEAMS License Key at privacyscrubber.com/pricing
๐ After Purchase: Activate PRO in Your MCP Client
After purchasing a PRO license at privacyscrubber.com/pricing, you will receive a license key. Add it to your MCP client config as an environment variable: PRIVACYSCRUBBER_KEY.
Claude Desktop
Edit ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or %APPDATA%\Claude\claude_desktop_config.json (Windows):
{
"mcpServers": {
"privacyscrubber": {
"command": "npx",
"args": ["-y", "@privacyscrubber/mcp-server"],
"env": {
"PRIVACYSCRUBBER_KEY": "YOUR_LICENSE_KEY_HERE"
}
}
}
}Restart Claude Desktop after saving.
Cursor
Go to Settings โ Features โ MCP Servers.
Find
privacyscrubberand click Edit.Add the environment variable:
PRIVACYSCRUBBER_KEY=YOUR_LICENSE_KEY_HERE.Restart Cursor.
Alternatively, export it system-wide so all tools pick it up:
# macOS / Linux โ add to ~/.zshrc or ~/.bashrc
export PRIVACYSCRUBBER_KEY="YOUR_LICENSE_KEY_HERE"Windsurf
Edit ~/.codeium/windsurf/mcp_config.json:
{
"mcpServers": {
"privacyscrubber": {
"command": "npx",
"args": ["-y", "@privacyscrubber/mcp-server"],
"env": {
"PRIVACYSCRUBBER_KEY": "YOUR_LICENSE_KEY_HERE"
}
}
}
}Verify Activation
After adding the key, ask your AI agent to call check_status:
Use the check_status tool from PrivacyScrubber MCP๐ Ecosystem & Production Architecture Guides
๐ Web App: https://privacyscrubber.com
๐ฆ Node.js / TypeScript SDK: @privacyscrubber/sdk
๐งฉ Chrome Extension: Chrome Web Store
๐ก๏ธ RAG & Vector Databases: Sanitizing PII Before Vector DB Ingestion (Pinecone, Chroma, Qdrant)
๐ค LangChain & LlamaIndex: In-Memory PII Middleware for AI Agent Pipelines
โก AWS Comprehend Alternative: In-Memory Redaction Without Egress or Cloud Overhead
๐ Academic Foundations & Regulatory Verification
PrivacyScrubber and the Zero-Trust Data Sanitization (ZTDS) protocol are backed by published scientific, clinical, and legal treatises:
Repository / Archive | DOI / Identifier | Focus Area | Regulatory & Compliance Scope |
IETF Specification | The ZTDS Protocol for Frontier AI Ingestion (Internet-Draft) | Global AI Privacy, Zero-Egress Architecture | |
Zenodo / CERN | Zero-Trust Data Sanitization (ZTDS) Protocol Foundation | Cross-Border AI Privacy, ISO 27001 A.8.11 | |
OSF (Center for Open Science) | Empirical Latency Benchmark & Memory Profiling (<2ms RAM) | Performance vs Cloud DLP Proxies | |
SSRN / Elsevier | Enterprise Generative AI Governance | EU AI Act, UK GDPR, US State Privacy | |
medRxiv (Cold Spring Harbor) | Multi-Center Clinical Trial De-Identification | HIPAA Safe Harbor Section 164.514(b) | |
Law Archive / OSF | Preserving Attorney-Client Privilege in AI Workflows | ABA Model Rules & Legal Ethics |
Citing PrivacyScrubber in Research & Audits
@software{sibiryakov2026privacyscrubber,
author = {Sibiryakov, Ilya},
title = {PrivacyScrubber: Zero-Trust Data Sanitization (ZTDS) Engine & MCP Server},
year = {2026},
publisher = {Zenodo},
doi = {10.5281/zenodo.22058770},
url = {https://github.com/moxno/privacyscrubber-mcp}
}๐ License
MIT ยฉ Ilya Sibiryakov (BrandMeWeb)
โ๏ธ Intellectual Property & Virtual Patent Marking
The Zero-Trust Data Sanitization (ZTDS) architecture, in-memory deterministic tokenization, cryptographic session handoff, and stdio execution methods implemented in this package are proprietary technology of Ilya Sibiryakov (BrandMeWeb) and are protected under Patent Pending status:
Patent Office: State of Israel Ministry of Justice, Patent Office (ILPO)
Application Number:
331905(Tracking ID:94221)Filing / Priority Date: September 14, 2026 (Paris Convention Art. 4 & 35 U.S.C. ยง 119 Priority)
Official Title: SYSTEM AND METHOD FOR CLIENT-SIDE ZERO-TRUST DATA SANITIZATION AND CRYPTOGRAPHIC SESSION HANDOFF IN ARTIFICIAL INTELLIGENCE WORKFLOWS
Virtual Patent Marking: privacyscrubber.com/patents/ in accordance with 35 U.S.C. ยง 287(a).
๐ Internet Engineering Task Force (IETF) Specification
Specification Title: The Zero-Trust Data Sanitization (ZTDS) Protocol for Frontier Artificial Intelligence Ingestion
Document Category: IETF Internet-Draft (Individual Submission)
IETF Datatracker: https://datatracker.ietf.org/doc/draft-sibiryakov-ztds-protocol/
Archive Plaintext: https://www.ietf.org/archive/id/draft-sibiryakov-ztds-protocol-01.txt
Available Tools
3 toolsreveal_textA
Replaces masked tokens (e.g., [EMAIL_1], [API_KEY_1]) in the LLM's response back with the original private data from the local volatile RAM-only session map.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | The AI generated response containing placeholders to restore. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | Yes | The detokenized text with original values restored. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full responsibility for behavioral disclosure. It reveals that the tool accesses a 'local volatile RAM-only session map' and performs replacement, but it does not mention authentication requirements, rate limits, side effects on the session map, or error handling for missing tokens. The description adds moderate transparency beyond schema fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence of 25 words with no redundant information. It is front-loaded with the core action and provides precise details efficiently. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (single parameter, no nested objects, has output schema), the description covers the input, process, and data source adequately. It does not describe the output schema contents or edge cases like missing tokens, but these are not critical due to the existence of an output schema. The description is sufficiently complete for an AI agent to use correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage for the sole parameter 'text' (description: 'The AI generated response containing placeholders to restore.'). The tool description adds value by specifying the format of placeholders (e.g., [EMAIL_1], [API_KEY_1]) and the source of original data (local volatile RAM-only session map), which enhances understanding beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific verb 'replaces' and identifies the resource: masked tokens like [EMAIL_1], [API_KEY_1] in the LLM's response, using original private data from a local session map. This distinguishes it from sibling tools sanitize_file and sanitize_text, which perform the inverse operation (masking).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is used after an LLM response contains masked tokens, but it does not explicitly state when to use it versus the sibling tools, nor does it provide when-not-to-use guidance or prerequisites. The usage context is clear but not formally outlined.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sanitize_fileA
Reads a local file, sanitizes its contents using the selected profile, and outputs the safe version for AI analysis. Securely keeps original identifiers in memory.
| Name | Required | Description | Default |
|---|---|---|---|
| profile | No | The detection profile to use. Available: 'General' (Free), or PRO profiles: 'Dev' (Engineering/Code), 'Medical', 'Pharma', 'Legal', 'Compliance', 'CCPA', 'Finance', 'Bizops', 'Sales', 'WealthMgmt', 'Insurance', 'Accounting', 'HR', 'Security', 'Marketing', 'Support', 'RealEstate', 'Agents', 'Academic', 'Creative', 'Tech', 'Personal'. Defaults to 'General'. | |
| file_path | Yes | Absolute path to the file to sanitize. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | Yes | The sanitized content of the file. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Without annotations, the description adds value by stating that original identifiers are securely kept in memory. However, it omits details on read behavior, file modification, error handling, or concurrency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no wasted words. The first sentence states the core action, and the second adds a critical security detail. Efficient and front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool with an output schema, the description covers the essential operation. It could mention error conditions or output format, but the presence of an output schema reduces the need. Overall adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description does not add new parameter meaning beyond what is already in the schema; it only references 'selected profile' without elaboration.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb (reads and sanitizes), the resource (local file), and the outcome (outputs safe version for AI analysis). It distinguishes from sibling 'sanitize_text' by specifying file input.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like 'sanitize_text' or 'reveal_text'. The description lacks context on prerequisites or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sanitize_textA
Locally scrubs PII, secrets, and credentials (like API keys, passwords, emails, phones, names) from code, logs, or text. Replaces them with safe placeholders (e.g., [EMAIL_1], [API_KEY_1]). Keep your data secure before passing it to any LLM.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | The raw text, code, or logs to sanitize. | |
| profile | No | The detection profile to use. Available: 'General' (Free), or PRO profiles: 'Dev' (Engineering/Code), 'Medical', 'Pharma', 'Legal', 'Compliance', 'CCPA', 'Finance', 'Bizops', 'Sales', 'WealthMgmt', 'Insurance', 'Accounting', 'HR', 'Security', 'Marketing', 'Support', 'RealEstate', 'Agents', 'Academic', 'Creative', 'Tech', 'Personal'. Defaults to 'General'. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | Yes | The sanitized text output. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It discloses that scrubbing is local (no data leak) and replaces matches with safe placeholders. It does not detail performance or determinism but provides sufficient behavioral insight.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with zero waste. The first sentence states the action and the second provides a practical recommendation. Front-loaded with essential information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema exists, the description need not explain return values. It covers the core purpose and usage context. Could mention that output is also text with placeholders, but not necessary.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with both parameters documented. The description adds value by listing all available profiles (e.g., 'Dev', 'Medical') and noting defaults, but this is also partly in the schema's description. Minor extra context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool scrubs PII, secrets, and credentials from text and replaces them with placeholders. It specifies the action is local and differentiates from siblings like 'sanitize_file' (for files) and 'reveal_text' (presumably reverse).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description advises using the tool before passing data to an LLM to keep data secure. It implies when to use and emphasizes 'Locally' to indicate no data is sent externally, but does not explicitly mention when not to use or name alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v1.0.2- Changed
reveal_text1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "properties": { + "text": { + "description": "The detokenized text with original values restored.", + "type": "string" + } + }, + "required": [ + "text" + ], + "type": "object" +}
- Changed
sanitize_file4 fields changed- removed
Input schema / properties / filePathRemoved value: -{ - "description": "Absolute path to the file to sanitize.", - "type": "string" -} - added
Input schema / properties / file_pathAdded value: +{ + "description": "Absolute path to the file to sanitize.", + "type": "string" +} - changed
Input schema / requiredPrevious value: -[ - "filePath" -]New value: +[ + "file_path" +] - changed
Output schema / (root)Previous value: -nullNew value: +{ + "properties": { + "text": { + "description": "The sanitized content of the file.", + "type": "string" + } + }, + "required": [ + "text" + ], + "type": "object" +}
- Changed
sanitize_text1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "properties": { + "text": { + "description": "The sanitized text output.", + "type": "string" + } + }, + "required": [ + "text" + ], + "type": "object" +}
3 tool updates
v1.0.1- First observed
reveal_text - First observed
sanitize_file - First observed
sanitize_text
TDQS
Scored across 3 tools
Each tool has a distinct purpose: sanitizing text, sanitizing files, and reversing the sanitization. No overlap in functionality.
All tool names follow a consistent verb_noun pattern with underscores (reveal_text, sanitize_file, sanitize_text).
With 3 tools, the server covers the core workflow of sanitization and reversal. Could benefit from an explicit session management tool, but is well-scoped overall.
The set covers sanitization for both text and files, plus the necessary reveal operation. Missing a tool to manage profiles or clear the session, but the core lifecycle is covered.
Maintenance
Related MCP Connectors
Redact PII from text before it reaches a model. Nothing stored, no third-party AI.
Detects and redacts PII (emails, phones, SSNs, names, addresses) from text. $0.02/call via x402.
Stateless PII redaction over MCP/REST. Free โค1000 words or $0.01/call; file upload supported.
Scan configs, files, or text for leaked secrets and obvious misconfigurations. Nothing stored.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceLocal-first CLI and MCP server for redacting sensitive text before sharing logs, configs, and errors with AI tools.MIT
- FlicenseAqualityBmaintenancePII sanitization layer for autonomous AI agent pipelines. Detects and redacts emails, phone numbers, national IDs, private keys, and financial data before text reaches LLMs. Supports EN, ES (LATAM), PT (BR/PT), DE, JA.12-

classifinder-mcpofficial
AlicenseAqualityBmaintenanceEnables AI agents to scan text for leaked secrets and prompt injection markers, and redact them before reaching an LLM.21MIT- AlicenseAqualityDmaintenanceScans prompts for PII and masks or redacts sensitive data locally before sending to an LLM, supporting multiple anonymization modes.1MIT