privacyscrubber-mcp
This server provides local, in-memory PII/secret redaction for AI workflows so sensitive data never reaches the LLM, plus a way to restore the originals afterward.
sanitize_text— Scrubs PII, secrets, and credentials (API keys, passwords, emails, phones, names) from raw text, code, or logs, replacing them with safe placeholders like[EMAIL_1]or[API_KEY_1].sanitize_file— Reads a file from an absolute path, sanitizes its contents, and returns a safe version for AI analysis while keeping original identifiers only in local memory.reveal_text— Replaces masked tokens in the LLM's response back with the original private values from the volatile RAM-only session map.Detection profiles — Choose a profile to tune detection:
General(free), or PRO profiles such asDev,Medical,Pharma,Legal,Compliance,CCPA,Finance,HR,Security,Marketing,Support,RealEstate,Academic, and more.
Together these let you send code/logs/files to an LLM and read back answers with real values, without exposing secrets or personal data to the model provider.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@privacyscrubber-mcpscrub this text: 'Name: Alice, SSN: 123-45-6789'"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
@privacyscrubber/mcp-server
CISO-Approved Zero-Trust PII & Secrets Redaction MCP Server for Cursor, Windsurf, and Claude Desktop. Locally scrubs PII, secrets, credentials, and custom regex rules from files and text contexts before they reach remote LLM providers to prevent API leaks and ensure HIPAA/SOC 2 compliance at the developer endpoint.
Support Zero-Trust Open Source: If PrivacyScrubber protects your API keys, source code, and developer credentials from model leaks, please Star this repository on GitHub or execute
gh repo star moxno/privacyscrubber-mcp.
Direct Commercial Licensing & Production Infrastructure:
Looking to automate zero-trust PII redaction inside backend microservices, ETL pipelines, or LangChain/LlamaIndex agents?
Developer SDK ($299/mo flat or $2,990/yr): Unlimited internal backend nodes, zero network egress, <1ms in-memory latency. Subscribe to Developer SDK | Technical Specs
TEAMS ($99/mo flat): Unlimited team seats across your entire domain with cryptographic P2P session handoff. Deploy TEAMS Plan
PRO ($15/mo or $110 Lifetime): Single-seat unlimited throughput across all 30 vertical profiles. Get PRO License
🔒 Zero-Trust Data Flow
All sensitive parameters, identifiers, and variables are intercepted locally inside your machine's RAM. They are replaced by tokens (e.g. [EMAIL_1]) before being sent to the AI. Once the AI responds, the tokens are safely swapped back to original values in your local context.
[Raw Input / Files] ──> [MCP sanitize_text] ──> [Masked Tokens] ──> [LLM API]
│ │
(In-Memory Map) (Result)
│ │
[Original Output] <─── [MCP reveal_text] <─────────────────────────────┘Related MCP server: trustboost-pii-sanitizer
🚀 Installation
1. Install via Smithery
To automatically configure and run with your preferred client, install using Smithery:
npx -y @smithery/cli install @privacyscrubber/mcp-server --write-to-clients2. Instant Run with NPX
Run the server directly without local installation:
npx -y @privacyscrubber/mcp-server3. Instant Run with UVX (Python / PyPI)
Run the server directly inside Python agent workflows (CrewAI, LangChain, AutoGen, LlamaIndex):
uvx privacyscrubber-mcpOr via pip:
pip install privacyscrubber-mcp
privacyscrubber-mcp4. Programmatic Node.js / TypeScript SDK (Lightweight Presidio Alternative)
Need direct, in-memory zero-trust PII sanitization in your backend microservice, Next.js app, or RAG vector pipeline rather than an MCP server? Use our official zero-dependency SDK:
npm install @privacyscrubber/sdkimport OpenAI from 'openai';
import { wrapOpenAI } from '@privacyscrubber/sdk';
// Transparently masks PII before sending to LLM and rehydrates responses:
const openai = wrapOpenAI(new OpenAI({ apiKey: process.env.OPENAI_API_KEY }));
// Outbound prompt is sanitized in local RAM before leaving your machine:
// "Schedule a call with [NAME_1] at [EMAIL_1] regarding API key [AWS_KEY_1]."
const completion = await openai.chat.completions.create({
model: 'gpt-4o',
messages: [{ role: 'user', content: 'Schedule a call with Alice Smith at alice@acme.com with AKIAIOSFODNN7EXAMPLE.' }]
});
// Incoming LLM answer is automatically rehydrated with "Alice Smith (alice@acme.com)":
console.log(completion.choices[0].message.content);⚡ Why @privacyscrubber/sdk vs Microsoft Presidio?
Microsoft Presidio is the Python standard, but deploying it in a Node.js / TypeScript stack requires running heavy Python microservices, Docker containers, and 500MB+ spaCy NLP models with 35–120ms latency. @privacyscrubber/sdk runs 100% in-process with zero dependencies:
Feature / Metric |
| Microsoft Presidio | AWS Comprehend / Google Cloud DLP |
Runtime & Dependencies | 0 dependencies (~150KB) | Python + Docker + spaCy (~500MB) | Heavy Cloud SDKs |
Execution Latency | Sub-millisecond in RAM (<1ms) | 35–120ms HTTP/gRPC roundtrip | 180–400ms external cloud roundtrip |
Docker / Sidecar Needed | None (Pure in-process Node/WASM) | Mandatory Docker container | VPC endpoints & IAM configurations |
Network Egress | 0 Bytes (100% Air-gapped) | Local internal network hop | Full unencrypted payload to cloud |
Data Loss & Reversibility | 0 Loss (RAM-reversible via | Manual token vault configuration | Irreversible masking / hashing |
OpenAI / LangChain 1-Liner | Built-in ( | Complex custom pipeline glue | Custom proxy architecture |
DevOps Secrets Interception | Built-in (AWS, JWT, DB URIs, GitHub PAT) | Custom regex recognizers needed | Cloud-specific classifiers |
Streaming Rehydration | Built-in ( | Buffering / chunk split failures | Not supported in real-time streams |
👉 View @privacyscrubber/sdk on NPM | Full documentation, Express middleware, LangChain transforms, and 25 compliance profiles.
🤖 AI Orchestrator Recipes: LangChain, LlamaIndex, CrewAI & AutoGPT
If you are building autonomous agents, RAG vector pipelines, or backend services rather than single-user IDE prompts, use @privacyscrubber/sdk to sanitize data in-memory:
1. LangChain.js (LCEL & RAG Document Ingestion)
import { ChatOpenAI } from '@langchain/openai';
import { createLangChainTransform, createDocumentTransformer } from '@privacyscrubber/sdk';
// A. RAG Pre-Ingestion: sanitize documents before embedding into vector stores
const docTransformer = createDocumentTransformer({
profile: 'General',
sensitiveMetadataKeys: ['account_owner', 'submitter_email'],
attachTelemetry: true
});
const sanitizedDocs = await docTransformer.transformDocuments(rawDocs);
// B. LCEL Runtime Chains: sanitize prompts and restore responses in local RAM
const transform = createLangChainTransform({ defaultProfile: 'Dev' });
const { scrubbedText, tokenMap } = transform.preprocess(
"Deploying database with user admin and pwd postgresql://user:SecretPass123@db.internal:5432/prod"
);
const model = new ChatOpenAI({ model: 'gpt-4o' });
const response = await model.invoke(scrubbedText);
const finalOutput = transform.postprocess(response.content, tokenMap);2. LlamaIndex.TS & Universal Vector DB Ingestion (Chroma, Pinecone, Qdrant)
import { createVectorIngestionGuard, createLlamaIndexTransform } from '@privacyscrubber/sdk';
// Initialize in-memory Vector Ingestion Guard (<1ms per batch, zero network egress)
const guard = createVectorIngestionGuard({
profile: 'Finance',
sensitiveMetadataKeys: ['contractor_email', 'billing_contact']
});
// A. Sanitize Chroma columnar batches ({ ids, documents, metadatas })
const cleanChroma = guard.sanitizeRecords(chromaBatch);
// B. Sanitize Pinecone / Qdrant record arrays ({ id, text/payload, metadata })
const cleanRecords = guard.sanitizeRecords(pineconeRecords);
// C. Align search query tokens with sanitized vector space before embedding
const { query: alignedQuery } = guard.sanitizeQuery("Search user john@example.com records");
// D. Transparent Vector Store Proxy: auto-sanitizes on write, re-hydrates on query
const guardedStore = guard.wrapVectorStore(nativeVectorStore);Complete runnable recipes are included inside the npm package under
@privacyscrubber/sdk/examples/(or runnpx @privacyscrubber/sdk).
3. CrewAI & Multi-Agent Swarms (Sidecar Mode / Python Interop)
# Run lightweight zero-dependency local daemon (<1ms in-memory, 0 external network calls):
npx @privacyscrubber/sdk sidecar --port 3100# In your Python CrewAI / AutoGen Agent tools:
import requests
def sanitize_agent_input(prompt: str) -> tuple[str, dict]:
res = requests.post("http://127.0.0.1:3100/sanitize", json={"text": prompt, "profile": "Dev"}).json()
return res["scrubbedText"], res["tokenMap"]
def restore_agent_output(ai_response: str, token_map: dict) -> str:
res = requests.post("http://127.0.0.1:3100/restore", json={"text": ai_response, "tokenMap": token_map}).json()
return res["restoredText"]4. AutoGPT / Multi-Turn Autonomous Agents
import { PrivacyScrubberEngine, createGuardedTools } from '@privacyscrubber/sdk';
const engine = new PrivacyScrubberEngine({ defaultProfile: 'Dev' });
// Equips agent with tools that intercept commands, file reads, and git diffs before model exposure:
const guardedTools = createGuardedTools(engine);Enterprise & Production Licensing:
Self-Serve Developer SDK ($299/mo flat or $2,990/yr): Unlimited internal backend nodes, microservices, and RAG pipelines. Subscribe to Developer SDK | Developer SDK Specs
PrivacyScrubber TEAMS ($99/mo flat): Unlimited team seats, centralized policy enforcement, encrypted session handoff. Deploy TEAMS Plan
Enterprise Air-Gapped License: On-premise source code distribution, zero-network custom models. Contact Enterprise
🛡️ Architecture & Security Deep-Dive (Zero-Trust vs Cloud DLP)
When AI IDEs (Cursor, Claude Desktop, Windsurf) connect to model providers, developer credentials, database connection strings, and internal customer PII are at continuous risk of prompt exfiltration. The @privacyscrubber/mcp-server enforces four immutable architectural guarantees:
Stdio Air-Gapped Transport: The MCP server communicates exclusively over local standard input/output (
stdio) child processes spawned by your IDE. It opens zero external listening ports and initiates zero remote network requests.Volatile RAM-Only Token Map: The mapping table between synthetic tokens (
[AWS_KEY_1],[EMAIL_1]) and raw cleartext is maintained exclusively in ephemeral node memory and is wiped the moment your IDE session closes.Deterministic AST Lookarounds vs Cloud Proxy Overhead: Unlike cloud DLP gateways (Nightfall, Skyflow) that add 200–400ms latency and transmit unencrypted code to third parties,
@privacyscrubber/mcp-serverruns locally in <2ms with zero data egress.Permanent Standards & Academic Validation:
IETF Specification: draft-sibiryakov-ztds-protocol-01 — Revision 01 · 2026-09-23 · 13 pages
CERN / Zenodo Foundation: DOI 10.5281/zenodo.22058770
Center for Open Science (OSF): DOI 10.17605/OSF.IO/5BYJF
Patent Pending: U.S. & International Patents Pending (Paris Convention / PCT Priority)
⚙️ Client Integrations
Claude Desktop
Add this to your Claude Desktop config file:
macOS:
~/Library/Application Support/Claude/claude_desktop_config.jsonWindows:
%APPDATA%\Claude\claude_desktop_config.json
{
"mcpServers": {
"privacyscrubber": {
"command": "npx",
"args": ["-y", "@privacyscrubber/mcp-server"],
"env": {
"PRIVACYSCRUBBER_KEY": "YOUR_OPTIONAL_PRO_LICENSE_KEY"
}
}
}
}Cursor / Windsurf
Navigate to Settings -> Features -> MCP.
Add new MCP server:
Name:
privacyscrubberType:
commandCommand:
npx -y @privacyscrubber/mcp-server
Optional: Set
PRIVACYSCRUBBER_KEYas an environment variable in your system shell.
Cline / Roo Code
Add to cline_mcp_settings.json:
{
"mcpServers": {
"privacyscrubber": {
"command": "npx",
"args": ["-y", "@privacyscrubber/mcp-server"],
"env": {
"PRIVACYSCRUBBER_KEY": "YOUR_OPTIONAL_PRO_LICENSE_KEY"
}
}
}
}Claude Code CLI
Add directly from your terminal:
claude mcp add privacyscrubber -- npx -y @privacyscrubber/mcp-server🛠️ Provided Tools & JSON-RPC Specifications
1. sanitize_text
Redacts PII, secrets, API keys, and credentials from a text block and populates the volatile local replacement mapping.
Arguments:
text(string, required): The raw content or logs to sanitize.profile(string, optional): Gated industry detection profile (e.g., 'General', 'Dev', 'Medical', 'Legal', 'Compliance'). Defaults to 'General'.
JSON-RPC Call Example:
{ "method": "tools/call", "params": { "name": "sanitize_text", "arguments": { "text": "Contact me at dev-key-1234 or jane.doe@company.com", "profile": "General" } } }Response Example:
{ "content": [ { "type": "text", "text": "Contact me at [SECRET_1] or [EMAIL_1]" } ] }
2. reveal_text
Detokenizes the AI response back to the original values locally.
Arguments:
text(string, required): The response from the LLM containing tokenized placeholders.
JSON-RPC Call Example:
{ "method": "tools/call", "params": { "name": "reveal_text", "arguments": { "text": "Please reach out to [EMAIL_1] regarding the update." } } }Response Example:
{ "content": [ { "type": "text", "text": "Please reach out to jane.doe@company.com regarding the update." } ] }
3. sanitize_file
Reads a local file, extracts text, sanitizes it, and returns the redacted template for LLM analysis.
Supported Formats: Plain text (source code, logs, CSV, JSON, markdown) and Microsoft Word (
.docx) documents.Arguments:
filePath(string, required): Absolute file path to read and sanitize.profile(string, optional): The industry detection profile.
4. guard_exec (Command Execution Firewall)
Safely executes terminal commands in an isolated child process, masking stdout/stderr PII, database credentials, and API keys in local RAM before passing them to the AI agent. Includes a CISO audit receipt in stderr.
Arguments:
command(string, required): The shell command to execute (e.g.cat .env,docker logs web,git diff).cwd(string, optional): Working directory.profile(string, optional): Detection profile (defaults toDev).timeout_ms(number, optional): Timeout in ms (defaults to15000).
5. guard_read_file (Credential-Masking File Reader)
Reads files (.env, configs, source code, database dumps) and tokenizes all passwords, JWTs, and PII in volatile memory, returning safe redacted content for AI reasoning.
Arguments:
file_path(string, required): Path to file.profile(string, optional): Detection profile (defaults toDev).max_lines(number, optional): Line cap for large files (defaults to500).
6. guard_git_diff (Pre-Commit & Diff Sanitizer)
Inspects staged (--cached) or unstaged repository diffs, redacting any newly introduced secrets or PII in local RAM before AI code review or commit message generation.
Arguments:
staged(boolean, optional): Iftrue, inspects staged changes (git diff --cached). Defaults tofalse.cwd(string, optional): Working directory.profile(string, optional): Detection profile (defaults toDev).
7. guard_apply_patch (Safe Patch Applicator)
Reverses token placeholders ([API_KEY_1], [SECRET_1]) in AI-generated code or text by looking up the local RAM session map, creating a .bak backup, and writing authentic cleartext directly to disk. The remote LLM never sees real secrets.
Arguments:
file_path(string, required): Path to target file.content(string, required): Content containing tokens to restore on disk.create_backup(boolean, optional): Backup existing file before write (defaults totrue).
8. create_agent_rules (1-Click Agent Rule Injection)
Automatically scaffolds CISO-grade Zero-Trust directives into .cursorrules, .windsurfrules, CLAUDE.md, .github/copilot-instructions.md, or .clinerules.
Arguments:
agent_types(array, optional):["all"],["cursor"],["windsurf"],["claude_code"],["copilot"],["cline"]. Defaults to["all"].workspace_dir(string, optional): Workspace directory.
9. check_status
Returns a visual dashboard showing your current tier, session request count, active profiles, and upgrade instructions. Use it at any time to check your license status or get setup help.
Arguments: (none required)
JSON-RPC Call Example:
{ "method": "tools/call", "params": { "name": "check_status", "arguments": {} } }Response Example (Free Tier):
╔══════════════════════════════════════════════════╗ ║ PrivacyScrubber MCP Server v2.2.8 ║ ╠══════════════════════════════════════════════════╣ ║ 🔓 Tier: FREE ║ ║ 📊 Session requests: 5 ║ ║ 📁 Input size limit: 15,000 characters/request ║ ╠══════════════════════════════════════════════════╣ ║ 🏷️ Profiles: General only — PRO unlocks 25 more ║ ║ 📋 Custom rules: 🔒 Locked — requires PRO ║ ╠══════════════════════════════════════════════════╣ ║ 💳 Upgrade to PRO — $110 Lifetime ║ ║ https://privacyscrubber.com/pricing?utm_source=npm&utm_medium=readme&utm_campaign=mcp_server ║ ╠══════════════════════════════════════════════════╣ ║ After purchase, add your key to MCP config: ║ ║ "PRIVACYSCRUBBER_KEY": "<your-key-here>" ║ ║ Full setup guide: ║ ║ https://privacyscrubber.com/pii-mcp/?utm_source=npm&utm_medium=readme&utm_campaign=mcp_server ║ ╚══════════════════════════════════════════════════╝
🛡️ Standalone CLI: ps-guard
PrivacyScrubber bundles ps-guard for Unix pipe, pre-commit, and agentic workflows:
# Pipe any output through RAM redaction
cat .env | npx ps-guard --profile dev
# Execute commands through the ZTDS safety wrapper
npx ps-guard -- npm test
# Review git diff with secrets redacted
npx ps-guard --diff --staged
# Generate rules for all AI IDEs (.cursorrules, .windsurfrules, CLAUDE.md, copilot, cline)
npx ps-guard --rules🌐 Browser Extension & Web Client
Looking for real-time protection directly inside your web browser?
Chrome Extension: Get the PrivacyScrubber Chrome Extension to sanitize prompts directly inside ChatGPT, Claude, and Gemini in real-time.
Web Sandbox: Use the zero-server browser sanitization tools at PrivacyScrubber Homepage.
📄 License & Commercial Upgrade
By default, the server runs under the Free Tier (restricted to 15,000 characters per request and the basic General PII profile). To unlock 30 specialized engineering, medical, legal, and financial PII profiles, as well as team-wide custom rules, you can purchase a commercial license.
Commercial Plan & Feature Comparison
Feature | Community Free | PRO Tier | TEAMS Tier | Developer SDK |
Execution Architecture | Local Stdio RAM | Local Stdio RAM | Local Stdio RAM | In-Process Node/WASM |
Volatile Tokenization | Yes (RAM-only) | Yes (RAM-only) | Yes (RAM-only) | Yes (RAM-only) |
Standard PII Masking | Yes | Yes | Yes | Yes |
Max Character Throughput | 15,000 chars / call | Unlimited | Unlimited | Unlimited |
30 Industry Profiles | General Only (5k trial) | All 30 Profiles | All 30 Profiles | All 30 Profiles |
Custom Regex Rules | Locked | Unlimited | Unlimited | Unlimited |
Team Rules Sync (GPO) | No | No | Yes (Shared Link) | Yes (Configurable) |
Microservice / RAG Export | No | No | No | Yes ( |
Subprocessor Liability | 0 (Client-side) | 0 (Client-side) | 0 (Client-side) | 0 (In-Process Node/WASM) |
Licensing Cost | $0 | $15/mo or $110 Lifetime | $99/mo Flat Rate | $299/mo or $2,990/yr |
Direct Activation | Default included |
👉 Acquire a Commercial License Key at privacyscrubber.com/pricing
🔐 After Purchase: Activate PRO in Your MCP Client
After purchasing a PRO license at privacyscrubber.com/pricing, you will receive a license key. Add it to your MCP client config as an environment variable: PRIVACYSCRUBBER_KEY.
Claude Desktop
Edit ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or %APPDATA%\Claude\claude_desktop_config.json (Windows):
{
"mcpServers": {
"privacyscrubber": {
"command": "npx",
"args": ["-y", "@privacyscrubber/mcp-server"],
"env": {
"PRIVACYSCRUBBER_KEY": "YOUR_LICENSE_KEY_HERE"
}
}
}
}Restart Claude Desktop after saving.
Cursor
Go to Settings → Features → MCP Servers.
Find
privacyscrubberand click Edit.Add the environment variable:
PRIVACYSCRUBBER_KEY=YOUR_LICENSE_KEY_HERE.Restart Cursor.
Alternatively, export it system-wide so all tools pick it up:
# macOS / Linux — add to ~/.zshrc or ~/.bashrc
export PRIVACYSCRUBBER_KEY="YOUR_LICENSE_KEY_HERE"Windsurf
Edit ~/.codeium/windsurf/mcp_config.json:
{
"mcpServers": {
"privacyscrubber": {
"command": "npx",
"args": ["-y", "@privacyscrubber/mcp-server"],
"env": {
"PRIVACYSCRUBBER_KEY": "YOUR_LICENSE_KEY_HERE"
}
}
}
}Verify Activation
After adding the key, ask your AI agent to call check_status:
Use the check_status tool from PrivacyScrubber MCP🔗 Ecosystem & Production Architecture Guides
🌐 Web App: https://privacyscrubber.com
📦 Node.js / TypeScript SDK: @privacyscrubber/sdk (npm)
🧩 Chrome Extension: Chrome Web Store
🛡️ RAG & Vector Databases: Sanitizing PII Before Vector DB Ingestion (Pinecone, Chroma, Qdrant)
🤖 LangChain & LlamaIndex: In-Memory PII Middleware for AI Agent Pipelines
⚡ AWS Comprehend Alternative: In-Memory Redaction Without Egress or Cloud Overhead
📚 Academic Foundations & Regulatory Verification
PrivacyScrubber and the Zero-Trust Data Sanitization (ZTDS) protocol are backed by published scientific, clinical, and legal treatises:
Repository / Archive | DOI / Identifier | Focus Area | Regulatory & Compliance Scope |
IETF Specification | The ZTDS Protocol for Frontier AI Ingestion (Internet-Draft) | Global AI Privacy, Zero-Egress Architecture | |
Zenodo / CERN | Zero-Trust Data Sanitization (ZTDS) Protocol Foundation | Cross-Border AI Privacy, ISO 27001 A.8.11 | |
OSF (Center for Open Science) | Empirical Latency Benchmark & Memory Profiling (<2ms RAM) | Performance vs Cloud DLP Proxies | |
SSRN / Elsevier | Enterprise Generative AI Governance | EU AI Act, UK GDPR, US State Privacy | |
medRxiv (Cold Spring Harbor) | Multi-Center Clinical Trial De-Identification | HIPAA Safe Harbor Section 164.514(b) | |
Law Archive / OSF | Preserving Attorney-Client Privilege in AI Workflows | ABA Model Rules & Legal Ethics |
Citing PrivacyScrubber in Research & Audits
@software{sibiryakov2026privacyscrubber,
author = {Sibiryakov, Ilya},
title = {PrivacyScrubber: Zero-Trust Data Sanitization (ZTDS) Engine & MCP Server},
year = {2026},
publisher = {Zenodo},
doi = {10.5281/zenodo.22058770},
url = {https://github.com/moxno/privacyscrubber-mcp}
}📄 License
MIT © Ilya Sibiryakov (BrandMeWeb)
⚖️ Intellectual Property & Patent Disclosures
The Zero-Trust Data Sanitization (ZTDS) architecture, in-memory deterministic tokenization, cryptographic session handoff, and stdio execution methods implemented in this package are proprietary technology of Ilya Sibiryakov (BrandMeWeb) and are protected under Patent Pending status:
Legal Status: U.S. & International Patents Pending
Priority Framework: Paris Convention Art. 4 & 35 U.S.C. § 119 Priority (locked through September 14, 2027)
Filing / Priority Date: September 14, 2026
Official Title: SYSTEM AND METHOD FOR CLIENT-SIDE ZERO-TRUST DATA SANITIZATION AND CRYPTOGRAPHIC SESSION HANDOFF IN ARTIFICIAL INTELLIGENCE WORKFLOWS
Patent Disclosures Hub: privacyscrubber.com/patents/ (formal verification via counsel NDA).
🌐 Internet Engineering Task Force (IETF) Specification
Specification Title: The Zero-Trust Data Sanitization (ZTDS) Protocol for Frontier Artificial Intelligence Ingestion
Document Category: IETF Internet-Draft (Individual Submission)
IETF Datatracker: https://datatracker.ietf.org/doc/draft-sibiryakov-ztds-protocol/
Archive Plaintext: https://www.ietf.org/archive/id/draft-sibiryakov-ztds-protocol-01.txt
Available Tools
3 toolsreveal_textA
Replaces masked tokens (e.g., [EMAIL_1], [API_KEY_1]) in the LLM's response back with the original private data from the local volatile RAM-only session map.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | The AI generated response containing placeholders to restore. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | Yes | The detokenized text with original values restored. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full responsibility for behavioral disclosure. It reveals that the tool accesses a 'local volatile RAM-only session map' and performs replacement, but it does not mention authentication requirements, rate limits, side effects on the session map, or error handling for missing tokens. The description adds moderate transparency beyond schema fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence of 25 words with no redundant information. It is front-loaded with the core action and provides precise details efficiently. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (single parameter, no nested objects, has output schema), the description covers the input, process, and data source adequately. It does not describe the output schema contents or edge cases like missing tokens, but these are not critical due to the existence of an output schema. The description is sufficiently complete for an AI agent to use correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage for the sole parameter 'text' (description: 'The AI generated response containing placeholders to restore.'). The tool description adds value by specifying the format of placeholders (e.g., [EMAIL_1], [API_KEY_1]) and the source of original data (local volatile RAM-only session map), which enhances understanding beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific verb 'replaces' and identifies the resource: masked tokens like [EMAIL_1], [API_KEY_1] in the LLM's response, using original private data from a local session map. This distinguishes it from sibling tools sanitize_file and sanitize_text, which perform the inverse operation (masking).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is used after an LLM response contains masked tokens, but it does not explicitly state when to use it versus the sibling tools, nor does it provide when-not-to-use guidance or prerequisites. The usage context is clear but not formally outlined.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sanitize_fileA
Reads a local file, sanitizes its contents using the selected profile, and outputs the safe version for AI analysis. Securely keeps original identifiers in memory.
| Name | Required | Description | Default |
|---|---|---|---|
| profile | No | The detection profile to use. Available: 'General' (Free), or PRO profiles: 'Dev' (Engineering/Code), 'Medical', 'Pharma', 'Legal', 'Compliance', 'CCPA', 'Finance', 'Bizops', 'Sales', 'WealthMgmt', 'Insurance', 'Accounting', 'HR', 'Security', 'Marketing', 'Support', 'RealEstate', 'Agents', 'Academic', 'Creative', 'Tech', 'Personal'. Defaults to 'General'. | |
| file_path | Yes | Absolute path to the file to sanitize. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | Yes | The sanitized content of the file. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Without annotations, the description adds value by stating that original identifiers are securely kept in memory. However, it omits details on read behavior, file modification, error handling, or concurrency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no wasted words. The first sentence states the core action, and the second adds a critical security detail. Efficient and front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool with an output schema, the description covers the essential operation. It could mention error conditions or output format, but the presence of an output schema reduces the need. Overall adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description does not add new parameter meaning beyond what is already in the schema; it only references 'selected profile' without elaboration.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb (reads and sanitizes), the resource (local file), and the outcome (outputs safe version for AI analysis). It distinguishes from sibling 'sanitize_text' by specifying file input.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like 'sanitize_text' or 'reveal_text'. The description lacks context on prerequisites or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sanitize_textA
Locally scrubs PII, secrets, and credentials (like API keys, passwords, emails, phones, names) from code, logs, or text. Replaces them with safe placeholders (e.g., [EMAIL_1], [API_KEY_1]). Keep your data secure before passing it to any LLM.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | The raw text, code, or logs to sanitize. | |
| profile | No | The detection profile to use. Available: 'General' (Free), or PRO profiles: 'Dev' (Engineering/Code), 'Medical', 'Pharma', 'Legal', 'Compliance', 'CCPA', 'Finance', 'Bizops', 'Sales', 'WealthMgmt', 'Insurance', 'Accounting', 'HR', 'Security', 'Marketing', 'Support', 'RealEstate', 'Agents', 'Academic', 'Creative', 'Tech', 'Personal'. Defaults to 'General'. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | Yes | The sanitized text output. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It discloses that scrubbing is local (no data leak) and replaces matches with safe placeholders. It does not detail performance or determinism but provides sufficient behavioral insight.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with zero waste. The first sentence states the action and the second provides a practical recommendation. Front-loaded with essential information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema exists, the description need not explain return values. It covers the core purpose and usage context. Could mention that output is also text with placeholders, but not necessary.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with both parameters documented. The description adds value by listing all available profiles (e.g., 'Dev', 'Medical') and noting defaults, but this is also partly in the schema's description. Minor extra context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool scrubs PII, secrets, and credentials from text and replaces them with placeholders. It specifies the action is local and differentiates from siblings like 'sanitize_file' (for files) and 'reveal_text' (presumably reverse).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description advises using the tool before passing data to an LLM to keep data secure. It implies when to use and emphasizes 'Locally' to indicate no data is sent externally, but does not explicitly mention when not to use or name alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v1.0.2- Changed
reveal_text1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "properties": { + "text": { + "description": "The detokenized text with original values restored.", + "type": "string" + } + }, + "required": [ + "text" + ], + "type": "object" +}
- Changed
sanitize_file4 fields changed- removed
Input schema / properties / filePathRemoved value: -{ - "description": "Absolute path to the file to sanitize.", - "type": "string" -} - added
Input schema / properties / file_pathAdded value: +{ + "description": "Absolute path to the file to sanitize.", + "type": "string" +} - changed
Input schema / requiredPrevious value: -[ - "filePath" -]New value: +[ + "file_path" +] - changed
Output schema / (root)Previous value: -nullNew value: +{ + "properties": { + "text": { + "description": "The sanitized content of the file.", + "type": "string" + } + }, + "required": [ + "text" + ], + "type": "object" +}
- Changed
sanitize_text1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "properties": { + "text": { + "description": "The sanitized text output.", + "type": "string" + } + }, + "required": [ + "text" + ], + "type": "object" +}
3 tool updates
v1.0.1- First observed
reveal_text - First observed
sanitize_file - First observed
sanitize_text
TDQS
Scored across 3 tools
Each tool has a distinct purpose: sanitizing text, sanitizing files, and reversing the sanitization. No overlap in functionality.
All tool names follow a consistent verb_noun pattern with underscores (reveal_text, sanitize_file, sanitize_text).
With 3 tools, the server covers the core workflow of sanitization and reversal. Could benefit from an explicit session management tool, but is well-scoped overall.
The set covers sanitization for both text and files, plus the necessary reveal operation. Missing a tool to manage profiles or clear the session, but the core lifecycle is covered.
Maintenance
Related MCP Connectors
Detect and redact PII and secrets before text reaches an LLM, with reversible placeholders.
Redact PII from text before it reaches a model. Nothing stored, no third-party AI.
Detects and redacts PII (emails, phones, SSNs, names, addresses) from text. $0.02/call via x402.
Stateless PII redaction over MCP/REST. Free ≤1000 words or $0.01/call; file upload supported.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceLocal-first CLI and MCP server for redacting sensitive text before sharing logs, configs, and errors with AI tools.MIT
- FlicenseAqualityBmaintenancePII sanitization layer for autonomous AI agent pipelines. Detects and redacts emails, phone numbers, national IDs, private keys, and financial data before text reaches LLMs. Supports EN, ES (LATAM), PT (BR/PT), DE, JA.12-
- AlicenseAqualityDmaintenanceScans prompts for PII and masks or redacts sensitive data locally before sending to an LLM, supporting multiple anonymization modes.1MIT
- AlicenseNot gradedqualityAmaintenanceScans text and files for common secrets (AWS, GitHub, etc.) and redacts them to prevent credential leakage in AI-assisted development. Runs entirely locally with no telemetry.MIT