eu-pii-redact
Provides a LangChain integration that wraps models or chains to redact PII before LLM processing and restore original values in responses, and includes a transformer for redacting documents before indexing for RAG.
Integrates with the EU PII Redaction API hosted on RapidAPI to detect and redact European personal data such as names, national IDs, IBANs, cards, VAT numbers, emails, phones, and IPs, with restoration support.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@eu-pii-redactredact ./support/ticket.txt and save the redacted copy next to it"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
eu-pii-redact
Keep European personal data out of LLM prompts and logs, and put it back in the answer. Python client, LangChain integration and MCP server for the EU PII Redaction API.
Hi, my name is Karina Dahl. I was charged twice for order 88412.
Please refund to DK50 0040 0440 1162 43 or call me on +45 32 12 34 56.
Karinabecomes
Hi, my name is [PERSON_1]. I was charged twice for order 88412.
Please refund to [IBAN_1] or call me on [PHONE_1].
[PERSON_1]Names from evidence, not capital letters: titles, greetings, roles,
Name:labels, sign-offs, or a first name from 140,000 European names. "Best Practices" and "Court of Justice" stay as they are.National ID numbers from 21 EU/EEA countries, checked by check digit and birth date (personnummer, HETU, DNI/NIE, codice fiscale, NIR, PESEL, BSN, CPR, ...).
IBANs (mod-97), cards (Luhn), EU VAT numbers, emails, phones (with + prefix), IPs.
Every mention of a person shares one placeholder; order and invoice numbers stay untouched.
You need a RapidAPI key with a plan for the API. The
free plan gives 500 requests a month.
Set it as RAPIDAPI_KEY.
Python
pip install eu-pii-redactfrom eu_pii_redact import Redactor
with Redactor() as r: # reads RAPIDAPI_KEY
redaction = r.redact(ticket)
reply = ask_llm(redaction.text) # the model only sees placeholders
print(redaction.restore(reply)) # real values back in the answerredaction.entities lists each finding with its type, offsets in the original text and placeholder
(national IDs also carry country and scheme). Texts longer than 50,000 characters are sent in pieces
with consistent placeholders. r.redact(text, types=["EMAIL", "IBAN"], mode="mask") masks with asterisks.
Related MCP server: PII Redaction MCP Server
LangChain
pip install "eu-pii-redact[langchain]"from eu_pii_redact.langchain import protect, PIIRedactionTransformer
safe_llm = protect(llm) # any model or chain that takes a string
safe_llm.invoke("Draft a reply to Karina Dahl about her refund to DK50 0040 0440 1162 43")
# the model sees [PERSON_1] and [IBAN_1]; the returned message has the real values again
docs = PIIRedactionTransformer().transform_documents(docs) # e.g. before indexing for RAGMCP server
Lets an AI assistant redact text and, more importantly, files without reading them:
redact_file reads and writes the file itself and only reports counts, so the personal data
never enters the conversation. restore_file puts the values back locally.
Tool | What it does |
| Redact a text; returns the redacted text and entities (no original values) |
| Redact a UTF-8 file into a new file; writes a restore mapping (owner-only permissions) |
| Put the original values back into a file using that mapping, offline |
Claude Desktop (one click): download eu-pii-redact-0.1.0.mcpb from the
latest release, open it, and paste
your RapidAPI key when asked. The bundle sources are in mcpb/.
Claude Code
claude mcp add eu-pii-redact -e RAPIDAPI_KEY=your-key -- uvx eu-pii-redactClaude Desktop / Cursor / other hosts (mcpServers config)
{
"mcpServers": {
"eu-pii-redact": {
"command": "uvx",
"args": ["eu-pii-redact"],
"env": { "RAPIDAPI_KEY": "your-key" }
}
}
}Postman
Import postman/eu-pii-redaction.postman_collection.json, set the collection variable
rapidapi_key, and send: four requests with example responses and built-in tests.
Limits
Rule-based, which keeps it fast (about 0.1 s for 50,000 characters) and predictable: no street addresses yet, phone numbers need an international prefix, and a name without any cue whose surname is an English word ("Will Smith") is missed on purpose. Use it as a strong first layer, not as a compliance guarantee.
License
MIT for this client package. The API itself is a hosted service.
Available Tools
3 toolsredact_fileA
Redact a text file into a new file without showing its contents. Returns only counts per type, so the personal data never enters the conversation.
| Name | Required | Description | Default |
|---|---|---|---|
| types | No | Only these types; default all. | |
| overwrite | No | ||
| input_path | Yes | UTF-8 text file to redact (.txt, .md, .csv, .json, .log, .eml...). | |
| output_path | Yes | Where to write the redacted copy. | |
| save_mapping | No | Also write <output_path>.map.json (owner-only permissions) so restore_file can put the original values back later. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false and openWorldHint=true, so write behavior is already known. The description adds genuine value beyond them: it writes to a NEW file and returns only per-type counts so raw PII never enters context. It omits overwrite semantics and the save_mapping side effect, but those gaps are covered in the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly packed sentences with no filler; the core action and the privacy-preserving return behavior are both front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly explains what is returned (counts per type) and that raw contents are withheld. The main omission is the relationship to sibling tools and the save_mapping/restore_file workflow, which the schema only partially hints at.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 80%, with descriptions for types, input_path, output_path, and save_mapping, so the schema does the heavy lifting. The description adds no parameter-level detail (only 'counts per type' loosely gestures at the types param), so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: redact a text file into a new file, plus the key constraint that contents are not shown. However, it never explicitly distinguishes itself from the sibling redact_text, leaving the agent to infer the file vs. string boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The clause 'so the personal data never enters the conversation' implies when this tool is preferable, but there is no explicit when-to-use, when-not-to-use, or naming of redact_text/restore_file as alternatives. Usage is only indirectly implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
redact_textARead-only
Redact personal data in a text and return the redacted text plus the detected entities (type, position, placeholder). Original values are not returned.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | "label" replaces with [TYPE_n] placeholders, "mask" with asterisks. | label |
| text | Yes | Text to redact. | |
| types | No | Only these types; default all. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true. The description adds valuable behavioral context beyond that: it discloses the return shape (redacted text plus detected entities) and explicitly states that original values are not returned, which is important privacy and safety information.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences with zero waste. The purpose and output behavior are front-loaded, and every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the annotations, 100% schema coverage, and lack of an output schema, the description is complete enough for correct invocation. It explains what is returned and what is not returned, covering the most important behavioral gap for this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters are already documented, including the enum behavior for 'mode' and the type filtering for 'types'. The description adds no parameter-level detail beyond what the schema provides, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Redact') and resource ('personal data in a text'), and distinguishes itself from the sibling redact_file by scoping to text input. The return contents are also previewed, so an agent knows what it will get.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by 'in a text' versus the sibling redact_file, but the description never explicitly says when to choose this tool over alternatives or when not to use it. No exclusions or routing guidance are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
restore_fileA
Put the original values back into a file, locally (no API call). Returns only how many placeholders were restored, not the values.
| Name | Required | Description | Default |
|---|---|---|---|
| overwrite | No | ||
| input_path | Yes | File containing placeholders, e.g. an edited redacted file or an LLM answer saved to disk. | |
| output_path | Yes | Where to write the restored file. | |
| mapping_path | Yes | The .map.json written by redact_file. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false and openWorldHint=false; the description adds real behavioral value beyond that: the operation is entirely local ('no API call') and the return value is only a restore count, not the restored values. It does not cover overwrite/conflict behavior, but the output disclosure is a genuine addition.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, with the local/no-API constraint and the return-value caveat both front-loaded. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, the description supplies the missing return-value semantics (count only) and the execution model (local, no API), which is exactly the information an agent needs that structured fields do not carry.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 75%, so the schema already explains input_path, output_path, and mapping_path. The description adds nothing about parameter meaning and leaves 'overwrite' to the schema (where it has no description). Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Put the original values back into a file') with the inverse-of-redaction semantics clear from context. It does not explicitly name or differentiate itself from redact_file/redact_text, but the operation is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied (restore after a redacted file was edited or an LLM answer saved to disk) but never stated as when-to-use guidance, and no exclusions or prerequisites are given. Adequate but leaves the routing decision to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.1.0- First observed
redact_file - First observed
redact_text - First observed
restore_file
TDQS
Scored across 3 tools
Each tool has a clearly distinct target and operation: redact_text for text, redact_file for files, restore_file for reversing file redaction. There is no overlap in purpose; an agent can easily select the correct tool.
All tool names follow a consistent snake_case verb_noun pattern (redact_text, redact_file, restore_file). The convention is predictable and uniform.
Three tools are well-scoped for a focused PII redaction server: one for text, one for files, one for restoration. Each tool earns its place without redundancy.
Core redaction and file restoration are covered, but redact_text lacks a corresponding restore operation (unlike redact_file), and no batch or structured-data redaction exists. These are minor gaps given the narrow scope.
Maintenance
Related MCP Connectors
Detect and redact PII and secrets before text reaches an LLM, with reversible placeholders.
Redact PII from text before it reaches a model. Nothing stored, no third-party AI.
Detects and redacts PII (emails, phones, SSNs, names, addresses) from text. $0.02/call via x402.
Stateless PII redaction over MCP/REST. Free ≤1000 words or $0.01/call; file upload supported.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceLocal-first CLI and MCP server for redacting sensitive text before sharing logs, configs, and errors with AI tools.MIT
- FlicenseNot gradedqualityDmaintenanceEnables AI agents to redact PII from text, summarize redacted content, and manage custom redaction patterns across multiple languages.-
- AlicenseAqualityAmaintenanceSanitizes text and files by removing PII, secrets, and custom patterns locally before sending to LLMs, with optional reverse-scrubbing.3328 npm2Cryptographic Autonomy 1.0 (Combined Work Exception)
- AlicenseBqualityCmaintenanceEnables AI assistants to mask personal data in local Turkish documents by accepting a local file path (TXT, MD, DOCX, or text PDF) and returning pseudonymized text such as [KISI-1] and [TCKN-1] instead of the original values. Tools scan for PII types, produce masked output, and restore real values into a local file without ever returning plaintext to the model.41MIT