Skip to main content
Glama

eu-pii-redact

Keep European personal data out of LLM prompts and logs, and put it back in the answer. Python client, LangChain integration and MCP server for the EU PII Redaction API.

Hi, my name is Karina Dahl. I was charged twice for order 88412.
Please refund to DK50 0040 0440 1162 43 or call me on +45 32 12 34 56.
Karina

becomes

Hi, my name is [PERSON_1]. I was charged twice for order 88412.
Please refund to [IBAN_1] or call me on [PHONE_1].
[PERSON_1]
  • Names from evidence, not capital letters: titles, greetings, roles, Name: labels, sign-offs, or a first name from 140,000 European names. "Best Practices" and "Court of Justice" stay as they are.

  • National ID numbers from 21 EU/EEA countries, checked by check digit and birth date (personnummer, HETU, DNI/NIE, codice fiscale, NIR, PESEL, BSN, CPR, ...).

  • IBANs (mod-97), cards (Luhn), EU VAT numbers, emails, phones (with + prefix), IPs.

  • Every mention of a person shares one placeholder; order and invoice numbers stay untouched.

You need a RapidAPI key with a plan for the API. The free plan gives 500 requests a month. Set it as RAPIDAPI_KEY.

Python

pip install eu-pii-redact
from eu_pii_redact import Redactor

with Redactor() as r:                       # reads RAPIDAPI_KEY
    redaction = r.redact(ticket)
    reply = ask_llm(redaction.text)         # the model only sees placeholders
    print(redaction.restore(reply))         # real values back in the answer

redaction.entities lists each finding with its type, offsets in the original text and placeholder (national IDs also carry country and scheme). Texts longer than 50,000 characters are sent in pieces with consistent placeholders. r.redact(text, types=["EMAIL", "IBAN"], mode="mask") masks with asterisks.

Related MCP server: PII Redaction MCP Server

LangChain

pip install "eu-pii-redact[langchain]"
from eu_pii_redact.langchain import protect, PIIRedactionTransformer

safe_llm = protect(llm)   # any model or chain that takes a string
safe_llm.invoke("Draft a reply to Karina Dahl about her refund to DK50 0040 0440 1162 43")
# the model sees [PERSON_1] and [IBAN_1]; the returned message has the real values again

docs = PIIRedactionTransformer().transform_documents(docs)   # e.g. before indexing for RAG

MCP server

Lets an AI assistant redact text and, more importantly, files without reading them: redact_file reads and writes the file itself and only reports counts, so the personal data never enters the conversation. restore_file puts the values back locally.

Tool

What it does

redact_text

Redact a text; returns the redacted text and entities (no original values)

redact_file

Redact a UTF-8 file into a new file; writes a restore mapping (owner-only permissions)

restore_file

Put the original values back into a file using that mapping, offline

Claude Desktop (one click): download eu-pii-redact-0.1.0.mcpb from the latest release, open it, and paste your RapidAPI key when asked. The bundle sources are in mcpb/.

Claude Code

claude mcp add eu-pii-redact -e RAPIDAPI_KEY=your-key -- uvx eu-pii-redact

Claude Desktop / Cursor / other hosts (mcpServers config)

{
  "mcpServers": {
    "eu-pii-redact": {
      "command": "uvx",
      "args": ["eu-pii-redact"],
      "env": { "RAPIDAPI_KEY": "your-key" }
    }
  }
}

Postman

Import postman/eu-pii-redaction.postman_collection.json, set the collection variable rapidapi_key, and send: four requests with example responses and built-in tests.

Limits

Rule-based, which keeps it fast (about 0.1 s for 50,000 characters) and predictable: no street addresses yet, phone numbers need an international prefix, and a name without any cue whose surname is an English word ("Will Smith") is missed on purpose. Use it as a strong first layer, not as a compliance guarantee.

License

MIT for this client package. The API itself is a hosted service.

Available Tools

3 tools
redact_fileA

Redact a text file into a new file without showing its contents. Returns only counts per type, so the personal data never enters the conversation.

ParametersJSON Schema
NameRequiredDescriptionDefault
typesNoOnly these types; default all.
overwriteNo
input_pathYesUTF-8 text file to redact (.txt, .md, .csv, .json, .log, .eml...).
output_pathYesWhere to write the redacted copy.
save_mappingNoAlso write <output_path>.map.json (owner-only permissions) so restore_file can put the original values back later.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false and openWorldHint=true, so write behavior is already known. The description adds genuine value beyond them: it writes to a NEW file and returns only per-type counts so raw PII never enters context. It omits overwrite semantics and the save_mapping side effect, but those gaps are covered in the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly packed sentences with no filler; the core action and the privacy-preserving return behavior are both front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly explains what is returned (counts per type) and that raw contents are withheld. The main omission is the relationship to sibling tools and the save_mapping/restore_file workflow, which the schema only partially hints at.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 80%, with descriptions for types, input_path, output_path, and save_mapping, so the schema does the heavy lifting. The description adds no parameter-level detail (only 'counts per type' loosely gestures at the types param), so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: redact a text file into a new file, plus the key constraint that contents are not shown. However, it never explicitly distinguishes itself from the sibling redact_text, leaving the agent to infer the file vs. string boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The clause 'so the personal data never enters the conversation' implies when this tool is preferable, but there is no explicit when-to-use, when-not-to-use, or naming of redact_text/restore_file as alternatives. Usage is only indirectly implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

redact_textA
Read-only

Redact personal data in a text and return the redacted text plus the detected entities (type, position, placeholder). Original values are not returned.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNo"label" replaces with [TYPE_n] placeholders, "mask" with asterisks.label
textYesText to redact.
typesNoOnly these types; default all.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true. The description adds valuable behavioral context beyond that: it discloses the return shape (redacted text plus detected entities) and explicitly states that original values are not returned, which is important privacy and safety information.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly written sentences with zero waste. The purpose and output behavior are front-loaded, and every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the annotations, 100% schema coverage, and lack of an output schema, the description is complete enough for correct invocation. It explains what is returned and what is not returned, covering the most important behavioral gap for this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters are already documented, including the enum behavior for 'mode' and the type filtering for 'types'. The description adds no parameter-level detail beyond what the schema provides, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Redact') and resource ('personal data in a text'), and distinguishes itself from the sibling redact_file by scoping to text input. The return contents are also previewed, so an agent knows what it will get.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by 'in a text' versus the sibling redact_file, but the description never explicitly says when to choose this tool over alternatives or when not to use it. No exclusions or routing guidance are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

restore_fileA

Put the original values back into a file, locally (no API call). Returns only how many placeholders were restored, not the values.

ParametersJSON Schema
NameRequiredDescriptionDefault
overwriteNo
input_pathYesFile containing placeholders, e.g. an edited redacted file or an LLM answer saved to disk.
output_pathYesWhere to write the restored file.
mapping_pathYesThe .map.json written by redact_file.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false and openWorldHint=false; the description adds real behavioral value beyond that: the operation is entirely local ('no API call') and the return value is only a restore count, not the restored values. It does not cover overwrite/conflict behavior, but the output disclosure is a genuine addition.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero filler, with the local/no-API constraint and the return-value caveat both front-loaded. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, the description supplies the missing return-value semantics (count only) and the execution model (local, no API), which is exactly the information an agent needs that structured fields do not carry.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, so the schema already explains input_path, output_path, and mapping_path. The description adds nothing about parameter meaning and leaves 'overwrite' to the schema (where it has no description). Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Put the original values back into a file') with the inverse-of-redaction semantics clear from context. It does not explicitly name or differentiate itself from redact_file/redact_text, but the operation is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied (restore after a redacted file was edited or an LLM answer saved to disk) but never stated as when-to-use guidance, and no exclusions or prerequisites are given. Adequate but leaves the routing decision to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv0.1.0
    • First observedredact_file
    • First observedredact_text
    • First observedrestore_file

TDQS

A4.2/5.0

Scored across 3 tools

Disambiguation5/5

Each tool has a clearly distinct target and operation: redact_text for text, redact_file for files, restore_file for reversing file redaction. There is no overlap in purpose; an agent can easily select the correct tool.

Naming Consistency5/5

All tool names follow a consistent snake_case verb_noun pattern (redact_text, redact_file, restore_file). The convention is predictable and uniform.

Tool Count5/5

Three tools are well-scoped for a focused PII redaction server: one for text, one for files, one for restoration. Each tool earns its place without redundancy.

Completeness4/5

Core redaction and file restoration are covered, but redact_text lacks a corresponding restore operation (unlike redact_file), and no batch or structured-data redaction exists. These are minor gaps given the narrow scope.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    A
    maintenance
    Sanitizes text and files by removing PII, secrets, and custom patterns locally before sending to LLMs, with optional reverse-scrubbing.
    3
    328 npm
    2
    Cryptographic Autonomy 1.0 (Combined Work Exception)
  • A
    license
    B
    quality
    C
    maintenance
    Enables AI assistants to mask personal data in local Turkish documents by accepting a local file path (TXT, MD, DOCX, or text PDF) and returning pseudonymized text such as [KISI-1] and [TCKN-1] instead of the original values. Tools scan for PII types, produce masked output, and restore real values into a local file without ever returning plaintext to the model.
    4
    1
    MIT