anonymize-mcp
The anonymize-mcp server provides production-grade NLP tools wrapping LINDAT/ÚFAL services, specializing in anonymization, named entity recognition, morphological analysis, readability checking, text correction, and translation — primarily for Czech and 35+ other languages.
Anonymization & Pseudonymization: An 8-step pipeline covering 80+ PII patterns (phone numbers, IDs/RČ, IČO, IBAN for 30+ countries, EU VAT, emails, URLs, crypto addresses, API tokens, and international IDs). Includes regex pre-pass, NLP-based NER pre-pass for companies/institutions, MasKIT engine, stop-list false-positive filtering, and optional
placeholder_modefor deterministic, auditable replacements (e.g., OSOBA1, MESTO1). A zero-egress local mode (ANONYMIZE_MCP_LOCAL=1) enables fully offline processing.Named Entity Recognition (NER): Identifies persons, organizations, locations, dates, and more in Czech (rich CNEC 2.0 tagset) and 34+ other languages (PER/ORG/LOC via multilingual UNER), with automatic language detection and optional XML/vertical output.
Morphological Analysis: Tokenization, lemmatization, POS tagging, and dependency parsing via UDPipe 2 across 961 models and 35+ languages, with auto language detection.
Readability Checking: Czech text analysis via PONK with four feature sets — overall metrics (ARI, lexical diversity, etc.), active grammatical rules, lexical surprise distribution, and speech act classification. Can also generate highlighted HTML.
Text Correction: Czech spell checking (standard or strict), diacritics addition (useful for OCR/mobile text), or diacritics removal via Korektor.
Translation: Translates between 8 languages (Czech, English, French, German, Polish, Russian, Ukrainian, Hindi) via Charles Translator — 17 direct pairs with automatic English-pivot for indirect pairs, and a document mode for CS↔EN to preserve structure.
Detects and anonymizes Bitcoin addresses (Legacy/P2SH/Bech32/Taproot) for crypto/Web3 use cases
Provides machine translation between 8 languages via Charles Translator API
Detects and anonymizes Ethereum addresses as part of crypto PII detection
Detects and anonymizes GitHub Personal Access Tokens (PAT) in text
Detects and anonymizes Google API keys in text
Detects and anonymizes Monero addresses as part of crypto PII detection
Detects and anonymizes OpenAI API tokens in text as part of PII detection
Detects and anonymizes ORCID identifiers in academic/research contexts
Detects and anonymizes Slack tokens in text
Detects and anonymizes Stripe API tokens in text
Detects and anonymizes XRP addresses as part of crypto PII detection
anonymize-mcp
MCP server obalující NLP nástroje LINDAT / ÚFAL MFF UK — multilingvální NER + morfologie (35 jazyků auto-detect), production-grade anonymizace s 80+ PII patterny napříč 9 sektory + mezinárodním pokrytím (US/UK/DE/FR/IT/ES/PL/RU/IN, EU VAT 28 zemí, IBAN 30+ zemí, crypto, API tokeny), překlad mezi 8 jazyky (17 přímých párů + auto EN-pivot), čitelnost a korektura.
🔒 Nově (v0.10.0): zero-egress lokální mód — plně offline anonymizace, žádný text neopustí stroj. Pro GDPR / právní / zdravotnická data. Zapnutí:
ANONYMIZE_MCP_LOCAL=1.
🌐 Nechcete nic instalovat? Vyzkoušejte to online zdarma → anonymizace.js.org — anonymizace, NER, morfologie, korektor a překlad češtiny přímo v prohlížeči. Doprovodný web k tomuto MCP serveru.
Pouze pro nekomerční použití. Modely NameTag a UDPipe jsou pod CC BY-NC-SA. LINDAT API je bezplatné pro akademické a osobní použití. Pro komerční nasazení kontaktujte autory nástrojů a
ufal@ufal.mff.cuni.cz.
Neoficiální komunitní projekt — není provozován ani schválen ÚFAL MFF UK; wrapper kolem veřejných LINDAT API od nezávislého vývojáře. Historie názvů:
ufal-mcp→wrapper-mcp(v0.8.0, na žádost ÚFAL) →anonymize-mcp(v0.9.0). Pokud máte nainstalovaný deprecated balíčekwrapper-mcp, přejděte napip install anonymize-mcp— je to tentýž projekt.
Co umí
Tool | Backend | K čemu |
| NER pro CZ (bohatý CNEC 2.0 tagset) + 34 dalších jazyků (UNER PER/ORG/LOC) s auto-detekcí | |
| MasKIT + NameTag | Production-grade pseudonymizace: regex pre-pass přes 80+ PII patternů — CZ + international (IBAN 30+ zemí, EU VAT 28, US SSN/EIN, DE/UK/FR/IT/ES/PL/RU/IN ID, crypto, API tokeny). Opt-in |
| Tokenizace, lemmatizace, POS tagging, závislostní parse — auto-detect 35 jazyků | |
| Čitelnost CZ — 4 feature sety: metrics + rules + lexical surprise + speech acts | |
| CZ spell checker + auto-doplnění/odstranění diakritiky | |
| Překlad mezi 8 jazyky (CZ/EN/FR/DE/PL/RU/UK/HI), 17 přímých párů + auto EN-pivot |
Podporované jazyky — NER + morfologie (35 jazyků, auto-detect)
🇨🇿 CZ · 🇸🇰 SK · 🇬🇧 EN · 🇩🇪 DE · 🇫🇷 FR · 🇮🇹 IT · 🇪🇸 ES · 🇵🇹 PT · 🇳🇱 NL
🇵🇱 PL · 🇭🇺 HU · 🇷🇴 RO · 🇸🇮 SL · 🇧🇬 BG · 🇬🇷 EL · 🇭🇷 HR · 🇷🇸 SR · 🇺🇦 UK · 🇷🇺 RU
🇫🇮 FI · 🇱🇹 LT · 🇱🇻 LV · 🇪🇪 ET · 🇩🇰 DA · 🇸🇪 SV · 🇳🇴 NO (Bokmål + Nynorsk)
🇨🇳 ZH · 🇦🇪 AR · 🇹🇷 TR · 🇻🇳 VI · 🇮🇳 HI · 🇮🇱 HE · 🇯🇵 JA · 🇰🇷 KO · 🇹🇭 TH
Related MCP server: pii-anonymizer
Pro koho je tohle (sektory + use cases)
Stress-tested napříč 9 sektory na 12.7KB cross-sektorovém spisu — výsledek 94/94 unique PII chyceno v jednom volání. Plus international corpus 17/17 (US/UK/DE/FR/IT/ES/PL/RU/IN + crypto + akademické + fleet):
Sektor | Use case | PII které MCP zvládne |
⚖️ Právo | Anonymizace spisu před AI review, GDPR compliance | Jména, RČ, adresy, č.j., sp.zn., IBAN, OP, datovky |
🏥 Medicína | Propouštěcí zprávy pro výzkum, statistika hospitalizací | RČ, IČZ, č. pojištěnce, kontakty lékaře — klinické kódy MKN-10 zachované |
🎓 Věda / akademie | Peer review, citace v publikaci | ORCID, Researcher ID, e-maily kolegů, granty |
💳 Bankovnictví | Compliance, výpisy do AI, vykazování | Č.ú., karta, IBAN, VS/KS/SS, header výpisu |
🏠 Reality / katastr | Anonymizace výpisů z KN, smluv | LV, parcely, k.ú., vlastník + RČ + adresa |
🚗 Pojišťovny | Likvidace škod, AI analýza | VIN, SPZ, č. pojistky, TP, OP, RČ pojištěného |
📜 Notáři | Notářské zápisy pro AI summary | NZ, OP, datovka notáře, sp. zn. |
📚 Studijní oddělení | Potvrzení o studiu, statistika studentů | UČO, studijní č., ISIC, kontakty studenta |
🔬 Výzkum / NGO | Anonymizace korpusu pro etiku výzkumu | Vše výše + zachování klinických/právních kódů |
Plus 35 jazyků v multilingvální stack (legal docs SK/EN/DE/PL/UK/RU/FR/HI/ES/IT/AR + 24 dalších otestovány na NER+morfologii, auto EN-pivot pro překlad mimo přímé Charles páry). CJK jména (čínská/japonská) maskována od v0.8.4.
Sektor #10 — International
Use case | PII které MCP zvládne |
🌍 US/UK/DE/FR/IT/ES/PL/RU/IN dokumenty | SSN, NIN, Steuer-ID, NIR, Codice Fiscale, DNI, PESEL, Aadhaar, PAN, cestovní pasy (8 jazyků) — auto bez |
💰 Crypto/Web3 outreach, smart contracts | Bitcoin (Legacy/P2SH/Bech32/Taproot), Ethereum, Monero, XRP, TRON |
🔐 DevOps logs / API key leak detection | OpenAI, Anthropic, OpenRouter, GitHub PAT, AWS, Google, Slack, Stripe tokeny |
🏢 Cross-border B2B | Foreign companies (SARL/SAS/GmbH/AG/Ltd/LLC/Inc/SpA/SL/Sp. z o.o.) + EU VAT (28 zemí) + IBAN (30+ zemí) |
Instalace
Z PyPI (doporučeno):
pip install anonymize-mcpNebo ze source:
git clone https://github.com/Buggy1111/anonymize-mcp.git
cd anonymize-mcp
pip install -e .Registrace v MCP klientovi
anonymize-mcp je standardní MCP server (stdio transport). Po registraci a restartu klienta máš k dispozici 6 nástrojů:
mcp__anonymize__extract_entities— multilingvální NER (35 jazyků auto-detect)mcp__anonymize__anonymize— production-grade pseudonymizace CZ (regex pre-pass + stop-list + placeholder mode)mcp__anonymize__analyze_morphology— morfologie 35 jazyků auto-detect (UDPipe 961 modelů)mcp__anonymize__check_readability— čitelnost CZ (4 feature sety)mcp__anonymize__correct_text— spell check + diakritika CZmcp__anonymize__translate_text— překlad mezi 8 jazyky
Claude Code (terminál)
claude mcp add anonymize -s user -- anonymize-mcpClaude Desktop
Starší Claude Desktop (Mac .app z anthropic.com, Windows .exe installer):
Edituj ~/Library/Application Support/Claude/claude_desktop_config.json (Mac)
nebo %APPDATA%\Claude\claude_desktop_config.json (Windows):
{
"mcpServers": {
"anonymize": {
"command": "anonymize-mcp"
}
}
}Nová Claude Desktop (Microsoft Store / appx package, "Cowork" UI): k 05/2026 podporuje pouze remote MCP servery přes HTTP URL. Lokální stdio MCP servery jako anonymize-mcp zde přidat nelze.
Na Windows může být
anonymize-mcp.exemimo PATH (typickyC:\Python\Python3xx\Scripts\anonymize-mcp.exe). V configu pak použij plnou cestu.
OpenAI Codex CLI (autorem netestováno)
Edituj ~/.codex/config.toml:
[mcp_servers.anonymize]
command = "anonymize-mcp"Cursor (autorem netestováno)
Edituj .cursor/mcp.json v projektu (nebo globálně ~/.cursor/mcp.json):
{
"mcpServers": {
"anonymize": {
"command": "anonymize-mcp"
}
}
}Windsurf, Cline, Zed, VS Code Copilot Agent (autorem netestováno)
Stejný mcpServers JSON formát — viz dokumentace daného klienta. command: "anonymize-mcp" (případně absolutní cesta).
Použití
V Claude Code stačí napsat například:
Anonymizuj text z
dokument.mdv placeholder_mode a vrať mi čistou verzi.
Vytáhni z dokumentu všechny osoby, soudy a č.j.
Klient přinesl ukrajinský dokument — přelož mi ho do češtiny, najdi entity a zanalyzuj morfologii.
Projeď moje podání přes PONK — vrať aktivovaná gramatická pravidla.
Klient mi posílá text bez diakritiky z mobilu — doplň diakritiku přes Korektor.
Autor
anonymize-mcp napsal Michal Bürgermeister (@Buggy1111, michalbugy12@gmail.com) — nezávislý vývojář z ČR.
Wrapper kolem skvělých nástrojů ÚFAL MFF UK — bez NameTag, MasKIT, UDPipe, PONK, Korektor a Charles Translator by tenhle MCP server neexistoval. Díky celému ÚFAL týmu (Jana Straková, Milan Straka, Jiří Mírovský, Barbora Hladká, Silvie Cinková a další) za roky práce na production-grade NLP nástrojích pro češtinu.
Issues, PR a feedback jsou vítané na github.com/Buggy1111/anonymize-mcp.
Licence
Tento nástroj má MIT licenci (viz LICENSE).
Pod ním jsou čtyři samostatné nástroje, každý s vlastní licencí:
Komponenta | Autoři | Licence software | Licence modelů |
NameTag 3 | Jana Straková, Milan Straka | MPL 2.0 | CC BY-NC-SA (NON-commercial) |
UDPipe | Milan Straka, Jana Straková | MPL 2.0 | CC BY-NC-SA (NON-commercial) |
MasKIT | Jiří Mírovský, Barbora Hladká | MPL 2.0 | (rule-based) |
PONK | Jiří Mírovský, Silvie Cinková, Barbora Hladká + autoři podaplikací: Ivan Kraus, Arnold Stanovský, Jan Černý, Ivana Kvapilíková, Tomáš Polák, Silvie Cinková | MPL 2.0 | (rule-based + UDPipe → CC BY-NC-SA) |
Důležité: tento nástroj nevolá lokální instalaci, ale veřejné API služby (lindat.mff.cuni.cz, quest.ms.mff.cuni.cz). Bezplatné pro akademické a osobní použití. Hromadný / placený / produkční traffic vyžaduje explicitní souhlas autorů a provozovatele API.
Bezpečnost
V cloudovém módu posíláš text na externí server (
quest.ms.mff.cuni.cz,lindat.mff.cuni.cz). Před odesláním citlivých dat nejdřív projeď text přesanonymize.Pro plně privátní zpracování použij zero-egress lokální mód (níže) — žádný text neopustí tvůj stroj.
Zero-egress / on-prem mód 🔒
Anonymizace kompletně lokálně — žádné volání externího API, žádný text neopustí stroj. Pro GDPR / právní / zdravotnická data, kde citlivý obsah nesmí ven.
pip install "anonymize-mcp[local]" # přidá ufal.nametag (lokální NER)
python -m anonymize_mcp.local_backend # jednorázově stáhne model (~31 MB)
ANONYMIZE_MCP_LOCAL=1 anonymize-mcp # spusť server v lokálním móduV Claude Code stačí přidat env proměnnou k registraci:
claude mcp add anonymize -s user -e ANONYMIZE_MCP_LOCAL=1 -- anonymize-mcpJak to funguje: anonymize přeskočí MasKIT API a anonymizuje přes lokální regex pre-pass (80+ vzorů) + NameTag NER běžící v procesu (ufal.nametag + CNEC 2.0 model). Jména, města, instituce, telefony, IČO, RČ, č.j. atd. se nahradí placeholdery (OSOBA1, MESTO1, TELEFON1…) bez jediného síťového volání.
Konfigurace (env):
Proměnná | Význam |
| Zapne zero-egress mód |
| Vlastní cesta k modelu (jinak auto-download do |
| Zakáže auto-download (model musíš dodat ručně) |
| Vědomě povolí cloudové tooly i v lokálním módu (jinak odmítnuté, viz níže) |
Co je v lokálním módu lokální (v0.10.1): anonymize i extract_entities běží plně offline (lokální CNEC 2.0 NER; multilingvální model vyžaduje API, tool na to upozorní warningem). Ostatní tooly (translate_text, correct_text, check_readability, analyze_morphology) by text poslaly na ÚFAL API — proto jsou v zero-egress módu odmítnuté s vysvětlující chybou; vědomě je povolíš přes ANONYMIZE_MCP_LOCAL_ALLOW_CLOUD=1.
Hardened setup (air-gapped/auditované stroje): model si předstáhni předem (python -m anonymize_mcp.local_backend — ověřuje se SHA-256) a server nasaď s ANONYMIZE_MCP_NO_DOWNLOAD=1 — pak proces nikdy neotevře žádné síťové spojení.
Tradeoff: lokální NameTag 1 (CNEC 2.0, CC BY-NC-SA, non-commercial) je o něco jednodušší než cloudový NameTag 3 a vynechává MasKIT rule-engine — výměnou za nulový egress.
Použité API (6 LINDAT REST endpointů)
POST https://lindat.mff.cuni.cz/services/nametag/api/recognize— NERPOST https://lindat.mff.cuni.cz/services/udpipe/api/process— morfologiePOST https://lindat.mff.cuni.cz/services/korektor/api/correct— spell checkPOST https://lindat.mff.cuni.cz/services/translation/api/v2/models/{src-tgt}— překladPOST https://quest.ms.mff.cuni.cz/maskit/api/process— anonymizacePOST https://quest.ms.mff.cuni.cz/ponk/api/process— čitelnost
Vývoj
python3 -m venv .venv
source .venv/bin/activate
pip install -e .
# Testy (272 offline testů; síťové: pytest -m network)
pip install -e ".[test]"
pytest -m "not network"Release proces
PyPI publish je automatický přes Trusted Publisher (OIDC).
# Bump version v pyproject.toml a src/anonymize_mcp/__init__.py
git commit -am "release: v0.X.0"
git tag v0.X.0
git push origin main --tagsAvailable Tools
6 toolsanalyze_morphologyA
Tokenizuje, lemmatizuje a označuje slovní druhy pomocí UDPipe 2.
Pro každý token vrací **lemma** (základní tvar), **UPOS** (universal POS tag),
**morphological features** (pád, rod, číslo, čas...) a volitelně závislostní
parse (head + deprel) nebo character ranges (offsety do originálu).
UDPipe 2 podporuje **961 modelů** pro téměř všechny jazyky světa.
Auto-detect (default) rozezná: czech, slovak, ukrainian, russian, polish,
german, english, french (via heuristics).
Hodí se pro:
- Fulltextové vyhledávání v právních textech (lemma "soud" matchuje "soudu/soudem/soudy")
- Filtrování podle slovních druhů (jen substantiva, jen verba)
- Detekce pasivních konstrukcí (Voice=Pass)
- Vícejazyčné dokumenty (UA legal aid, EN smlouvy, DE Klage…)
Args:
text: Vstupní text.
model: UDPipe model alias. ``auto`` (default) detekuje jazyk podle markerů.
Explicit: ``czech``, ``slovak``, ``english``, ``ukrainian``, ``russian``,
``polish``, ``german``, ``french``, atd. — 961 modelů celkem.
include_parse: True = vrátí závislostní parse (head, deprel).
include_ranges: True = vrátí ``token_range`` (char offsets do originálu).
Užitečné pro inline highlighting nebo mapování token → text position.
Returns:
``sentences``, ``model``, ``token_count``, ``sentence_count``,
``detected_language`` (jen u auto).
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| model | No | auto | |
| include_parse | No | ||
| include_ranges | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations present, so description fully handles transparency. Discloses UDPipe 2 usage, 961 model support, auto-detect language list, and behavior of optional parameters (include_parse, include_ranges) with practical examples.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Well-structured with bullet points and clear sections. Slightly lengthy due to language list, but overall concise and front-loaded with core purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given presence of output schema, description need not detail return values. Covers input, use cases, and optional outputs adequately. Mentions output fields (sentences, model, token_count, etc.) for agent to infer structure.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage 0% requires description to explain parameters. Description explains 'text', 'model' (with defaults and examples), and both boolean flags with usage context. Lacks constraints (e.g., text length), but otherwise sufficient.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states tool performs tokenization, lemmatization, and POS tagging using UDPipe 2. Lists specific output fields (lemma, UPOS, morphological features, optional parse, ranges). Distinguishes from sibling tools like extract_entities and translate_text by focusing on morphological analysis.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit use cases (fulltext search, POS filtering, passive detection, multilingual documents) and mentions language auto-detection. Does not explicitly state when not to use, but the use cases are clear and differentiate from siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
anonymizeA
Production-grade pseudonymizace českých právních textů (v0.6.0).
Pipeline (8 kroků):
1. **Regex pre-pass** (`regex_pre_pass=True`) — strukturovaná PII
(telefon, IČO, RČ, č.j., sp. zn., e-mail, URL, PSČ, SPZ, IBAN, DIČ,
OP, datovka) se anonymizuje **PŘED** MasKITem, aby nebyly fragmentovány.
Telefon "777 123 456" se anonymizuje **celý** jako jeden blok TELEFON1.
2. **Strict wrapper pre-pass** (`strict=True`) — NameTag najde
firmy/úřady/instituce, které MasKIT vynechává nebo fragmentuje,
a anonymizuje je sentinely → FIRMA1, INSTITUCE1.
3. **MasKIT** — pseudonymizace zbývajících PII (jména, adresy, ...).
4. **Stop-list filter** (`stop_list_filter=True`) — MasKIT občas
chybně nahrazuje běžná slova ("stát" → "UniAgentury", "sporu" →
"Pardubic"). Wrapper detekuje a vrátí originál, přidá warning.
5. **Restore sentinely** → finální placeholdery (TELEFON1, FIRMA1, ...).
6. **Fragmentation warnings** — detekce známých MasKIT problémů.
7. **Type classification** — NameTag dohledá typ entity pro každou náhradu.
8. **Placeholder mode** (`placeholder_mode=True`) — místo MasKIT náhodných
fake names (`Jan Novák`) použij deterministické `OSOBA1`, `OSOBA2`...,
`MESTO1`, `ULICE1`, ... S dedupingem: Jiří × 15× v textu → OSOBA1 × 15×.
**Reprodukovatelné** (stejný vstup → stejný výstup) a **auditovatelné**.
Args:
text: Vstupní text (čeština).
output: Formát výstupu — ``txt`` (default), ``html``, ``conllu``.
keep_mapping: Když True, vrátí mapping. **POZOR**: pokud má text
dál opustit důvěrné prostředí, mapping vypni!
classify_types: NameTag dohledá typ entity. Default ``True``.
strict: Wrapper pre-pass na firmy/úřady. Default ``True``.
placeholder_mode: ⭐ **NEW v0.6.0** — deterministic placeholdery
místo MasKIT fake names. Pro reprodukovatelnost a auditovatelnost.
regex_pre_pass: Default ``True``. Strukturovaná PII regexem PŘED MasKITem.
stop_list_filter: Default ``True``. Rollback MasKIT false positives.
Returns:
``anonymized`` (čistý text), ``raw`` (MasKIT raw), ``replacements``
(list s ``original``, ``placeholder``, ``type``, ``source``),
``warnings``, ``sources`` ({maskit, wrapper-regex, wrapper-strict,
wrapper-placeholder}).
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| output | No | txt | |
| keep_mapping | No | ||
| classify_types | No | ||
| strict | No | ||
| placeholder_mode | No | ||
| regex_pre_pass | No | ||
| stop_list_filter | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It details the 8-step pipeline, each parameter's effect (e.g., placeholder_mode for reproducibility, stop_list_filter for false positives), and return structure. This is highly transparent about the tool's behavior and side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with numbered steps and bullet points, but it is verbose. Every sentence provides value, and it is front-loaded with the tool's purpose. However, it could be slightly more concise without losing clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (8 parameters, multi-step pipeline, return types), the description is remarkably complete. It explains pipeline stages, parameter interactions, output fields, and even version-specific features. The output schema is described in the Returns section, fulfilling completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description compensates by explaining each parameter (text, output, keep_mapping, etc.) in plain language, including defaults and behavioral impact. This adds significant meaning beyond the schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it is a 'Production-grade pseudonymizace českých právních textů' (pseudonymization of Czech legal texts). It provides a specific verb (pseudonymize), resource (Czech legal texts), and detailed pipeline. This distinguishes it from sibling tools like analyze_morphology or translate_text.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the pipeline and parameter behaviors, implying usage for Czech legal text anonymization. It mentions caveats like turning off keep_mapping if text leaves confidential environment. However, it does not explicitly state when not to use this tool or provide alternatives among siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_readabilityA
Analyzuje čitelnost českého textu pomocí PONK — 4 feature sety (v0.7.0).
PONK byl navržen pro úřední komunikaci s občany. V0.7.0 wrapper vystavuje
všechny 4 jeho feature sety, ne jen metriky:
1. **Overall metrics** — ARI (years of education needed), Verb Distance,
Activity, Lexical diversity. (Always returned.)
2. **Grammatical rules** (``include_rules=True``) — list pravidel které se
v textu aktivovala. Každé pravidlo má český název a popis. Aktuálně PONK
detekuje: Nedostatek sloves, Přemíra podstatných jmen, Dlouhé věty,
Sloveso příliš daleko v klauzi, ...
3. **Lexical surprise** (``include_lexical_surprise=True``) — distribuce
sémantické překvapivosti slov (1=běžné, 16=velmi vzácné/odborné).
Vrátí summary: kolik slov je common / surprising / very_surprising.
4. **Speech acts** (``include_speech_acts=True``) — typy vět (Situace,
Kontext, Postup, Proces, Podmínky, Doporučení, Odkazy, Prameny).
Args:
text: Vstupní text.
input_format: ``txt`` (default), ``md``, ``docx``.
include_rules: Default ``True``. List aktivovaných gramatických pravidel.
include_lexical_surprise: Default ``True``. Distribuce vzácnosti slov.
include_speech_acts: Default ``True``. Typy vět/řečové akty.
include_highlighted_html: Default ``False`` (úspora bandwidthu — HTML
má 100+ KB). Zapni pro vizualizační report/PDF.
Returns:
``metrics``, ``counts``, ``version``, ``processing_time_s``,
+ volitelné ``rules`` (list), ``lexical_surprise`` (dict),
``speech_acts`` (dict), ``highlighted_html`` (str).
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| input_format | No | txt | |
| include_rules | No | ||
| include_lexical_surprise | No | ||
| include_speech_acts | No | ||
| include_highlighted_html | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses behavioral traits: it mentions PONK is designed for official communication, notes the highlight HTML is large (100+ KB) for bandwidth consideration, and explains the optional feature sets. No annotations exist, so the description carries the burden; it is fairly transparent about what is returned and the cost of certain options.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with numbered lists and bullet points, front-loading the purpose. It is reasonably concise for the complexity, though slightly verbose with repeated enumeration of feature sets.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (6 parameters, 1 required, and a rich output schema), the description covers the returned fields and optional components. It is complete enough for an agent to understand outputs, though it could mention that the output schema exists but doesn't need to detail it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description compensates by explaining each parameter's purpose, default values, and the effect of include flags. It also warns about the highlight HTML size. This adds significant meaning beyond the raw schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Analyzuje čitelnost českého textu pomocí PONK' and lists four specific feature sets (overall metrics, grammatical rules, lexical surprise, speech acts). It distinguishes itself from sibling tools like analyze_morphology, anonymize, etc., by focusing on readability analysis.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains what the tool does but does not provide explicit guidance on when to use it versus siblings or when not to use it. Usage is implied through the purpose, but no when/when-not or alternative recommendations are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
correct_textA
Opraví český text pomocí Korektor — pravopis nebo doplnění diakritiky.
Use cases pro legal-tech:
- **spellcheck** (default) — kontrola pravopisu před odesláním podání
- **spellcheck_strict** — agresivnější (až 2 edits/word)
- **diacritics** — doplnění diakritiky do textu bez ní
(OCR výstupy, emaily, mobilní zprávy: ``Jan Vzorek bez hacku`` → ``Jan Vzorek bez háčků``)
- **strip** — odstranění diakritiky (např. pro URL slugy nebo legacy systémy)
Pozor: CZ-only. Modely jsou z roku 2013, vlastní jména a nová slova mohou
mít omezenou přesnost.
Args:
text: Vstupní český text.
mode: Operace — ``spellcheck`` (default), ``spellcheck_strict``,
``diacritics``, ``strip``.
Returns:
``corrected`` (upravený text), ``model``, ``mode``, ``changed`` (bool).
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| mode | No | spellcheck |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of behavioral disclosure. It reveals the tool is CZ-only, models are from 2013, and accuracy may be limited for proper names and new words. It also lists return fields (corrected, model, mode, changed), providing transparency about output.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with sections for use cases, args, and returns. It is front-loaded with the main purpose and every sentence adds value without waste. The length is appropriate for the complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-parameter tool with no annotations, the description covers purpose, parameters, modes, return values, and limitations. It lacks details on error handling or idempotency, but overall it is sufficiently complete for an AI agent to understand and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema has 0% description coverage; the description compensates for the 'mode' parameter by detailing each enum value's purpose and examples. However, for 'text', it only states 'Vstupní český text' without additional semantics, which is minimal but clear given the tool name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool corrects Czech text using Korektor for spelling or diacritics. It distinguishes from sibling tools by listing specific modes (spellcheck, diacritics, strip) and use cases relevant to legal-tech, with examples like OCR output enhancement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly explains each mode's purpose and when to use them (e.g., diacritics for OCR outputs, strip for URL slugs). It also cautions that the tool is CZ-only and notes model limitations from 2013, providing clear context for proper usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extract_entitiesA
Rozpozná pojmenované entity pomocí NameTag 3 — CZ i 30+ dalších jazyků.
Pro **češtinu** používá bohatý CNEC 2.0 tagset (osoba/firma/instituce/
PSČ/telefon/datum/…). Pro ostatní jazyky (SK, EN, DE, FR, IT, ES, PT,
NL, PL, HU, UK, RU, RO, SL, BG, EL, HR, SR, FI, LT, LV, ET, DA, SV,
NO, ZH, AR, TR, VI, HI a další) přepne na multilingvální UNER model
s tagsetem PER/ORG/LOC.
Args:
text: Vstupní text (UTF-8).
model: ``auto`` (default) — automatická detekce CZ vs non-CZ.
``czech`` vynutí CNEC 2.0 (bohatý CZ tagset). ``multilingual``
vynutí UNER PER/ORG/LOC pro non-CZ. Lze zadat i plné jméno
modelu (např. ``nametag3-multilingual-onto-250203``).
fix_romance: Default True. Pro PT/ES texty oprava typického
UNER bugu, kdy se "X de Place" zaeviduje celé jako PER —
wrapper rozdělí na PER + LOC a generuje warning.
include_xml: Default ``False``. Inline XML s ``<ne type="...">`` tagy
pro HTML highlighting (extra API call).
include_vertical: Default ``False``. Tabulkový formát ``id\ttype\ttext``
(extra API call).
Returns:
``entities`` (list s ``type``, ``label``, ``text``, ``tokens``,
``nested``), ``model``, ``count``, ``warnings``,
``detected_language`` (jen u ``auto``),
``xml`` (jen pokud ``include_xml``),
``vertical`` (jen pokud ``include_vertical``).
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| model | No | auto | |
| fix_romance | No | ||
| include_xml | No | ||
| include_vertical | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It discloses that include_xml and include_vertical cause extra API calls, and fix_romance generates warnings. It does not mention rate limits or auth, but the behavioral traits are adequately described.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with an overview, parameter details, and return values. It is informative without being overly verbose, though a slightly more streamlined presentation could improve conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (5 parameters, 1 required, output schema present), the description covers all parameters and return values comprehensively. It also notes extra API calls for certain options, ensuring the agent has sufficient context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema coverage, the description fully explains each parameter: text, model (with values and defaults), fix_romance (function and default), include_xml, include_vertical. It provides concrete details beyond the schema, such as model name examples and the effect of fix_romance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it recognizes named entities using NameTag 3, supporting Czech with a rich tagset and over 30 other languages with a multilingual model. It distinguishes itself from sibling tools like anonymize or correct_text by focusing specifically on entity extraction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use different model options (auto, czech, multilingual) and mentions fix_romance for specific languages. However, it does not explicitly state when not to use this tool or compare it to alternatives, leaving some ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
translate_textA
Přeloží text přes Charles Translator (LINDAT) — 8 jazyků, 17 přímých párů + auto EN-pivot pro nepřímé páry.
Podporované jazyky: ``cs`` (čeština), ``en``, ``fr``, ``de``, ``pl``,
``ru``, ``uk`` (ukrajinština), ``hi`` (hindština).
**Přímé páry** (17): cs↔en (+doc), cs↔uk, cs↔ru, en↔fr, en↔de, en↔ru,
en↔pl, en→hi (jednosměrně).
**EN-pivot** (auto): pro páry mimo seznam (typicky de→cs, pl→cs, fr→cs,
fr→de) wrapper provede 2 volání ``src→en→tgt`` a vrátí finální překlad
+ warning + ``pivot=True``. Doc-mode v pivotu nepodporován.
Klíčové páry pro legal-tech:
- ``cs-en`` / ``en-cs`` — anglické sumáře, mezinárodní komunikace
- ``doc-cs-en`` / ``doc-en-cs`` (s ``document_mode=True``) — celé dokumenty
se zachovanou strukturou odstavců
- ``cs-uk`` / ``uk-cs`` — ukrajinští klienti / legal aid pro UA migranty
- ``cs-ru`` / ``ru-cs`` — ruskojazyční klienti
- ``de-cs`` / ``pl-cs`` / ``fr-cs`` — automatický EN-pivot pro EU sousedy
Pozor: SK ↔ CZ pár v Charles Translatoru chybí. SK je auto-alias na CS
(mutual intelligibility). HI lze jen jako tgt (en→hi), ne jako src.
Charles Translator umí vlastní jména zachovat v originále — užitečné
pro legal: *"Jan Vzorek podal žalobu u Krajského soudu v Ostravě."*
→ *"Jan Vzorek filed a lawsuit at the Krajský soud v Ostrava."*
Args:
text: Text k překladu (UTF-8).
src: Zdrojový jazyk (default ``cs``).
tgt: Cílový jazyk (default ``en``).
document_mode: True pro doc mode (cs↔en only). Zachová strukturu.
Returns:
``translated`` (přeložený text), ``src``, ``tgt``, ``pair``
(skutečně použitý model name), ``document_mode``, ``input_chars``,
``output_chars``.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| src | No | cs | |
| tgt | No | en | |
| document_mode | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description carries full burden. It discloses EN-pivot behavior (two calls, warning, pivot flag), document mode constraints, and preservation of proper names. The return structure is described. Mutation aspect is not explicit but translation is inherently non-destructive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is verbose (over 300 words) with extensive lists and examples. While well-structured with Markdown and headings, it could be shortened without losing essential information. Some detail is redundant for an AI agent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of language pairs and pivot logic, the description covers all necessary behavioral and constraint details. Output schema exists, so return values are documented separately. Missing an explicit example call, but overall complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema has 0% description coverage, so description must compensate. It explains each parameter: 'text' (UTF-8), 'src' (default cs), 'tgt' (default en), and 'document_mode' (only for cs↔en, preserves structure). This adds significant meaning beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool translates text using Charles Translator, specifies 8 languages and 17 direct pairs, and distinguishes it from sibling tools like analyze_morphology or anonymize. The verb 'přeloží' (translate) and resource are explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides extensive usage guidance: supported languages, direct vs. pivot pairs, document mode limitations, and key pairs for legal-tech. It implicitly advises when to avoid pivot (no doc-mode) and mentions missing SK↔CZ. However, it doesn't explicitly contrast with sibling tools, though purpose is distinct enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v0.1.0- First observed
analyze_morphology - First observed
anonymize - First observed
check_readability - First observed
correct_text - First observed
extract_entities - First observed
translate_text
TDQS
Scored across 6 tools
Each tool has a unique, well-defined purpose: morphological analysis, anonymization, readability checking, text correction, entity extraction, and translation. There is no overlap in functionality, and descriptions clearly differentiate them.
Most tools follow a verb_noun pattern (analyze_morphology, check_readability, correct_text, extract_entities, translate_text), while 'anonymize' is a standalone verb. This minor inconsistency does not hinder understanding.
With 6 tools, the server covers essential NLP tasks for Czech legal texts without being over- or under-scoped. Each tool earns its place.
The toolset provides a comprehensive pipeline for processing legal texts (analysis, correction, anonymization, translation, entity extraction, readability). Minor gaps like summarization exist, but core workflows are well covered.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Stateless PII redaction over MCP/REST. Free ≤1000 words or $0.01/call; file upload supported.
Toxicity, sentiment, NER, PII detection, and language identification tools
Translate MCP — wraps LibreTranslate API (https://libretranslate.com/)
Detect and redact Norwegian PII (fodselsnummer, names, health data) in text and PDFs.
Related MCP Servers
- AlicenseAqualityDmaintenanceProvides local anonymization of Czech legal documents by replacing sensitive entities with pseudonyms to ensure privacy during LLM interactions. It allows users to safely process documents like contracts and judgments by keeping original data offline and facilitating local deanonymization.55MIT
- FlicenseNot gradedqualityDmaintenanceMCP server for automatic detection and redaction of PII in text, with anonymization and deanonymization capabilities, all local processing.1-
- FlicenseNot gradedqualityCmaintenanceMCP server for anonymizing and deanonymizing PII through the Pseudora API, enabling safe sharing of sensitive text with AI assistants.-
- AlicenseAqualityAmaintenanceLocal pseudonymisation MCP server that detects PII in text, replaces it with opaque tokens before sending to cloud LLMs, and restores tokens afterward.2891MIT