Skip to main content
Glama

obsify

Lassen Sie einen KI-Assistenten an sensiblen Dateien arbeiten, ohne dass deren Rohwerte jemals in den Kontext des Modells gelangen.

obsify ist ein lokaler, deterministischer MCP-Server. Das Frontier-Modell denkt über die Struktur nach – Schemata, synthetische Zwillinge, maskiertes Feedback – während deterministischer lokaler Code die Substanz berührt und nur maskierte, aggregierte Ergebnisse zurückgibt. Keine LLM-Aufrufe, kein Netzwerk zur Laufzeit: Erkennung erfolgt per Regex + Prüfsummen + Wörterbücher + Presidio's lokalem NER.

Es wird mit australischer Entitätsunterstützung (ABN / ACN / TFN, prüfsummenvalidiert) und einer labelgesteuerten Routing-Ebene ausgeliefert, die „Wann sollte der Assistent Rohdaten vermeiden“ zu einer deterministischen, erzwungenen Entscheidung und nicht zu einer Ermessensentscheidung macht.

Ehrlicher Geltungsbereich: run_on_real führt modellgeschriebenen Code in einer Best-Effort-lokalen Sandbox aus und maskiert dessen Ausgabe Best-Effort. Es ist kein Gefängnis. Lesen Sie SECURITY.md, bevor Sie es auf etwas richten, dessen Verlust Sie sich nicht leisten können. Geben Sie Aggregate zurück.

Warum

Die Einspeisung vertraulicher Dokumente in ein gehostetes LLM bedeutet, dass die Substanz Ihre Grenzen verlässt. Die üblichen Antworten sind „Verwenden Sie das LLM nicht“ oder „Vertrauen Sie dem Anbieter.“ obsify geht einen dritten Weg — Compute-to-Data: Bringen Sie den Code zu den Daten, nicht die Daten zum Modell.

  • Das Modell sieht das Schema einer Tabelle, nicht ihre Zeilen.

  • Das Modell entwickelt anhand eines synthetischen Zwillings (gefälschte Werte, echte Struktur).

  • Der Analysecode des Modells wird lokal ausgeführt; nur maskierte, aggregierte Ausgabe wird zurückgegeben.

Das Denken des Frontier-Modells bleibt erhalten. Nur seine Augen auf Rohwerte werden entfernt.

Related MCP server: MCP DB Results Anonymizer

Werkzeuge

Werkzeug

Was es tut

Rückgabe

scan_pii(path)

Durchsucht eine Datei/einen Ordner nach PII

Typen, Orte, Anzahlen — niemals Werte

make_synthetic_twin(path, out)

Originalgetreue Fälschung einer Excel-Arbeitsmappe

Schema-Zusammenfassung; Zwilling nach out geschrieben (Werte gefälscht, leak-verifiziert)

run_on_real(code, data_path)

Compute-to-Data: Führen Sie Ihren Code lokal gegen die echte Datei aus (gebunden an DATA_PATH)

Nur PII-maskierte, größenbegrenzte stdout/stderr — Aggregate zurückgeben

redact_text(text)

Maskiert PII in einem String zu <TYPE>-Token

Der geschwärzte String

verify_value_free(text, terms)

Fail-Closed-Prüfung, dass text keines von terms (oder deren Varianten) preisgibt

{"value_free": bool}

Unterstützte Dokumente: PDF (Text + Tabellen; komplexe Tabellen-Fallback via obsify[tables]), Excel .xlsx/.xlsm und Word .docx (Absätze + Tabellen). Nicht lesbare oder nicht unterstützte Dateien werden als explizite Hinweise/blinde Flecken ausgewiesen, niemals stillschweigend verworfen. (Noch keine OCR — gescannte/Bild- Seiten werden als abdeckungsarm gekennzeichnet, nicht transkribiert.)

Maskierung bekannter Entitäten (optional). Stellen Sie eine lokale .obsify.entities-Liste mit zu verbergenden Namen bereit; scan_pii / redact_text fangen sie deterministisch ab – sowie die Suffix-/Abkürzungsvarianten, die NER übersieht (BRIGHTWATER HLDGS P/L für Brightwater Holdings Pty Ltd) – als KNOWN_ENTITY. Die Liste bleibt lokal und gelangt niemals in den Kontext des Modells. Siehe docs/known_entities.md.

Demo

Testen Sie alle fünf Werkzeuge live gegen synthetische Daten mit dem offiziellen MCP Inspector:

python -m obsify.make_corpus --out ./corpus_demo
npx @modelcontextprotocol/inspector obsify-mcp

Rufen Sie scan_pii für ./corpus_demo/ledger.xlsx auf und bestätigen Sie, dass es nur Typen / Anzahlen / Orte zurückgibt – niemals Werte. Siehe docs/verifying.md.

Testen Sie es – synthetisches Korpus

Generieren Sie ein gefälschtes, aber realistisches Korpus (alles synthetisch; ABN/ACN/TFN sind prüfsummenvalidiert), das alle drei Formate umfasst, und richten Sie dann ein Werkzeug darauf:

pip install "obsify[demo]"                 # reportlab, for the sample PDFs
python -m obsify.make_corpus --out ./corpus_demo

Es schreibt eine mehrblättrige Excel-Tabelle (ein numerisches False-Positive-Minenfeld), ein PDF-Anschreiben (Prosatext + Probebilanz-Tabelle) und ein DOCX-Prüfvermerk (Absätze + Lieferantentabelle). Großartig, um die Reifen von scan_pii / make_synthetic_twin zu testen, ohne echte Daten zu berühren.

Installation und Ausführung als MCP-Server

Erfordert Python 3.11+. obsify spricht MCP über stdio – der Client startet es als lokalen Unterprozess; nichts wird remote gehostet. Registrieren Sie es bei jedem MCP-fähigen Client (Claude Desktop, Claude Code, Cursor, VS Code, …), indem Sie einen Block zur Konfiguration dieses Clients hinzufügen.

Empfohlen – Installation ohne Installation via uvx:

{ "mcpServers": { "obsify": { "command": "uvx", "args": ["obsify-mcp"] } } }

uvx holt obsify von PyPI und führt es bei Bedarf aus – keine dauerhafte Installation. Beim ersten Start laden Sie das spaCy NER-Modell (en_core_web_lg, ~560 MB) einmal herunter und zwischenspeichern es; dies holt ein öffentliches Modell und sendet keine Benutzerdaten (setzen Sie OBSIFY_AUTO_DOWNLOAD=0, um es zu verbieten und das Modell selbst zu installieren). Spätere Ausführungen sind sofort und vollständig offline.

Oder installieren Sie es (pip / pipx):

pipx install obsify        # isolated, on PATH  (or: pip install obsify)

Dann weisen Sie den Client auf den installierten Befehl:

{ "mcpServers": { "obsify": { "command": "obsify-mcp" } } }

Starten Sie den Client neu und die Werkzeuge erscheinen. Optionale Extras: obsify[tables] (komplexe Tabellen-PDF- Fallback via camelot + Ghostscript), obsify[compute] (pandas, praktisch innerhalb von run_on_real-Code).

PATH-Falle (Hauptgrund Nr. 1 für „Server verbindet nicht“): Der command muss im PATH aufgelöst werden können, den der Client sieht. Ein GUI-Client teilt möglicherweise nicht den PATH Ihrer venv. Lösungen: Verwenden Sie uvx/pipx (global auflösbar) oder geben Sie einen absoluten Pfad an – "/pfad/zu/.venv/bin/obsify-mcp" (macOS/Linux) oder "C:\\pfad\zu\\.venv\\Scripts\\obsify-mcp.exe" (Windows).

Aus diesem Repository (bevor es auf PyPI ist):

pip install "git+https://github.com/Formative-Sum41/obsify.git"   # gets `obsify-mcp` + `obsify`

Die Routing-Ebene – deterministisch, keine Ermessensentscheidung

Der schwierige Teil von „Hilf mir, aber lies die vertrauliche Datei nicht“ ist die Entscheidung, wann geschützt werden muss. obsify verlagert diese Entscheidung aus dem Modell in die Umgebung:

  1. .obsify.json – ein Label-Manifest, das Pfade klassifiziert (public / confidential / restricted).

  2. obsify.guard (ausgeführt als python -m obsify.guard) – ein PreToolUse-Guard, der das direkte Lesen einer gelabelten Datei blockiert (Exit 2) und den Assistenten an scan_pii / make_synthetic_twin / run_on_real weiterleitet.

  3. Eine Konvention (in CLAUDE.md), sodass der Assistent obsify bevorzugt, bevor er überhaupt auf den Guard trifft.

Richten Sie es mit einem Befehl ein:

obsify init [--dir PATH] [--with-claude-md]

obsify init ist von Grund auf nicht destruktiv – es besitzt genau eine Datei und gibt Ihnen Schnipsel für den Rest:

  • .obsify.json – obsify besitzt dies; init schreibt es (niemals ohne --force überschrieben).

  • .claude/settings.json – Ihre Datei: init gibt den PreToolUse-Hook-Block zum Einfügen aus, bearbeitet ihn nie (da er Code ausführt, ist die Registrierung Ihre Entscheidung).

  • CLAUDE.md – Ihre Datei: die Konvention ist opt-in. Standardmäßig wird sie ausgegeben; --with-claude-md hängt einen markierungsumschlossenen, idempotenten Block an, der Ihren Inhalt niemals überschreibt.

Vollständige Konvention: docs/obsify_routing.md.

Wie die Erkennung präzise bleibt

  • Prüfsummenvalidierte Identifikatoren. ABN/ACN/TFN-Kandidaten werden per Regex vorgeschlagen und durch ihre offiziellen Prüfsummen bestätigt, sodass eine zufällige Zahl niemals als Identifikator gemeldet wird.

  • Kontextabhängige IDs. Eine bloße Zahl wird nur dann als ABN/ACN/TFN akzeptiert, wenn ein Label-Wort („TFN“, „ABN“, „BSB“, …) in der Nähe ist – dies unterdrückt die False-Positive-Flut durch sequenzielle Journal-IDs in numerischen Tabellen.

  • Buchstabenlose / NER-mit-Ziffern-Unterdrückung. Reine Zahlen, Beträge, Daten und alphanumerische Codes werden nicht als Namen/Organisationen gekennzeichnet; echte Namen, E-Mails und Adressen (die Buchstaben enthalten) sind nicht betroffen. Validierte buchstabenlose PII bleibt ausgenommen: Prüfsummen-IDs (ABN/ACN/TFN/Medicare), Luhn-Karten, gültige IPs, BSB-bezogene Konten und Telefone (via Kontext oder Telefonform) – während ein Dezimalpunkt immer noch einen Betrag und kein Telefon kennzeichnet.

Gemessene Genauigkeit

obsify wird mit einem bewerteten Evaluierungs-Framework ausgeliefert (eval/ – beschriftetes synthetisches Korpus + Antwortschlüssel + Bewerter gegen den ausgelieferten Detektor, plus eine unabhängige Drittanbieter-Kreuzprüfung). Überschrift auf dem synthetischen Korpus: 100 % Recall bei erwarteten Erkennungselementen, 0 False Positives auf einer numerischen FP-Foltertabelle (mit einer gruppierten Zahlensperre), bloße kontextabhängige IDs korrekt unterdrückt. Unabhängige Kreuzprüfung vs. Microsoft presidio-research: EMAIL/IBAN 100 %, PERSON 94 %.

Das Framework hat sich bezahlt gemacht – es fand echte Fehler, die dann behoben wurden: Kreditkarten und Telefonnummern wurden stillschweigend durch den numerischen Rauschfilter unterdrückt (jetzt ausgenommen via Prüfsummen- validierung / Telefonform), und Medicare, IP, Geburtsdatum, AU-Pass und Führerschein hatten keinen Erkenner (jetzt hinzugefügt, prüfsummen- oder kontextabhängig). Vollständige Methode, Zahlen und verbleibende dokumentierte Lücken (SWIFT/BIC, Nicht-Geburtsdaten): eval/README.md.

Tests

pip install -e ".[dev]"
pytest tests/            # or run any file directly: python tests/test_obsify.py

Zwölf Suiten (73 Tests), ausgeführt in CI unter Linux + Windows / Python 3.11 + 3.12:

  • mcp-protocol – startet den echten Server über stdio und spricht MCP mit ihm (derselbe Weg, den ein Client wie Claude verwendet): bestätigt, dass alle fünf Werkzeuge mit gültigen Schemata registriert sind und dass Aufrufe durch JSON-RPC hin und zurück gehen – einschließlich scan_pii, das nur die Struktur, Ende-zu-Ende zurückgibt.

  • checksums – verankert an extern veröffentlichten ABN/ACN/TFN-Arbeitsbeispielen (gültig und korrupt), was die Generator↔Validator-Zirkularität durchbricht.

  • obsify / twin / redaction – die Datenschutzinvarianten: nur Strukturausgabe, leak-freie Zwillinge und eine Fail-Closed-Selbstprüfung.

  • precision – die False-Positive-Unterdrücker töten numerisches Tabellenrauschen, während echte Namen erhalten bleiben.

  • routing – die Block-/Erlaubnis-Klassifizierung des Guards und der nicht-destruktive Vertrag von obsify init.

  • corpus – das synthetische PDF+Excel+DOCX-Korpus Ende-zu-Ende: formatbezogene Erkennung, DOCX- Absatz+Tabellen-Extraktion und nur Strukturausgabe über jedes Format.

  • evaluation – das bewertete Framework als Regressionstor (Recall, Unterdrückung, FP-Folter, Lücken).

  • robustness – Graceful Degradation: korrupte/überdimensionierte/leere/verschachtelte/nicht unterstützte Eingaben stürzen nie ab und werden immer als Hinweise ausgewiesen.

  • model / variants – Logik zum automatischen Herunterladen des Modells beim ersten Start; Variantennormalisierung hinter verify_value_free.

Für interaktive Verifizierung (MCP Inspector) und die letzte Meile des Live-Clients siehe docs/verifying.md.

Lizenz

MIT – siehe LICENSE.

Available Tools

5 tools
make_synthetic_twinA

Generate a SYNTHETIC TWIN of a real Excel workbook at path, written to out. Schema (sheets, headers, column types, true row counts) is preserved; every data value is freshly FAKED — no real value is copied. Reason and write your analysis code against the twin; then run it on the real file with run_on_real. Returns the schema summary (safe shape).

ParametersJSON Schema
NameRequiredDescriptionDefault
outYes
pathYes
cap_rowsNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well. It discloses key behavioral traits: schema is preserved, every data value is freshly FAKED, no real value is copied, and it returns a safe schema summary. This gives the agent essential safety and data-handling context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three dense sentences, front-loaded with the core action and then efficient supplementary detail. There is no fluff or redundancy; every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the core function, the workflow pairing with run_on_real, and the return value (schema summary). It lacks any explanation of cap_rows and its potential effect on 'true row counts,' which would be a notable gap for a tool of moderate complexity. Overall, it is quite complete but not flawless.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does explain path and out (': path', 'written to out'), but it does not mention cap_rows at all. This leaves one of three parameters semantically opaque, so the compensation is partial.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb+resource: 'Generate a SYNTHETIC TWIN of a real Excel workbook at `path`, written to `out`.' It clearly differentiates from siblings by framing this as the twin-creation step and explicitly mentions run_on_real as the subsequent step for real-file execution.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit workflow guidance: 'Reason and write your analysis code against the twin; then run it on the real file with run_on_real.' This tells exactly when to use this tool and names the alternative (run_on_real) for the next phase, satisfying the 'when/when-not/alternatives' criterion.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

redact_textA

Return text with detected PII replaced by placeholders (e.g. , ). Deterministic; checksum-validated identifiers and context/precision rules apply so bare numbers are not over-masked.

entities is an optional PATH to a local .obsify.entities file of KNOWN names to hide; matches (incl. variants) are masked as . If omitted, a nearby .obsify.entities is auto-used.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYes
entitiesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It does well by disclosing determinism, checksum-validated identifiers, over-masking avoidance, and the entities-file auto-use behavior. Minor gaps remain around error handling or what happens when no PII is detected, but transparency is strong overall.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: the first sentence states the core operation, followed by key behavioral constraints and then the optional parameter explanation. Every sentence earns its place without unnecessary verbiage.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simple 2-parameter shape and the presence of an output schema, the description is highly complete. It covers the main transformation, important edge-case prevention (bare numbers), and the optional entities file behavior. The description is sufficient for an agent to invoke the tool correctly without needing further clarification.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does. It explains that `entities` is a path to a local `.obsify.entities` file, that matched names are masked as `<KNOWN_ENTITY>`, and that a nearby file is auto-used if omitted. This adds substantial meaning beyond the bare schema fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence clearly states the verb (redact), the resource (text), and the output format (PII replaced by placeholders), making the purpose immediately obvious. It also distinguishes itself from sibling tools like scan_pii and verify_value_free by explicitly conveying the redaction operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context about how the tool behaves and when the optional entities file applies, but it does not explicitly state when to prefer redact_text over sibling tools or when not to use it. There are no alternative tool comparisons or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_on_realA

COMPUTE-TO-DATA: execute your Python code LOCALLY against the real file at data_path (bound to the variable DATA_PATH in your code); the returned output is size-capped and best-effort PII-masked. The data never enters your context; substance never leaves. Return AGGREGATES (counts/sums/summaries) via print() — output masking is defense-in-depth, NOT a guarantee (NER can miss a name in a raw record), so never print raw records or identifiers. The masking field carries this caveat with the result. Network is disabled and a timeout applies.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYes
timeoutNo
data_pathYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and handles it well: output is size-capped, PII masking is best-effort and explicitly not a guarantee, network is disabled, a timeout applies, and execution is local. It also warns that raw records/identifiers should never be printed, adding important safety context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well organized: concept label, action, safety constraints, and usage guidance. Bolded callouts ('Return AGGREGATES...', 'best-effort') make key instructions easy to parse, and no sentence is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a code-execution tool with no output schema and no annotations, the description covers the essential operational surface: local execution, data binding, output size, masking caveat, aggregate printing, network isolation, and timeout. It is sufficient for an agent to invoke the tool safely and interpret the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It adds strong semantics for code (Python executed locally) and data_path (bound to DATA_PATH), but timeout is only indirectly covered by 'a timeout applies' and the schema's default, not explained as a configurable parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'COMPUTE-TO-DATA' and clearly states the tool executes Python code locally against a real file at data_path, binding it to DATA_PATH. This is a specific verb+resource pairing and is distinct from siblings like make_synthetic_twin, which implies synthetic data operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description strongly implies when to use it: when you need to compute over real data without pulling raw data into context ('data never enters your context'). It gives actionable guidance to print aggregates and avoid raw records, but it does not explicitly name alternative tools or state when-not-to-use conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scan_piiA

Scan a file or folder for PII and return TYPES + LOCATIONS + COUNTS only — never the detected values. Safe to surface to an LLM: it learns what PII exists and where, without the substance entering context. Recurses into subfolders; skips unreadable files and caps very large sheets, reporting both as notes.

entities is an optional PATH to a local .obsify.entities file (one name per line) of KNOWN names to hide; matches (incl. suffix/abbreviation variants) are reported as KNOWN_ENTITY. If omitted, a nearby .obsify.entities is auto-used. The names are read locally and never returned.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes
entitiesNo
max_cellsNo

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description provides thorough behavioral details: it returns only metadata (not values), skips unreadable files, caps very large sheets, and reads the entities file locally without returning the names. It also explains the automatic fallback for the entities file. This gives a clear picture of side effects and limitations, exceeding the typical level of transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with a clear first paragraph on functionality and a second on the entities parameter. It is concise enough to convey necessary details without fluff, though the entities explanation could be slightly tighter. The information is relevant and not redundant, earning a high score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description provides a high-level overview of the return value (types, locations, counts) without specifying the exact output format, which is acceptable given no output schema. It covers main behaviors (recursion, skipping, capping) and the entities file. It lacks explicit error handling or return structure details, but for a scan tool, the description sufficiently completes the context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds significant meaning for the 'entities' parameter by explaining its purpose, format, and default behavior. It indirectly touches on 'max_cells' by mentioning capping large sheets, but does not explicitly link it to the parameter. The 'path' parameter is self-explanatory given the context. Overall, it compensates for the lack of schema descriptions, though not perfectly for max_cells.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool scans files or folders for PII and returns only types, locations, and counts, never the values. It also mentions recursion, skipping unreadable files, and capping large sheets, which fully specifies the tool's function. This distinguishes it from sibling tools like redact_text or make_synthetic_twin.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not explicitly explain when to use this tool over its siblings. It implies usage for scanning and reporting PII metadata, and the safety note ('Safe to surface to an LLM') hints at a use case, but there is no direct comparison or guidance on choosing between tools. The behavior details (recursion, skipping) could inform usage, but explicit 'use when' instructions are absent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_value_freeA

Fail-closed check that text contains NONE of terms (nor their suffix-normalized / distinctive-token variants). Returns {"value_free": bool} with zero detail on what matched — for verifying an artifact before it leaves the perimeter.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYes
termsYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of behavioral clarity. It discloses the fail-closed behavior, the variant-matching behavior, and the deliberately detail-poor return shape. It does not explicitly state there are no side effects, but the read-only check nature is strongly implied.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded, stating the check first, then the return contract, then the intended target scenario. Every sentence contributes useful information without repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter boolean verification tool, the description is nearly complete: it names inputs, behavior, return value, and intended boundary context. It leaves minor edge-case behavior unspecified, but this does not materially hamper selection or invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has no parameter descriptions, but the description defines the core semantics: `text` is the artifact being verified and `terms` are the prohibited strings matched directly or through normalized variants. It adds meaningful algorithmic context beyond the bare schema, though it omits edge cases like empty terms behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a fail-closed verification check that `text` contains none of `terms` or their variants, giving a specific verb, resource, and scope. It distinguishes itself from sibling tools by being a boolean verification gate rather than a scanning or redaction operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly identifies the intended use case: verifying an artifact before it leaves the perimeter. It implies this is a pre-release/compliance gate rather than a diagnostic tool, and the zero-detail return further signals it is not for troubleshooting that needs matched context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 5 tool updatesv0.2.0
    • First observedmake_synthetic_twin
    • First observedredact_text
    • First observedrun_on_real
    • First observedscan_pii
    • First observedverify_value_free

TDQS

A4.4/5.0

Scored across 5 tools

Disambiguation5/5

Each tool has a clear, distinct purpose: make_synthetic_twin creates a fake dataset, run_on_real executes code against real data, scan_pii identifies PII locations, redact_text masks PII in text, and verify_value_free checks for forbidden terms. No two tools overlap in what they accomplish.

Naming Consistency4/5

Most tools follow a verb_noun pattern (make_synthetic_twin, scan_pii, redact_text, verify_value_free), but run_on_real breaks the pattern with a prepositional phrase. The style is consistent (all snake_case, verbs first) but the deviation is noticeable.

Tool Count4/5

With 5 tools, the count is well-scoped for a focused PII-handling server. Each tool covers a necessary step in the workflow, and the count is within the typical 3-15 range, though a few additional helpers could be justified (e.g., a check for twin accuracy).

Completeness4/5

The tool surface covers the core lifecycle: protect data (scan, redact, verify) and enable safe analysis (twin, run on real). Minor gaps exist, such as no tool to validate the synthetic twin's fidelity against the real file, and verify_value_free lacks a positive counterpart, but agents can work around these.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    Let LLMs analyze sensitive data safely by querying a tokenized, join-preserving copy of the database, with fail-closed PII scanning and provable numeric equivalence.
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Acts as an anonymizing proxy between AI agents and databases, detecting PII and replacing it with realistic fake data so agents never see real data.
    Apache 2.0
  • A
    license
    A
    quality
    B
    maintenance
    Self-hosted governance layer between an AI assistant and your data: allow/deny policy, deterministic PII masking, row caps, and a hash-chained audit log with an Ed25519-signed receipt for every access, verifiable offline.
    4
    481 npm
    3
    MIT