obsify
obsify
AIアシスタントが機密ファイルを扱う際、生の値がモデルのコンテキストに入ることを防ぎます。
obsifyはローカルで動作する決定論的なMCPサーバーです。フロンティアモデルは形状(スキーマ、合成ツイン、マスクされたフィードバック)について推論し、決定論的なローカルコードが実体に触れ、マスクされた集計結果のみを返します。LLM呼び出しはなく、実行時にネットワークも不要です。検出は正規表現+チェックサム+辞書+PresidioのローカルNERを使用します。
オーストラリアのエンティティサポート(ABN / ACN / TFN、チェックサム検証済み)と、ラベル駆動のルーティングレイヤーを備えており、「アシスタントがいつ生データを避けるべきか」を判断ではなく、決定論的で強制された決定にします。
正直な範囲:
run_on_realはモデルが書いたコードをベストエフォートのローカルサンドボックスで実行し、その出力をベストエフォートでマスクします。これは刑務所ではありません。漏洩させられないものを指定する前にSECURITY.mdをお読みください。集計結果を返してください。
なぜ
機密文書をホスト型LLMに送信すると、実体が境界外に出てしまいます。通常の答えは「LLMを使わない」か「プロバイダーを信頼する」です。obsifyは第三の道を取ります——計算からデータへ:データをモデルに持ってくるのではなく、コードをデータに持って行きます。
モデルはスプレッドシートのスキーマを見ますが、行は見ません。
モデルは合成ツイン(偽の値、実際の構造)に対して開発します。
モデルの分析コードはローカルで実行され、マスクされた集計出力のみが返ります。
フロンティアモデルの推論は保持されます。生の値に対する目だけが取り除かれます。
Related MCP server: MCP DB Results Anonymizer
ツール
ツール | 機能 | 戻り値 |
| ファイル/フォルダのPIIをスキャン | 種類、場所、カウント — 値は決して返さない |
| Excelワークブックの忠実な偽造 | スキーマの概要; |
| 計算からデータへ:実際のファイルに対してローカルでコードを実行( | PIIマスク済み、サイズ制限付きのstdout/stderrのみ — 集計結果を返す |
| 文字列内のPIIを | 編集済みの文字列 |
|
|
|
対応ドキュメント: PDF(テキスト+表;複雑な表はobsify[tables]でフォールバック)、Excel .xlsx/.xlsm、Word .docx(段落+表)。読み取り不能または未対応のファイルは、明示的な注記/盲点として表示され、静かに破棄されることはありません。(OCRは未対応 — スキャン/画像ページは低カバレッジとしてフラグされ、文字起こしはされません。)
既知エンティティのマスキング(オプション)。 ローカルの.obsify.entitiesリストに隠したい名前を指定すると、scan_pii / redact_textがそれらを決定論的に捕捉します — NERが見逃す接尾辞/略語のバリアント(BRIGHTWATER HLDGS P/L for Brightwater Holdings Pty Ltd)も — KNOWN_ENTITYとして。リストはローカルに留まり、モデルのコンテキストに入ることはありません。詳細はdocs/known_entities.mdを参照。
デモ
公式のMCP Inspectorを使って、5つのツールすべてを合成データに対してライブで試せます:
python -m obsify.make_corpus --out ./corpus_demo
npx @modelcontextprotocol/inspector obsify-mcp./corpus_demo/ledger.xlsxに対してscan_piiを呼び出し、種類/カウント/場所のみを返し、値は決して返さないことを確認してください。docs/verifying.mdを参照。
試してみる — 合成コーパス
偽造だが現実的なコーパス(すべて合成;ABN/ACN/TFNはチェックサム検証済み)を生成し、3つの形式すべてにわたって、ツールを指定します:
pip install "obsify[demo]" # reportlab, for the sample PDFs
python -m obsify.make_corpus --out ./corpus_demoこれにより、マルチシートのExcel台帳(数値の偽陽性の地雷原)、PDFのエンゲージメントレター(散文+試算表)、DOCXの監査メモ(段落+ベンダー表)が書き込まれます。実際のデータに触れずにscan_pii / make_synthetic_twinを試すのに最適です。
MCPサーバーとしてインストールして実行
Python 3.11+が必要です。obsifyはstdio上でMCPを話します — クライアントがローカルサブプロセスとして起動し、リモートでホストされるものはありません。MCP対応クライアント(Claude Desktop、Claude Code、Cursor、VS Code、…)に、そのクライアントの設定に1ブロック追加して登録します。
推奨 — uvxによるゼロインストール:
{ "mcpServers": { "obsify": { "command": "uvx", "args": ["obsify-mcp"] } } }uvxはPyPIからobsifyを取得し、オンデマンドで実行します — 恒久的なインストールは不要です。初回実行時、obsifyはspaCy NERモデル(en_core_web_lg、約560 MB)を一度ダウンロードしてキャッシュします。これは公開モデルを取得し、ユーザーデータは送信しません(OBSIFY_AUTO_DOWNLOAD=0を設定して禁止し、自分でモデルをインストールすることもできます)。以降の実行は瞬時で完全にオフラインです。
またはインストール(pip / pipx):
pipx install obsify # isolated, on PATH (or: pip install obsify)次に、クライアントをインストールされたコマンドに向けます:
{ "mcpServers": { "obsify": { "command": "obsify-mcp" } } }クライアントを再起動するとツールが表示されます。オプションの追加機能:obsify[tables](camelot + Ghostscriptによる複雑な表のPDFフォールバック)、obsify[compute](pandas、run_on_realコード内で便利)。
PATHの落とし穴(「サーバーが接続できない」原因の第1位):
commandはクライアントが見るPATHで解決可能でなければなりません。GUIクライアントはvenvのPATHを共有しない場合があります。修正方法:uvx/pipxを使用する(グローバルに解決可能)、または絶対パスを指定する —"/path/to/.venv/bin/obsify-mcp"(macOS/Linux)または"C:\\path\\.venv\\Scripts\\obsify-mcp.exe"(Windows)。
このリポジトリから(PyPIに公開される前):
pip install "git+https://github.com/Formative-Sum41/obsify.git" # gets `obsify-mcp` + `obsify`ルーティングレイヤー — 判断ではなく決定論的
「助けてくれ、でも機密ファイルは読まないで」という難しい部分は、いつ保護するかを決めることです。obsifyはその決定をモデルから環境に移します:
.obsify.json— パスを分類するラベルマニフェスト(public/confidential/restricted)。obsify.guard(python -m obsify.guardとして実行) — ラベル付きファイルの直接読み取りをブロックし(終了コード2)、アシスタントをscan_pii/make_synthetic_twin/run_on_realにリダイレクトするPreToolUseガード。規約(
CLAUDE.md内) — アシスタントがガードに引っかかる前にobsifyを優先するようにします。
1つのコマンドでセットアップ:
obsify init [--dir PATH] [--with-claude-md]obsify initは設計上非破壊的です — 正確に1つのファイルを所有し、残りのスニペットを提供します:
.obsify.json— obsifyが所有;initが書き込みます(--forceなしでは上書きされません)。.claude/settings.json— あなたのファイル:initはPreToolUseフックブロックを表示して貼り付け用に出力し、編集はしません(コードを実行するため、登録はあなたの判断です)。CLAUDE.md— あなたのファイル:規約はオプトインです。デフォルトでは表示します;--with-claude-mdはマーカーで囲まれた冪等なブロックを追加し、あなたのコンテンツを決して上書きしません。
完全な規約:docs/obsify_routing.md。
検出の精度を保つ方法
チェックサム検証済み識別子。 ABN/ACN/TFNの候補は正規表現で提案され、公式のチェックサムで確認されるため、ランダムな数字が識別子として報告されることはありません。
コンテキストが必要なID。 裸の数字は、ラベルワード(「TFN」、「ABN」、「BSB」、…)が近くにある場合にのみABN/ACN/TFNとして受け入れられます — これにより、数値台帳での連続仕訳IDの偽陽性の洪水を防ぎます。
文字なし / NER+数字の抑制。 純粋な数字、金額、日付、英数字コードは名前/組織としてフラグされません。実際の名前、メール、住所(文字を含む)は影響を受けません。検証済みの文字なしPIIは例外のまま:チェックサムID(ABN/ACN/TFN/Medicare)、Luhnカード、有効なIP、BSBに隣接する口座、電話番号(コンテキストまたは電話の形状による) — 小数点は依然として金額を示し、電話番号ではありません。
測定された精度
obsifyにはスコア付き評価ハーネス(eval/ — ラベル付き合成コーパス+解答キー+出荷時の検出器に対するスコアラー、さらに独立した第三者によるクロスチェック)が含まれています。合成コーパスの見出し:期待検出項目に対して100%再現率、数値FP拷問シート(グループ化数字ガード付き)で0偽陽性、裸のコンテキストゲートIDは正しく抑制。Microsoft presidio-researchとの独立クロスチェック:EMAIL/IBAN 100%、PERSON 94%。
ハーネスはその価値を証明しました — 実際の欠陥を発見し、その後修正されました: クレジットカードと電話番号が数値ノイズフィルターによって静かに抑制されていました(現在はチェックサム検証/電話形状によって例外)、Medicare、IP、生年月日、オーストラリアのパスポートと運転免許証には認識機能がありませんでした(現在は追加され、チェックサムまたはコンテキストでゲート)。完全な方法、数値、残っている文書化されたギャップ(SWIFT/BIC、非DOB日付):eval/README.md。
テスト
pip install -e ".[dev]"
pytest tests/ # or run any file directly: python tests/test_obsify.py12スイート(73テスト)、CIでLinux + Windows / Python 3.11 + 3.12で実行:
mcp-protocol — 実際のサーバーをstdioで起動し、MCPを話します(Claudeのようなクライアントが使用するのと同じパス):5つのツールすべてが有効なスキーマで登録され、呼び出しがJSON-RPCを介してラウンドトリップすることを確認 —
scan_piiが形状のみをエンドツーエンドで返すことを含む。checksums — 外部公開されたABN/ACN/TFNの実例(有効および破損)に固定され、ジェネレーター↔バリデーターの循環性を断ち切ります。
obsify / twin / redaction — プライバシーの不変条件:形状のみの出力、漏洩のないツイン、フェイルクローズの自己チェック。
precision — 偽陽性抑制機能が数値台帳ノイズを除去しつつ、実際の名前を保持することを確認。
routing — ガードのブロック/許可分類と
obsify initの非破壊的契約。corpus — 合成PDF+Excel+DOCXコーパスのエンドツーエンド:形式ごとの検出、DOCXの段落+表抽出、すべての形式での形状のみの出力。
evaluation — スコア付きハーネスを回帰ゲートとして(再現率、抑制、FP拷問、ギャップ)。
robustness — グレースフルデグラデーション:破損/巨大/空/ネスト/未対応の入力が決してクラッシュせず、常に注記として表示される。
model / variants — 初回実行時のモデル自動ダウンロードロジック;
verify_value_freeの背後にあるバリアント正規化。
インタラクティブな検証(MCP Inspector)とライブクライアントの最終確認については、docs/verifying.mdを参照。
ライセンス
MIT — LICENSEを参照。
Available Tools
5 toolsmake_synthetic_twinA
Generate a SYNTHETIC TWIN of a real Excel workbook at path, written to
out. Schema (sheets, headers, column types, true row counts) is preserved;
every data value is freshly FAKED — no real value is copied. Reason and write
your analysis code against the twin; then run it on the real file with
run_on_real. Returns the schema summary (safe shape).
| Name | Required | Description | Default |
|---|---|---|---|
| out | Yes | ||
| path | Yes | ||
| cap_rows | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well. It discloses key behavioral traits: schema is preserved, every data value is freshly FAKED, no real value is copied, and it returns a safe schema summary. This gives the agent essential safety and data-handling context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three dense sentences, front-loaded with the core action and then efficient supplementary detail. There is no fluff or redundancy; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the core function, the workflow pairing with run_on_real, and the return value (schema summary). It lacks any explanation of cap_rows and its potential effect on 'true row counts,' which would be a notable gap for a tool of moderate complexity. Overall, it is quite complete but not flawless.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does explain path and out (': path', 'written to out'), but it does not mention cap_rows at all. This leaves one of three parameters semantically opaque, so the compensation is partial.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb+resource: 'Generate a SYNTHETIC TWIN of a real Excel workbook at `path`, written to `out`.' It clearly differentiates from siblings by framing this as the twin-creation step and explicitly mentions run_on_real as the subsequent step for real-file execution.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit workflow guidance: 'Reason and write your analysis code against the twin; then run it on the real file with run_on_real.' This tells exactly when to use this tool and names the alternative (run_on_real) for the next phase, satisfying the 'when/when-not/alternatives' criterion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
redact_textA
Return text with detected PII replaced by placeholders (e.g.
, ). Deterministic; checksum-validated identifiers and
context/precision rules apply so bare numbers are not over-masked.
entities is an optional PATH to a local .obsify.entities file of KNOWN names
to hide; matches (incl. variants) are masked as . If omitted, a
nearby .obsify.entities is auto-used.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| entities | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It does well by disclosing determinism, checksum-validated identifiers, over-masking avoidance, and the entities-file auto-use behavior. Minor gaps remain around error handling or what happens when no PII is detected, but transparency is strong overall.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: the first sentence states the core operation, followed by key behavioral constraints and then the optional parameter explanation. Every sentence earns its place without unnecessary verbiage.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simple 2-parameter shape and the presence of an output schema, the description is highly complete. It covers the main transformation, important edge-case prevention (bare numbers), and the optional entities file behavior. The description is sufficient for an agent to invoke the tool correctly without needing further clarification.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does. It explains that `entities` is a path to a local `.obsify.entities` file, that matched names are masked as `<KNOWN_ENTITY>`, and that a nearby file is auto-used if omitted. This adds substantial meaning beyond the bare schema fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence clearly states the verb (redact), the resource (text), and the output format (PII replaced by placeholders), making the purpose immediately obvious. It also distinguishes itself from sibling tools like scan_pii and verify_value_free by explicitly conveying the redaction operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context about how the tool behaves and when the optional entities file applies, but it does not explicitly state when to prefer redact_text over sibling tools or when not to use it. There are no alternative tool comparisons or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_on_realA
COMPUTE-TO-DATA: execute your Python code LOCALLY against the real file at
data_path (bound to the variable DATA_PATH in your code); the returned output
is size-capped and best-effort PII-masked. The data never enters your context;
substance never leaves. Return AGGREGATES (counts/sums/summaries) via print() —
output masking is defense-in-depth, NOT a guarantee (NER can miss a name in a raw
record), so never print raw records or identifiers. The masking field carries this
caveat with the result. Network is disabled and a timeout applies.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | ||
| timeout | No | ||
| data_path | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and handles it well: output is size-capped, PII masking is best-effort and explicitly not a guarantee, network is disabled, a timeout applies, and execution is local. It also warns that raw records/identifiers should never be printed, adding important safety context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well organized: concept label, action, safety constraints, and usage guidance. Bolded callouts ('Return AGGREGATES...', 'best-effort') make key instructions easy to parse, and no sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a code-execution tool with no output schema and no annotations, the description covers the essential operational surface: local execution, data binding, output size, masking caveat, aggregate printing, network isolation, and timeout. It is sufficient for an agent to invoke the tool safely and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds strong semantics for code (Python executed locally) and data_path (bound to DATA_PATH), but timeout is only indirectly covered by 'a timeout applies' and the schema's default, not explained as a configurable parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'COMPUTE-TO-DATA' and clearly states the tool executes Python code locally against a real file at data_path, binding it to DATA_PATH. This is a specific verb+resource pairing and is distinct from siblings like make_synthetic_twin, which implies synthetic data operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description strongly implies when to use it: when you need to compute over real data without pulling raw data into context ('data never enters your context'). It gives actionable guidance to print aggregates and avoid raw records, but it does not explicitly name alternative tools or state when-not-to-use conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scan_piiA
Scan a file or folder for PII and return TYPES + LOCATIONS + COUNTS only — never the detected values. Safe to surface to an LLM: it learns what PII exists and where, without the substance entering context. Recurses into subfolders; skips unreadable files and caps very large sheets, reporting both as notes.
entities is an optional PATH to a local .obsify.entities file (one name per
line) of KNOWN names to hide; matches (incl. suffix/abbreviation variants) are
reported as KNOWN_ENTITY. If omitted, a nearby .obsify.entities is auto-used.
The names are read locally and never returned.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| entities | No | ||
| max_cells | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description provides thorough behavioral details: it returns only metadata (not values), skips unreadable files, caps very large sheets, and reads the entities file locally without returning the names. It also explains the automatic fallback for the entities file. This gives a clear picture of side effects and limitations, exceeding the typical level of transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with a clear first paragraph on functionality and a second on the entities parameter. It is concise enough to convey necessary details without fluff, though the entities explanation could be slightly tighter. The information is relevant and not redundant, earning a high score.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description provides a high-level overview of the return value (types, locations, counts) without specifying the exact output format, which is acceptable given no output schema. It covers main behaviors (recursion, skipping, capping) and the entities file. It lacks explicit error handling or return structure details, but for a scan tool, the description sufficiently completes the context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds significant meaning for the 'entities' parameter by explaining its purpose, format, and default behavior. It indirectly touches on 'max_cells' by mentioning capping large sheets, but does not explicitly link it to the parameter. The 'path' parameter is self-explanatory given the context. Overall, it compensates for the lack of schema descriptions, though not perfectly for max_cells.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool scans files or folders for PII and returns only types, locations, and counts, never the values. It also mentions recursion, skipping unreadable files, and capping large sheets, which fully specifies the tool's function. This distinguishes it from sibling tools like redact_text or make_synthetic_twin.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not explicitly explain when to use this tool over its siblings. It implies usage for scanning and reporting PII metadata, and the safety note ('Safe to surface to an LLM') hints at a use case, but there is no direct comparison or guidance on choosing between tools. The behavior details (recursion, skipping) could inform usage, but explicit 'use when' instructions are absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_value_freeA
Fail-closed check that text contains NONE of terms (nor their
suffix-normalized / distinctive-token variants). Returns {"value_free": bool}
with zero detail on what matched — for verifying an artifact before it leaves
the perimeter.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| terms | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of behavioral clarity. It discloses the fail-closed behavior, the variant-matching behavior, and the deliberately detail-poor return shape. It does not explicitly state there are no side effects, but the read-only check nature is strongly implied.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded, stating the check first, then the return contract, then the intended target scenario. Every sentence contributes useful information without repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter boolean verification tool, the description is nearly complete: it names inputs, behavior, return value, and intended boundary context. It leaves minor edge-case behavior unspecified, but this does not materially hamper selection or invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has no parameter descriptions, but the description defines the core semantics: `text` is the artifact being verified and `terms` are the prohibited strings matched directly or through normalized variants. It adds meaningful algorithmic context beyond the bare schema, though it omits edge cases like empty terms behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a fail-closed verification check that `text` contains none of `terms` or their variants, giving a specific verb, resource, and scope. It distinguishes itself from sibling tools by being a boolean verification gate rather than a scanning or redaction operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly identifies the intended use case: verifying an artifact before it leaves the perimeter. It implies this is a pre-release/compliance gate rather than a diagnostic tool, and the zero-detail return further signals it is not for troubleshooting that needs matched context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
v0.2.0- First observed
make_synthetic_twin - First observed
redact_text - First observed
run_on_real - First observed
scan_pii - First observed
verify_value_free
TDQS
Scored across 5 tools
Each tool has a clear, distinct purpose: make_synthetic_twin creates a fake dataset, run_on_real executes code against real data, scan_pii identifies PII locations, redact_text masks PII in text, and verify_value_free checks for forbidden terms. No two tools overlap in what they accomplish.
Most tools follow a verb_noun pattern (make_synthetic_twin, scan_pii, redact_text, verify_value_free), but run_on_real breaks the pattern with a prepositional phrase. The style is consistent (all snake_case, verbs first) but the deviation is noticeable.
With 5 tools, the count is well-scoped for a focused PII-handling server. Each tool covers a necessary step in the workflow, and the count is within the typical 3-15 range, though a few additional helpers could be justified (e.g., a check for twin accuracy).
The tool surface covers the core lifecycle: protect data (scan, redact, verify) and enable safe analysis (twin, run on real). Minor gaps exist, such as no tool to validate the synthetic twin's fidelity against the real file, and verify_value_free lacks a positive counterpart, but agents can work around these.
Maintenance
Related MCP Connectors
Deterministic runtime safety for AI agents: scan PII, gate tool actions, verify LLM output.
- SnipgetOAuthai.snipget
300+ deterministic data utilities for AI agents: validate, normalize, parse, match, redact.
Security gateway for AI agents: policy, approval, and audited execution, no secrets shared.
Safe, read-only Postgres and MySQL access for AI agents. Audit log + column-level controls.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceLet LLMs analyze sensitive data safely by querying a tokenized, join-preserving copy of the database, with fail-closed PII scanning and provable numeric equivalence.MIT
- AlicenseNot gradedqualityDmaintenanceActs as an anonymizing proxy between AI agents and databases, detecting PII and replacing it with realistic fake data so agents never see real data.Apache 2.0
- AlicenseAqualityBmaintenanceSelf-hosted governance layer between an AI assistant and your data: allow/deny policy, deterministic PII masking, row caps, and a hash-chained audit log with an Ed25519-signed receipt for every access, verifiable offline.4481 npm3MIT
- AlicenseAqualityBmaintenanceEnables safe interaction with cloud LLMs by redacting sensitive entities into reversible placeholders, enforcing deterministic egress policies with human approval, and rehydrating responses so real data never leaves the process.4MIT