Skip to main content
Glama

claude-pii-guard

Keep raw personal data out of the Anthropic API while keeping Claude Code's Auto mode hands-off. Instead of detecting and masking on the wire, this repo (1) denies Claude Code access to where personal data lives, and (2) runs the work that needs that data locally through a compute-to-data MCP server (safe-data) that returns only aggregates, schemas, and pseudonymized results. Docs are in Japanese; the code and config are generic.

Claude Code の Auto モードをそのままに、生の個人情報が Anthropic API に届かない構成を作るためのツール群。骨格は「検出して消す」ではなく「Claude が到達できる場所に生データを置かない」+「個人情報が要る処理は手元で実行して結果だけ返す」。

構成

層

成果物

強制力

届く経路を限定する / 読ませない(P0)

settings/phase0.settings.json + merge_settings.py

あり(deny / sandbox)

見せずに処理する(P1)

src/safe_data/ — MCP サーバー safe-data

あり(唯一の口)

ソース側で断つ(P1)

db/*.sql(ビュー + 専用ロール)、safeify-csv(ETL)

あり

使い方を教える

skills/safe-analysis/SKILL.md、settings/CLAUDE.snippet.md

なし(体験のため)

最後の網(P2a・実装済み)

mods/pii-guard/(Claude Code Mod)+ src/pii_guard/(ローカルデーモン:辞書 + 規則 + vault) — docs/phase2-mod-design.md

あり(fail-closed)

Phase 2b/2c(未実装)

NER、ui.render 復元、support-intake / slack-safe

—

Related MCP server: MCP PowerBI Controller

safe-data のツール

ツール

返すもの

返さないもの

schema_describe(view?)

claude スキーマのビューの列 + 合成サンプル

実データ

files_describe()

~/PII/inbox のファイル id・推定スキーマ・PII フラグ

値、ファイル名

sql_run(sql, max_rows)

1 文の SELECT の結果(行数上限・n<11 抑止)

SELECT *、書込み、設定読取

py_run(script, inputs)

ネットワーク無しコンテナで実行した結果 JSON(8KB)

保護対象の値を含む結果、print 出力

fixture_make(source, n)

ビュー/ファイルと同じ形の合成行

—

read_masked(file_id, limit)

inbox のテキストを pii-guard で擬似化した先頭 N 行

デーモン停止時は何も返さない

pii-guard(Mod + デーモン)

# 辞書: 自社の顧客 CSV か DB から(氏名・かな・メール・電話・ID の列を指定)
uv run pii-guard-dict --from-csv ~/PII/exports/customers.csv --id user_id --name name --kana kana --email mail --phone tel
# デーモンを LaunchAgent として常駐
scripts/pii-guard-launchd.sh install
# Mod を読み込んで起動(常用するなら ~/.claude/settings.json の env に CLAUDE_CODE_PLUGIN_DIRS を置く)
claude --plugin-dir ~/work/claude-pii-guard/mods/pii-guard

イベント

すること

失敗時

prompt.submit

打ち込んだ本文と context を /mask

{ drop }(送信しない)

session.append

保存される全行(ツール結果・失敗結果・添付)を /mask、画像/文書は除外

行の本文を伏せ字に差し替え

prompt.context

CLAUDE.md・メモリ・git 状態を /mask

context を空にする

tool.call(Bash と MCP)

引数に辞書一致・患者 ID があれば { deny }

{ deny }

turn.step

汚染中はモデルを呼ばない

—

ui.render

画面に描くときだけ札を実値に戻す(モデル・転記は札のまま)

札のまま表示

/pii

status / untaint / reload / off / on

—

vault は Fernet で暗号化(鍵は ~/.config/safe-data/vault.key、0600)。pii-guard-scan が毎晩 3:15 に転記・メモリ・tool-results を走査し、保護対象の値が残っていれば種類と件数を ~/.config/safe-data/scan.log に出す(値は出さない)。検出は辞書(自社 DB 由来、異体字・かな・ローマ字の揺れを吸収)→ 規則(メール・電話・〒・文脈付きマイナンバー・文脈付き生年月日・住所・患者 ID)。site で強さを変える(コード行は辞書と ID のみ)。claude plugin test(17 件)と pytest(49 件)で検証。GitHub Actions が push ごとに両方を回す。

セットアップ

scripts/install.sh            # uv sync、鍵生成、inbox 作成、claude mcp add、Skill リンク、runner イメージ
scripts/day0.sh               # 環境の記録とカナリア作成
python3 settings/merge_settings.py          # dry run
python3 settings/merge_settings.py --apply  # ~/.claude/settings.json に合流(バックアップあり)

その後 scripts/phase0_acceptance.md を通し、DB は db/README.md。

元に戻すときは scripts/uninstall.sh(launchd・MCP 登録・Skill リンク・settings.json のバックアップ復元・~/.config/safe-data・Docker イメージを確認しながら外す。~/PII の中身・DB・claude.ai のコネクタは触らず、手順を表示する)。

開発

uv sync
uv run pytest -q
uv run safe-data-mcp --check   # 実効設定

自分の環境に合わせる

  • settings/phase0.settings.json の deny には Claude Desktop 内蔵ツール(mcp__computer-use、mcp__Claude_Browser__*、mcp__ccd_*)や GitKraken / AWS プラグインのツール名が入っている。存在しないツールへの deny は何にも一致しないだけで害はない。自分の環境の名前は /mcp で確認し、settings/connectors.local.json に写す。

  • src/safe_data/config.py の PII 列パターンは日本語のサポート業務向け。自分のスキーマに合わせて config.toml の pii_columns.patterns で上書きする。

  • db/03_views.sql は例。列名を自分のテーブルに合わせる。

  • docs/design.md は筆者環境を前提にした設計記録(Skill 名・ドメインは例)。

注意(colima / Docker Desktop)

py_run はホストのファイルをコンテナに bind mount する。VM 型の Docker(colima)はホームディレクトリしか VM に共有していないので、pii_inbox と analysis_dir は $HOME 配下に置く(/tmp 配下は空ディレクトリとしてマウントされ、スクリプトが読めない)。safe-data は --mount type=bind を使うので、共有されていないパスは明示的なエラーになる。

設計上の約束

  • ツールはエラーを isError ではなく {ok:false, error} で返す(失敗出力が文脈に流れない)

  • 戻り値は 40,000 字未満(Claude Code の 50,000 字ディスク退避を発生させない)

  • 理由文に値を書かない(種類と件数だけ)

  • hook は fail-open なので主役にしない。この層(deny / sandbox / MCP / DB)が主役

Available Tools

5 tools
files_describeA

PII inbox(Claude が直接読めない場所)にあるファイルの一覧と推定スキーマを返す。値は返さない。 ファイルは id で参照する(py_run の inputs に渡す)。

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose key traits: it returns only metadata (no values), the source is a PII inbox Claude cannot read directly, and outputs are id references. It omits error behavior and any access/permission caveats, so it is strong but not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short, front-loaded sentences with no filler. The most important constraint (no values returned) and the downstream usage hint are both stated compactly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return shape need not be explained, and there are no parameters to document. The description covers what it returns, what it withholds, and how ids are consumed, making it essentially complete for a zero-arg discovery tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to compensate for; the baseline of 4 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: returns a list of files plus estimated schemas from the PII inbox. It also clarifies the negative scope ('does not return values') and distinguishes itself from siblings by noting files are referenced by id for py_run's inputs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the workflow — describe files here, then pass ids to py_run's inputs — which is useful implied guidance. However, it never explicitly says when to use this versus schema_describe or when it is not needed, leaving the choice to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fixture_makeA

ビュー名またはファイル id と同じ形の合成データ(Faker)を n 行返す。テストやスクリプト開発用。

Args:
    source: claude スキーマのビュー名、または files_describe の id。
    n: 行数(最大 50)。
ParametersJSON Schema
NameRequiredDescriptionDefault
nNo
sourceYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose a real constraint (n is capped at 50) and the generation mechanism (Faker), which is useful beyond the schema. It says nothing about side effects, permissions, or what happens on an invalid source, so coverage is only partial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences plus a tight Args block; the core behavior and the testing purpose are front-loaded with zero filler. Every line earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and both parameters are documented. What remains thin is error/edge behavior (invalid source, exceeding the cap) and confirmation of the read-only, side-effect-free nature of a generator.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate and largely does: it defines 'source' as either a claude schema view name or a files_describe id, and 'n' as the row count with a 50-row maximum. It adds meaningful semantics the bare schema (string/integer) lacks, though it omits the default of 5 that the schema carries.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: returns n rows of synthetic Faker data matching the shape of a given view or file id. The shape-matching behavior is distinctive and lets an agent distinguish it from siblings like sql_run or py_run, though it never explicitly names an alternative tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'for tests and script development' gives a clear use context, so usage is implied rather than left blank. However, there is no when-not guidance and no explicit routing to schema_describe or files_describe even though those tools produce the 'source' inputs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

py_runA

analysis/ 配下のスクリプトを、ネットワーク無しのコンテナで inbox ファイルに対して実行し、小さな JSON 結果だけ返す。

スクリプト内では `import sd` して `df = sd.load("<file id>")` で読み、`sd.result({...})` で結果を返す。
print 出力は返らない。結果に保護対象の値(PII 列の値)が含まれると拒否される。

Args:
    script: analysis ディレクトリからの相対パス(例: retention.py)。
    inputs: files_describe で得た id のリスト。
ParametersJSON Schema
NameRequiredDescriptionDefault
inputsNo
scriptYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does a lot: network-disabled execution, print output discarded, only a small JSON result returned, and hard rejection when PII column values appear in the result. It omits execution limits (timeouts/quotas) and whether the run is side-effect-free with respect to the input files, but the safety-relevant traits are unusually well disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first line, followed by the script contract (sd.load / sd.result), then the two arguments. Dense but well-ordered; the inline code examples earn their space by defining the required calling convention.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no further explanation, and the description still clarifies that only sd.result payloads survive and that print output is dropped. For a sandboxed code-execution tool the main residual gaps are error/failure modes and any timeout behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does for both parameters: script is a path relative to the analysis directory (with an example), and inputs is a list of ids produced by files_describe. Both semantics are supplied even though the schema itself documents neither.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: it runs scripts from analysis/ against inbox files inside a network-disabled container and returns a small JSON result. It is clearly separable from sql_run and files_describe by the 'execute a Python script' framing, though it never names the sibling it is not.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implicitly shows the workflow (obtain ids via files_describe, then pass them as inputs) and that scripts must live under analysis/, which is useful context. However it gives no explicit when-to-use/when-not guidance and never contrasts itself with sql_run, so an agent must infer the boundary between running SQL and running a Python script.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

schema_describeA

PII を除いたビュー(claude スキーマ)の列定義と、Faker で作ったダミー行を返す。実データは返さない。

Args:
    view: 特定のビュー名。省略時は全ビュー。
ParametersJSON Schema
NameRequiredDescriptionDefault
viewNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose the important safety traits (PII removed, data is Faker-generated dummy rows, no real data returned), which is genuinely useful context, but says nothing about permissions, result size, or pagination.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core behavior is front-loaded in the first sentence, the no-real-data caveat follows immediately, and the single parameter is documented compactly. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and the single parameter is covered. The only real gap is routing guidance against the sql_run/fixture_make siblings, which is minor for such a narrow tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and there is one parameter, yet the description fully explains it: 'view' selects a specific view and defaults to all views when omitted. That compensates well for the undocumented schema, though it adds no format or naming-convention detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: it returns column definitions for the PII-stripped (claude) views plus Faker-generated dummy rows, and explicitly notes it does not return real data. This separates it from sql_run, though it never names that sibling directly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'does not return real data' clause implies this is for schema inspection rather than querying, but there is no explicit when-to-use/when-not guidance and no mention of the obvious alternative sql_run or fixture_make. Usage is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sql_runA

claude スキーマの PII 無しビューに対して SELECT を1本実行する(読み取り専用・行数上限・少数セル抑止)。

Args:
    sql: SELECT または WITH で始まる1文。SELECT * は不可。
    max_rows: 返す最大行数(上限は設定値)。
ParametersJSON Schema
NameRequiredDescriptionDefault
sqlYes
max_rowsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose the key behaviors: read-only, a row-count cap, and small-cell suppression (a privacy safeguard an agent should know about). It omits error/timeout behavior and the exact meaning of the max_rows ceiling, so it is not fully transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded one-line summary followed by structured Args entries; every sentence carries constraint information with no filler. The mixed-language summary slightly hurts scannability for a non-Japanese reader but does not waste space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. The description covers what is executed, the safety profile, and both parameters' constraints; only error handling and the concrete row ceiling are absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: it constrains sql (SELECT/WITH prefix, no SELECT *) and explains max_rows as a maximum with a configured upper bound. It still doesn't state the default value or the actual ceiling, leaving some ambiguity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (execute one SELECT) and a specific resource (the PII-free view of the claude schema), with scope qualifiers (single statement). It does not explicitly distinguish itself from py_run, the other execution sibling, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives constraints (must start with SELECT or WITH, SELECT * disallowed, read-only) but never says when to choose this over py_run or schema_describe, nor what to do if the query is rejected. Usage is implied by the constraints rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 5 tool updatesv0.1.0
    • First observedfiles_describe
    • First observedfixture_make
    • First observedpy_run
    • First observedschema_describe
    • First observedsql_run

TDQS

A4/5.0

Scored across 5 tools

Disambiguation4/5

Each tool targets a distinct resource or operation (views, SQL, inbox files, synthetic fixtures, sandbox scripts). The main overlap is schema_describe and fixture_make, which both produce Faker dummy rows, though schema_describe is for schema inspection and fixture_make for generating test data.

Naming Consistency5/5

All names use consistent snake_case with a clear object_action pattern: schema_describe, sql_run, files_describe, fixture_make, py_run. Minor abbreviations (sql, py) are readable and do not break the convention.

Tool Count5/5

Five tools is well-scoped for a privacy-safe data access server. Each tool has a clear role in the workflow: schema discovery, querying, file cataloging, fixture generation, and sandbox execution.

Completeness4/5

The surface covers the core read-only analysis lifecycle: discover schemas, query safe views, inspect PII inbox metadata, generate synthetic data, and run sandboxed scripts. A minor gap is the lack of a tool to list available analysis scripts or otherwise help discover script paths for py_run.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers