Skip to main content
Glama

obsify

민감한 파일의 원시 값이 모델의 컨텍스트에 들어가지 않도록 AI 어시스턴트가 작업할 수 있게 합니다.

obsify는 로컬에서 작동하는 결정론적 MCP 서버입니다. 프론티어 모델은 형태(shape) — 스키마, 합성 트윈, 마스킹된 피드백 — 에 대해 추론하고, 결정론적 로컬 코드는 실체(substance)를 다루며 마스킹되고 집계된 결과만 반환합니다. LLM 호출이 없고, 런타임에 네트워크도 없습니다: 탐지는 정규식 + 체크섬 + 사전 + Presidio의 로컬 NER을 사용합니다.

호주 엔터티 지원(ABN / ACN / TFN, 체크섬 검증)과 레이블 기반 라우팅 계층을 제공하여 "어시스턴트가 원시 데이터를 피해야 하는 시점"을 판단 호출이 아닌 결정론적이고 강제된 결정으로 만듭니다.

정직한 범위: run_on_real은 모델이 작성한 코드를 최선의 노력 로컬 샌드박스에서 실행하고 그 출력을 최선의 노력으로 마스킹합니다. 감옥이 아닙니다. 유출을 감당할 수 없는 항목에 사용하기 전에 SECURITY.md를 읽으십시오. 집계 결과를 반환합니다.

이유

기밀 문서를 호스팅된 LLM에 제공하면 실체가 경계를 벗어납니다. 일반적인 대답은 "LLM을 사용하지 마라" 또는 "제공자를 신뢰하라"입니다. obsify는 세 번째 경로를 취합니다 — 데이터로의 컴퓨팅(compute-to-data): 데이터를 모델로 가져오지 않고 코드를 데이터로 가져옵니다.

  • 모델은 스프레드시트의 스키마를 보고, 행을 보지 않습니다.

  • 모델은 합성 트윈(가짜 값, 실제 구조)에 대해 개발합니다.

  • 모델의 분석 코드는 로컬에서 실행됩니다. 마스킹되고 집계된 출력만 반환됩니다.

프론티어 모델의 추론은 보존됩니다. 원시 값에 대한 눈만 제거됩니다.

Related MCP server: MCP DB Results Anonymizer

도구

도구

기능

반환값

scan_pii(path)

파일/폴더에서 PII 스캔

유형, 위치, 개수 — 값은 절대 반환하지 않음

make_synthetic_twin(path, out)

Excel 워크북의 충실한 가짜 복사본

스키마 요약; 트윈이 out에 기록됨 (값은 가짜, 유출 검증됨)

run_on_real(code, data_path)

데이터로의 컴퓨팅: 실제 파일에 대해 로컬에서 코드 실행 (DATA_PATH에 바인딩됨)

PII 마스킹되고 크기 제한된 stdout/stderr만 — 집계 결과 반환

redact_text(text)

문자열에서 PII를 <TYPE> 토큰으로 마스킹

수정된 문자열

verify_value_free(text, terms)

text가 terms(또는 그 변형) 중 어느 것도 누출하지 않는지 실패-폐쇄 확인

{"value_free": bool}

지원 문서: PDF (텍스트 + 표; 복잡한 표는 obsify[tables]를 통한 폴백), Excel .xlsx/.xlsm, Word .docx (문단 + 표). 읽을 수 없거나 지원되지 않는 파일은 명시적 메모/사각지대로 표시되며, 조용히 버려지지 않습니다. (OCR 없음 — 스캔/이미지 페이지는 저범위로 표시되며, 변환되지 않습니다.)

알려진 엔터티 마스킹 (선택 사항). 숨길 이름의 로컬 .obsify.entities 목록을 제공하면 scan_pii / redact_text가 결정론적으로 이를 잡아내며 — NER이 놓치는 접미사/약어 변형(BRIGHTWATER HLDGS P/L for Brightwater Holdings Pty Ltd)도 — KNOWN_ENTITY로 처리합니다. 목록은 로컬에 남아 모델의 컨텍스트에 절대 들어가지 않습니다. docs/known_entities.md 참조.

데모

공식 MCP Inspector를 사용하여 합성 데이터에 대해 다섯 가지 도구를 모두 실시간으로 시험해 보세요:

python -m obsify.make_corpus --out ./corpus_demo
npx @modelcontextprotocol/inspector obsify-mcp

./corpus_demo/ledger.xlsx에 대해 scan_pii를 호출하고 유형/개수/위치만 반환하고 값은 절대 반환하지 않는지 확인하세요. docs/verifying.md 참조.

사용해 보기 — 합성 코퍼스

세 가지 형식을 모두 포괄하는 가짜이지만 현실적인 코퍼스(모두 합성; ABN/ACN/TFN은 체크섬 검증됨)를 생성한 다음 도구를 가리키세요:

pip install "obsify[demo]"                 # reportlab, for the sample PDFs
python -m obsify.make_corpus --out ./corpus_demo

다중 시트 Excel 원장(숫자 오탐 지뢰밭), PDF 위임장(산문 + 시산표), DOCX 감사 메모(문단 + 공급업체 표)를 작성합니다. 실제 데이터를 건드리지 않고 scan_pii / make_synthetic_twin을 시험해 보기에 좋습니다.

설치 및 MCP 서버로 실행

Python 3.11+ 필요. obsify는 stdio를 통해 MCP와 통신합니다 — 클라이언트가 로컬 하위 프로세스로 실행합니다. 원격으로 호스팅되지 않습니다. MCP 호환 클라이언트(Claude Desktop, Claude Code, Cursor, VS Code 등)에 등록하려면 해당 클라이언트의 설정에 블록 하나를 추가하세요.

권장 — uvx를 통한 제로 설치:

{ "mcpServers": { "obsify": { "command": "uvx", "args": ["obsify-mcp"] } } }

uvx는 PyPI에서 obsify를 가져와 요청 시 실행합니다 — 영구 설치 불필요. 첫 실행 시 obsify는 spaCy NER 모델(en_core_web_lg, ~560 MB)을 한 번 다운로드하여 캐시합니다. 이는 공개 모델을 가져오며 사용자 데이터를 보내지 않습니다(금지하려면 OBSIFY_AUTO_DOWNLOAD=0 설정하고 직접 모델 설치). 이후 실행은 즉시 이루어지며 완전히 오프라인입니다.

또는 설치 (pip / pipx):

pipx install obsify        # isolated, on PATH  (or: pip install obsify)

그런 다음 클라이언트를 설치된 명령어로 지정하세요:

{ "mcpServers": { "obsify": { "command": "obsify-mcp" } } }

클라이언트를 다시 시작하면 도구가 나타납니다. 선택적 추가 기능: obsify[tables] (camelot + Ghostscript를 통한 복잡한 표 PDF 폴백), obsify[compute] (pandas, run_on_real 코드 내에서 유용).

PATH 문제 (서버 연결 실패의 #1 원인): command는 클라이언트가 보는 PATH에서 확인 가능해야 합니다. GUI 클라이언트는 가상 환경의 PATH를 공유하지 않을 수 있습니다. 해결책: uvx/pipx 사용 (전역적으로 확인 가능), 또는 절대 경로 제공 — "/path/to/.venv/bin/obsify-mcp" (macOS/Linux) 또는 "C:\path\to\.venv\Scripts\obsify-mcp.exe" (Windows).

이 저장소에서 (PyPI에 올리기 전):

pip install "git+https://github.com/Formative-Sum41/obsify.git"   # gets `obsify-mcp` + `obsify`

라우팅 계층 — 결정론적, 판단 호출이 아님

"도와주지만 기밀 파일을 읽지 마세요"의 어려운 부분은 보호 시점을 결정하는 것입니다. obsify는 그 결정을 모델에서 환경으로 옮깁니다:

  1. .obsify.json — 경로를 분류하는 레이블 매니페스트 (public / confidential / restricted).

  2. obsify.guard (python -m obsify.guard로 실행) — 레이블이 지정된 파일의 직접 읽기를 차단(종료 코드 2)하고 어시스턴트를 scan_pii / make_synthetic_twin / run_on_real로 리디렉션하는 PreToolUse 가드.

  3. 관례 (CLAUDE.md에 있음) — 어시스턴트가 가드에 도달하기 전에 obsify를 선호하도록 함.

한 명령으로 설정:

obsify init [--dir PATH] [--with-claude-md]

obsify init은 비파괴적으로 설계되었습니다 — 정확히 하나의 파일을 소유하고 나머지에 대한 스니펫을 제공합니다:

  • .obsify.json — obsify가 소유; init이 작성함 (--force 없이 덮어쓰지 않음).

  • .claude/settings.json — 사용자 파일: init은 PreToolUse 훅 블록을 출력하여 붙여넣도록 하며, 직접 편집하지 않음 (코드를 실행하므로 등록은 사용자 결정).

  • CLAUDE.md — 사용자 파일: 관례는 옵트인. 기본적으로 출력; --with-claude-md는 마커로 감싸진 멱등 블록을 추가하여 사용자 콘텐츠를 덮어쓰지 않음.

전체 관례: docs/obsify_routing.md.

탐지 정밀도 유지 방법

  • 체크섬 검증 식별자. ABN/ACN/TFN 후보는 정규식으로 제안되고 공식 체크섬으로 확인되므로, 임의의 숫자가 식별자로 보고되지 않습니다.

  • 컨텍스트 필요 ID. 레이블 단어("TFN", "ABN", "BSB", …)가 근처에 있을 때만 숫자가 ABN/ACN/TFN으로 허용됩니다 — 이는 숫자 원장에서 순차 저널 ID 오탐 홍수를 차단합니다.

  • 문자 없는 / 숫자 포함 NER 억제. 순수 숫자, 금액, 날짜 및 영숫자 코드는 이름/조직으로 플래그되지 않습니다. 실제 이름, 이메일 및 주소(문자를 포함)는 영향을 받지 않습니다. 검증된 문자 없는 PII는 예외로 유지: 체크섬 ID(ABN/ACN/TFN/Medicare), Luhn 카드, 유효 IP, BSB 인접 계좌, 전화(컨텍스트 또는 전화 형태를 통해) — 소수점은 여전히 금액을 표시하며 전화가 아닙니다.

측정된 정확도

obsify는 점수 평가 하네스(eval/ — 레이블이 지정된 합성 코퍼스 + 정답 키 + 배송되는 탐지기에 대한 스코어러, 추가로 독립적인 제3자 교차 확인)를 제공합니다. 합성 코퍼스에 대한 헤드라인: 예상 탐지 항목에 대해 100% 재현율, 숫자 FP-고문 시트(그룹화된 숫자 가드 포함)에서 0 오탐, 맥락 게이트 ID가 올바르게 억제됨. Microsoft presidio-research에 대한 독립 교차 확인: EMAIL/IBAN 100%, PERSON 94%.

하네스는 그 가치를 입증했습니다 — 실제 결함을 발견했고, 이후 수정되었습니다: 신용 카드와 전화번호가 숫자 노이즈 필터에 의해 조용히 억제되고 있었음(이제 체크섬 검증/전화 형태를 통해 예외 처리됨), Medicare, IP, 생년월일, 호주 여권 및 운전면허증에는 인식기가 없었음(이제 추가됨, 체크섬 또는 컨텍스트 게이트). 전체 방법, 숫자 및 남은 문서화된 격차(SWIFT/BIC, 생년월일이 아닌 날짜): eval/README.md.

테스트

pip install -e ".[dev]"
pytest tests/            # or run any file directly: python tests/test_obsify.py

12개 스위트(73개 테스트), Linux + Windows / Python 3.11 + 3.12에서 CI 실행:

  • mcp-protocol — 실제 서버를 stdio로 실행하고 MCP와 통신(Claude와 같은 클라이언트가 사용하는 동일한 경로): 다섯 도구 모두 유효한 스키마로 등록되고 호출이 JSON-RPC를 통해 왕복하는지 확인 — scan_pii가 형태만, 종단 간 반환하는지 포함.

  • checksums — 외부에 게시된 ABN/ACN/TFN 작업 예제(유효 및 손상)에 고정되어 생성기↔검증기 순환성을 깨뜨림.

  • obsify / twin / redaction — 개인정보 보호 불변성: 형태만 출력, 누출 없는 트윈, 실패-폐쇄 자체 점검.

  • precision — 오탐 억제기가 숫자 원장 노이즈를 죽이면서 실제 이름을 유지하는지.

  • routing — 가드의 차단/허용 분류 및 obsify init의 비파괴 계약.

  • corpus — 합성 PDF+Excel+DOCX 코퍼스 종단 간: 형식별 탐지, DOCX 문단+표 추출, 모든 형식에서 형태만 출력.

  • evaluation — 회귀 게이트로서 점수 하네스(재현율, 억제, FP-고문, 격차).

  • robustness — 우아한 저하: 손상/초대형/빈/중첩/지원되지 않는 입력이 절대 충돌하지 않고 항상 메모로 표시됨.

  • model / variants — 첫 실행 모델 자동 다운로드 로직; verify_value_free 뒤의 변형 정규화.

대화형 검증(MCP Inspector) 및 라이브 클라이언트 최종 점검은 docs/verifying.md 참조.

라이선스

MIT — LICENSE 참조.

Available Tools

5 tools
make_synthetic_twinA

Generate a SYNTHETIC TWIN of a real Excel workbook at path, written to out. Schema (sheets, headers, column types, true row counts) is preserved; every data value is freshly FAKED — no real value is copied. Reason and write your analysis code against the twin; then run it on the real file with run_on_real. Returns the schema summary (safe shape).

ParametersJSON Schema
NameRequiredDescriptionDefault
outYes
pathYes
cap_rowsNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well. It discloses key behavioral traits: schema is preserved, every data value is freshly FAKED, no real value is copied, and it returns a safe schema summary. This gives the agent essential safety and data-handling context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three dense sentences, front-loaded with the core action and then efficient supplementary detail. There is no fluff or redundancy; every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the core function, the workflow pairing with run_on_real, and the return value (schema summary). It lacks any explanation of cap_rows and its potential effect on 'true row counts,' which would be a notable gap for a tool of moderate complexity. Overall, it is quite complete but not flawless.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does explain path and out (': path', 'written to out'), but it does not mention cap_rows at all. This leaves one of three parameters semantically opaque, so the compensation is partial.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb+resource: 'Generate a SYNTHETIC TWIN of a real Excel workbook at `path`, written to `out`.' It clearly differentiates from siblings by framing this as the twin-creation step and explicitly mentions run_on_real as the subsequent step for real-file execution.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit workflow guidance: 'Reason and write your analysis code against the twin; then run it on the real file with run_on_real.' This tells exactly when to use this tool and names the alternative (run_on_real) for the next phase, satisfying the 'when/when-not/alternatives' criterion.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

redact_textA

Return text with detected PII replaced by placeholders (e.g. , ). Deterministic; checksum-validated identifiers and context/precision rules apply so bare numbers are not over-masked.

entities is an optional PATH to a local .obsify.entities file of KNOWN names to hide; matches (incl. variants) are masked as . If omitted, a nearby .obsify.entities is auto-used.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYes
entitiesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It does well by disclosing determinism, checksum-validated identifiers, over-masking avoidance, and the entities-file auto-use behavior. Minor gaps remain around error handling or what happens when no PII is detected, but transparency is strong overall.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: the first sentence states the core operation, followed by key behavioral constraints and then the optional parameter explanation. Every sentence earns its place without unnecessary verbiage.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simple 2-parameter shape and the presence of an output schema, the description is highly complete. It covers the main transformation, important edge-case prevention (bare numbers), and the optional entities file behavior. The description is sufficient for an agent to invoke the tool correctly without needing further clarification.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does. It explains that `entities` is a path to a local `.obsify.entities` file, that matched names are masked as `<KNOWN_ENTITY>`, and that a nearby file is auto-used if omitted. This adds substantial meaning beyond the bare schema fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence clearly states the verb (redact), the resource (text), and the output format (PII replaced by placeholders), making the purpose immediately obvious. It also distinguishes itself from sibling tools like scan_pii and verify_value_free by explicitly conveying the redaction operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context about how the tool behaves and when the optional entities file applies, but it does not explicitly state when to prefer redact_text over sibling tools or when not to use it. There are no alternative tool comparisons or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_on_realA

COMPUTE-TO-DATA: execute your Python code LOCALLY against the real file at data_path (bound to the variable DATA_PATH in your code); the returned output is size-capped and best-effort PII-masked. The data never enters your context; substance never leaves. Return AGGREGATES (counts/sums/summaries) via print() — output masking is defense-in-depth, NOT a guarantee (NER can miss a name in a raw record), so never print raw records or identifiers. The masking field carries this caveat with the result. Network is disabled and a timeout applies.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYes
timeoutNo
data_pathYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and handles it well: output is size-capped, PII masking is best-effort and explicitly not a guarantee, network is disabled, a timeout applies, and execution is local. It also warns that raw records/identifiers should never be printed, adding important safety context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well organized: concept label, action, safety constraints, and usage guidance. Bolded callouts ('Return AGGREGATES...', 'best-effort') make key instructions easy to parse, and no sentence is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a code-execution tool with no output schema and no annotations, the description covers the essential operational surface: local execution, data binding, output size, masking caveat, aggregate printing, network isolation, and timeout. It is sufficient for an agent to invoke the tool safely and interpret the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It adds strong semantics for code (Python executed locally) and data_path (bound to DATA_PATH), but timeout is only indirectly covered by 'a timeout applies' and the schema's default, not explained as a configurable parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'COMPUTE-TO-DATA' and clearly states the tool executes Python code locally against a real file at data_path, binding it to DATA_PATH. This is a specific verb+resource pairing and is distinct from siblings like make_synthetic_twin, which implies synthetic data operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description strongly implies when to use it: when you need to compute over real data without pulling raw data into context ('data never enters your context'). It gives actionable guidance to print aggregates and avoid raw records, but it does not explicitly name alternative tools or state when-not-to-use conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scan_piiA

Scan a file or folder for PII and return TYPES + LOCATIONS + COUNTS only — never the detected values. Safe to surface to an LLM: it learns what PII exists and where, without the substance entering context. Recurses into subfolders; skips unreadable files and caps very large sheets, reporting both as notes.

entities is an optional PATH to a local .obsify.entities file (one name per line) of KNOWN names to hide; matches (incl. suffix/abbreviation variants) are reported as KNOWN_ENTITY. If omitted, a nearby .obsify.entities is auto-used. The names are read locally and never returned.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes
entitiesNo
max_cellsNo

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description provides thorough behavioral details: it returns only metadata (not values), skips unreadable files, caps very large sheets, and reads the entities file locally without returning the names. It also explains the automatic fallback for the entities file. This gives a clear picture of side effects and limitations, exceeding the typical level of transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with a clear first paragraph on functionality and a second on the entities parameter. It is concise enough to convey necessary details without fluff, though the entities explanation could be slightly tighter. The information is relevant and not redundant, earning a high score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description provides a high-level overview of the return value (types, locations, counts) without specifying the exact output format, which is acceptable given no output schema. It covers main behaviors (recursion, skipping, capping) and the entities file. It lacks explicit error handling or return structure details, but for a scan tool, the description sufficiently completes the context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds significant meaning for the 'entities' parameter by explaining its purpose, format, and default behavior. It indirectly touches on 'max_cells' by mentioning capping large sheets, but does not explicitly link it to the parameter. The 'path' parameter is self-explanatory given the context. Overall, it compensates for the lack of schema descriptions, though not perfectly for max_cells.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool scans files or folders for PII and returns only types, locations, and counts, never the values. It also mentions recursion, skipping unreadable files, and capping large sheets, which fully specifies the tool's function. This distinguishes it from sibling tools like redact_text or make_synthetic_twin.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not explicitly explain when to use this tool over its siblings. It implies usage for scanning and reporting PII metadata, and the safety note ('Safe to surface to an LLM') hints at a use case, but there is no direct comparison or guidance on choosing between tools. The behavior details (recursion, skipping) could inform usage, but explicit 'use when' instructions are absent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_value_freeA

Fail-closed check that text contains NONE of terms (nor their suffix-normalized / distinctive-token variants). Returns {"value_free": bool} with zero detail on what matched — for verifying an artifact before it leaves the perimeter.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYes
termsYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of behavioral clarity. It discloses the fail-closed behavior, the variant-matching behavior, and the deliberately detail-poor return shape. It does not explicitly state there are no side effects, but the read-only check nature is strongly implied.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded, stating the check first, then the return contract, then the intended target scenario. Every sentence contributes useful information without repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter boolean verification tool, the description is nearly complete: it names inputs, behavior, return value, and intended boundary context. It leaves minor edge-case behavior unspecified, but this does not materially hamper selection or invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has no parameter descriptions, but the description defines the core semantics: `text` is the artifact being verified and `terms` are the prohibited strings matched directly or through normalized variants. It adds meaningful algorithmic context beyond the bare schema, though it omits edge cases like empty terms behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a fail-closed verification check that `text` contains none of `terms` or their variants, giving a specific verb, resource, and scope. It distinguishes itself from sibling tools by being a boolean verification gate rather than a scanning or redaction operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly identifies the intended use case: verifying an artifact before it leaves the perimeter. It implies this is a pre-release/compliance gate rather than a diagnostic tool, and the zero-detail return further signals it is not for troubleshooting that needs matched context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 5 tool updatesv0.2.0
    • First observedmake_synthetic_twin
    • First observedredact_text
    • First observedrun_on_real
    • First observedscan_pii
    • First observedverify_value_free

TDQS

A4.4/5.0

Scored across 5 tools

Disambiguation5/5

Each tool has a clear, distinct purpose: make_synthetic_twin creates a fake dataset, run_on_real executes code against real data, scan_pii identifies PII locations, redact_text masks PII in text, and verify_value_free checks for forbidden terms. No two tools overlap in what they accomplish.

Naming Consistency4/5

Most tools follow a verb_noun pattern (make_synthetic_twin, scan_pii, redact_text, verify_value_free), but run_on_real breaks the pattern with a prepositional phrase. The style is consistent (all snake_case, verbs first) but the deviation is noticeable.

Tool Count4/5

With 5 tools, the count is well-scoped for a focused PII-handling server. Each tool covers a necessary step in the workflow, and the count is within the typical 3-15 range, though a few additional helpers could be justified (e.g., a check for twin accuracy).

Completeness4/5

The tool surface covers the core lifecycle: protect data (scan, redact, verify) and enable safe analysis (twin, run on real). Minor gaps exist, such as no tool to validate the synthetic twin's fidelity against the real file, and verify_value_free lacks a positive counterpart, but agents can work around these.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    Let LLMs analyze sensitive data safely by querying a tokenized, join-preserving copy of the database, with fail-closed PII scanning and provable numeric equivalence.
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Acts as an anonymizing proxy between AI agents and databases, detecting PII and replacing it with realistic fake data so agents never see real data.
    Apache 2.0
  • A
    license
    A
    quality
    B
    maintenance
    Self-hosted governance layer between an AI assistant and your data: allow/deny policy, deterministic PII masking, row caps, and a hash-chained audit log with an Ed25519-signed receipt for every access, verifiable offline.
    4
    481 npm
    3
    MIT