pc2e-pii-shield
pc2e-pii-shield
보안이 강화된 프로덕션 등급의 Model Context Protocol (MCP) 서버로, 자동, 클라이언트 측, 엣지 개인 식별 정보(PII) 마스킹과 함께 읽기 전용 PostgreSQL 쿼리 실행을 제공합니다. LLM 에이전트(예: Cursor, Cline, Claude Code)가 GDPR, PDPA 및 데이터 프라이버시 원칙을 엄격히 준수하면서 데이터베이스에서 SQL 쿼리를 실행할 수 있게 합니다.
재사용 가능한 보안 미들웨어 제품으로 설계 및 개발된 이 서버는 데이터베이스 쿼리 결과를 가로채 민감한 데이터 유출을 방지합니다.
기술 아키텍처
flowchart TD
Client["AI Agent / Client (Cursor/Cline)"]
Proxy["Nginx Reverse Proxy"]
App["pc2e-pii-shield (Express)"]
DB["Postgres Database (Tailscale-Only)"]
Client ==>|HTTPS / SSE Request| Proxy
Proxy ==>|x-api-key Authentication| App
App ==>|Regex Read-Only Validation| DB
DB ==>|Raw SQL Results| App
App ==>|PII Tokenization & Masking| Proxy
Proxy ==>|Sanitized Event Stream| Client핵심 구성 요소
자동 마스킹 인터셉터 (
masking.ts): SQL 결과 집합을 동적으로 스캔합니다. 하이브리드 접근 방식을 사용하여 열 스키마 매칭(예:name,email,phone을 포함하는 필드)과 정규식 기반 콘텐츠 스캐닝을 결합해 데이터가 서버를 떠나기 전에 민감한 식별자를 감지하고 마스킹합니다.가명화 캐시 (
cache.ts): 원시 값을 임시 자리 표시자(예:__PERSON_A__,__EMAIL_1__)에 매핑하는 인메모리 TTL 기반 캐시(기본값: 30분)입니다. 무제한 메모리 소비를 방지하면서 양방향 복원을 허용합니다.AST 수준 변이 가드 (
db.ts): 원시 SQL 입력을 가로채는 엄격한 정규식 검증기입니다. SELECT가 아닌 모든 명령을 차단하고DROP,ALTER,DELETE,TRUNCATE,CREATE,GRANT와 같은 금지 키워드가 포함된 쿼리를 거부하여 애플리케이션 계층에서 엄격한 읽기 전용 경계를 보장합니다.동시 세션 관리자 (
index.ts): 기본 단일 연결 템플릿과 달리 이 서버는 연결sessionId를 키로 하는SSEServerTransport인스턴스의 활성 맵을 유지하여 여러 원격 개발자 또는 에이전트가 상태 충돌 없이 동시에 연결하고 스트리밍할 수 있게 합니다.텔레메트리 및 메트릭 엔드포인트 (
/stats): 연결 수, 고유 클라이언트 IP 추적, 집계 쿼리 실행 통계를 노출하여 설치 및 활성 사용량을 실시간으로 모니터링합니다.
Related MCP server: PostgreSQL MCP Server
보안 모델 및 위협 완화
제로 트러스트 데이터베이스 연결: 자격 증명 노출을 방지하도록 설계되었습니다. 데이터베이스는 격리된 Tailscale 전용 네트워크 인터페이스(예:
100.92.174.76)에서 실행되어 데이터베이스 포트가 공개 인터넷에 노출되지 않도록 보장합니다.암호화된 전송 및 API 키 보안: 서버 앞단에는 와일드카드 SSL 인증서를 사용하는 Nginx가 HTTPS(포트 443)로 배치되어, 요청을 전달하기 전에 보안 API 키 인증 게이트(
x-api-key)를 적용합니다.인메모리 수명 주기: 가명화 매핑은 엄격한 TTL과 함께 메모리에 저장되어 마스킹된 PII의 영구 디스크 흔적을 남기지 않습니다.
설치 및 배포
1. 사전 요구 환경 설정
환경 템플릿을 복사하세요:
cp .env.example .env.env 안에 데이터베이스 자격 증명을 구성하고 보안 API 키를 생성하세요.
2. 네이티브 빌드
Node.js(v18+)가 설치되어 있는지 확인하세요:
npm install
npm run build
npm start3. 컨테이너화된 배포
Docker Compose를 사용하여 배포하세요:
docker compose up -d --build이것은 호스트 포트 3088을 컨테이너의 내부 포트 3000에 매핑하여 SSE 서버를 자동으로 실행합니다.
4. 직접 실행 (NPX)
코드를 수동으로 다운로드하지 않고 Stdio 전송을 통해 서버를 즉시 실행할 수 있습니다:
npx -y mcp-pii-shield --db-uri "postgresql://username:password@localhost:5432/your_database"또는 SSE 전송을 통해 서버를 실행하세요:
npx -y mcp-pii-shield --sse --port 3000 --db-uri "postgresql://username:password@localhost:5432/your_database" --api-key "your_secret_key"클라이언트 통합
A. 로컬 클라이언트 통합 (Stdio를 통한 NPX)
로컬 AI 클라이언트가 npx를 사용하여 서버를 직접 실행하도록 구성하세요.
Claude Desktop (config.json)
다음 블록을 ~/Library/Application Support/Claude/claude_desktop_config.json(macOS) 또는 %APPDATA%\Claude\claude_desktop_config.json(Windows)에 추가하세요:
{
"mcpServers": {
"pc2e-pii-shield": {
"command": "npx",
"args": [
"-y",
"mcp-pii-shield",
"--db-uri",
"postgresql://username:password@localhost:5432/your_database"
]
}
}
}Cursor (설정 → 기능 → MCP)
+ 새 MCP 서버 추가를 클릭하세요.
이름을
pc2e-pii-shield로 설정하세요.유형을
command로 설정하세요.명령을 다음으로 설정하세요:
npx -y mcp-pii-shield --db-uri "postgresql://username:password@localhost:5432/your_database"
VS Code (Cline / Roo Code)
클라이언트 설정 JSON에 다음을 추가하세요:
{
"mcpServers": {
"pc2e-pii-shield": {
"command": "npx",
"args": [
"-y",
"mcp-pii-shield",
"--db-uri",
"postgresql://username:password@localhost:5432/your_database"
]
}
}
}B. 원격 클라이언트 통합 (SSE를 통한 HTTPS)
호스팅된 서버(예: 공용 NAS 인스턴스)에 연결하는 경우 SSE 전송 URL을 통해 연결하세요.
VS Code (Cline / Roo Code)
{
"mcpServers": {
"pc2e-pii-shield": {
"sseUrl": "https://pii-shield.thegeekybeng.com/sse?api_key=your_api_key_here"
}
}
}Cursor
+ 새 MCP 서버 추가를 클릭하세요.
이름을
pc2e-pii-shield로 설정하세요.유형을
SSE로 설정하세요.URL을 다음으로 설정하세요:
https://pii-shield.thegeekybeng.com/sse?api_key=your_api_key_here
프로젝트 배경 및 기술 책임자
이 프로젝트는 Andrew Yeo가 설계, 구축, 오픈소스화했습니다.
기술 책임자 소개
Andrew는 싱가포르에 기반을 둔 시니어 시스템 아키텍트이자 AI 엔지니어로, 다음과 같은 경력을 보유하고 있습니다:
25년의 전문 경력을 APAC에서 보유하며 프로그램 전달, 고객 온보딩, 기술 벤더 관리를 담당했습니다.
16년 이상의 시스템 아키텍처 및 기술 리더십을 보유하며 강력한 엔터프라이즈 인프라와 마이크로서비스 플랫폼을 설계하고 배포했습니다.
2년 이상의 전담 실무 AI/ML 엔지니어링 경험을 보유하며 AI 안전, LLM 메트릭, 보안 에이전틱 워크플로우를 전문으로 합니다.
검증된 실적
보안 시민 플랫폼: MPS-Connect(시민 선거구 케이스워크 플랫폼)와 **Case-Writer-Intelligence (CWI)**를 설계 및 배포했으며, 3단계 인과관계 엔진과 7개의 인간 개입 승인 게이트를 통합하여 문서 분류 시간을 40% 단축했습니다.
AI 계측 및 테스트: **Portable Continuous Context Engine (PC2E)**를 설계하여 6개 LLM 제공업체에 걸쳐 50,000건의 사례를 체계적이고 경험적으로 평가하여 모델 정렬 및 규정 준수를 벤치마킹했습니다.
기술 전문 분야: CI/CD 및 DevSecOps (GitHub Actions, Docker), 컨테이너화된 배포, 제로 트러스트 네트워크 토폴로지, 로컬/엣지 SLM 오케스트레이션 전문가입니다.
Available Tools
3 toolsadd_to_rosterA
Register new names to the active regex scan roster for local name-matching detection.
| Name | Required | Description | Default |
|---|---|---|---|
| names | Yes | An array of names to be dynamically added to the scanner roster. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are not provided, so the description carries the burden, but it is minimal. It clarifies the scope (local name-matching detection) but does not disclose behavioral traits such as whether the roster is persistent, how additions affect existing entries, or any potential side effects (e.g., deduplication). It goes beyond a simple 'Add' but lacks substantial behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence that packs essential information: action, target, and purpose. It is front-loaded with the verb. No filler or redundant content. Five is appropriate for its brevity and efficiency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with one parameter and no output schema. The description covers the purpose and target, but lacks details about behavior (e.g., duplicates, confirmation) and does not mention return values. Given the low complexity, this is acceptable but not fully complete; a 3 is appropriate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% (the parameter 'names' is documented as 'An array of names to be dynamically added to the scanner roster'). The description adds value by clarifying that the names are 'new' and for 'local name-matching detection', which enhances the schema's meaning. With full coverage, baseline is 3; the added specificity justifies a 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Register new names to the active regex scan roster for local name-matching detection' clearly states the action (register names), the resource (active regex scan roster), and the purpose (local name-matching detection). It distinguishes from siblings (unmask_text, run_secure_query) by specifying the roster for name-matching, which is specific enough.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context (for local name-matching detection) but does not explicitly specify when to use this tool versus alternatives, nor any exclusions (e.g., when to prefer unmask_text). Sibling tools exist but are not referenced or contrasted. Adequate but lacks explicit guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_secure_queryA
Execute a read-only SELECT database query. All PII values (names, emails, phones, NRIC/IDs) in the results will be automatically masked before being returned.
| Name | Required | Description | Default |
|---|---|---|---|
| sql_query | Yes | The read-only SQL SELECT query to run (e.g. SELECT name, email FROM contacts LIMIT 5) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it does so well by disclosing: (1) the operation is read-only, and (2) all PII values in results will be automatically masked. This gives the agent critical behavioral expectations (e.g., don't expect unmasked PII in results). It does not cover edge cases like error handling or large result pagination, but for the information provided, this is a strong disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences (33 words) with a clear action-first structure. Front-loads the primary purpose ('Execute a read-only SELECT database query') and follows with the key behavioral differentiator (PII masking). Every word contributes meaning; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple 1-parameter tool with no output schema, the description covers all essential aspects: the operation, the constraint on input, and a key output transformation (masking). Additional details like error messages for invalid queries or rate limiting would be nice but are not critical for this complexity, and the behavioral notes alone elevate it above the norm.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds value by qualifying the query as 'read-only' and emphasizing the PII masking behavior, which affects result processing semantics beyond what the schema example shows. It could have gone further by specifying what happens with non-SELECT input (error vs. rejection).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Execute a read-only SELECT database query') with a specific verb and resource, and the PII masking note explains what makes it 'secure.' This effectively differentiates it from sibling tools (unmask_text, add_to__roster) by making clear this is the querying tool that returns masked data.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool (read-only data retrieval) but does not explicitly state alternatives or exclusions (e.g., 'for write operations use X'). The sibling tools could offer more context, but no explicit comparison is provided. The read-only and SELECT constraints give some usage guardrails.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
unmask_textA
Restore the original raw PII values in a text payload by replacing placeholders (e.g. PERSON_A, EMAIL_1) with their original values cached during this session.
| Name | Required | Description | Default |
|---|---|---|---|
| masked_text | Yes | The text containing placeholders to be restored. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description must convey behavior. It mentions the session-cached values but does not disclose what happens if the cache is missing, whether the operation is reversible, or any side effects (e.g., does it mutate input or return a new string?). It provides some context but lacks critical behavioral details for a tool with no annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the core action and provides examples. It contains no redundant or tangential information, making it optimally concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with one parameter and no output schema, the description covers the main mechanism but omits the return value and potential error conditions (e.g., missing cache entries). While the session dependency is mentioned, a mention of expected output or failure handling would enhance completeness. Still, it is adequate for a simple tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides a basic description of 'masked_text.' The tool description adds value by giving concrete examples of placeholder formats and explaining that they are replaced with original values. This goes beyond the schema's simple definition, enriching parameter understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: restoring original PII values by replacing placeholders like __PERSON_A__ and __EMAIL_1__ with cached values. It uses a specific verb and resource, making it unmistakable. Although siblings are unrelated, the purpose is distinct and well-defined.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies usage context by mentioning 'cached during this session,' which tells the agent when the tool is applicable (after a prior masking operation). It does not explicitly list alternatives or exclusions, but given the unrelated siblings, this is not a significant gap. The context is clear enough for selecting this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v1.0.0- First observed
add_to_roster - First observed
run_secure_query - First observed
unmask_text
TDQS
Scored across 3 tools
Each tool addresses a distinct concern: one unmask text, one manage the name roster, and one execute queries with automatic masking. There is no overlap that would cause an agent to misselect.
Most tools follow a verb_noun pattern (unmask_text, run_secure_query), but add_to_roster breaks the pattern with an intervening preposition. This is a minor deviation and the intent remains clear.
Three tools is a reasonable, focused set for a PII-shielding server. It is slightly lean but each tool serves a clear purpose without unnecessary bloat.
The core masking lifecycle is covered—query masking, unmasking, and roster management—but obvious gaps exist: no tool for masking non-query text, no roster removal or listing, and no way to manage the cached placeholders beyond unmasking. These gaps could force workarounds.
Maintenance
Related MCP Connectors
Guard AI agents' PostgreSQL/MySQL access via MCP: SQL audit, auth, masking, write approval
Paid remote MCP for governed database query review, SQL simulation, approvals, and audits.
Hosted MCP server for PostgreSQL diagnostics: slow queries, missing indexes, connection pressure.
Draxlr's remote MCP server connects AI assistants to your SQL databases and dashboards. Explore schemas, run read-only queries, manage saved queries and dashboards, and export results, all with row-level security so each user sees only their own data.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceA secure MCP server that enables querying PostgreSQL databases through an SSH tunnel with enforced read-only access, connection pooling, and comprehensive data exploration tools.-
- AlicenseNot gradedqualityDmaintenanceA production-ready MCP server that enables safe, read-only SQL SELECT queries against PostgreSQL databases with built-in security validation. It features connection pooling, automatic row limits, and structured logging to ensure secure and reliable database interactions.31 npmISC
- AlicenseNot gradedqualityDmaintenanceRead-only PostgreSQL MCP server that enables running SELECT queries, listing tables and schemas, and describing columns, with built-in protection against writes and malicious SQL attacks.476 npmMIT
- AlicenseAqualityDmaintenanceA secure, read-only PostgreSQL MCP server that provides safe database introspection and querying capabilities.1419 npmMIT