pc2e-pii-shield
pc2e-pii-shield
一个安全、生产级的 Model Context Protocol(MCP)服务器,提供只读 PostgreSQL 查询执行,并具备自动、客户端侧和边缘端的个人可识别信息(PII)脱敏能力。它允许 LLM 智能体(例如 Cursor、Cline、Claude Code)在数据库上执行 SQL 查询,同时确保严格符合 GDPR、PDPA 和数据隐私原则。
该服务器作为可复用的安全中间件产品进行设计与工程化实现,拦截数据库查询结果,以防止敏感数据外泄。
技术架构
flowchart TD
Client["AI Agent / Client (Cursor/Cline)"]
Proxy["Nginx Reverse Proxy"]
App["pc2e-pii-shield (Express)"]
DB["Postgres Database (Tailscale-Only)"]
Client ==>|HTTPS / SSE Request| Proxy
Proxy ==>|x-api-key Authentication| App
App ==>|Regex Read-Only Validation| DB
DB ==>|Raw SQL Results| App
App ==>|PII Tokenization & Masking| Proxy
Proxy ==>|Sanitized Event Stream| Client核心组件
自动脱敏拦截器(
masking.ts): 动态扫描 SQL 结果集。它采用混合方法:列模式匹配(例如,包含name、email、phone的字段)结合基于正则表达式的内容扫描,在数据离开服务器之前检测并脱敏敏感标识符。假名化缓存(
cache.ts): 一个内存中的、基于 TTL 的缓存(默认:30 分钟),将原始值映射到临时占位符(例如__PERSON_A__、__EMAIL_1__)。这支持双向还原,同时防止无限制的内存消耗。AST 级变更防护器(
db.ts): 一个严格的正则表达式验证器,用于拦截原始 SQL 输入。它阻止任何非 SELECT 命令,并拒绝包含DROP、ALTER、DELETE、TRUNCATE、CREATE或GRANT等禁止关键字的查询,从而在应用层确保严格的只读边界。并发会话管理器(
index.ts): 与基本的单连接模板不同,该服务器维护一个以连接sessionId为键的SSEServerTransport实例活动映射,允许多个远程开发人员或智能体并发连接和流式传输,而不会发生状态冲突。遥测与指标端点(
/stats): 公开连接计数、唯一客户端 IP 跟踪和聚合查询执行统计信息,以实时监控安装和活跃使用情况。
Related MCP server: PostgreSQL MCP Server
安全模型与威胁缓解
零信任数据库连接: 旨在防止凭据泄露。数据库运行在仅限 Tailscale 的隔离网络接口上(例如
100.92.174.76),确保数据库端口永远不会暴露到公共互联网。加密传输与 API 密钥安全: 服务器由 Nginx 通过 HTTPS(端口 443)前置代理,并使用通配符 SSL 证书,在转发请求之前强制执行安全的 API 密钥身份验证门禁(
x-api-key)。内存生命周期: 假名化映射以严格的 TTL 存储在内存中,不会在磁盘上留下已脱敏 PII 的任何持久痕迹。
安装与部署
1. 前置环境设置
复制环境模板:
cp .env.example .env在 .env 中配置您的数据库凭据并生成一个安全的 API 密钥。
2. 原生构建
确保已安装 Node.js(v18+):
npm install
npm run build
npm start3. 容器化部署
使用 Docker Compose 部署:
docker compose up -d --build此配置将主机端口 3088 映射到容器内部端口 3000,自动运行 SSE 服务器。
4. 直接执行(NPX)
您可以立即通过 Stdio 传输运行服务器,无需手动下载代码:
npx -y mcp-pii-shield --db-uri "postgresql://username:password@localhost:5432/your_database"或者通过 SSE 传输运行服务器:
npx -y mcp-pii-shield --sse --port 3000 --db-uri "postgresql://username:password@localhost:5432/your_database" --api-key "your_secret_key"客户端集成
A. 本地客户端集成(通过 NPX 经 Stdio 传输)
配置您的本地 AI 客户端,使用 npx 直接启动服务器。
Claude Desktop(config.json)
将以下代码块添加到您的 ~/Library/Application Support/Claude/claude_desktop_config.json(macOS)或 %APPDATA%\Claude\claude_desktop_config.json(Windows):
{
"mcpServers": {
"pc2e-pii-shield": {
"command": "npx",
"args": [
"-y",
"mcp-pii-shield",
"--db-uri",
"postgresql://username:password@localhost:5432/your_database"
]
}
}
}Cursor(设置 → 功能 → MCP)
单击 + 添加新的 MCP 服务器。
将 名称 设置为
pc2e-pii-shield。将 类型 设置为
command。将 命令 设置为:
npx -y mcp-pii-shield --db-uri "postgresql://username:password@localhost:5432/your_database"
VS Code(Cline / Roo Code)
将以下内容添加到您的客户端设置 JSON:
{
"mcpServers": {
"pc2e-pii-shield": {
"command": "npx",
"args": [
"-y",
"mcp-pii-shield",
"--db-uri",
"postgresql://username:password@localhost:5432/your_database"
]
}
}
}B. 远程客户端集成(通过 HTTPS 经 SSE 传输)
如果您要连接到托管服务器(例如您的公共 NAS 实例),请通过 SSE 传输 URL 连接。
VS Code(Cline / Roo Code)
{
"mcpServers": {
"pc2e-pii-shield": {
"sseUrl": "https://pii-shield.thegeekybeng.com/sse?api_key=your_api_key_here"
}
}
}Cursor
单击 + 添加新的 MCP 服务器。
将 名称 设置为
pc2e-pii-shield。将 类型 设置为
SSE。将 URL 设置为:
https://pii-shield.thegeekybeng.com/sse?api_key=your_api_key_here
项目背景与技术负责人
本项目由 Andrew Yeo 架构、构建并开源。
关于首席架构师
Andrew 是常驻新加坡的高级系统架构师和 AI 工程师,具备:
25 年专业经验,覆盖亚太地区,管理项目交付、客户引导和技术供应商管理。
16 年以上系统架构和技术领导经验,设计并部署了稳健的企业基础设施和微服务平台。
2 年以上专注的 AI/ML 实操工程经验,专精于 AI 安全、LLM 指标和安全智能体工作流。
已验证的工作成果
安全公民服务平台: 架构并部署了 MPS-Connect(公民选区个案处理平台)和 Case-Writer-Intelligence(CWI),集成了 3 阶段因果引擎和 7 个人工在环审批关卡,将文档分拣时间减少了 40%。
AI 度量与测试: 设计了 Portable Continuous Context Engine(PC2E),运行了一项跨六个 LLM 提供商的 50,000 个案例的系统性实证评估,以对模型对齐和合规性进行基准测试。
技术专长: 精通 CI/CD 与 DevSecOps(GitHub Actions、Docker)、容器化部署、零信任网络拓扑以及本地/边缘 SLM 编排。
Available Tools
3 toolsadd_to_rosterA
Register new names to the active regex scan roster for local name-matching detection.
| Name | Required | Description | Default |
|---|---|---|---|
| names | Yes | An array of names to be dynamically added to the scanner roster. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are not provided, so the description carries the burden, but it is minimal. It clarifies the scope (local name-matching detection) but does not disclose behavioral traits such as whether the roster is persistent, how additions affect existing entries, or any potential side effects (e.g., deduplication). It goes beyond a simple 'Add' but lacks substantial behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence that packs essential information: action, target, and purpose. It is front-loaded with the verb. No filler or redundant content. Five is appropriate for its brevity and efficiency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with one parameter and no output schema. The description covers the purpose and target, but lacks details about behavior (e.g., duplicates, confirmation) and does not mention return values. Given the low complexity, this is acceptable but not fully complete; a 3 is appropriate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% (the parameter 'names' is documented as 'An array of names to be dynamically added to the scanner roster'). The description adds value by clarifying that the names are 'new' and for 'local name-matching detection', which enhances the schema's meaning. With full coverage, baseline is 3; the added specificity justifies a 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Register new names to the active regex scan roster for local name-matching detection' clearly states the action (register names), the resource (active regex scan roster), and the purpose (local name-matching detection). It distinguishes from siblings (unmask_text, run_secure_query) by specifying the roster for name-matching, which is specific enough.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context (for local name-matching detection) but does not explicitly specify when to use this tool versus alternatives, nor any exclusions (e.g., when to prefer unmask_text). Sibling tools exist but are not referenced or contrasted. Adequate but lacks explicit guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_secure_queryA
Execute a read-only SELECT database query. All PII values (names, emails, phones, NRIC/IDs) in the results will be automatically masked before being returned.
| Name | Required | Description | Default |
|---|---|---|---|
| sql_query | Yes | The read-only SQL SELECT query to run (e.g. SELECT name, email FROM contacts LIMIT 5) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it does so well by disclosing: (1) the operation is read-only, and (2) all PII values in results will be automatically masked. This gives the agent critical behavioral expectations (e.g., don't expect unmasked PII in results). It does not cover edge cases like error handling or large result pagination, but for the information provided, this is a strong disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences (33 words) with a clear action-first structure. Front-loads the primary purpose ('Execute a read-only SELECT database query') and follows with the key behavioral differentiator (PII masking). Every word contributes meaning; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple 1-parameter tool with no output schema, the description covers all essential aspects: the operation, the constraint on input, and a key output transformation (masking). Additional details like error messages for invalid queries or rate limiting would be nice but are not critical for this complexity, and the behavioral notes alone elevate it above the norm.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds value by qualifying the query as 'read-only' and emphasizing the PII masking behavior, which affects result processing semantics beyond what the schema example shows. It could have gone further by specifying what happens with non-SELECT input (error vs. rejection).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Execute a read-only SELECT database query') with a specific verb and resource, and the PII masking note explains what makes it 'secure.' This effectively differentiates it from sibling tools (unmask_text, add_to__roster) by making clear this is the querying tool that returns masked data.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool (read-only data retrieval) but does not explicitly state alternatives or exclusions (e.g., 'for write operations use X'). The sibling tools could offer more context, but no explicit comparison is provided. The read-only and SELECT constraints give some usage guardrails.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
unmask_textA
Restore the original raw PII values in a text payload by replacing placeholders (e.g. PERSON_A, EMAIL_1) with their original values cached during this session.
| Name | Required | Description | Default |
|---|---|---|---|
| masked_text | Yes | The text containing placeholders to be restored. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description must convey behavior. It mentions the session-cached values but does not disclose what happens if the cache is missing, whether the operation is reversible, or any side effects (e.g., does it mutate input or return a new string?). It provides some context but lacks critical behavioral details for a tool with no annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the core action and provides examples. It contains no redundant or tangential information, making it optimally concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with one parameter and no output schema, the description covers the main mechanism but omits the return value and potential error conditions (e.g., missing cache entries). While the session dependency is mentioned, a mention of expected output or failure handling would enhance completeness. Still, it is adequate for a simple tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides a basic description of 'masked_text.' The tool description adds value by giving concrete examples of placeholder formats and explaining that they are replaced with original values. This goes beyond the schema's simple definition, enriching parameter understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: restoring original PII values by replacing placeholders like __PERSON_A__ and __EMAIL_1__ with cached values. It uses a specific verb and resource, making it unmistakable. Although siblings are unrelated, the purpose is distinct and well-defined.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies usage context by mentioning 'cached during this session,' which tells the agent when the tool is applicable (after a prior masking operation). It does not explicitly list alternatives or exclusions, but given the unrelated siblings, this is not a significant gap. The context is clear enough for selecting this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v1.0.0- First observed
add_to_roster - First observed
run_secure_query - First observed
unmask_text
TDQS
Scored across 3 tools
Each tool addresses a distinct concern: one unmask text, one manage the name roster, and one execute queries with automatic masking. There is no overlap that would cause an agent to misselect.
Most tools follow a verb_noun pattern (unmask_text, run_secure_query), but add_to_roster breaks the pattern with an intervening preposition. This is a minor deviation and the intent remains clear.
Three tools is a reasonable, focused set for a PII-shielding server. It is slightly lean but each tool serves a clear purpose without unnecessary bloat.
The core masking lifecycle is covered—query masking, unmasking, and roster management—but obvious gaps exist: no tool for masking non-query text, no roster removal or listing, and no way to manage the cached placeholders beyond unmasking. These gaps could force workarounds.
Maintenance
Related MCP Connectors
Guard AI agents' PostgreSQL/MySQL access via MCP: SQL audit, auth, masking, write approval
Paid remote MCP for governed database query review, SQL simulation, approvals, and audits.
Hosted MCP server for PostgreSQL diagnostics: slow queries, missing indexes, connection pressure.
Draxlr's remote MCP server connects AI assistants to your SQL databases and dashboards. Explore schemas, run read-only queries, manage saved queries and dashboards, and export results, all with row-level security so each user sees only their own data.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceA secure MCP server that enables querying PostgreSQL databases through an SSH tunnel with enforced read-only access, connection pooling, and comprehensive data exploration tools.-
- AlicenseNot gradedqualityDmaintenanceA production-ready MCP server that enables safe, read-only SQL SELECT queries against PostgreSQL databases with built-in security validation. It features connection pooling, automatic row limits, and structured logging to ensure secure and reliable database interactions.31 npmISC
- AlicenseNot gradedqualityDmaintenanceRead-only PostgreSQL MCP server that enables running SELECT queries, listing tables and schemas, and describing columns, with built-in protection against writes and malicious SQL attacks.476 npmMIT
- AlicenseAqualityDmaintenanceA secure, read-only PostgreSQL MCP server that provides safe database introspection and querying capabilities.1419 npmMIT