Prompt Shield & AI Safety MCP
<p align="center">
<img src="./assets/logo.png" width="130" height="130" alt="Prompt Shield & AI Safety MCP Logo" />
</p>
# Prompt Shield & AI Safety MCP
[](https://smithery.ai)
[](https://modelcontextprotocol.io)
[](https://base.org)
[](#included-tools)
[](LICENSE)
**Deterministic prompt injection classification, instruction stripping, secret entropy detection, PII obfuscation, and dataset bias auditing.**
Built specifically for LLM application builders, LangChain/LlamaIndex engineers, enterprise AI developers, and AI red-teamers.
---
## ⚡ Quickstart
### Smithery Install
```bash
smithery skill add whambammy/prompt-shield-security-mcp
```
### Claude Desktop / Cursor (`claude_desktop_config.json`)
```json
{
"mcpServers": {
"prompt-shield-security-mcp": {
"command": "npx",
"args": ["-y", "@whambammy/prompt-shield-security-mcp"],
"env": {
"PAYMENT_WALLET": "0x9793E7269b3301893318dEa8338576Ba612F39B3",
"BASE_RPC_URL": "https://mainnet.base.org"
}
}
}
}
```
---
## 🛠️ Included Tools
| Tool Name | Price (USDC) | Capability |
| :--- | :---: | :--- |
| `prompt_injection_jailbreak_classifier` | $0.045 | High-speed deterministic classifier detecting indirect prompt injections, delimiter hijacking, and system role impersonation attacks in user inputs. |
| `strip_prompt_injection` | $0.02 | Neutralizes indirect prompt injections, system prompt leak probes, jailbreak tokens, and hidden instruction tags in untrusted text. |
| `secrets_entropy_scanner` | $0.030 | Shannon entropy and pattern analyzer scanning source code for leaked private keys, AWS access secrets, JWTs, and database connection strings. |
| `obfuscate_pii_entities` | $0.01 | Detects and masks Personally Identifiable Information (SSNs, credit cards, emails, phone numbers, API keys) prior to model ingestion. |
| `synthetic_dataset_bias_auditor` | $0.040 | Audits synthetic agent training data distributions for demographic bias, representation skew, and label drift across sensitive attribute categories. |
---
## 🔄 End-to-End Workflow
An enterprise agent receives untrusted user uploads -> runs secrets entropy scan to catch leaked tokens -> obfuscates PII -> passes through the prompt injection classifier -> strips residual jailbreak tags before invoking downstream frontier models.
---
## 💰 The x402 Base L2 Micropayment Protocol
When an agent invokes a tool without payment, the server responds with a deterministic `HTTP 402 Payment Required` challenge containing:
- Target tool price in USDC
- Base Native USDC Contract: `0x833589fCD6eDb6E08f4c7C32D4f71b54bdA02913`
- Recipient payout wallet address
- Single-use cryptographic nonce
Once broadcasted on Base L2, resubmitting with `paymentSignature` unlocks deterministic execution.
---
## 📄 License
MIT License. Created by [Whambammy](https://github.com/Whambammy).
TDQS
Scored across 5 tools
Most tools target clearly distinct safety functions: PII masking, bias auditing, injection detection, injection neutralization, and secrets scanning. The two prompt-injection tools could be confused at a glance, but their descriptions distinguish detection/classification from stripping/neutralization.
All names use snake_case and are descriptive, but the structural pattern is mixed: some are verb-first (obfuscate_pii_entities, strip_prompt_injection) while others are noun phrases ending in auditor/classifier/scanner. This is readable but not a predictable convention.
Five tools is well within the sensible range for a focused AI safety/prompt shield server. Each tool covers a distinct guardrail capability, so none feels redundant or extraneous.
The surface covers key input-side guardrails: PII, bias, prompt injection detection/removal, and secrets scanning. Minor gaps remain around output-side moderation, toxicity classification, or policy enforcement, but the core lifecycle for a prompt safety server is reasonably covered.