Skip to main content
Glama
Whambammy

Prompt Shield & AI Safety MCP

README.md
<p align="center">
  <img src="./assets/logo.png" width="130" height="130" alt="Prompt Shield & AI Safety MCP Logo" />
</p>

# Prompt Shield & AI Safety MCP

[![Smithery Compatible](https://img.shields.io/badge/Smithery-Compatible-blue.svg)](https://smithery.ai)
[![Model Context Protocol](https://img.shields.io/badge/MCP-Standard%20v1.0-emerald.svg)](https://modelcontextprotocol.io)
[![Base L2 Settlement](https://img.shields.io/badge/Base%20L2-USDC%20x402-blue.svg)](https://base.org)
[![Tools](https://img.shields.io/badge/Tools-5%20Curated-purple.svg)](#included-tools)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)

**Deterministic prompt injection classification, instruction stripping, secret entropy detection, PII obfuscation, and dataset bias auditing.**

Built specifically for LLM application builders, LangChain/LlamaIndex engineers, enterprise AI developers, and AI red-teamers.

---

## ⚡ Quickstart

### Smithery Install
```bash
smithery skill add whambammy/prompt-shield-security-mcp
```

### Claude Desktop / Cursor (`claude_desktop_config.json`)
```json
{
  "mcpServers": {
    "prompt-shield-security-mcp": {
      "command": "npx",
      "args": ["-y", "@whambammy/prompt-shield-security-mcp"],
      "env": {
        "PAYMENT_WALLET": "0x9793E7269b3301893318dEa8338576Ba612F39B3",
        "BASE_RPC_URL": "https://mainnet.base.org"
      }
    }
  }
}
```

---

## 🛠️ Included Tools

| Tool Name | Price (USDC) | Capability |
| :--- | :---: | :--- |
| `prompt_injection_jailbreak_classifier` | $0.045 | High-speed deterministic classifier detecting indirect prompt injections, delimiter hijacking, and system role impersonation attacks in user inputs. |
| `strip_prompt_injection` | $0.02 | Neutralizes indirect prompt injections, system prompt leak probes, jailbreak tokens, and hidden instruction tags in untrusted text. |
| `secrets_entropy_scanner` | $0.030 | Shannon entropy and pattern analyzer scanning source code for leaked private keys, AWS access secrets, JWTs, and database connection strings. |
| `obfuscate_pii_entities` | $0.01 | Detects and masks Personally Identifiable Information (SSNs, credit cards, emails, phone numbers, API keys) prior to model ingestion. |
| `synthetic_dataset_bias_auditor` | $0.040 | Audits synthetic agent training data distributions for demographic bias, representation skew, and label drift across sensitive attribute categories. |


---

## 🔄 End-to-End Workflow

An enterprise agent receives untrusted user uploads -> runs secrets entropy scan to catch leaked tokens -> obfuscates PII -> passes through the prompt injection classifier -> strips residual jailbreak tags before invoking downstream frontier models.

---

## 💰 The x402 Base L2 Micropayment Protocol

When an agent invokes a tool without payment, the server responds with a deterministic `HTTP 402 Payment Required` challenge containing:
- Target tool price in USDC
- Base Native USDC Contract: `0x833589fCD6eDb6E08f4c7C32D4f71b54bdA02913`
- Recipient payout wallet address
- Single-use cryptographic nonce

Once broadcasted on Base L2, resubmitting with `paymentSignature` unlocks deterministic execution.

---

## 📄 License
MIT License. Created by [Whambammy](https://github.com/Whambammy).

TDQS

A3.5/5.0

Scored across 5 tools

Disambiguation4/5

Most tools target clearly distinct safety functions: PII masking, bias auditing, injection detection, injection neutralization, and secrets scanning. The two prompt-injection tools could be confused at a glance, but their descriptions distinguish detection/classification from stripping/neutralization.

Naming Consistency3/5

All names use snake_case and are descriptive, but the structural pattern is mixed: some are verb-first (obfuscate_pii_entities, strip_prompt_injection) while others are noun phrases ending in auditor/classifier/scanner. This is readable but not a predictable convention.

Tool Count5/5

Five tools is well within the sensible range for a focused AI safety/prompt shield server. Each tool covers a distinct guardrail capability, so none feels redundant or extraneous.

Completeness4/5

The surface covers key input-side guardrails: PII, bias, prompt injection detection/removal, and secrets scanning. Minor gaps remain around output-side moderation, toxicity classification, or policy enforcement, but the core lifecycle for a prompt safety server is reasonably covered.

Maintenance

ActivityMaintained
ResponsivenessNo issues