Prompt Shield & AI Safety MCP
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Prompt Shield & AI Safety MCPScan this user message for prompt injection and strip any jailbreak attempts."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Prompt Shield & AI Safety MCP
Deterministic prompt injection classification, instruction stripping, secret entropy detection, PII obfuscation, and dataset bias auditing.
Built specifically for LLM application builders, LangChain/LlamaIndex engineers, enterprise AI developers, and AI red-teamers.
⚡ Quickstart
Smithery Install
smithery skill add whambammy/prompt-shield-security-mcpClaude Desktop / Cursor (claude_desktop_config.json)
{
"mcpServers": {
"prompt-shield-security-mcp": {
"command": "npx",
"args": ["-y", "@whambammy/prompt-shield-security-mcp"],
"env": {
"PAYMENT_WALLET": "0x9793E7269b3301893318dEa8338576Ba612F39B3",
"BASE_RPC_URL": "https://mainnet.base.org"
}
}
}
}Related MCP server: zentric-protocol-mcp
🛠️ Included Tools
Tool Name | Price (USDC) | Capability |
| $0.045 | High-speed deterministic classifier detecting indirect prompt injections, delimiter hijacking, and system role impersonation attacks in user inputs. |
| $0.02 | Neutralizes indirect prompt injections, system prompt leak probes, jailbreak tokens, and hidden instruction tags in untrusted text. |
| $0.030 | Shannon entropy and pattern analyzer scanning source code for leaked private keys, AWS access secrets, JWTs, and database connection strings. |
| $0.01 | Detects and masks Personally Identifiable Information (SSNs, credit cards, emails, phone numbers, API keys) prior to model ingestion. |
| $0.040 | Audits synthetic agent training data distributions for demographic bias, representation skew, and label drift across sensitive attribute categories. |
🔄 End-to-End Workflow
An enterprise agent receives untrusted user uploads -> runs secrets entropy scan to catch leaked tokens -> obfuscates PII -> passes through the prompt injection classifier -> strips residual jailbreak tags before invoking downstream frontier models.
💰 The x402 Base L2 Micropayment Protocol
When an agent invokes a tool without payment, the server responds with a deterministic HTTP 402 Payment Required challenge containing:
Target tool price in USDC
Base Native USDC Contract:
0x833589fCD6eDb6E08f4c7C32D4f71b54bdA02913Recipient payout wallet address
Single-use cryptographic nonce
Once broadcasted on Base L2, resubmitting with paymentSignature unlocks deterministic execution.
📄 License
MIT License. Created by Whambammy.
Available Tools
5 toolsobfuscate_pii_entitiesA
Detects and masks Personally Identifiable Information (SSNs, credit cards, emails, phone numbers, API keys) prior to model ingestion. (0.01 USDC on Base L2)
| Name | Required | Description | Default |
|---|---|---|---|
| payload | Yes | Input parameters or JSON string payload for the tool execution | |
| paymentSignature | No | Base L2 USDC micropayment signature or transaction hash for x402 settlement |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It does reveal the core behavior (detection and masking), the PII types handled, and the associated cost ('0.01 USDC on Base L2'). However, it does not disclose what the output looks like, how original data is handled, or the mechanics of the payment requirement, leaving meaningful gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single focused sentence plus a parenthetical cost note. It front-loads the core operation and PII types without redundant filler, and every part contributes useful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema and the description does not explain the return value, which is a notable gap for a tool that transforms input data. It also leaves the payment flow ambiguous: the description mentions a 0.01 USDC charge, yet paymentSignature is optional in the schema and no guidance is given on how to supply it. These omissions make the tool incompletely specified for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds some meaning beyond the schema by specifying the PII categories the payload should be checked for and the payment cost, but it does not clarify the exact expected payload format or whether paymentSignature is mandatory despite the cost note.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the operation ('Detects and masks') and the resource ('Personally Identifiable Information'), listing concrete entity types such as SSNs, credit cards, emails, phone numbers, and API keys. It also gives context ('prior to model ingestion') that helps distinguish it from mere detection tools, though it does not explicitly name sibling alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'prior to model ingestion' provides a clear, actionable usage context: use this tool before sending data to a model. It does not explicitly state when not to use it or name alternative sibling tools, but the intended scenario is reasonably clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prompt_injection_jailbreak_classifierA
High-speed deterministic classifier detecting indirect prompt injections, delimiter hijacking, and system role impersonation attacks in user inputs. (0.045 USDC on Base L2)
| Name | Required | Description | Default |
|---|---|---|---|
| payload | Yes | Input parameters or JSON string payload for the tool execution | |
| paymentSignature | No | Base L2 USDC micropayment signature or transaction hash for x402 settlement |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It does disclose 'high-speed deterministic classifier', which tells the agent the tool is deterministic and fast, and notes a 0.045 USDC cost on Base L2. However, it omits return format (what a classification result looks like), behavior when paymentSignature is absent, and error handling—gaps that matter for a paid tool with no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with the purpose front-loaded and the cost as a compact parenthetical. Every clause earns its place and there is no filler or repetition of schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must explain what the classifier returns, and it does not—an agent cannot tell whether it gets a label, a confidence score, or a boolean. The payment flow is also ambiguous: the cost is stated but the schema marks paymentSignature optional, leaving unclear whether the tool executes without payment. These omissions leave an agent under-informed for a paid security tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3. The description reinforces that payload is user input to be analyzed, but adds no format guidance (plain text vs JSON string) beyond the generic schema wording, and the paymentSignature semantics are only covered by the schema's boilerplate. The description does not meaningfully extend the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('detecting') with a clear resource (prompt injection/jailbreak attacks) and enumerates three concrete attack types (indirect prompt injections, delimiter hijacking, system role impersonation). It distinguishes itself from the sibling strip_prompt_injection by being a classifier/detector rather than a sanitizer, so an agent can tell them apart from the description alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The purpose implies when to use it (classify user inputs for injection attacks), but it never explicitly states when to prefer this over the sibling strip_prompt_injection (which removes rather than detects), nor does it give exclusion criteria or mention whether the payment requirement gates usage. The cost note is contextual but not a usage guideline.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
secrets_entropy_scannerB
Shannon entropy and pattern analyzer scanning source code for leaked private keys, AWS access secrets, JWTs, and database connection strings. (0.030 USDC on Base L2)
| Name | Required | Description | Default |
|---|---|---|---|
| payload | Yes | Input parameters or JSON string payload for the tool execution | |
| paymentSignature | No | Base L2 USDC micropayment signature or transaction hash for x402 settlement |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It usefully discloses the analysis approach and the 0.030 USDC cost, but it does not state whether the tool is read-only, what it returns, or whether source code is transmitted externally. These are notable gaps, though the cost disclosure adds some transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that clearly communicates the tool's purpose and includes the relevant cost. There is no redundancy, and every clause contributes useful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description leaves important invocation details unstated: how to structure the payload, what the response contains, and how the payment signature relates to the 0.030 USDC fee. An agent can infer the general purpose but not fully understand the execution contract.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description does not add meaningful parameter-level detail beyond the schema; the payload is described generically as 'Input parameters or JSON string payload' even though the tool's purpose implies it should contain source code.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: analyzing Shannon entropy and patterns in source code to find leaked private keys, AWS access secrets, JWTs, and database connection strings. However, it does not differentiate itself from the closely related sibling detect_secrets_leak, so an agent may struggle to choose between them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance is given about when to use this tool over detect_secrets_leak or other secret-related siblings. The phrase 'scanning source code' implies a context, but there are no conditions, exclusions, or alternative routing instructions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
strip_prompt_injectionA
Neutralizes indirect prompt injections, system prompt leak probes, jailbreak tokens, and hidden instruction tags in untrusted text. (0.02 USDC on Base L2)
| Name | Required | Description | Default |
|---|---|---|---|
| payload | Yes | Input parameters or JSON string payload for the tool execution | |
| paymentSignature | No | Base L2 USDC micropayment signature or transaction hash for x402 settlement |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden, and it does well by enumerating exactly what it neutralizes and adding an explicit cost note ('0.02 USDC on Base L2'). It does not state the shape of the returned sanitized text or failure behavior, a minor gap for an otherwise clear transformation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler; every clause carries content, and the parenthetical pricing is compact and useful for selection. It is appropriately sized for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity sanitization tool with fully described parameters, the description covers the core behavior and cost. It slightly under-specifies the return value and whether paymentSignature is mandatory in practice, but an agent has enough to select and invoke it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the structured fields already document both parameters. The description adds no meaning about how payload should be shaped or how paymentSignature relates to the 0.02 USDC fee beyond what the schema already states, warranting the baseline score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb ('Neutralizes') and a well-defined resource: indirect prompt injections, system prompt leak probes, jailbreak tokens, and hidden instruction tags in untrusted text. The action clearly marks it as a sanitizer rather than a detector, distinguishing it from the sibling prompt_injection_jailbreak_classifier.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended use is implied by 'in untrusted text': an agent can infer this is the tool to strip adversarial LLM instructions before further processing. But the description never names the classifier sibling or states when not to use it, leaving routing partially to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
synthetic_dataset_bias_auditorB
Audits synthetic agent training data distributions for demographic bias, representation skew, and label drift across sensitive attribute categories. (0.040 USDC on Base L2)
| Name | Required | Description | Default |
|---|---|---|---|
| payload | Yes | Input parameters or JSON string payload for the tool execution | |
| paymentSignature | No | Base L2 USDC micropayment signature or transaction hash for x402 settlement |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations available, the description carries the full burden of behavioral disclosure, but it only states the audit scope and a price. It does not disclose whether the operation is read-only, whether it requires a payment signature for execution, how the input payload should be structured, what report format is returned, or any side effects. The pricing note is useful but not a behavioral trait.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is minimal and front-loaded with the core purpose in the first sentence. The second sentence adds the cost detail, which is decision-relevant but arguably belongs in annotations. Overall, there is no fluff, and it earns conciseness credit, though it lacks any structural breakdown for the agent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no output schema, no annotations, and a generic payload parameter, so the description needs to explain what the agent should pass and what it will receive. It explains the audit dimensions but omits the expected payload format, the return shape, and whether payment is mandatory. An agent would likely struggle to construct a correct call without additional assumptions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the two declared parameters, so the baseline is 3. The description adds general context that the payload relates to synthetic training data and sensitive attribute categories, but it does not provide concrete parameter-level guidance beyond the schema's generic 'Input parameters or JSON string payload' text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description specifies a concrete verb ('Audits'), a distinct target ('synthetic agent training data distributions'), and enumerates the exact audit dimensions (demographic bias, representation skew, label drift, sensitive attribute categories). This clearly differentiates it from sibling audit tools like smart_contract_reentrancy_auditor or mesh_watertight_manifold_auditor without needing to open the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to choose this tool over alternatives, nor does it mention any exclusions or prerequisites. It implies the use case by its name and purpose, but there is no explicit context such as 'use when you need to evaluate a synthetic dataset for fairness' or comparison to any sibling tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
v1.0.0- First observed
obfuscate_pii_entities - First observed
prompt_injection_jailbreak_classifier - First observed
secrets_entropy_scanner - First observed
strip_prompt_injection - First observed
synthetic_dataset_bias_auditor
TDQS
Scored across 5 tools
Most tools target clearly distinct safety functions: PII masking, bias auditing, injection detection, injection neutralization, and secrets scanning. The two prompt-injection tools could be confused at a glance, but their descriptions distinguish detection/classification from stripping/neutralization.
All names use snake_case and are descriptive, but the structural pattern is mixed: some are verb-first (obfuscate_pii_entities, strip_prompt_injection) while others are noun phrases ending in auditor/classifier/scanner. This is readable but not a predictable convention.
Five tools is well within the sensible range for a focused AI safety/prompt shield server. Each tool covers a distinct guardrail capability, so none feels redundant or extraneous.
The surface covers key input-side guardrails: PII, bias, prompt injection detection/removal, and secrets scanning. Minor gaps remain around output-side moderation, toxicity classification, or policy enforcement, but the core lifecycle for a prompt safety server is reasonably covered.
Maintenance
Related MCP Connectors
Deterministic runtime safety for AI agents: scan PII, gate tool actions, verify LLM output.
Deterministic trust gate for AI output: leaked-secret, prompt-injection & PII in one call.
The WAF for agents. Pattern-based + heuristic firewall scans prompts, RAG documents, tool argume...
Deterministic prompt-injection detector; signed, offline-verifiable verdicts. Not an LLM.
Related MCP Servers
- AlicenseAqualityBmaintenanceProtects AI agents from threats like prompt injection, jailbreaks, and SQL injection through a multi-layer scanning pipeline. It also enables PII redaction and rehydration to ensure data privacy during LLM interactions.12318 npm2Apache 2.0
- FlicenseNot gradedqualityDmaintenanceSecurity middleware for LLM apps and AI agent pipelines. Detects prompt injection attacks (22 signatures, 7 languages) and anonymizes PII (17 entity types). Deterministic, sub-25ms, GDPR Art.30 compliant.-
- AlicenseNot gradedqualityDmaintenanceProvides a pre-flight/post-flight firewall for LLM calls with comprehensive detection, classification, policy enforcement, reversible redaction, output safety, and immutable audit logging.1MIT
- AlicenseNot gradedqualityDmaintenanceAnalyzes inputs and outputs in real-time to protect against prompt injections, data leaks, secrets exposure, and phishing URLs.10 npm3MIT