Sandbox Code Auditor MCP
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Sandbox Code Auditor MCPScan this untrusted code for sandbox escapes and generate a Seccomp BPF filter"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Sandbox Code Auditor MCP
Python sandbox escape detection, SQL injection taint analysis, Seccomp BPF filter generation, and ReDoS regex scanning.
Built specifically for Coding agents (Claude Code, Cursor, AutoGen), CI/CD security pipelines, and automated code review bots.
⚡ Quickstart
Smithery Install
smithery skill add whambammy/sandbox-code-auditor-mcpClaude Desktop / Cursor (claude_desktop_config.json)
{
"mcpServers": {
"sandbox-code-auditor-mcp": {
"command": "npx",
"args": ["-y", "@whambammy/sandbox-code-auditor-mcp"],
"env": {
"PAYMENT_WALLET": "0x9793E7269b3301893318dEa8338576Ba612F39B3",
"BASE_RPC_URL": "https://mainnet.base.org"
}
}
}
}Related MCP server: vibecheck
🛠️ Included Tools
Tool Name | Price (USDC) | Capability |
| $0.045 | Audits Python ASTs for dangerous builtins, |
| $0.040 | Parses SQL query ASTs to verify parameterized binding, flagging raw string concatenations that lead to second-order SQL injection vulnerabilities. |
| $0.035 | Audits regular expressions for catastrophic polynomial and exponential backtracking (ReDoS) vulnerabilities using NFA/DFA cycle decomposition. |
| $0.040 | Generates minimal Seccomp BPF syscall filter profiles for sandboxing untrusted agent processes, blocking ptrace, fork, and raw socket creation. |
| $0.025 | Audits CORS response headers for dangerous wildcards with credentials ( |
🔄 End-to-End Workflow
An autonomous coding agent receives a script from an external PR -> analyzes the Python AST for sandbox escapes -> audits SQL queries for injection taint -> tests regexes for ReDoS -> outputs a Seccomp BPF filter to execute the code securely.
💰 The x402 Base L2 Micropayment Protocol
When an agent invokes a tool without payment, the server responds with a deterministic HTTP 402 Payment Required challenge containing:
Target tool price in USDC
Base Native USDC Contract:
0x833589fCD6eDb6E08f4c7C32D4f71b54bdA02913Recipient payout wallet address
Single-use cryptographic nonce
Once broadcasted on Base L2, resubmitting with paymentSignature unlocks deterministic execution.
📄 License
MIT License. Created by Whambammy.
Available Tools
5 toolscors_policy_preflight_misconfig_checkerA
Audits CORS response headers for dangerous wildcards with credentials (Access-Control-Allow-Origin: * + Credentials: true). (0.025 USDC on Base L2)
| Name | Required | Description | Default |
|---|---|---|---|
| payload | Yes | Input parameters or JSON string payload for the tool execution | |
| paymentSignature | No | Base L2 USDC micropayment signature or transaction hash for x402 settlement |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the behavioral burden. It implies a read-only audit and discloses the 0.025 USDC cost, but does not explain whether the tool fetches a URL, analyzes supplied headers, or what happens on failure or when payment is absent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence with no filler; the core check is front-loaded and the cost is tucked in parentheses. Every element earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and a vague payload definition, the description is too thin. An agent still cannot tell what input format to use or what result shape to expect, leaving a significant gap for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The schema's 'payload' description is generic, and the tool description does not clarify what payload should contain (e.g., URL vs. header JSON). The cost mention does hint that paymentSignature is relevant, but without saying when it is required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Audits') and a precise target: CORS response headers, with the exact dangerous condition spelled out (Access-Control-Allow-Origin: * + Credentials: true). This clearly separates it from the general verify_cors_headers sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended use is implied: use it when you need to audit a CORS configuration for wildcard-with-credentials misconfigurations. However, it never explicitly states when not to use it or contrasts it with verify_cors_headers or other security checkers.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
linux_seccomp_bpf_filter_generatorA
Generates minimal Seccomp BPF syscall filter profiles for sandboxing untrusted agent processes, blocking ptrace, fork, and raw socket creation. (0.040 USDC on Base L2)
| Name | Required | Description | Default |
|---|---|---|---|
| payload | Yes | Input parameters or JSON string payload for the tool execution | |
| paymentSignature | No | Base L2 USDC micropayment signature or transaction hash for x402 settlement |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full disclosure burden. It does add value by enumerating the blocked syscalls and disclosing the 0.040 USDC cost, but it does not state what the output looks like (e.g., BPF bytecode vs. source), whether it merely generates profiles or applies them, or any platform/kernel prerequisites.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single efficient sentence that front-loads the core purpose followed by the concrete syscall block list. The price parenthetical earns its place as cost context, though it is slightly tangential to the tool's function.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a security-sensitive tool with no output schema and no annotations, the description is incomplete: it never explains the return value format (critical for consuming the generated filter), prerequisites, or side effects. An agent cannot fully predict invocation behavior beyond the basic purpose.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description hints at what the payload relates to (syscall filtering) but does not clarify the expected payload format or how to specify the filter policy, leaving agents to infer the input structure.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it generates minimal Seccomp BPF syscall filter profiles, and names the exact syscalls blocked (ptrace, fork, raw socket creation). The sandboxing purpose and target (untrusted agent processes) differentiate it clearly from all siblings, none of which perform seccomp/BPF work.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'for sandboxing untrusted agent processes' gives a clear when-to-use context. It does not name alternatives or exclusions, but since no sibling tool overlaps functionally, the context is adequate for selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
python_ast_sandbox_escape_detectorA
Audits Python ASTs for dangerous builtins, __subclasses__ gadget chains, importlib, and bytecode compilation tricks used in sandbox escapes. (0.045 USDC on Base L2)
| Name | Required | Description | Default |
|---|---|---|---|
| payload | Yes | Input parameters or JSON string payload for the tool execution | |
| paymentSignature | No | Base L2 USDC micropayment signature or transaction hash for x402 settlement |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. The verb 'Audits' and the analysis-focused wording convey a read-only, non-destructive operation, and the cost note '(0.045 USDC on Base L2)' adds a payment constraint. However, it does not disclose potential side effects like external data transmission, failure modes, or whether results are persisted. Basic safety is implied but not detailed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, focused sentence that front-loads the core action and target, then lists the audited threat categories, and appends the cost. No filler, no repetition of schema fields. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is a specialized security analyzer with no output schema and no annotations, so the description must compensate. It does not specify what the tool returns (verdict, list of found issues, score, etc.), nor the exact input format for 'payload' (raw source code, serialized AST, JSON). The payment mechanism is also ambiguous: the description mentions a price but does not explain whether paymentSignature is required or how settlement works. These gaps leave an agent unsure how to invoke the tool correctly or interpret results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents both parameters (payload and paymentSignature) with 100% coverage, giving a baseline of 3. The description adds meaningful context by specifying that the payload is a Python AST and that the tool examines it for specific escape vectors. This goes beyond the generic schema description ('Input parameters or JSON string payload'), clarifying the tool's expected input domain.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Audits') and resource ('Python ASTs') and enumerates concrete threat patterns (dangerous builtins, __subclasses__ chains, importlib, bytecode tricks). This clearly distinguishes it from siblings like validate_code_syntax, diff_ast_trees, and safe_eval_math_expression, which address syntax validity, AST diffing, and safe expression evaluation respectively.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The purpose strongly implies when to use the tool: when a Python AST needs security auditing for sandbox escapes. However, there is no explicit mention of when not to use it or which sibling alternatives to prefer. The guidance is thus implied rather than stated, meeting the 'implied usage' level but not going further.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
regex_redos_exponential_scannerA
Audits regular expressions for catastrophic polynomial and exponential backtracking (ReDoS) vulnerabilities using NFA/DFA cycle decomposition. (0.035 USDC on Base L2)
| Name | Required | Description | Default |
|---|---|---|---|
| payload | Yes | Input parameters or JSON string payload for the tool execution | |
| paymentSignature | No | Base L2 USDC micropayment signature or transaction hash for x402 settlement |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so description carries the burden. It discloses the method (NFA/DFA cycle decomposition) and cost (0.035 USDC on Base L2), but does not describe the output format or any side effects or prerequisites. It adds some value but lacks critical behavioral details like return value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence plus cost note. Efficient and front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity and lack of output schema, the description is incomplete. It does not specify what the payload should contain (e.g., the regex) or what the tool returns. Missing critical usage context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description does not add any parameter semantics; the payload description is generic and the description does not clarify what to pass. It doesn't compensate for the vague schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (audits) and resource (regular expressions) with a specific vulnerability type (ReDoS). Distinguishes from siblings like generate_regex_dfa by focusing on vulnerability scanning rather than DFA generation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for auditing regexes for ReDoS but does not explicitly compare with alternatives or state when not to use. There is no mention of exclusions or alternative tools, so guidance is limited.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sql_ast_sqli_taint_analyzerA
Parses SQL query ASTs to verify parameterized binding, flagging raw string concatenations that lead to second-order SQL injection vulnerabilities. (0.040 USDC on Base L2)
| Name | Required | Description | Default |
|---|---|---|---|
| payload | Yes | Input parameters or JSON string payload for the tool execution | |
| paymentSignature | No | Base L2 USDC micropayment signature or transaction hash for x402 settlement |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full transparency burden. It reasonably conveys a read-only analysis behavior through 'parses', 'verify', and 'flagging', and it discloses the execution cost. It does not, however, describe failure behavior, input expectations, or whether it modifies anything.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, focused sentence that front-loads the core behavior and appends only the relevant cost. There is no repetition of the tool name and no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex analyzer with no output schema and no annotations, an agent still lacks critical invocation details: what exactly the payload should contain (SQL text vs AST JSON), how to represent the query, and what the output report looks like. The description explains the high-level purpose well but not enough to confidently construct a valid call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline 3 applies. The parameter descriptions are generic boilerplate, and the tool description only broadly hints that the payload relates to SQL queries and ASTs. It does not specify the exact payload structure, but the schema already provides some description for both parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific operation ('Parses SQL query ASTs'), the intended goal ('verify parameterized binding'), and the vulnerability it targets ('raw string concatenations that lead to second-order SQL injection'). This clearly distinguishes it from generic validation or sanitization siblings like validate_code_syntax and sanitize_sql_query.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended use is implied by the function statement: use this when you need to verify SQL parameterized binding via AST analysis. However, there are no explicit conditions, exclusions, or references to alternatives such as sanitize_sql_query, so an agent is left to infer when this tool should be preferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
v1.0.0- First observed
cors_policy_preflight_misconfig_checker - First observed
linux_seccomp_bpf_filter_generator - First observed
python_ast_sandbox_escape_detector - First observed
regex_redos_exponential_scanner - First observed
sql_ast_sqli_taint_analyzer
TDQS
Scored across 5 tools
Each tool targets a completely distinct artifact and vulnerability class (Python AST escape, SQL taint, regex ReDoS, Seccomp BPF, CORS headers). There is no realistic scenario where an agent would confuse one for another, and the descriptions reinforce the boundaries.
All five names follow a strict snake_case pattern of <domain>_<subject>_<agent-noun>, e.g. python_ast_sandbox_escape_detector, regex_redos_exponential_scanner. The convention is applied uniformly with no stylistic deviations.
Five tools is a well-scoped set for a focused security-audit server, with each tool covering a distinct check. No tool feels redundant or missing for the stated surface.
The surface covers several sandbox-relevant classes (escape, SQLi, ReDoS, seccomp, CORS) but omits common adjacent checks like shell injection, deserialization/pickle, and path traversal. Agents can work around these gaps, and notably one tool generates filters while the rest audit, a minor asymmetry.
Maintenance
Related MCP Connectors
Security reviews for coding agents: diffs checked against your org policy and live infrastructure.
Deterministic runtime safety for AI agents: scan PII, gate tool actions, verify LLM output.
The WAF for agents. Pattern-based + heuristic firewall scans prompts, RAG documents, tool argume...
Zero-config MCP security scanner for AI-generated apps. 25K+ vulnerability patterns.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceFull-stack security for AI agents — static analysis + MCP runtime interception. 31 rules detect prompt injection, data exfiltration, backdoors, tool poisoning, and cross-file attack chains. Includes MCP proxy for real-time blocking, Python AST taint tracking, multi-language injection detection (8 languages), and AI-powered deep analysis. Free, offline, zero-config.83 npm15MIT
- AlicenseAqualityDmaintenanceAgent-native "safe to ship?" security gate for AI-generated code. Uses real parsers and inter-rocedural taint analysis (JS/TS, Python, Go) to flag the classes AI coding agents get wrong — secrets, SQL injection, SS, SSRF, path traversal, command injection, weak JWT/CORS — and ranks findings by confidence. Exposes a scan tool over MCP.110 npm2MIT
- AlicenseNot gradedqualityCmaintenanceReal security scanners for AI coding agents — SAST (441 rules), secret detection (419+ patterns), dependency CVEs (OSV.dev), MCP/skill vetting, MITRE ATT&CK. Open-source, Rust, free14 npmApache 2.0
- AlicenseBqualityCmaintenanceProvides local, dependency-free security scanning tools for LLM configurations, prompts, RAG sources, and more, enabling AI coding agents to detect prompt injections and other vulnerabilities without external network access.8MIT