BlastRadius MCP
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@BlastRadius MCPscore this command before it runs: rm -rf / --no-preserve-root"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
BlastRadius MCP
A zero-trust security proxy for Model Context Protocol servers. It resolves each agent tool call the way the shell will, scores it 0-100 before it runs, blocks the dangerous ones, redacts credentials flowing in either direction, requires a signed human approval token for high-impact actions, detects multi-step attack sequences, and leaves a tamper-evident audit trail.
Apache-2.0. Two runtime dependencies. Runs locally, opens no sockets, phones nothing home, no license key and no paid tier.
The problem
Give an agent bash, filesystem, SQL, or cloud access through MCP and you have
handed a language model the ability to run rm -rf, DROP TABLE, or
aws s3 rb. Those are not hypotheticals, they are the first thing anyone writes
in a demo. Meanwhile:
Tool inputs and outputs carry unmasked AWS keys, GitHub PATs, and database passwords.
Nothing simulates the consequences before the call lands.
Nothing records that the call happened, so nothing can be reviewed afterwards.
Existing options ask you to route every call through a hosted control plane, or to rewrite your agent loop. BlastRadius sits in front of the MCP server instead. Your agent's loop does not change.
And the second problem: matching the text is not the same as matching the command
This is the one that gets products like this wrong.
A guardrail matches the submitted string. The shell executes the resolved string. Bash rewrites the text first, through quote removal, parameter expansion, command substitution, and pipeline assembly. So these are the same command to the kernel and completely different strings to a regex:
D=/; rm -rf $D* text: harmless-looking runs: rm -rf /*
rm -rf $(echo /) text: no slash runs: rm -rf /
bash -c "rm -rf /" text: a bash invocation runs: rm -rf /
cat payload.sh | sh text: reading a file runs: an unreviewed scriptA guard evaluating the pre-transformation string and a shell executing the post-transformation string are evaluating two different commands.
This is not a hypothetical concern. The GuardFall survey, published 30 June 2026, tested the pre-execution command guards in eleven actively maintained coding agents spanning roughly 548,000 GitHub stars. Agents whose guard regexes the raw command string leaked on the large majority of bypass cases attempted against them. Goose leaked on 22 of 23. OpenCode on 16 of 16. Only Continue, which tokenizes with shell-aware parsing and resolves expansion and substitution before matching the fully resolved command, blocked all twenty-one cases.
So this engine resolves first, then matches. See Seeing past the bypass.
Related MCP server: agent-trust-firewall
Install
npx -y blastradius-mcpRequires Node 22+. That is the whole install.
Standalone
Run it as its own MCP server and call its tools from your agent.
{
"mcpServers": {
"blastradius": { "command": "npx", "args": ["-y", "blastradius-mcp"] }
}
}Proxy mode
Wrap an existing server. Nothing downstream changes.
npx blastradius-mcp proxy --command "npx -y @modelcontextprotocol/server-filesystem /workspace"{
"mcpServers": {
"guarded-filesystem": {
"command": "npx",
"args": [
"-y", "blastradius-mcp", "proxy",
"--command", "npx -y @modelcontextprotocol/server-filesystem /workspace"
]
}
}
}Working configs for Cursor, Cline, and VS Code are in examples/.
Installable as a one-click local MCP bundle via manifest.json.
Real output
Every block below was produced by running the engine, not written by hand.
Regenerate them yourself with node scripts/generate-readme-transcripts.mjs.
Scoring a destructive action
// simulate_action("rm -rf / --no-preserve-root")
{
"dangerScore": 100,
"severity": "CRITICAL",
"category": "SHELL",
"reasons": [
"Recursive root, home, or wildcard filesystem deletion detected (rm -rf / or *)"
],
"destructive": true,
"irreversible": true,
"rollbackFeasible": false,
"recommendedMitigations": [
"Target specific named directories only",
"Use trash-cli or create a verified dry-run backup first"
]
}The same command, different environment
Score 95 on staging. 100 when the context says production.
// simulate_action("aws s3 rb s3://prod-bucket --force", context={environment:"production"})
{
"dangerScore": 100,
"severity": "CRITICAL",
"category": "CLOUD",
"reasons": [
"Object storage mass deletion (bucket removal or a syncing delete) destroys data and versions",
"Target environment is marked as PRODUCTION (production), escalating blast radius."
],
"destructive": true,
"irreversible": true,
"rollbackFeasible": false
}A read is not a risk
// simulate_action("SELECT id, email FROM users LIMIT 10")
{
"dangerScore": 5,
"severity": "SAFE",
"category": "SQL",
"destructive": false,
"irreversible": false,
"rollbackFeasible": true,
"reasons": ["Standard read-only or low-impact routine action."]
}Redacting credentials on the way in
// inspect_payload_dlp({authorization:"Bearer AKIAIOSFODNN7EXAMPLE",
// dsn:"postgres://admin:hunter2@db.internal:5432/prod",
// card:"4111 1111 1111 1111", note:"owner joe@acme.com"})
{
"hasFindings": true,
"findingsCount": 5,
"findings": [
{ "type": "AWS_ACCESS_KEY", "category": "CREDENTIAL", "severity": "CRITICAL", "preview": "AKIA...MPLE" },
{ "type": "DATABASE_CONN_URI_PASSWORD", "category": "CREDENTIAL", "severity": "HIGH", "preview": "post...er2@" },
{ "type": "CREDIT_CARD_NUMBER", "category": "PII", "severity": "HIGH", "preview": "4111...1111" },
{ "type": "EMAIL_ADDRESS", "category": "PII", "severity": "MEDIUM", "preview": "joe@....com" }
],
"sanitizedContent": "{\"authorization\":\"Bearer [REDACTED_AWS_KEY:MPLE]\",\"dsn\":\"postgres://admin:[REDACTED_PASSWORD]@db.internal:5432/prod\",\"card\":\"[REDACTED_CREDIT_CARD:1111]\",\"note\":\"owner [REDACTED_EMAIL:.com]\"}"
}The policy gate blocks
// enforce_policy("bash", {command: "rm -rf /var/lib/postgresql"})
{
"decision": "BLOCK",
"ruleMatched": "ZT-001",
"reasons": [
"Violates rule [Block Root and Wildcard Deletion]: Matches strictly forbidden pattern 'rm -rf /'"
],
"requiresToken": false
}Step-up approval, and the replay it prevents
The same call three times: no token, valid token, same token again.
// 1. no token
{
"decision": "REQUIRE_CONFIRMATION",
"ruleMatched": "ZT-003",
"reasons": ["Elevated risk (Severity: CRITICAL, Danger Score: 95). Cryptographic confirmation token required to proceed."],
"requiresToken": true,
"tokenValid": false
}
// request_confirmation_token(...).payload
{
"tokenId": "br-tok-62b3da7b-67c9-48b4-b766-f1a12d4a25bf",
"toolName": "bash",
"actionFingerprint": "95:terraform-destroy",
"requestedBy": "jimil@acme.dev",
"issuedAt": 1791190531,
"expiresAt": 1791190831,
"reason": "Approved in ticket OPS-4471"
}
// 2. with the token
{
"decision": "ALLOW",
"ruleMatched": "ZT-003",
"reasons": ["Action elevated and approved via valid token: br-tok-62b3da7b-..."],
"requiresToken": false,
"tokenValid": true
}
// 3. same token again
{
"decision": "REQUIRE_CONFIRMATION",
"reasons": [
"Provided confirmation token is invalid or expired: Token has already been consumed (replay prevention)",
"Elevated risk (Severity: CRITICAL, Danger Score: 95). Cryptographic confirmation token required to proceed."
],
"requiresToken": true,
"tokenValid": false
}That third block is the whole point. An approval token authorises exactly one action.
The six tools
Tool | What it does |
| Scores an action 0-100 before execution, reports severity, affected entities, and whether rollback is possible |
| Finds and redacts secrets and PII in any text |
| The decision point. Returns |
| Issues a signed, single-use, tool-scoped approval token |
| Re-walks the hash chain and its signatures to detect tampering |
| Live metrics, active policy, build capabilities |
Seeing past the bypass
src/analyzer/shellResolver.ts resolves a command
the way a POSIX shell would, to the extent that is possible without executing
anything. The rule patterns then run against what the shell would actually run.
The pass performs, in order:
Quote removal.
rm -rf "/"andrm -rf /are the same command.Variable assignment tracking.
D=/; rm -rf $D*resolves$Dfrom the assignment earlier in the same command, because the shell would.Parameter expansion.
$HOME,$ROOT,$PWDand friends expand. When the value is not statically knowable it is substituted as a markedUNKNOWN, which is treated as a root-or-home target rather than an opaque token, so a recursive delete aimed at an unseen variable still matches.Command substitution.
$(...)and backticks are resolved to the literal content inside them. This is deliberately conservative:rm -rf $(echo /)resolves to a token containing a slash, which is enough to match. Nothing is ever executed to find out what a substitution would produce.Pipeline splitting. Stages are separated at the top level only, so
echo 'a | b'stays one stage.Inline program extraction.
sh -ctakes the whole remainder of the line as its program, sobash -c "rm -rf /"resolves to the destructive command it runs.Interpreter detection. Any stage that receives content from a pipe and is itself an interpreter is flagged, because the program that runs is not in the request.
On top of that the engine scores structural facts no regex can express:
a destructive binary (
rm,unlink,shred,srm) aimed at a root, home, or wildcard targetshredandsrmagainst any path, because they overwrite in place and so cannot be recovered even without a recursive flaga recursive
rmagainst a system path such as/etcdestruction delegated to
findvia-delete,-exec rm, or-okcontent piped into an interpreter, with or without a substitution
arguments assembled at runtime by
xargsa destructive command composed inline inside
sh -c
Every report carries the resolution block, so a reviewer can see the exact text
that was matched:
// simulate_action("D=/; rm -rf $D*")
{
"dangerScore": 100,
"severity": "CRITICAL",
"reasons": [
"Recursive rm against a broad target (-rf /*), resolved through variable or substitution expansion."
],
"resolution": {
"resolvedCommand": "D=/; rm -rf /*",
"traits": ["assigned:D", "unresolved-variable:D"],
"hasPipeline": false,
"interpreters": [],
"unresolvedVariables": []
}
}Coverage
tests/shellResolver.test.ts pins the GuardFall bypass classes. 23 of 23 dangerous
commands that a raw-string guard would miss are scored HIGH or above:
Class | Examples |
Variable expansion |
|
Command substitution |
|
Quoting |
|
Alternate destructive binaries |
|
Indirect execution |
|
The same file pins the other direction: 18 routine commands that must not trip,
including rm -rf node_modules, rm -rf ./build, kubectl apply -f deployment.yaml,
and terraform plan. A guard that blocks everything is not a guard.
What resolution still does not catch
Stated plainly, because this is the tier the literature says gets bypassed and pretending otherwise would repeat the mistake this project already fixed once:
Encoding. Base64, hex, and ROT-encoded payloads are not decoded.
Multi-hop indirection.
bash /tmp/x.shwherex.shwas written three calls ago is opaque to a single-call analysis.A destructive payload assembled character by character at runtime.
Anything outside a shell. MCP tool arguments that are not command strings go through scoring, but not through a resolver.
Resolution closes the specific gap GuardFall documented for shell commands. It is not general program analysis, and no honest claim should say it is.
Detection coverage
24 signatures in src/analyzer/rules.ts, scored by rule ID.
Scoring takes the max matching rule, not a sum, so overlapping patterns do not
inflate a score. All 24 are matched against the resolved command, so the
alternative destructive binaries and indirections listed above are reachable
without a dedicated rule.
ID | Scores | Covers |
| 100 |
|
| 100 |
|
| 95 | Fork bombs |
| 90 |
|
| 80 |
|
| 85 |
|
| 85 |
|
| 95 |
|
| 100 |
|
| 90 |
|
| 90 |
|
| 85 |
|
| 80 |
|
| 95 |
|
| 95 |
|
| 85 |
|
| 95 |
|
| 95 |
|
| 85 |
|
| 80 |
|
| 80 |
|
| 90 | Mass deletion by synchronisation and pruning: |
| 95 | Destructive database operations issued through a client binary: |
| 90 | Encoded content piped to a decoder that then executes: |
Bands: SAFE <15, LOW 15-39, MEDIUM 40-69, HIGH 70-89, CRITICAL 90+.
Categories: SHELL, SQL, CLOUD, FILESYSTEM, NETWORK, GENERIC.
11 DLP detectors in src/analyzer/dlpScanner.ts:
Credentials: AWS access keys, GitHub PATs (
ghp_andgithub_pat_), OpenAI, Stripe secret keys, Slack tokens, SSH private keys, database URI passwords, JWTsPII: credit cards (Visa, Mastercard, Amex, Discover, Diners, UnionPay, in contiguous, spaced, or hyphenated form), US SSNs, email addresses
What it does not catch
Stated up front, because a rule set that pretends to be complete is worse than one that does not:
Encoded payloads. Base64, hex, and string concatenation defeat both the resolver and the rules. See what resolution still does not catch.
Multi-step and multi-call attacks within one call. Sequence detection below covers the audited cases, not arbitrary workflows.
Anything semantic. It does not know that
orders_archive_2019matters less thanorders.Agent-supplied identity.
requestedByon an approval token is free text from the calling agent. See Step-up approval tokens.
Open an issue if you have a dangerous pattern that is not covered. The rule set is the part most worth improving.
Default policy
Four rules, active out of the box, in src/policy/defaultPolicies.ts.
ID | Action | Detail |
| BLOCK |
|
| BLOCK | Fork bombs, |
| REQUIRE_CONFIRMATION | Any |
| BLOCK |
|
Override it with your own JSON via --policy, see examples/custom_policy.json:
npx blastradius-mcp --policy ./my-policy.jsonZT-004 matches with a plain includes() on the serialized parameters. That is
deliberately conservative and it does over-block: any path containing .env
matches, including .envrc. Tune it with severityThreshold rather than by
loosening the pattern.
Sequence detection — the part a single call cannot see
Per-call allow/deny misses workflows assembled from individually benign steps. Microsoft's own analysis of an MCP control plane states the gap directly:
"Per-call allow still misses malicious workflows built from individually permitted steps. Sequence policy is the next real control."
src/analyzer/sequenceDetector.ts is that
control. It runs over the hash-chained decision log after the per-call verdict and
detects five workflow shapes:
Signal | Fires when | Default escalation |
| A credential-adjacent read is followed by a destructive or irreversible decision |
|
| A run of ≥5 reads across ≥3 distinct payloads, terminated by a network or DLP-positive sink |
|
| ≥3 distinct privilege families approved in a non-decreasing severity run, or a strictly rising score on one tool |
|
| ≥3 refusals against one tool family |
|
| ≥4 destructive decisions inside 8 decisions, boosted when the span is under 120 seconds |
|
Three properties that matter:
It can only tighten a decision, never loosen one. If the per-call rule already blocked or demanded a token, that stands. Sequence analysis never downgrades.
Every signal cites its evidence. The reason string names the entry indices that triggered it, so a reviewer can reconstruct the finding rather than trust a score.
It reports its own blind spot.
truncated: trueis returned when the analysis window excludes earlier history, because a pattern beginning before the window cannot be found inside it.
Thresholds are tunable through the same policy document as everything else:
{
"sequence": {
"enabled": true,
"windowSize": 20,
"destructiveBurstThreshold": 4,
"probingThreshold": 3
}
}What sequence detection still misses
Documented deliberately, because a detector with unstated limits is not a detector you can rely on:
Sub-window attacks. Analysis sees only the last
windowSizedecisions. An agent that reads a secret, runs 21 benign calls, then destroys will settruncated: trueand produce no signal.Cross-caller laundering. Nothing correlates by
callerId. An attack split across two agent identities resets both the ordering and the burst span.Slow, patient exfiltration.
reconnaissance-then-exfiltrationrequires a contiguous read run immediately before the sink. Five reads interleaved with writes across ten minutes is not a run. The detector is tuned for agent-speed attacks, which is the same thing that makes it weak against a patient one.Destructive classification is reconstructed from severity, score, and reason wording, because the ledger does not persist the engine's
destructiveflag. Changing reason wording without updating the detector narrows it silently.
Audit ledger
Append-only JSONL, SHA-256 hash chain, one HMAC-SHA256 signature per entry.
Payloads are hashed, not stored, so secrets never land in the log.
verify_audit_logre-computes every hash and re-computes every signature. The hash chain alone does not coverseverity,category,reasons, ordlpFindingsCount; the signature is what covers those, which is why it is verified rather than merely written.Editing any field of any entry breaks verification and reports the index.
The signing key defaults to a random per-process value, so set it explicitly or previously written entries will not verify after a restart:
export BLAST_RADIUS_AUDIT_KEY="a-long-random-secret"
export BLAST_RADIUS_SIGNING_KEY="a-different-long-random-secret"The ledger defaults to process.cwd(), which under a GUI client is wherever the
app happened to launch from. Point it somewhere deliberate:
export BLAST_RADIUS_AUDIT_PATH="/var/log/blastradius/audit.jsonl"What the chain does not do: detect deletion of entries (the ledger is verified from the in-memory copy), detect truncation, or survive a full rewrite with recomputed hashes. There is no external or Merkle anchoring. Treat it as tamper-evident, not tamper-proof.
Step-up approval tokens
HMAC-SHA256, JWT-shaped payload.signature, 300-second default TTL, hard-capped at
one hour.
Signature compared with
crypto.timingSafeEqual, never===.Scoped to a single tool name, or
*for any tool.Single-use. The token is consumed the moment it authorises an action, so replaying it returns
REQUIRE_CONFIRMATIONwithToken has already been consumed (replay prevention).Consumed records are pruned once the token's own expiry passes, which bounds memory without reopening the hole.
One honest limitation: requestedBy is a free-text string supplied by the calling
agent. There is no out-of-band human channel and no identity provider behind it.
The token proves an approval was issued and spent once. It does not prove a
specific human pressed a button. Wiring it to a real approval service is the
natural next step, and it is not built.
Proxy mode
npx blastradius-mcp proxy --command "<downstream mcp server command>"Inspects inbound tools/call requests and, more usefully, scans downstream
responses and redacts credentials before they reach the client. A tool that
echoes a key back at your model is a routine leak path, and nothing upstream of it
usually cleans up.
A step-up token is accepted as a confirmationToken argument on the call. It is
stripped before forwarding, so an approval credential never reaches a server with
no business holding it.
Known limitations, stated plainly:
The downstream command is split on whitespace, so a path containing a space breaks.
It is a line-oriented JSON-RPC passthrough, not a full MCP SDK client.
The proxy exposes the downstream server's tools, not BlastRadius's own six.
Testing
npm run test:srcℹ tests 145
ℹ pass 145
ℹ fail 0145 tests across nine files, covering the scoring engine, the DLP scanner, the policy engine, the token manager, audit-ledger tamper detection, shell-aware command resolution with the GuardFall bypass classes, false-positive pinning on routine commands, sequence and workflow detection, and a full JSON-RPC lifecycle over a real stdio subprocess. CI runs the suite on Node 22 and 24, with CodeQL on javascript-typescript.
The build is tsc under strict: true. No bundler, no framework, no test-runner
dependency.
Three probes are kept in scripts/ because they are more useful as runnable output
than as assertions:
node scripts/probe-guardfall.mjs # 23 bypass classes, one line each
node scripts/probe-false-positives.mjs # 39 routine commands that must not trip
node scripts/probe-sequence.mjs # workflow patterns firing, and staying quietConfiguration
Variable | Default | Purpose |
| random per process | HMAC secret for approval tokens |
| random per process | HMAC secret for audit entries |
|
| Where the ledger is written |
Set the first two or tokens and audit entries stop verifying across restarts.
Flag | Purpose |
| Load a custom policy JSON |
| Guard an existing MCP server |
|
A note on how this shipped
Worth reading, because it is the honest version of what this project is.
The first public version of this repo carried a tests 24/24 passing badge. There
were 11 tests, in 2 files. Four of the six test files were 0 bytes: the audit
ledger, the policy engine, the token manager, and the end-to-end server had no
coverage at all. It also claimed single-use tokens with replay protection, where
the function that enforced it existed and was never called. It shipped a 0-byte
LICENSE behind an Apache-2.0 badge. And its README advertised rate limits, x402
micropayments, and Splunk SIEM export, none of which existed.
All of it is fixed: 145 real tests, a real CI gate, tokens that are genuinely
single-use, an actual Apache-2.0 LICENSE, and a paid-tiers section deleted
rather than softened. An Apache-2.0 repository cannot enforce a paywall anyway.
Then there was the second one, which is more interesting.
On 30 June 2026, a paper named the exact design this project used and measured it as broken. GuardFall tested the command guards in eleven coding agents and found that matching the raw command string leaked on most bypass attempts. That was this project's engine. It shipped regexes over submitted text, exactly the tier the survey measured at 22-of-23 and 16-of-16 failure.
So the engine was rewritten to resolve commands before matching them: quote
removal, variable assignment tracking, parameter expansion, command substitution,
pipeline splitting, and inline program extraction, all without executing anything.
All 23 bypass cases now score HIGH or above, pinned by
tests/shellResolver.test.ts.
The reason this is in the README rather than a changelog entry is that a tool whose entire pitch is I stop agents doing dangerous things had a bug in the category it exists to prevent, twice. The first time I did not look. The second time I read the literature and found the paper that said so. If you are evaluating this, you should know that the person maintaining it now checks the claims, and how.
Contributing
See CONTRIBUTING.md. The rule set is the highest-value
contribution: a new signature in src/analyzer/rules.ts needs a test in the
matching tests/*.test.ts and a note about expected false positives.
Security issues: SECURITY.md, or use GitHub private vulnerability reporting.
License
Apache-2.0. See LICENSE.
Available Tools
6 toolsenforce_policyB
The central zero-trust decision point. Evaluates a prospective tool call against organizational security policies, scans parameters for DLP, and records a cryptographic audit trail.
| Name | Required | Description | Default |
|---|---|---|---|
| callerId | No | Unique identifier for the calling agent or user. | agent-client |
| toolName | Yes | The name of the target tool intended to be called. | |
| parameters | Yes | The parameters intended to be passed to the tool. | |
| confirmationToken | No | Optional cryptographically signed approval token for elevated destructive actions. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses a side effect ('records a cryptographic audit trail'), which tells the agent this is not a pure read, but it omits critical behavior: what a deny decision looks like, whether the call is blocked or merely reported, and how confirmationToken gates elevated destructive actions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with the most important framing front-loaded and zero filler. It could be marginally shorter, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a central security gate with no annotations and no output schema, the description is under-specified: it never says what the tool returns (allow/deny, reason codes) or explains the confirmation-token elevation flow described in the schema. The nested-object parameters argument is also left opaque.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters are already documented in the schema, establishing a baseline of 3. The description only indirectly references 'parameters' as being scanned for DLP and adds no format or constraint detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific function with strong framing ('central zero-trust decision point') plus three concrete actions: policy evaluation, DLP scanning, and audit recording. It is distinguishable from most siblings, though it overlaps conceptually with inspect_payload_dlp (DLP) and simulate_action (evaluation) without naming how it differs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'prospective tool call' implies this is a pre-execution gate, which is useful context, but there is no explicit when-to-use or when-not-to-use guidance and no mention of the sibling tools an agent should consider instead (e.g., simulate_action for dry runs). Usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_security_postureB
Returns current operational security metrics, the active policy profile, build edition and capabilities, blocked threat statistics, and DLP redaction totals.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It implies a read-only snapshot through the verb 'returns', but never states that it is non-mutating, whether it requires elevated privileges, whether results are cached or point-in-time, or how large/fresh the payload is. It adds return-content detail but no operational behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single well-formed sentence that front-loads the action ('Returns current operational security metrics') and then lists the payload contents. No filler or redundancy. Readability is slightly hurt by a long comma-delimited list, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must explain what comes back, and it does so item by item (metrics, policy profile, edition/capabilities, blocked threats, DLP totals). That covers the return contract well. The missing piece is the behavioral/usage layer — freshness, point-in-time semantics, and when to call it versus its siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, and the schema is trivially complete, so there is nothing for the description to disambiguate. Baseline for a parameterless tool is 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it returns current security posture data, and it enumerates the concrete contents (operational metrics, active policy profile, build edition/capabilities, blocked threat stats, DLP redaction totals). That is far more informative than a tautology. It does not explicitly contrast itself with siblings such as inspect_payload_dlp or enforce_policy, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool, when not to, or which sibling to prefer for adjacent questions. An agent must infer from the name alone that this is the read-side counterpart to enforce_policy or verify_audit_log. No prerequisites or contexts are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inspect_payload_dlpC
Deep Data Loss Prevention (DLP) scanner. Detects, reports, and redacts API keys, passwords, JWTs, cloud credentials, credit cards, SSNs, and PII from prompts, code, and logs.
| Name | Required | Description | Default |
|---|---|---|---|
| content | Yes | The raw text, code snippet, or log output to inspect. | |
| maskSensitive | No | Whether to return a sanitized version with credentials and PII masked. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It mentions detection, reporting, and redaction, but does not disclose whether the operation is read-only, whether it modifies the input, any authentication needs, or rate limits. The default maskSensitive=true behavior is only described in the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with the purpose front-loaded. The list of detectable data types is long but informative and earns its place. There is no wasted wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description should ideally explain the return value (findings format, redacted output) and safety posture. It covers purpose and inputs but leaves output expectations and behavioral constraints unclear, making it only minimally complete for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (content and maskSensitive) are fully documented in the schema. The description adds no additional meaning or syntax beyond what the schema already provides, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb (scanner) and resource (payload for DLP) and enumerates specific sensitive data types it handles. It does not explicitly differentiate from sibling tools like get_security_posture or verify_audit_log, but the payload-inspection focus is distinctive enough for an agent to identify it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use or when-not-to-use guidance, and no alternatives are named. The description implies usage for scanning prompts, code, and logs, but leaves the agent to infer when to call this versus the other security tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
request_confirmation_tokenA
Generates a time-bound, cryptographically signed (HMAC-SHA256) step-up approval token allowing an agent to execute high-impact actions when human consent is granted.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | Justification for elevation. | |
| toolName | Yes | The specific tool name to authorize. | |
| ttlSeconds | No | Token validity period in seconds (default: 300s). | |
| requestedBy | Yes | Identifier of the human supervisor or authorizing entity. | |
| actionFingerprint | Yes | Unique signature or hash of the approved payload. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses important traits such as HMAC-SHA256 signing, time-bound validity, and the step-up approval nature, but it omits authentication requirements, rate limits, failure modes, and whether token generation itself is gated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, dense sentence with no redundant phrasing. The key concept—generating a signed step-up token—is front-loaded, making it well-structured for quick scanning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and no annotations, the description is only moderately complete. It explains what the token is for generally, but does not describe the return value, token format, error cases, or how the human consent step is verified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters are already documented in the input schema. The description adds no parameter-level detail beyond what the schema provides, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: generates a time-bound, HMAC-SHA256-signed step-up approval token. It also explains the token's purpose for high-impact actions. However, it does not explicitly differentiate this tool from siblings such as enforce_policy or simulate_action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides a clear context for use: when an agent needs to execute high-impact actions and human consent is granted. It lacks explicit when-not-to-use guidance and does not name alternative tools for related authorization or policy tasks.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
simulate_actionA
Simulates and predicts the blast radius, danger score (0-100), and destructive reversibility of a shell command, SQL query, or cloud operation before execution.
| Name | Required | Description | Default |
|---|---|---|---|
| context | No | Execution context such as environment (prod/stage), working directory, or target host. | |
| actionType | No | Optional action category hint. | |
| commandOrQuery | Yes | The shell command, SQL query, or cloud API call to simulate. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses the return content (blast radius, 0-100 danger score, reversibility), but never states explicitly that the tool does NOT execute the command, whether it needs elevated permissions, or how it handles unknown/unparseable input — key facts for a safety-oriented simulator.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that packs the purpose, accepted input types, and expected outputs with zero filler. Nothing redundant or off-topic.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description correctly takes on the job of naming the return values, which it does. But for a pre-execution risk-assessment tool it omits whether the action is actually executed, the result format, and any annotation-level safety guarantees, leaving meaningful gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters (including the nested context object and the actionType enum) are already documented in the schema. The description only echoes the shell/SQL/cloud mapping, adding no syntax, format, or enum-value meaning beyond the schema; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (simulates/predicts) and resource (shell command, SQL query, cloud operation), plus the exact outputs produced (blast radius, danger score 0-100, reversibility). It is clearly distinct from the sibling tools like enforce_policy or request_confirmation_token, but it never explicitly names or contrasts an alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'before execution' implies the intended usage window (pre-flight dry run), which is meaningful for an agent deciding when to call this. However, there is no explicit when-not guidance, no prerequisites, and no routing between this and siblings such as enforce_policy or get_security_posture.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_audit_logB
Cryptographically verifies the immutable hash-chained audit ledger to detect tampering, deleted records, or unauthorized log modifications.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Number of recent audit records to verify. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses the verification mechanism (hash-chain) and the classes of anomalies it can detect, which is meaningful context, but it does not state that the operation is read-only/non-mutating, whether it requires elevated permissions, or whether cryptographic verification is expensive or long-running over a large ledger.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with the core action front-loaded and no filler. Every clause (mechanism, target, detected anomalies) earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool this is close to adequate, but with no output schema the description should hint at what a verification result contains (pass/fail, offending record index, proof details). It also omits the important scoping caveat implied by 'limit' — that only recent records are checked, so tampering further back could be missed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the single 'limit' parameter is fully documented in the schema as 'Number of recent audit records to verify.' The description adds no semantics about the parameter (e.g., that verifying only a subset leaves earlier ledger history unchecked), so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and mechanism ('cryptographically verifies') applied to a specific resource ('immutable hash-chained audit ledger') and states the outcome ('detect tampering, deleted records, or unauthorized log modifications'). This is far more than a restatement of the name, but it never references any sibling tool (e.g., get_security_posture, inspect_payload_dlp) to help the agent route between them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to invoke this versus alternatives such as get_security_posture or inspect_payload_dlp, and no mention of prerequisites or when verification is unnecessary. The description is purely declarative about what the tool does, leaving the selection decision entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v0.1.0- First observed
enforce_policy - First observed
get_security_posture - First observed
inspect_payload_dlp - First observed
request_confirmation_token - First observed
simulate_action - First observed
verify_audit_log
TDQS
Scored across 6 tools
Each tool has a distinct primary purpose spanning prediction, DLP scanning, posture reporting, policy enforcement, token issuance, and audit verification. There is mild overlap where enforce_policy also scans parameters for DLP and evaluates actions like simulate_action, but the descriptions clarify their respective roles well.
All six tools follow a strict verb_noun pattern (simulate_action, inspect_payload_dlp, get_security_posture, enforce_policy, request_confirmation_token, verify_audit_log). The convention is applied uniformly with no deviations.
Six tools is well-scoped for a security guardrail server, with each tool mapping to a discrete responsibility in the pre-execution safety workflow. No tool feels redundant or missing at the count level.
The surface covers a full guardrail lifecycle: simulation, DLP inspection, policy enforcement, step-up approval, posture reporting, and audit verification. Minor gaps exist such as no tool to view/configure policy definitions or revoke issued tokens, but core workflows are covered.
Maintenance
Related MCP Connectors
Deterministic runtime safety for AI agents: scan PII, gate tool actions, verify LLM output.
Zero-trust gateway for AI agents: score tool calls, verify agent cards, enforce policy, audit.
See, price, and control every tool call your AI agents make: policy checks, cost, and audit tools.
Security gateway for AI agents: policy, approval, and audited execution, no secrets shared.
Related MCP Servers
- FlicenseNot gradedqualityBmaintenanceProvides a secure MCP boundary for AI agents, intercepting and validating tool calls, redacting secrets, and requiring human approval for sensitive actions with a tamper-evident audit trail.-
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to securely invoke tools by enforcing identity proof, capability verification, and risk scoring on every request, blocking unsafe calls before they execute.MIT
- AlicenseNot gradedqualityBmaintenanceEnables autonomous agents to run shell and subprocess commands through a guardrail that blocks destructive system mutations, path traversals, and reverse shells. It also adds prompt-injection detection, spend circuit breakers, loop detection, and privacy sanitization so agent execution stays within safe, budgeted bounds.7MIT
- AlicenseNot gradedqualityBmaintenanceEnables autonomous agents to enforce financial guardrails by tracking real-time token burn and cost velocity, automatically throttling or freezing execution when budgets are exceeded. It also blocks destructive shell commands, prompt injections, and runaway planning loops while sanitizing outputs before they are returned.7MIT