Skip to main content
Glama
Jimil-Joshi

BlastRadius MCP

BlastRadius MCP

A zero-trust security proxy for Model Context Protocol servers. It resolves each agent tool call the way the shell will, scores it 0-100 before it runs, blocks the dangerous ones, redacts credentials flowing in either direction, requires a signed human approval token for high-impact actions, detects multi-step attack sequences, and leaves a tamper-evident audit trail.

Apache-2.0. Two runtime dependencies. Runs locally, opens no sockets, phones nothing home, no license key and no paid tier.

npm tests ci codeql license node types


The problem

Give an agent bash, filesystem, SQL, or cloud access through MCP and you have handed a language model the ability to run rm -rf, DROP TABLE, or aws s3 rb. Those are not hypotheticals, they are the first thing anyone writes in a demo. Meanwhile:

  • Tool inputs and outputs carry unmasked AWS keys, GitHub PATs, and database passwords.

  • Nothing simulates the consequences before the call lands.

  • Nothing records that the call happened, so nothing can be reviewed afterwards.

Existing options ask you to route every call through a hosted control plane, or to rewrite your agent loop. BlastRadius sits in front of the MCP server instead. Your agent's loop does not change.

And the second problem: matching the text is not the same as matching the command

This is the one that gets products like this wrong.

A guardrail matches the submitted string. The shell executes the resolved string. Bash rewrites the text first, through quote removal, parameter expansion, command substitution, and pipeline assembly. So these are the same command to the kernel and completely different strings to a regex:

D=/; rm -rf $D*          text: harmless-looking     runs: rm -rf /*
rm -rf $(echo /)         text: no slash             runs: rm -rf /
bash -c "rm -rf /"       text: a bash invocation    runs: rm -rf /
cat payload.sh | sh      text: reading a file       runs: an unreviewed script

A guard evaluating the pre-transformation string and a shell executing the post-transformation string are evaluating two different commands.

This is not a hypothetical concern. The GuardFall survey, published 30 June 2026, tested the pre-execution command guards in eleven actively maintained coding agents spanning roughly 548,000 GitHub stars. Agents whose guard regexes the raw command string leaked on the large majority of bypass cases attempted against them. Goose leaked on 22 of 23. OpenCode on 16 of 16. Only Continue, which tokenizes with shell-aware parsing and resolves expansion and substitution before matching the fully resolved command, blocked all twenty-one cases.

So this engine resolves first, then matches. See Seeing past the bypass.


Related MCP server: agent-trust-firewall

Install

npx -y blastradius-mcp

Requires Node 22+. That is the whole install.

Standalone

Run it as its own MCP server and call its tools from your agent.

{
  "mcpServers": {
    "blastradius": { "command": "npx", "args": ["-y", "blastradius-mcp"] }
  }
}

Proxy mode

Wrap an existing server. Nothing downstream changes.

npx blastradius-mcp proxy --command "npx -y @modelcontextprotocol/server-filesystem /workspace"
{
  "mcpServers": {
    "guarded-filesystem": {
      "command": "npx",
      "args": [
        "-y", "blastradius-mcp", "proxy",
        "--command", "npx -y @modelcontextprotocol/server-filesystem /workspace"
      ]
    }
  }
}

Working configs for Cursor, Cline, and VS Code are in examples/. Installable as a one-click local MCP bundle via manifest.json.


Real output

Every block below was produced by running the engine, not written by hand. Regenerate them yourself with node scripts/generate-readme-transcripts.mjs.

Scoring a destructive action

// simulate_action("rm -rf / --no-preserve-root")
{
  "dangerScore": 100,
  "severity": "CRITICAL",
  "category": "SHELL",
  "reasons": [
    "Recursive root, home, or wildcard filesystem deletion detected (rm -rf / or *)"
  ],
  "destructive": true,
  "irreversible": true,
  "rollbackFeasible": false,
  "recommendedMitigations": [
    "Target specific named directories only",
    "Use trash-cli or create a verified dry-run backup first"
  ]
}

The same command, different environment

Score 95 on staging. 100 when the context says production.

// simulate_action("aws s3 rb s3://prod-bucket --force", context={environment:"production"})
{
  "dangerScore": 100,
  "severity": "CRITICAL",
  "category": "CLOUD",
  "reasons": [
    "Object storage mass deletion (bucket removal or a syncing delete) destroys data and versions",
    "Target environment is marked as PRODUCTION (production), escalating blast radius."
  ],
  "destructive": true,
  "irreversible": true,
  "rollbackFeasible": false
}

A read is not a risk

// simulate_action("SELECT id, email FROM users LIMIT 10")
{
  "dangerScore": 5,
  "severity": "SAFE",
  "category": "SQL",
  "destructive": false,
  "irreversible": false,
  "rollbackFeasible": true,
  "reasons": ["Standard read-only or low-impact routine action."]
}

Redacting credentials on the way in

// inspect_payload_dlp({authorization:"Bearer AKIAIOSFODNN7EXAMPLE",
//                      dsn:"postgres://admin:hunter2@db.internal:5432/prod",
//                      card:"4111 1111 1111 1111", note:"owner joe@acme.com"})
{
  "hasFindings": true,
  "findingsCount": 5,
  "findings": [
    { "type": "AWS_ACCESS_KEY",              "category": "CREDENTIAL", "severity": "CRITICAL", "preview": "AKIA...MPLE" },
    { "type": "DATABASE_CONN_URI_PASSWORD",  "category": "CREDENTIAL", "severity": "HIGH",     "preview": "post...er2@" },
    { "type": "CREDIT_CARD_NUMBER",          "category": "PII",        "severity": "HIGH",     "preview": "4111...1111" },
    { "type": "EMAIL_ADDRESS",               "category": "PII",        "severity": "MEDIUM",   "preview": "joe@....com" }
  ],
  "sanitizedContent": "{\"authorization\":\"Bearer [REDACTED_AWS_KEY:MPLE]\",\"dsn\":\"postgres://admin:[REDACTED_PASSWORD]@db.internal:5432/prod\",\"card\":\"[REDACTED_CREDIT_CARD:1111]\",\"note\":\"owner [REDACTED_EMAIL:.com]\"}"
}

The policy gate blocks

// enforce_policy("bash", {command: "rm -rf /var/lib/postgresql"})
{
  "decision": "BLOCK",
  "ruleMatched": "ZT-001",
  "reasons": [
    "Violates rule [Block Root and Wildcard Deletion]: Matches strictly forbidden pattern 'rm -rf /'"
  ],
  "requiresToken": false
}

Step-up approval, and the replay it prevents

The same call three times: no token, valid token, same token again.

// 1. no token
{
  "decision": "REQUIRE_CONFIRMATION",
  "ruleMatched": "ZT-003",
  "reasons": ["Elevated risk (Severity: CRITICAL, Danger Score: 95). Cryptographic confirmation token required to proceed."],
  "requiresToken": true,
  "tokenValid": false
}

// request_confirmation_token(...).payload
{
  "tokenId": "br-tok-62b3da7b-67c9-48b4-b766-f1a12d4a25bf",
  "toolName": "bash",
  "actionFingerprint": "95:terraform-destroy",
  "requestedBy": "jimil@acme.dev",
  "issuedAt": 1791190531,
  "expiresAt": 1791190831,
  "reason": "Approved in ticket OPS-4471"
}

// 2. with the token
{
  "decision": "ALLOW",
  "ruleMatched": "ZT-003",
  "reasons": ["Action elevated and approved via valid token: br-tok-62b3da7b-..."],
  "requiresToken": false,
  "tokenValid": true
}

// 3. same token again
{
  "decision": "REQUIRE_CONFIRMATION",
  "reasons": [
    "Provided confirmation token is invalid or expired: Token has already been consumed (replay prevention)",
    "Elevated risk (Severity: CRITICAL, Danger Score: 95). Cryptographic confirmation token required to proceed."
  ],
  "requiresToken": true,
  "tokenValid": false
}

That third block is the whole point. An approval token authorises exactly one action.


The six tools

Tool

What it does

simulate_action

Scores an action 0-100 before execution, reports severity, affected entities, and whether rollback is possible

inspect_payload_dlp

Finds and redacts secrets and PII in any text

enforce_policy

The decision point. Returns ALLOW, BLOCK, or REQUIRE_CONFIRMATION, plus the rule that fired

request_confirmation_token

Issues a signed, single-use, tool-scoped approval token

verify_audit_log

Re-walks the hash chain and its signatures to detect tampering

get_security_posture

Live metrics, active policy, build capabilities


Seeing past the bypass

src/analyzer/shellResolver.ts resolves a command the way a POSIX shell would, to the extent that is possible without executing anything. The rule patterns then run against what the shell would actually run.

The pass performs, in order:

  1. Quote removal. rm -rf "/" and rm -rf / are the same command.

  2. Variable assignment tracking. D=/; rm -rf $D* resolves $D from the assignment earlier in the same command, because the shell would.

  3. Parameter expansion. $HOME, $ROOT, $PWD and friends expand. When the value is not statically knowable it is substituted as a marked UNKNOWN, which is treated as a root-or-home target rather than an opaque token, so a recursive delete aimed at an unseen variable still matches.

  4. Command substitution. $(...) and backticks are resolved to the literal content inside them. This is deliberately conservative: rm -rf $(echo /) resolves to a token containing a slash, which is enough to match. Nothing is ever executed to find out what a substitution would produce.

  5. Pipeline splitting. Stages are separated at the top level only, so echo 'a | b' stays one stage.

  6. Inline program extraction. sh -c takes the whole remainder of the line as its program, so bash -c "rm -rf /" resolves to the destructive command it runs.

  7. Interpreter detection. Any stage that receives content from a pipe and is itself an interpreter is flagged, because the program that runs is not in the request.

On top of that the engine scores structural facts no regex can express:

  • a destructive binary (rm, unlink, shred, srm) aimed at a root, home, or wildcard target

  • shred and srm against any path, because they overwrite in place and so cannot be recovered even without a recursive flag

  • a recursive rm against a system path such as /etc

  • destruction delegated to find via -delete, -exec rm, or -ok

  • content piped into an interpreter, with or without a substitution

  • arguments assembled at runtime by xargs

  • a destructive command composed inline inside sh -c

Every report carries the resolution block, so a reviewer can see the exact text that was matched:

// simulate_action("D=/; rm -rf $D*")
{
  "dangerScore": 100,
  "severity": "CRITICAL",
  "reasons": [
    "Recursive rm against a broad target (-rf /*), resolved through variable or substitution expansion."
  ],
  "resolution": {
    "resolvedCommand": "D=/; rm -rf /*",
    "traits": ["assigned:D", "unresolved-variable:D"],
    "hasPipeline": false,
    "interpreters": [],
    "unresolvedVariables": []
  }
}

Coverage

tests/shellResolver.test.ts pins the GuardFall bypass classes. 23 of 23 dangerous commands that a raw-string guard would miss are scored HIGH or above:

Class

Examples

Variable expansion

D=/; rm -rf $D*, rm -rf $HOME, TARGET=/etc; shred -u $TARGET/*

Command substitution

rm -rf $(echo /), rm -rf \echo /``

Quoting

rm -rf "/", rm -rf '/'*' '/"'

Alternate destructive binaries

unlink /etc/shadow, find / -type f -delete, find . -exec rm {} \;

Indirect execution

cat payload.sh | sh, bash -c "rm -rf /", sh -c "$(curl …)", xargs rm < targets.txt

The same file pins the other direction: 18 routine commands that must not trip, including rm -rf node_modules, rm -rf ./build, kubectl apply -f deployment.yaml, and terraform plan. A guard that blocks everything is not a guard.

What resolution still does not catch

Stated plainly, because this is the tier the literature says gets bypassed and pretending otherwise would repeat the mistake this project already fixed once:

  • Encoding. Base64, hex, and ROT-encoded payloads are not decoded.

  • Multi-hop indirection. bash /tmp/x.sh where x.sh was written three calls ago is opaque to a single-call analysis.

  • A destructive payload assembled character by character at runtime.

  • Anything outside a shell. MCP tool arguments that are not command strings go through scoring, but not through a resolver.

Resolution closes the specific gap GuardFall documented for shell commands. It is not general program analysis, and no honest claim should say it is.



Detection coverage

24 signatures in src/analyzer/rules.ts, scored by rule ID. Scoring takes the max matching rule, not a sum, so overlapping patterns do not inflate a score. All 24 are matched against the resolved command, so the alternative destructive binaries and indirections listed above are reachable without a dedicated rule.

ID

Scores

Covers

SH-001

100

rm -rf /, rm -rf *, ~, $HOME

SH-002

100

mkfs, dd if=… of=/dev/…, fdisk

SH-003

95

Fork bombs

SH-004

90

curl … | sh

SH-005

80

chmod 777 / 666

SH-006

85

shutdown, reboot, poweroff

SH-007

85

git push --force, reset --hard, clean -fdx

SH-008

95

reg delete, del /f /s /q C:\windows

SQL-001

100

DROP DATABASE, DROP SCHEMA

SQL-002

90

DROP TABLE, TRUNCATE

SQL-003

90

DELETE FROM with no WHERE

SQL-004

85

UPDATE … SET with no WHERE

SQL-005

80

GRANT ALL PRIVILEGES

CLD-001

95

terminate-instances, delete-cluster, delete-db-instance, delete-stack

CLD-002

95

kubectl delete namespace|all|nodes|pv|pvc, drain

CLD-003

85

delete-user, delete-role, delete-access-key

CLD-004

95

terraform destroy|apply, pulumi destroy, -auto-approve

CLD-005

95

s3 rb, gsutil rm -r, s3 sync … --delete

CLD-006

85

helm uninstall|delete, kubectl delete

CLD-007

80

npm publish, twine upload, docker push :latest

FS-001

80

.env, .aws/credentials, id_rsa, /etc/shadow, service account JSON

FS-002

90

Mass deletion by synchronisation and pruning: rsync --delete, docker prune, git worktree remove --force

DB-001

95

Destructive database operations issued through a client binary: redis-cli FLUSHALL, mongosh dropDatabase, psql TRUNCATE

SH-009

90

Encoded content piped to a decoder that then executes: base64 -d | sh

Bands: SAFE <15, LOW 15-39, MEDIUM 40-69, HIGH 70-89, CRITICAL 90+. Categories: SHELL, SQL, CLOUD, FILESYSTEM, NETWORK, GENERIC.

11 DLP detectors in src/analyzer/dlpScanner.ts:

  • Credentials: AWS access keys, GitHub PATs (ghp_ and github_pat_), OpenAI, Stripe secret keys, Slack tokens, SSH private keys, database URI passwords, JWTs

  • PII: credit cards (Visa, Mastercard, Amex, Discover, Diners, UnionPay, in contiguous, spaced, or hyphenated form), US SSNs, email addresses

What it does not catch

Stated up front, because a rule set that pretends to be complete is worse than one that does not:

  • Encoded payloads. Base64, hex, and string concatenation defeat both the resolver and the rules. See what resolution still does not catch.

  • Multi-step and multi-call attacks within one call. Sequence detection below covers the audited cases, not arbitrary workflows.

  • Anything semantic. It does not know that orders_archive_2019 matters less than orders.

  • Agent-supplied identity. requestedBy on an approval token is free text from the calling agent. See Step-up approval tokens.

Open an issue if you have a dangerous pattern that is not covered. The rule set is the part most worth improving.


Default policy

Four rules, active out of the box, in src/policy/defaultPolicies.ts.

ID

Action

Detail

ZT-001

BLOCK

rm -rf /, rm -rf *, DROP DATABASE, DROP SCHEMA, mkfs, format c:

ZT-002

BLOCK

Fork bombs, dd if=/dev/zero, del /f /s /q c:\windows

ZT-003

REQUIRE_CONFIRMATION

Any HIGH or CRITICAL action needs a signed token

ZT-004

BLOCK

.env, .env.production, id_rsa, id_ed25519, .aws/credentials, /etc/shadow, service-account.json

Override it with your own JSON via --policy, see examples/custom_policy.json:

npx blastradius-mcp --policy ./my-policy.json

ZT-004 matches with a plain includes() on the serialized parameters. That is deliberately conservative and it does over-block: any path containing .env matches, including .envrc. Tune it with severityThreshold rather than by loosening the pattern.


Sequence detection — the part a single call cannot see

Per-call allow/deny misses workflows assembled from individually benign steps. Microsoft's own analysis of an MCP control plane states the gap directly:

"Per-call allow still misses malicious workflows built from individually permitted steps. Sequence policy is the next real control."

src/analyzer/sequenceDetector.ts is that control. It runs over the hash-chained decision log after the per-call verdict and detects five workflow shapes:

Signal

Fires when

Default escalation

exfiltration-then-destruction

A credential-adjacent read is followed by a destructive or irreversible decision

BLOCK if critical, else CONFIRMATION

reconnaissance-then-exfiltration

A run of ≥5 reads across ≥3 distinct payloads, terminated by a network or DLP-positive sink

CONFIRMATION

privilege-ladder

≥3 distinct privilege families approved in a non-decreasing severity run, or a strictly rising score on one tool

CONFIRMATION

guardrail-probing

≥3 refusals against one tool family

CONFIRMATION, BLOCK at ≥5

destructive-burst

≥4 destructive decisions inside 8 decisions, boosted when the span is under 120 seconds

CONFIRMATION, BLOCK at ≥6

Three properties that matter:

  • It can only tighten a decision, never loosen one. If the per-call rule already blocked or demanded a token, that stands. Sequence analysis never downgrades.

  • Every signal cites its evidence. The reason string names the entry indices that triggered it, so a reviewer can reconstruct the finding rather than trust a score.

  • It reports its own blind spot. truncated: true is returned when the analysis window excludes earlier history, because a pattern beginning before the window cannot be found inside it.

Thresholds are tunable through the same policy document as everything else:

{
  "sequence": {
    "enabled": true,
    "windowSize": 20,
    "destructiveBurstThreshold": 4,
    "probingThreshold": 3
  }
}

What sequence detection still misses

Documented deliberately, because a detector with unstated limits is not a detector you can rely on:

  • Sub-window attacks. Analysis sees only the last windowSize decisions. An agent that reads a secret, runs 21 benign calls, then destroys will set truncated: true and produce no signal.

  • Cross-caller laundering. Nothing correlates by callerId. An attack split across two agent identities resets both the ordering and the burst span.

  • Slow, patient exfiltration. reconnaissance-then-exfiltration requires a contiguous read run immediately before the sink. Five reads interleaved with writes across ten minutes is not a run. The detector is tuned for agent-speed attacks, which is the same thing that makes it weak against a patient one.

  • Destructive classification is reconstructed from severity, score, and reason wording, because the ledger does not persist the engine's destructive flag. Changing reason wording without updating the detector narrows it silently.


Audit ledger

Append-only JSONL, SHA-256 hash chain, one HMAC-SHA256 signature per entry.

  • Payloads are hashed, not stored, so secrets never land in the log.

  • verify_audit_log re-computes every hash and re-computes every signature. The hash chain alone does not cover severity, category, reasons, or dlpFindingsCount; the signature is what covers those, which is why it is verified rather than merely written.

  • Editing any field of any entry breaks verification and reports the index.

The signing key defaults to a random per-process value, so set it explicitly or previously written entries will not verify after a restart:

export BLAST_RADIUS_AUDIT_KEY="a-long-random-secret"
export BLAST_RADIUS_SIGNING_KEY="a-different-long-random-secret"

The ledger defaults to process.cwd(), which under a GUI client is wherever the app happened to launch from. Point it somewhere deliberate:

export BLAST_RADIUS_AUDIT_PATH="/var/log/blastradius/audit.jsonl"

What the chain does not do: detect deletion of entries (the ledger is verified from the in-memory copy), detect truncation, or survive a full rewrite with recomputed hashes. There is no external or Merkle anchoring. Treat it as tamper-evident, not tamper-proof.


Step-up approval tokens

HMAC-SHA256, JWT-shaped payload.signature, 300-second default TTL, hard-capped at one hour.

  • Signature compared with crypto.timingSafeEqual, never ===.

  • Scoped to a single tool name, or * for any tool.

  • Single-use. The token is consumed the moment it authorises an action, so replaying it returns REQUIRE_CONFIRMATION with Token has already been consumed (replay prevention).

  • Consumed records are pruned once the token's own expiry passes, which bounds memory without reopening the hole.

One honest limitation: requestedBy is a free-text string supplied by the calling agent. There is no out-of-band human channel and no identity provider behind it. The token proves an approval was issued and spent once. It does not prove a specific human pressed a button. Wiring it to a real approval service is the natural next step, and it is not built.


Proxy mode

npx blastradius-mcp proxy --command "<downstream mcp server command>"

Inspects inbound tools/call requests and, more usefully, scans downstream responses and redacts credentials before they reach the client. A tool that echoes a key back at your model is a routine leak path, and nothing upstream of it usually cleans up.

A step-up token is accepted as a confirmationToken argument on the call. It is stripped before forwarding, so an approval credential never reaches a server with no business holding it.

Known limitations, stated plainly:

  • The downstream command is split on whitespace, so a path containing a space breaks.

  • It is a line-oriented JSON-RPC passthrough, not a full MCP SDK client.

  • The proxy exposes the downstream server's tools, not BlastRadius's own six.


Testing

npm run test:src
ℹ tests 145
ℹ pass 145
ℹ fail 0

145 tests across nine files, covering the scoring engine, the DLP scanner, the policy engine, the token manager, audit-ledger tamper detection, shell-aware command resolution with the GuardFall bypass classes, false-positive pinning on routine commands, sequence and workflow detection, and a full JSON-RPC lifecycle over a real stdio subprocess. CI runs the suite on Node 22 and 24, with CodeQL on javascript-typescript.

The build is tsc under strict: true. No bundler, no framework, no test-runner dependency.

Three probes are kept in scripts/ because they are more useful as runnable output than as assertions:

node scripts/probe-guardfall.mjs        # 23 bypass classes, one line each
node scripts/probe-false-positives.mjs  # 39 routine commands that must not trip
node scripts/probe-sequence.mjs        # workflow patterns firing, and staying quiet

Configuration

Variable

Default

Purpose

BLAST_RADIUS_SIGNING_KEY

random per process

HMAC secret for approval tokens

BLAST_RADIUS_AUDIT_KEY

random per process

HMAC secret for audit entries

BLAST_RADIUS_AUDIT_PATH

./blastradius-audit.jsonl

Where the ledger is written

Set the first two or tokens and audit entries stop verifying across restarts.

Flag

Purpose

--policy, -p <file>

Load a custom policy JSON

proxy --command "<cmd>"

Guard an existing MCP server

--help, --version


A note on how this shipped

Worth reading, because it is the honest version of what this project is.

The first public version of this repo carried a tests 24/24 passing badge. There were 11 tests, in 2 files. Four of the six test files were 0 bytes: the audit ledger, the policy engine, the token manager, and the end-to-end server had no coverage at all. It also claimed single-use tokens with replay protection, where the function that enforced it existed and was never called. It shipped a 0-byte LICENSE behind an Apache-2.0 badge. And its README advertised rate limits, x402 micropayments, and Splunk SIEM export, none of which existed.

All of it is fixed: 145 real tests, a real CI gate, tokens that are genuinely single-use, an actual Apache-2.0 LICENSE, and a paid-tiers section deleted rather than softened. An Apache-2.0 repository cannot enforce a paywall anyway.

Then there was the second one, which is more interesting.

On 30 June 2026, a paper named the exact design this project used and measured it as broken. GuardFall tested the command guards in eleven coding agents and found that matching the raw command string leaked on most bypass attempts. That was this project's engine. It shipped regexes over submitted text, exactly the tier the survey measured at 22-of-23 and 16-of-16 failure.

So the engine was rewritten to resolve commands before matching them: quote removal, variable assignment tracking, parameter expansion, command substitution, pipeline splitting, and inline program extraction, all without executing anything. All 23 bypass cases now score HIGH or above, pinned by tests/shellResolver.test.ts.

The reason this is in the README rather than a changelog entry is that a tool whose entire pitch is I stop agents doing dangerous things had a bug in the category it exists to prevent, twice. The first time I did not look. The second time I read the literature and found the paper that said so. If you are evaluating this, you should know that the person maintaining it now checks the claims, and how.


Contributing

See CONTRIBUTING.md. The rule set is the highest-value contribution: a new signature in src/analyzer/rules.ts needs a test in the matching tests/*.test.ts and a note about expected false positives.

Security issues: SECURITY.md, or use GitHub private vulnerability reporting.

License

Apache-2.0. See LICENSE.

Available Tools

6 tools
enforce_policyB

The central zero-trust decision point. Evaluates a prospective tool call against organizational security policies, scans parameters for DLP, and records a cryptographic audit trail.

ParametersJSON Schema
NameRequiredDescriptionDefault
callerIdNoUnique identifier for the calling agent or user.agent-client
toolNameYesThe name of the target tool intended to be called.
parametersYesThe parameters intended to be passed to the tool.
confirmationTokenNoOptional cryptographically signed approval token for elevated destructive actions.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses a side effect ('records a cryptographic audit trail'), which tells the agent this is not a pure read, but it omits critical behavior: what a deny decision looks like, whether the call is blocked or merely reported, and how confirmationToken gates elevated destructive actions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with the most important framing front-loaded and zero filler. It could be marginally shorter, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a central security gate with no annotations and no output schema, the description is under-specified: it never says what the tool returns (allow/deny, reason codes) or explains the confirmation-token elevation flow described in the schema. The nested-object parameters argument is also left opaque.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters are already documented in the schema, establishing a baseline of 3. The description only indirectly references 'parameters' as being scanned for DLP and adds no format or constraint detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific function with strong framing ('central zero-trust decision point') plus three concrete actions: policy evaluation, DLP scanning, and audit recording. It is distinguishable from most siblings, though it overlaps conceptually with inspect_payload_dlp (DLP) and simulate_action (evaluation) without naming how it differs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'prospective tool call' implies this is a pre-execution gate, which is useful context, but there is no explicit when-to-use or when-not-to-use guidance and no mention of the sibling tools an agent should consider instead (e.g., simulate_action for dry runs). Usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_security_postureB

Returns current operational security metrics, the active policy profile, build edition and capabilities, blocked threat statistics, and DLP redaction totals.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It implies a read-only snapshot through the verb 'returns', but never states that it is non-mutating, whether it requires elevated privileges, whether results are cached or point-in-time, or how large/fresh the payload is. It adds return-content detail but no operational behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single well-formed sentence that front-loads the action ('Returns current operational security metrics') and then lists the payload contents. No filler or redundancy. Readability is slightly hurt by a long comma-delimited list, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description must explain what comes back, and it does so item by item (metrics, policy profile, edition/capabilities, blocked threats, DLP totals). That covers the return contract well. The missing piece is the behavioral/usage layer — freshness, point-in-time semantics, and when to call it versus its siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, and the schema is trivially complete, so there is nothing for the description to disambiguate. Baseline for a parameterless tool is 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: it returns current security posture data, and it enumerates the concrete contents (operational metrics, active policy profile, build edition/capabilities, blocked threat stats, DLP redaction totals). That is far more informative than a tautology. It does not explicitly contrast itself with siblings such as inspect_payload_dlp or enforce_policy, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool, when not to, or which sibling to prefer for adjacent questions. An agent must infer from the name alone that this is the read-side counterpart to enforce_policy or verify_audit_log. No prerequisites or contexts are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

inspect_payload_dlpC

Deep Data Loss Prevention (DLP) scanner. Detects, reports, and redacts API keys, passwords, JWTs, cloud credentials, credit cards, SSNs, and PII from prompts, code, and logs.

ParametersJSON Schema
NameRequiredDescriptionDefault
contentYesThe raw text, code snippet, or log output to inspect.
maskSensitiveNoWhether to return a sanitized version with credentials and PII masked.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It mentions detection, reporting, and redaction, but does not disclose whether the operation is read-only, whether it modifies the input, any authentication needs, or rate limits. The default maskSensitive=true behavior is only described in the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences with the purpose front-loaded. The list of detectable data types is long but informative and earns its place. There is no wasted wording.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the description should ideally explain the return value (findings format, redacted output) and safety posture. It covers purpose and inputs but leaves output expectations and behavioral constraints unclear, making it only minimally complete for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (content and maskSensitive) are fully documented in the schema. The description adds no additional meaning or syntax beyond what the schema already provides, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb (scanner) and resource (payload for DLP) and enumerates specific sensitive data types it handles. It does not explicitly differentiate from sibling tools like get_security_posture or verify_audit_log, but the payload-inspection focus is distinctive enough for an agent to identify it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use or when-not-to-use guidance, and no alternatives are named. The description implies usage for scanning prompts, code, and logs, but leaves the agent to infer when to call this versus the other security tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

request_confirmation_tokenA

Generates a time-bound, cryptographically signed (HMAC-SHA256) step-up approval token allowing an agent to execute high-impact actions when human consent is granted.

ParametersJSON Schema
NameRequiredDescriptionDefault
reasonNoJustification for elevation.
toolNameYesThe specific tool name to authorize.
ttlSecondsNoToken validity period in seconds (default: 300s).
requestedByYesIdentifier of the human supervisor or authorizing entity.
actionFingerprintYesUnique signature or hash of the approved payload.

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses important traits such as HMAC-SHA256 signing, time-bound validity, and the step-up approval nature, but it omits authentication requirements, rate limits, failure modes, and whether token generation itself is gated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, dense sentence with no redundant phrasing. The key concept—generating a signed step-up token—is front-loaded, making it well-structured for quick scanning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and no annotations, the description is only moderately complete. It explains what the token is for generally, but does not describe the return value, token format, error cases, or how the human consent step is verified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters are already documented in the input schema. The description adds no parameter-level detail beyond what the schema provides, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: generates a time-bound, HMAC-SHA256-signed step-up approval token. It also explains the token's purpose for high-impact actions. However, it does not explicitly differentiate this tool from siblings such as enforce_policy or simulate_action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides a clear context for use: when an agent needs to execute high-impact actions and human consent is granted. It lacks explicit when-not-to-use guidance and does not name alternative tools for related authorization or policy tasks.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

simulate_actionA

Simulates and predicts the blast radius, danger score (0-100), and destructive reversibility of a shell command, SQL query, or cloud operation before execution.

ParametersJSON Schema
NameRequiredDescriptionDefault
contextNoExecution context such as environment (prod/stage), working directory, or target host.
actionTypeNoOptional action category hint.
commandOrQueryYesThe shell command, SQL query, or cloud API call to simulate.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses the return content (blast radius, 0-100 danger score, reversibility), but never states explicitly that the tool does NOT execute the command, whether it needs elevated permissions, or how it handles unknown/unparseable input — key facts for a safety-oriented simulator.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that packs the purpose, accepted input types, and expected outputs with zero filler. Nothing redundant or off-topic.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description correctly takes on the job of naming the return values, which it does. But for a pre-execution risk-assessment tool it omits whether the action is actually executed, the result format, and any annotation-level safety guarantees, leaving meaningful gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters (including the nested context object and the actionType enum) are already documented in the schema. The description only echoes the shell/SQL/cloud mapping, adding no syntax, format, or enum-value meaning beyond the schema; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb (simulates/predicts) and resource (shell command, SQL query, cloud operation), plus the exact outputs produced (blast radius, danger score 0-100, reversibility). It is clearly distinct from the sibling tools like enforce_policy or request_confirmation_token, but it never explicitly names or contrasts an alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'before execution' implies the intended usage window (pre-flight dry run), which is meaningful for an agent deciding when to call this. However, there is no explicit when-not guidance, no prerequisites, and no routing between this and siblings such as enforce_policy or get_security_posture.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_audit_logB

Cryptographically verifies the immutable hash-chained audit ledger to detect tampering, deleted records, or unauthorized log modifications.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoNumber of recent audit records to verify.

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses the verification mechanism (hash-chain) and the classes of anomalies it can detect, which is meaningful context, but it does not state that the operation is read-only/non-mutating, whether it requires elevated permissions, or whether cryptographic verification is expensive or long-running over a large ledger.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with the core action front-loaded and no filler. Every clause (mechanism, target, detected anomalies) earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter tool this is close to adequate, but with no output schema the description should hint at what a verification result contains (pass/fail, offending record index, proof details). It also omits the important scoping caveat implied by 'limit' — that only recent records are checked, so tampering further back could be missed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the single 'limit' parameter is fully documented in the schema as 'Number of recent audit records to verify.' The description adds no semantics about the parameter (e.g., that verifying only a subset leaves earlier ledger history unchecked), so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and mechanism ('cryptographically verifies') applied to a specific resource ('immutable hash-chained audit ledger') and states the outcome ('detect tampering, deleted records, or unauthorized log modifications'). This is far more than a restatement of the name, but it never references any sibling tool (e.g., get_security_posture, inspect_payload_dlp) to help the agent route between them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to invoke this versus alternatives such as get_security_posture or inspect_payload_dlp, and no mention of prerequisites or when verification is unnecessary. The description is purely declarative about what the tool does, leaving the selection decision entirely to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 6 tool updatesv0.1.0
    • First observedenforce_policy
    • First observedget_security_posture
    • First observedinspect_payload_dlp
    • First observedrequest_confirmation_token
    • First observedsimulate_action
    • First observedverify_audit_log

TDQS

A3.5/5.0

Scored across 6 tools

Disambiguation4/5

Each tool has a distinct primary purpose spanning prediction, DLP scanning, posture reporting, policy enforcement, token issuance, and audit verification. There is mild overlap where enforce_policy also scans parameters for DLP and evaluates actions like simulate_action, but the descriptions clarify their respective roles well.

Naming Consistency5/5

All six tools follow a strict verb_noun pattern (simulate_action, inspect_payload_dlp, get_security_posture, enforce_policy, request_confirmation_token, verify_audit_log). The convention is applied uniformly with no deviations.

Tool Count5/5

Six tools is well-scoped for a security guardrail server, with each tool mapping to a discrete responsibility in the pre-execution safety workflow. No tool feels redundant or missing at the count level.

Completeness4/5

The surface covers a full guardrail lifecycle: simulation, DLP inspection, policy enforcement, step-up approval, posture reporting, and audit verification. Minor gaps exist such as no tool to view/configure policy definitions or revoke issued tokens, but core workflows are covered.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    B
    maintenance
    Provides a secure MCP boundary for AI agents, intercepting and validating tool calls, redacting secrets, and requiring human approval for sensitive actions with a tamper-evident audit trail.
    -
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI agents to securely invoke tools by enforcing identity proof, capability verification, and risk scoring on every request, blocking unsafe calls before they execute.
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables autonomous agents to run shell and subprocess commands through a guardrail that blocks destructive system mutations, path traversals, and reverse shells. It also adds prompt-injection detection, spend circuit breakers, loop detection, and privacy sanitization so agent execution stays within safe, budgeted bounds.
    7
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables autonomous agents to enforce financial guardrails by tracking real-time token burn and cost velocity, automatically throttling or freezing execution when budgets are exceeded. It also blocks destructive shell commands, prompt injections, and runaway planning loops while sanitizing outputs before they are returned.
    7
    MIT