PromptBrake Free Tools
Server Details
Free AI security tools: injection payloads, OWASP LLM mapper, ADLC release planner, test builder.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-06-18
- URL
TDQS
Scored across 4 tools
Each tool targets a distinct action: building a test pack, fetching injection payloads, mapping OWASP risks, and planning a release. There is mild overlap between get_prompt_injection_payloads and map_owasp_llm_risk since both surface test prompts, but their framing (attack category vs. OWASP risk) keeps them separable.
All four tools follow a clean snake_case verb_noun pattern (build_test_pack, get_prompt_injection_payloads, map_owasp_llm_risk, plan_adlc_release). The convention is uniform and immediately readable.
Four tools is a small but reasonable set for a free/teaser toolset, and each has a clearly distinct role. It is slightly thin—one could imagine more coverage—but no tool feels redundant or out of place.
The set covers planning, test construction, payloads, and risk mapping, but has notable gaps: nothing runs or evaluates tests, and each payload/risk tool appears scoped to a single category per call. As a standalone domain surface it leaves obvious dead ends.
Available Tools
4 toolsbuild_test_packAI test case builderARead-onlyIdempotentInspect
Validate 1-20 response tests and return a tests.json file for PromptBrake CI. Checks are case-insensitive text comparisons; they do not judge meaning or verify backend actions. Runs statelessly: nothing is stored or executed.
| Name | Required | Description | Default |
|---|---|---|---|
| tests | Yes | 1-20 response tests with unique names. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive/closed-world, and the description adds genuinely new behavioral semantics on top: checks are case-insensitive text comparisons, they do not judge meaning or verify backend actions, and the call is stateless with nothing stored or executed. That prevents an agent from over-trusting the validation result.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action and output artifact, followed by the two caveats that matter most. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only validation tool with no output schema, the description covers the return artifact (tests.json), input bounds, comparison semantics, and statelessness. An agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single nested 'tests' array documents name/check/value/prompt, enums, max lengths and item bounds. The description reinforces the 1-20 limit and case-insensitive comparison but adds no syntax or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action and artifact: validates 1-20 response tests and returns a tests.json file for PromptBrake CI. The verb 'validate' differs slightly from the name 'build_test_pack', but the resource and output are unambiguous. Siblings (payloads, OWASP mapping, release planning) are unrelated enough that no explicit differentiation is needed.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'for PromptBrake CI' situates the tool in a clear workflow, and the 1-20 test range signals when the input is appropriate. It does not state when NOT to use it or name an alternative, but no sibling competes for the same job.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_prompt_injection_payloadsPrompt injection payload libraryARead-onlyIdempotentInspect
Fixed prompt injection test inputs for one attack category, to try against your own chatbot or LLM API.
| Name | Required | Description | Default |
|---|---|---|---|
| category | Yes | Attack category of the payloads to return. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and openWorldHint=false, so the safety profile is fully covered externally. The description adds one useful behavioral fact beyond that – the payloads are "fixed" (static/deterministic) rather than generated – but says nothing about return shape, size, or how many payloads come back.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence naming the artifact, its scope, and its intended use, with no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, deterministic lookup with a fully documented enum parameter and no output schema, the description gives an agent enough to call it correctly. The only minor gap is that it doesn't hint at the quantity or format of the returned payloads.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single parameter is a fully enumerated category, so the schema already carries the semantics. The description only reinforces that exactly one category's payloads are returned per call (implying multiple calls for full coverage), which is a marginal addition. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource (fixed prompt injection test inputs) and a precise scope (one attack category) plus the intended consumer (your own chatbot or LLM API). It is clear what the tool returns, though it never names or contrasts itself with the sibling build_test_pack, which sounds adjacent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"To try against your own chatbot or LLM API" implies the usage context, but there is no explicit when-to-use-this-instead guidance, no mention of when to prefer build_test_pack, and no statement about whether repeated calls are needed for multiple categories.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
map_owasp_llm_riskOWASP LLM test case mapperBRead-onlyIdempotentInspect
Concrete test prompts, signals to look for, and an owner for one OWASP LLM risk category.
| Name | Required | Description | Default |
|---|---|---|---|
| category | Yes | OWASP LLM Top 10 risk id, e.g. llm01. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds useful content context (what is returned: prompts, signals, owner) but does not disclose permissions, rate limits, or other operational behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single, front-loaded phrase with no wasted words. The fragment structure is terse but effective for a simple lookup tool, though it stops short of a complete action statement.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read-only lookup with no output schema, the description adequately says what the tool returns. However, it omits usage guidance against siblings and does not clarify return format or category handling beyond the enum in the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the single category parameter is fully documented with an enum and example. The description's phrase 'one OWASP LLM risk category' maps to the parameter but adds no syntax or format details beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific output bundle (test prompts, signals, owner) for one OWASP LLM risk category, so the resource and scope are clear. It does not explicitly distinguish itself from siblings such as get_prompt_injection_payloads or build_test_pack.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no exclusions, and no mention of alternatives. The agent must infer from the output description alone whether this tool is the right choice versus its siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
plan_adlc_releaseADLC release readiness plannerARead-onlyIdempotentInspect
Turn release decisions for an AI agent or chatbot into a Markdown release plan, a planning-completeness score, open decisions, and a starter PromptBrake gate policy. It is a planning aid, not security validation.
| Name | Required | Description | Default |
|---|---|---|---|
| rollout | Yes | How the release is deployed. | |
| blockers | No | Observed behaviors that must block release; choose at least three. | |
| rollback | Yes | State of the rollback or emergency kill-switch procedure. | |
| agent_name | Yes | Name of the AI system. | |
| test_scope | Yes | Planned behavioral validation before release. | |
| system_type | Yes | Kind of AI system being released. | |
| candidate_id | Yes | Release candidate, e.g. version or commit SHA. | |
| capabilities | No | Capabilities of the release candidate; choose all that apply. | |
| min_executed | No | Default 18. | |
| data_boundary | No | Optional sensitive-data and tenant boundary. | |
| release_owner | Yes | Person or role accountable for the release decision. | |
| human_approval | Yes | Whether high-impact actions require human approval and if it is implemented. | |
| warning_policy | Yes | How validation warnings are treated in the release gate. | |
| evidence_record | Yes | Where release evidence is kept. | |
| feedback_signal | Yes | Production signal that becomes the next regression test. | |
| authority_boundary | Yes | What the system may do, must never do, and which resources it may access. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds genuinely useful context beyond that: it enumerates what the tool returns and warns that its output is advisory, not security validation, which manages agent expectations about the artifact's authority.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences, front-loaded with the transformation and its outputs, followed by a single scoping caveat. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by listing the four return artifacts. Combined with a fully-covered 16-parameter schema and complete annotations, the definition is sufficient to invoke the tool correctly, though the boundary against the security siblings could be made explicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all 16 parameters are already documented in the schema. The description mentions no parameter semantics (required vs optional, enum meanings, the min_executed default), so it adds nothing beyond what the schema provides, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Turn release decisions ... into') and enumerates concrete outputs: a Markdown release plan, a planning-completeness score, open decisions, and a starter PromptBrake gate policy. The 'planning aid, not security validation' line implicitly separates it from the security-oriented siblings, but no sibling is named directly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The clause 'It is a planning aid, not security validation' gives partial when-to-use framing, implying it should not be reached for validation work. However, it never says when to prefer it over build_test_pack or map_owasp_llm_risk, so the routing guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
- First observed
build_test_pack - First observed
get_prompt_injection_payloads - First observed
map_owasp_llm_risk - First observed
plan_adlc_release
Related MCP Connectors
130+ QA & dev tools for AI agents: prompt injection, RAG testing, VLM eval, guardrails. Free.
AI-security knowledge as MCP: standards-mapped tools (OWASP, NIST, MITRE) for AI agents.
Japanese LLM security — prompt injection detection (jpi-guard) + PII masking (PII Guard). Free.
Zero-config MCP security scanner for AI-generated apps. 25K+ vulnerability patterns.
Related MCP Servers
- AlicenseBqualityCmaintenanceProvides local, dependency-free security scanning tools for LLM configurations, prompts, RAG sources, and more, enabling AI coding agents to detect prompt injections and other vulnerabilities without external network access.8MIT
- AlicenseAqualityDmaintenanceSecurity scanning for AI coding tools (Claude Code, Cursor, Windsurf) including secrets detection, MCP config vulnerabilities, agent instruction checks, threat modeling, prompt injection testing, pre-commit security checks, and dependency vulnerability scanning.730 npm1MIT
- AlicenseNot gradedqualityCmaintenanceReal security scanners for AI coding agents — SAST (441 rules), secret detection (419+ patterns), dependency CVEs (OSV.dev), MCP/skill vetting, MITRE ATT&CK. Open-source, Rust, free14 npmApache 2.0
- AlicenseNot gradedqualityDmaintenanceProvides AI agents with 25 security analysis tools including vulnerability scanning, package hallucination detection, prompt injection firewall, and CI/CD integration.1MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.