keel
Keel
将扫描器噪音转化为猎人级、非破坏性证据的 MCP 控制平面
十三个 MCP 工具。语义去重。按主机限速。展示猎人能力的安全证据——不损害目标。
为什么选择 Keel · 安装 · 客户端 · 证据 · 工具
向代理倾倒 150 个工具很容易。真正的难题是跨工具去重、可利用性与噪音的区分,以及不猛击目标。Keel 就是解决这三个问题的控制平面。
AI 客户端与 Keel 对话,而不是与 httpx、nuclei 或 shell 对话。Keel 一次起草一波请求,强制执行范围和速率限制,将扫描器命中合并为语义卡片,并在测试者拥有的数据上运行仅 GET 的剧本。当剧本返回 proven 时,你会得到一个猎人可跟随的 curl 重放——仍然没有写入、shell 或载荷轰炸。
仅在你获得授权测试的项目上使用。
为什么选择 Keel
难题 | 扫描器转储的做法 | Keel 的做法 |
跨工具去重 | 每行一个 Nuclei 模板 ID;同一个 IDOR 出现五次 | 基于漏洞类别 + 规范化路由 + 方法 + 参数的语义键。UUID/id/hex 令牌折叠。兼容的观察结果合并。 |
可利用性与噪音 | 高严重性 = "直接上报" | 卡片经历 |
不猛击目标 | 一次性触发所有模板,429 时重试 | 每主机一个活跃波次,令牌桶,Nuclei 并发为 1,无 OAST,无重定向,无未签名模板,排除 dos/fuzz/bruteforce/intrusive 标签。HTTP 429 变为冷却期。 |
一个包装了庞大工具箱的封装层没有这一层。Keel 有——在调度器、适配器和证据代理中。
Related MCP server: BountyProof MCP
架构
flowchart TD
A[AI coding client] -->|stdio MCP| B[Keel]
B --> C[Scope and rate gate]
C --> W[Background job and wave scheduler]
W --> H[httpx: one target]
W --> N[nuclei: HTTP templates, bounded]
C --> P[Proof broker: GET only]
P --> T[Tester-owned resource]
H --> S[Semantic card store]
N --> S
P --> S
S --> Q[Triage and evidence states]使用你获授权测试的主机名调用
begin_engagement。draft_waves提议可达性加模板微波次。此时不产生流量。execute_wave立即返回一个任务。轮询wave_status。cancel_wave终止扫描器。query_cards返回与猎人相关的卡片。assess_exploitability说明什么可以证明它。draft_proof然后execute_proof针对测试者数据运行仅 GET 的剧本。proven表示不变量成立。protected表示对照生效。
安装
macOS(Homebrew)。pipx 是独立工具——先安装它。Apple 的 /usr/bin/python3 通常是 3.9,无法安装 Keel。
brew install pipx python@3.12
pipx ensurepath
# open a new terminal, then:
pipx install keel-pentest
keel-pentest setup
keel-pentest doctor如果机器上已有 python3.12 且你不想用 Homebrew pipx:
python3.12 -m pip install --user pipx
python3.12 -m pipx ensurepath
python3.12 -m pipx install keel-pentestsetup 将 ProjectDiscovery 的 httpx 和 nuclei 下载到 ~/.keel/bin。即使 GUI 客户端的 PATH 很薄,Keel 也能在那里找到它们。首次扫描无需额外的 KEEL_HTTPX_BIN。
然后将你的 MCP 客户端指向 keel-pentest 可执行文件:
claude mcp add --scope user --transport stdio keel -- keel-pentest
codex mcp add keel -- keel-pentest
hermes mcp add keel --command keel-pentestOpenCode:"command": ["keel-pentest"]。
Python 3.10+。不要 pip install keel——那是另一个项目。操作系统说明和 pip/venv:INSTALL.md。客户端形态:clients/README.md。
可选后续:KEEL_APPROVAL_FILE 用于团队清单,固定范围、模板 ID 和证据目标。默认模式为自我声明——begin_engagement 即授权。速率限制、每主机一波、签名模板和脱敏证据仍然生效。
不会造成损害却仍能证明影响的安全证据
扫描器输出是一个假设。Keel 使用一次性测试者账户和唯一金丝雀来证明(或反驳)它。每个剧本都是仅 GET、有预算的,并返回一个 curl 重放。重放就是报告工件:如果此问题未修复,拥有普通账户的猎人就能做到这一点。
剧本 | 证明什么 | 如何做到,且不造成损害 |
| IDOR / BOLA | 测试者 A 读取其金丝雀;测试者 B 对同一 A 拥有的 URL 发起 GET。金丝雀相同且 2xx = |
| 反射型 XSS / HTML 注入 | 注入唯一标记加一个无害的 |
| 开放重定向 | 将重定向参数指向 |
| 缺少授权 | 测试者 A 基线必须显示金丝雀;同一 URL 在无凭据时必须不显示。2xx + 未认证金丝雀 = |
| 仅可达性 | A 读取自己的金丝雀。这是 |
execute_proof 存储状态码、金丝雀布尔值、截断标志、哈希、猎人影响文本和复现脚本。它不持久化响应体或机密。
在 cross_account_read / unauth_access_probe 之前,在测试者拥有的对象中植入一个非机密金丝雀。反射型 XSS 和开放重定向自行注入标记。
MCP 工具
工具 | 作用 |
| 注册范围和流量上限 |
| 提议可达性 + 模板微波次;不产生流量 |
| 排队一个后台任务 |
| 阶段、进度、结果;省略 |
| 停止排队中或运行中的扫描器 |
| 按优先级排序的语义卡片 |
| 仅重新运行原始 Nuclei 模板 |
| 候选影响、缺失证据、阴性对照、剧本 |
| 记录猎人假设 |
| 白名单证据计划;不产生流量 |
| 运行仅 GET 的剧本 |
| 冷却期、预算、待处理波次 |
| 仅追加的应用程序事件 |
begin_engagement 需要 engagement_id 和 scope_hosts(纯主机名,例如 target.example)。默认值:3 req/s,一次一个主机,每波 120 秒 / 120 个请求。allow_safe_proof=true 启用证据。仅传递测试者凭据名称;将机密放入 KEEL_CREDENTIALS_FILE。
示例提示词
Use only Keel MCP tools. Do not shell out to httpx, nuclei, curl, or exploit tools.
1. begin_engagement for bb-2026-01 with scope_hosts ["target.example"], 3 req/s.
Set allow_safe_proof true if I will run proofs.
2. draft_waves for https://target.example.
3. execute_wave for each wave. Poll wave_status until completed, retryable_failed,
terminal_failed, or cancelled.
4. query_cards (include_noise false), then assess_exploitability on candidates.
5. For a card with a safe playbook, draft_proof then execute_proof using tester
credential names and the canary I planted. Treat protected as refuted.
6. Summarize duplicates, evidence state, hunter_impact, and the repro_script.
Claim exploitable only when Keel reports proven.流量控制
在起草、准入、摄取和证据阶段严格执行范围和排除项
每主机一波;同主机任务等待
共享的全局和每主机令牌桶
持久化请求预留;重试消耗新的预留
Nuclei:签名 HTTP 模板,无 OAST,无重定向,无重试,排除 dos/fuzz/bruteforce/intrusive
隔离的空扫描器配置;剥离代理和 ProjectDiscovery 云环境变量
HTTP 429 停止波次并遵循 Retry-After
有界响应读取;证据不含原始响应体
故障排查
keel-pentest doctor
keel-pentest setup # if doctor reports missing httpx/nuclei客户端重启后调用 begin_engagement 可恢复 SQLite 参与记录。如果更改了范围,请使用新的 engagement_id。
证据需要 allow_safe_proof=true,对于会话剧本,还需要 KEEL_CREDENTIALS_FILE,将 tester-a 之类的名称映射到 Authorization 或 Cookie。
许可证
MIT。版权所有 (c) 2026 Lutfi Z.P.
PyPI:keel-pentest。MCP Registry:io.github.lutfizp/keel。源码:github.com/lutfizp/keel。
Available Tools
9 toolsbegin_engagementC
Register scope, rate limits, and proof flags for one engagement.
| Name | Required | Description | Default |
|---|---|---|---|
| scope_hosts | Yes | ||
| engagement_id | Yes | ||
| exclude_hosts | No | ||
| allow_safe_proof | No | ||
| tester_account_a | No | ||
| tester_account_b | No | ||
| operator_confirmed | No | ||
| requests_per_second | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description does not disclose any side effects, permissions, or error behaviors. Without annotations, the description is insufficient to understand what happens when the tool is invoked.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence without unnecessary words. It is well-structured and easy to read.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool lacks an output schema and annotations, and the description does not mention what the response contains, possible errors, or any other context. It is insufficient for an agent to understand the full behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description mentions 'scope,' 'rate limits,' and 'proof flags' which partially map to parameters like scope_hosts and requests_per_second, but it does not explain the meaning or format of each parameter. The schema has no parameter descriptions, so the description does not compensate adequately.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Register') and the resource ('one engagement'), distinguishing it from siblings that focus on proof execution or health checks.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool over alternatives. It does not mention any preconditions or scenarios where this tool is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
draft_proofD
Describe an allowlisted proof without sending traffic.
| Name | Required | Description | Default |
|---|---|---|---|
| card_id | Yes | ||
| playbook_id | Yes | ||
| engagement_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It hints at being non-destructive by saying 'without sending traffic,' but does not explain what drafting entails, side effects, or response behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is succinct but too sparse to be effective. It lacks necessary detail while also not being well-structured to convey core functionality.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, no parameter descriptions, and a vague purpose, the agent has insufficient information to determine when or how to invoke this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
All three parameters (engagement_id, card_id, playbook_id) have no descriptions in the schema or prose. Coverage is 0%, and the description adds no meaning beyond the parameter names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Describe an allowlisted proof' is vague; 'describe' is not a strong verb for the action, and 'allowlisted proof' is ambiguous. It does not clearly distinguish itself from sibling tools like execute_proof or draft_waves.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only usage hint is a negative constraint ('without sending traffic'), which is insufficient. No positive conditions or comparisons to alternatives (e.g., when to use draft_proof vs execute_proof) are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
draft_wavesA
Propose probe_alive then template_scan waves without executing them.
| Name | Required | Description | Default |
|---|---|---|---|
| seed_url | Yes | ||
| engagement_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of behavioral disclosure. It usefully states that the tool does not execute the waves and specifies the wave order. However, it does not disclose whether the proposal persists, requires permissions, or has any side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single short, front-loaded sentence with no filler. Every word contributes meaning, and the core distinction ('without executing them') is stated directly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with only two required scalar parameters and no output schema, so the description does not need much. It covers the main purpose and non-execution, but it omits what the proposal produces or returns and how the parameters relate to the waves.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides only names and types with no descriptions, and the description never mentions the parameters. 'seed_url' and 'engagement_id' are somewhat self-explanatory, but the 0% schema coverage is not compensated by any parameter-level guidance in the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('propose'), identifies the resource ('waves'), and names the exact wave sequence ('probe_alive then template_scan'). The phrase 'without executing them' clearly distinguishes this tool from the sibling execute_wave.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use when you want to preview or plan waves before execution, and 'without executing them' effectively rules out execute_wave. It does not explicitly name alternatives or state when to switch to execution, but the intended context is reasonably clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
engagement_healthC
Report registered engagements, cooldowns, and pending waves.
| Name | Required | Description | Default |
|---|---|---|---|
| engagement_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. 'Report' suggests a read-only style operation, but the description does not explicitly state that no state changes occur, does not mention auth requirements or side effects, and provides no detail about what 'registered' or 'pending' statuses mean. With no output schema, return behavior is also undisclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single compact sentence with no filler or repetition. It front-loads the verb and the key reported categories, which makes it easy to scan, although it is terse enough that it contributes to under-specification in other dimensions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and no parameter descriptions, this one-line description is not enough for a fully informed call. A no-argument health check is guessable, but the behavior of engagement_id, the meaning of 'registered,' and the relationship to sibling status tools like query_cards and state_impact are left unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description never mentions engagement_id. The agent cannot tell whether the optional parameter filters the report to one engagement, scopes the results, or is required for a valid call. The description adds no meaning beyond the raw property name in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'Report' and names concrete resources: registered engagements, cooldowns, and pending waves. This makes the tool's core purpose clear and distinguishes it from the execution-focused siblings like execute_wave and begin_engagement, though it does not explicitly differentiate it from query-oriented siblings like query_cards.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives. The description states what it reports but does not mention prerequisites, exclusions, or a preferred context such as 'check status before executing a wave.' An agent would have to infer usage from the tool name and sibling names.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
execute_proofC
Run an allowlisted proof. Requires allow_safe_proof and operator_confirmed.
| Name | Required | Description | Default |
|---|---|---|---|
| card_id | Yes | ||
| session_a | Yes | ||
| session_b | No | ||
| playbook_id | Yes | ||
| engagement_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description bears the full burden of behavioral disclosure, and it does reveal a meaningful precondition — an allowlisted proof and operator confirmation — implying an approval gate beyond the schema's surface, which is useful. However, it says nothing about side effects, return values, reversibility, or whether execution is long-running, a notable gap for an 'execute' tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no filler and the preconditions are stated directly. It is concise to the point of thinness — the efficiency is real, but the brevity reflects under-specification rather than disciplined economy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given zero annotations, no output schema, and 0% schema description coverage, the description leaves critical information uncovered: the meaning of a 'proof', expected parameter values, and the outcome of execution. For a 5-parameter (4 required) tool, this is incomplete and would leave an agent uncertain how to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description offers no compensatory explanation of card_id, session_a, session_b, playbook_id, or engagement_id, or how they interrelate. With five unannotated string parameters, the agent is left guessing at values and formats, which the description failed to address.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('run') and a resource ('allowlisted proof'), with the 'allowlisted' qualifier adding an authorization constraint that helps set context. However, it never defines what a 'proof' is or differentiates this from close siblings like execute_wave and draft_proof, leaving the agent to infer the distinction on its own.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description offers no guidance on when to choose this tool over its siblings, despite obvious ambiguity with execute_wave, draft_proof, and draft_waves. The 'Requires allow_safe_proof and operator_confirmed' line reads as a precondition rather than a usage context, and no alternatives or exclusions are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
execute_waveC
Run one admitted wave behind the per-host token bucket.
| Name | Required | Description | Default |
|---|---|---|---|
| wave_id | Yes | ||
| engagement_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of disclosing side effects. 'Run one admitted wave' hints at a mutating action but does not state whether it is idempotent, what happens to the wave, what errors occur, or what the rate limit entails. The token bucket reference suggests throttling but lacks concrete behavioral details expected for an execute-style operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, dense sentence that front-loads the core action with no filler words. Every word contributes to the intended meaning, and it is brief. However, its extreme brevity sacrifices clarity—conciseness is not a substitute for explaining 'admitted' or the token bucket without further context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Even though the tool is simple (2 string params, no output schema), the description fails to cover key aspects like return values, side effects, or the meaning of 'admitted' and 'per-host token bucket.' For a mutating tool with no annotations, more behavioral context is necessary. The lack of any output or error information makes it incomplete for an agent to call this safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not mention either 'wave_id' or 'engagement_id'. There is no explanation of how the parameters influence execution or what 'admitted' means for them. The description provides zero help in understanding parameter semantics, leaving the agent completely reliant on parameter names alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description provides a verb ('run') and a resource ('wave') with additional context about a token bucket, but the meaning of 'admitted wave' is jargon-heavy and unclear without domain knowledge. It does not clearly differentiate from the sibling 'execute_proof'—both suggest executing something. It is not a tautology, but it fails to concretely state what the tool does or what a 'wave' is.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus the siblings. It does not mention alternatives like 'execute_proof' or conditions under which a wave is 'admitted.' The token bucket hint implies rate limiting but does not explain when a user should call this versus other execution tools. No exclusions or prerequisites are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
query_cardsB
Return hunter-relevant cards. Informational and hardening are hidden by default.
| Name | Required | Description | Default |
|---|---|---|---|
| engagement_id | Yes | ||
| include_noise | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description must carry the full burden of behavioral disclosure. It does state a key behavior: informational and hardening cards are hidden by default, which tells the agent about default filtering. However, it does not mention whether the tool is read-only, any permission requirements, rate limits, or failure modes. The non-mutating nature of a 'query' is implied but not explicitly stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely compact at two short sentences, leading with the primary purpose. It avoids redundancy and wastes no words, though it could have used the available space to clarify parameters or usage since it is so brief.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and no annotations, the description is incomplete. It does not describe the return format, possible results, pagination, error cases, or what constitutes 'hunter-relevant'. The single behavioral note about default hiding is helpful but does not make the tool safely callable by an agent that needs to know what to expect or how to interpret the output.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain the parameters. It does not explain what engagement_id is or what include_noise does beyond default false. The text 'Informational and hardening are hidden by default' indirectly suggests include_noise might control showing those, but it never explicitly links the parameter to that behavior. The agent is left guessing about the meaning and usage of both parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool returns 'hunter-relevant cards', which is a clear verb (return/query) and resource (cards). It does not formally distinguish itself from sibling tools, but the action-oriented siblings (execute_proof, draft_waves, etc.) are clearly different, so the purpose is recognizable without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this is a query tool for retrieving cards, but it gives no explicit guidance on when to use it versus alternatives. The note 'Informational and hardening are hidden by default' hints at the include_noise parameter, but it does not explicitly say 'use include_noise when you need these types of cards' or provide any exclusions relative to other tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
second_lookC
Re-run a bounded template scan on a single card URL.
| Name | Required | Description | Default |
|---|---|---|---|
| card_id | Yes | ||
| engagement_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description must carry all behavioral information. It only says 're-run' and 'bounded template scan,' which hints at a read-only operation but does not disclose side effects, auth requirements, rate limits, or return behavior. This is sparse coverage that leaves significant behavioral uncertainty.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, tight sentence with no filler. It is front-loaded with the core action ('re-run') and scope ('bounded template scan'). While it is not verbose, its brevity comes at the cost of missing crucial details, so it earns a 4 for clarity of structure but not a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of annotations, output schema, and parameter descriptions, the description is severely under-informed. It does not explain what an 'engagement' or 'card' is, what a 'template scan' yields, or how to interpret results. For a tool with two required parameters and no output schema, this is insufficient for an agent to call it correctly without additional context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% (parameters have no descriptions), and the description does not compensate. It mentions 'single card URL,' implying card_id is a URL, but leaves engagement_id unexplained. The agent must infer parameter purpose and types from names alone, which is insufficient.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('re-run'), resource ('bounded template scan'), and object ('single card URL'), making the tool's core function clear. It distinguishes implicitly from siblings like execute_proof or execute_wave by emphasizing a 'second look' on a single card, but it does not explicitly name alternatives, so it falls short of full differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 're-run' implies a use case where a previous scan already occurred and a refresh is needed, offering some contextual guidance. However, there is no explicit mention of when to choose this tool over siblings, and no exclusions are stated. The guidance is present but implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
state_impactC
Record hunter impact_class and preconditions on a card.
| Name | Required | Description | Default |
|---|---|---|---|
| impact | Yes | ||
| card_id | Yes | ||
| hunter_why | Yes | ||
| engagement_id | Yes | ||
| preconditions | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses that this is a write operation ('Record'), but with no annotations and no output schema, that's all it reveals. It doesn't specify whether this creates a new record, updates an existing state, requires any authentication, or what happens on repeated calls. For a mutation tool, this is a substantial gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, about eight words, with no filler. It leads with the action and object, making it easy to parse and free of redundant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With five required parameters, no annotations, and no output schema, this description is too minimal to support correct invocation. It doesn't explain what a valid 'preconditions' string looks like, what 'hunter_why' is for, or what the tool returns. The agent would need to inspect external docs or guess.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, but it only mentions two of the five required parameters (impact, preconditions) and doesn't explain formats, constraints, or how they relate. engagement_id, card_id, and hunter_why are absent from the description, and the schema only labels them as strings. This leaves the agent to guess at their meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a concrete verb ('Record') with a specific resource ('a card') and identifies the payload ('hunter impact_class and preconditions'). It distinguishes itself from the sibling tools, none of which address recording impact state. However, it introduces the term 'impact_class' that doesn't appear in the schema ('impact'), and omits the other required fields from the description.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus the siblings. It doesn't state prerequisites, whether it should be called before/after other tools like execute_proof or draft_waves, or any alternative to use instead. The only inference is from the verb 'record', but that's not enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
v0.1.1- First observed
begin_engagement - First observed
draft_proof - First observed
draft_waves - First observed
engagement_health - First observed
execute_proof - First observed
execute_wave - First observed
query_cards - First observed
second_look - First observed
state_impact
TDQS
Scored across 9 tools
Each tool targets a distinct resource and action: engagements, waves, proofs, and cards are cleanly separated. The two execute tools are disambiguated by proof vs wave, and the two draft tools by waves vs proof, so an agent should not confuse them.
Most tools follow a clear verb_noun snake_case pattern like execute_proof, begin_engagement, draft_waves, and query_cards. engagement_health and second_look break the pattern by using noun phrases, but the overall convention remains readable and predictable.
Nine tools is a well-scoped size for this domain, covering engagement setup, wave and proof execution, and card interaction without unnecessary redundancy or bloat.
The core loop is represented, but there is no explicit admission step for drafted waves before execute_wave, and engagements lack update/close lifecycle operations. These are notable workflow gaps that could stall end-to-end operations.
Maintenance
Related MCP Connectors
MCP server for Pentest-Tools.com: run scans, manage findings and reports via your preffered LLM.
Find, vet, and run MCP tools through a secure audited gateway with prompt-injection risk scoring
AI pentesting: run scans, triage vulnerabilities, review PRs, manage schedules and assets.
Enrich, search, assess, and manage threat intelligence through 80+ typed MCP tools.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceAutonomous pentests from one command: real security tools, working PoCs, and audit-ready reports, all driven via MCP.271 PyPI1,716MIT
- AlicenseBqualityDmaintenanceAn MCP server for authorized bug bounty work that enforces an evidence-driven workflow with session management, preflight checks, surface discovery, and verified scanning.12MIT
- AlicenseNot gradedqualityCmaintenanceEnables automated bug bounty hunting and security research with tools for reconnaissance, web vulnerability scanning, API testing, binary analysis, and mobile app analysis through an MCP interface.MIT
- AlicenseNot gradedqualityCmaintenanceEnables authorized penetration testing through MCP, providing parallel reconnaissance, vulnerability scanning, attack path analysis, and self-contained HTML reporting with compliance tagging.MIT