Skip to main content
Glama

Keel

스캐너 노이즈를 헌터급 비파괴 증명으로 바꾸는 MCP 컨트롤 플레인

Python License MCP PyPI Registry Version

열세 개의 MCP 도구. 의미 기반 중복 제거. 호스트별 속도 제한. 헌터가 무엇을 할 수 있는지 보여주는 안전한 증명 — 대상에 피해를 주지 않으면서.

왜 Keel인가 · 설치 · 클라이언트 · 증명 · 도구


에이전트에 150개의 도구를 던지는 것은 쉽습니다. 진짜 어려운 문제는 도구 간 중복 제거, 실제 악용 가능 여부 vs 노이즈, 그리고 대상을 두드리지 않는 것입니다. Keel은 이 세 가지를 위한 컨트롤 플레인입니다.

AI 클라이언트는 httpx, nuclei, 또는 셸이 아니라 Keel과 대화합니다. Keel은 한 번에 하나의 웨이브를 설계하고, 범위와 속도 제한을 적용하며, 스캐너 히트를 의미 기반 카드로 병합하고, 테스터 소유 데이터에 대해 GET 전용 플레이북을 실행합니다. 플레이북이 proven을 반환하면 헌터가 따라 할 수 있는 curl 재현 명령을 얻습니다 — 여전히 쓰기, 셸, 페이로드 스팸 없이.

테스트 권한이 있는 프로그램에만 사용하세요.

왜 Keel인가

어려운 문제

스캐너 덤프가 하는 일

Keel이 하는 일

도구 간 중복 제거

행마다 Nuclei 템플릿 ID 하나; 동일한 IDOR이 다섯 번 나타남

취약점 클래스 + 정규화된 경로 + 메서드 + 파라미터에서 의미 키 생성. UUID/id/hex 토큰은 축소됨. 호환 가능한 관측값은 병합됨.

악용 가능 vs 노이즈

높은 심각도 = "바로 적용"

카드는 observed → hypothesis → corroborated → proven / refuted 단계를 거침. 정보성 및 강화 항목은 숨겨짐. 유형화된 악용 가능성은 누락된 증거와 음성 대조군을 명명함.

대상을 두드리지 않기

모든 템플릿을 한 번에 발사, 429에서 재시도

호스트당 활성 웨이브 하나, 토큰 버킷, Nuclei 동시성 1, OAST 없음, 리다이렉트 없음, 서명되지 않은 템플릿 없음, dos/fuzz/bruteforce/intrusive 태그 제외. HTTP 429는 쿨다운이 됨.

거대한 도구 상자에 셸로 연결하는 래퍼에는 그런 계층이 없습니다. Keel에는 있습니다 — 스케줄러, 어댑터, 증명 브로커에.

Related MCP server: BountyProof MCP

아키텍처

flowchart TD
    A[AI coding client] -->|stdio MCP| B[Keel]
    B --> C[Scope and rate gate]
    C --> W[Background job and wave scheduler]
    W --> H[httpx: one target]
    W --> N[nuclei: HTTP templates, bounded]
    C --> P[Proof broker: GET only]
    P --> T[Tester-owned resource]
    H --> S[Semantic card store]
    N --> S
    P --> S
    S --> Q[Triage and evidence states]
  1. 테스트 권한이 있는 호스트 이름으로 begin_engagement을 시작합니다.

  2. draft_waves가 도달 가능성과 템플릿 마이크로 웨이브를 제안합니다. 아직 트래픽은 없습니다.

  3. execute_wave가 즉시 작업을 반환합니다. wave_status로 폴링합니다. cancel_wave가 스캐너를 중단합니다.

  4. query_cards가 헌터 관련 카드를 반환합니다. assess_exploitability가 무엇이 증명이 될지 말해줍니다.

  5. draft_proof 후 execute_proof가 테스터 데이터에 대해 GET 전용 플레이북을 실행합니다. proven은 불변식이 유지되었음을 의미합니다. protected는 대조군이 작동했음을 의미합니다.

설치

macOS (Homebrew). pipx는 별도 도구입니다 — 먼저 설치하세요. Apple의 /usr/bin/python3은 종종 3.9이며 Keel을 설치할 수 없습니다.

brew install pipx python@3.12
pipx ensurepath
# open a new terminal, then:
pipx install keel-pentest
keel-pentest setup
keel-pentest doctor

python3.12이 이미 머신에 있고 Homebrew pipx를 원하지 않는다면:

python3.12 -m pip install --user pipx
python3.12 -m pipx ensurepath
python3.12 -m pipx install keel-pentest

setup은 ProjectDiscovery httpx와 nuclei를 ~/.keel/bin에 다운로드합니다. GUI 클라이언트가 얇은 PATH를 가져도 Keel은 그곳에서 찾습니다. 첫 스캔에 추가 KEEL_HTTPX_BIN은 필요 없습니다.

그런 다음 MCP 클라이언트를 keel-pentest 실행 파일에 연결합니다:

claude mcp add --scope user --transport stdio keel -- keel-pentest
codex mcp add keel -- keel-pentest
hermes mcp add keel --command keel-pentest

OpenCode: "command": ["keel-pentest"].

Python 3.10+. pip install keel은 하지 마세요 — 그것은 다른 프로젝트입니다. OS 참고 사항 및 pip/venv: INSTALL.md. 클라이언트 형태: clients/README.md.

선택 사항(나중에): 팀 매니페스트로 범위, 템플릿 ID, 증명 대상을 고정하는 KEEL_APPROVAL_FILE. 기본 모드는 자기 증명(self-attested)입니다 — begin_engagement이 인가입니다. 속도 제한, 호스트당 웨이브 하나, 서명된 템플릿, 정화된 증거는 여전히 적용됩니다.

여전히 영향력을 증명하는 안전한 증명

스캐너 출력은 가설입니다. Keel은 일회용 테스터 계정과 고유 카나리로 그것을 증명(또는 반증)합니다. 모든 플레이북은 GET 전용, 예산 제한, curl 재현 명령 반환입니다. 재현 명령이 보고서 산출물입니다: 이것이 수정되지 않으면, 일반 계정을 가진 헌터가 이것을 할 수 있습니다.

플레이북

증명하는 것

피해 없이 증명하는 방법

cross_account_read

IDOR / BOLA

테스터 A가 자신의 카나리를 읽음; 테스터 B가 동일한 A 소유 URL에 GET 요청. 동일한 카나리 + 2xx = proven. 401/403/404 또는 카나리 없는 2xx = protected (반증됨).

reflected_marker

반사형 XSS / HTML 인젝션

고유 마커와 무해한 <keel> 프로브를 주입. 이스케이프되지 않은 반사 = proven. HTML 인코딩된 반사 = protected. 대조군 GET에는 프로브가 이미 포함되어 있지 않아야 함.

open_redirect_canary

오픈 리다이렉트

리다이렉트 파라미터를 https://keel-proof.invalid/<marker>로 지정. Location 호스트가 그 카나리인 3xx = proven.

unauth_access_probe

인증 부재

테스터 A 기준선에 카나리가 표시되어야 함; 자격 증명 없는 동일 URL에는 표시되지 않아야 함. 인증 없는 2xx + 카나리 = proven.

own_session_marker

도달 가능성만

A가 자신의 카나리를 읽음. 이것은 corroborated이며, 취약점 증명이 절대 아님.

execute_proof는 상태 코드, 카나리 불리언, 잘림 플래그, 해시, 헌터 영향 텍스트, 재현 스크립트를 저장합니다. 응답 본문이나 비밀은 영속화하지 않습니다.

cross_account_read / unauth_access_probe 전에 테스터 소유 객체에 비밀 아닌 카나리를 심으세요. 반사형 XSS와 오픈 리다이렉트는 마커를 스스로 주입합니다.

MCP 도구

도구

역할

begin_engagement

범위와 트래픽 상한 등록

draft_waves

도달 가능성 + 템플릿 마이크로 웨이브 제안; 트래픽 없음

execute_wave

백그라운드 작업 하나 대기열에 추가

wave_status

단계, 진행률, 결과; job_id 생략 시 목록 표시

cancel_wave

대기 중이거나 실행 중인 스캐너 중지

query_cards

우선순위가 매겨진 의미 카드

second_look

원래 Nuclei 템플릿만 재실행

assess_exploitability

후보 영향, 누락된 증거, 음성 대조군, 플레이북

state_impact

헌터 가설 기록

draft_proof

허용 목록 증명 계획; 트래픽 없음

execute_proof

GET 전용 플레이북 실행

engagement_health

쿨다운, 예산, 대기 중인 웨이브

engagement_audit

추가 전용 애플리케이션 이벤트

begin_engagement에는 engagement_id와 scope_hosts(일반 호스트 이름, 예: target.example)가 필요합니다. 기본값: 3 req/s, 한 번에 호스트 하나, 웨이브당 120초 / 120개 요청. allow_safe_proof=true가 증명을 활성화합니다. 테스터 자격 증명 이름만 전달하세요. 비밀은 KEEL_CREDENTIALS_FILE에 넣으세요.

예시 프롬프트

Use only Keel MCP tools. Do not shell out to httpx, nuclei, curl, or exploit tools.

1. begin_engagement for bb-2026-01 with scope_hosts ["target.example"], 3 req/s.
   Set allow_safe_proof true if I will run proofs.
2. draft_waves for https://target.example.
3. execute_wave for each wave. Poll wave_status until completed, retryable_failed,
   terminal_failed, or cancelled.
4. query_cards (include_noise false), then assess_exploitability on candidates.
5. For a card with a safe playbook, draft_proof then execute_proof using tester
   credential names and the canary I planted. Treat protected as refuted.
6. Summarize duplicates, evidence state, hunter_impact, and the repro_script.
   Claim exploitable only when Keel reports proven.

트래픽 제어

  • 초안, 승인, 수집, 증명 시 정확한 범위 및 제외

  • 호스트당 웨이브 하나; 동일 호스트 작업은 대기

  • 공유 전역 및 호스트별 토큰 버킷

  • 지속적 요청 예약; 재시도는 새 예약을 소비

  • Nuclei: 서명된 HTTP 템플릿, OAST 없음, 리다이렉트 없음, 재시도 없음, dos/fuzz/bruteforce/intrusive 제외

  • 격리된 빈 스캐너 구성; 프록시 및 ProjectDiscovery-cloud 환경 변수 제거

  • HTTP 429는 웨이브를 중지하고 Retry-After를 준수

  • 제한된 응답 읽기; 원시 본문 없는 증거

문제 해결

keel-pentest doctor
keel-pentest setup    # if doctor reports missing httpx/nuclei

클라이언트 재시작 후 begin_engagement은 SQLite 참여 기록을 복원합니다. 범위를 변경했다면 새 engagement_id를 사용하세요.

증명에는 allow_safe_proof=true가 필요하며, 세션 플레이북의 경우 KEEL_CREDENTIALS_FILE에 tester-a 같은 이름을 Authorization 또는 Cookie에 매핑해야 합니다.

keel MCP server

라이선스

MIT. Copyright (c) 2026 Lutfi Z.P.

PyPI: keel-pentest. MCP 레지스트리: io.github.lutfizp/keel. 소스: github.com/lutfizp/keel.

keel MCP server

Available Tools

9 tools
begin_engagementC

Register scope, rate limits, and proof flags for one engagement.

ParametersJSON Schema
NameRequiredDescriptionDefault
scope_hostsYes
engagement_idYes
exclude_hostsNo
allow_safe_proofNo
tester_account_aNo
tester_account_bNo
operator_confirmedNo
requests_per_secondNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description does not disclose any side effects, permissions, or error behaviors. Without annotations, the description is insufficient to understand what happens when the tool is invoked.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence without unnecessary words. It is well-structured and easy to read.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool lacks an output schema and annotations, and the description does not mention what the response contains, possible errors, or any other context. It is insufficient for an agent to understand the full behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description mentions 'scope,' 'rate limits,' and 'proof flags' which partially map to parameters like scope_hosts and requests_per_second, but it does not explain the meaning or format of each parameter. The schema has no parameter descriptions, so the description does not compensate adequately.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Register') and the resource ('one engagement'), distinguishing it from siblings that focus on proof execution or health checks.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool over alternatives. It does not mention any preconditions or scenarios where this tool is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

draft_proofD

Describe an allowlisted proof without sending traffic.

ParametersJSON Schema
NameRequiredDescriptionDefault
card_idYes
playbook_idYes
engagement_idYes

TDQS

D1.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It hints at being non-destructive by saying 'without sending traffic,' but does not explain what drafting entails, side effects, or response behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is succinct but too sparse to be effective. It lacks necessary detail while also not being well-structured to convey core functionality.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, no parameter descriptions, and a vague purpose, the agent has insufficient information to determine when or how to invoke this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

All three parameters (engagement_id, card_id, playbook_id) have no descriptions in the schema or prose. Coverage is 0%, and the description adds no meaning beyond the parameter names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Describe an allowlisted proof' is vague; 'describe' is not a strong verb for the action, and 'allowlisted proof' is ambiguous. It does not clearly distinguish itself from sibling tools like execute_proof or draft_waves.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only usage hint is a negative constraint ('without sending traffic'), which is insufficient. No positive conditions or comparisons to alternatives (e.g., when to use draft_proof vs execute_proof) are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

draft_wavesA

Propose probe_alive then template_scan waves without executing them.

ParametersJSON Schema
NameRequiredDescriptionDefault
seed_urlYes
engagement_idYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of behavioral disclosure. It usefully states that the tool does not execute the waves and specifies the wave order. However, it does not disclose whether the proposal persists, requires permissions, or has any side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short, front-loaded sentence with no filler. Every word contributes meaning, and the core distinction ('without executing them') is stated directly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple with only two required scalar parameters and no output schema, so the description does not need much. It covers the main purpose and non-execution, but it omits what the proposal produces or returns and how the parameters relate to the waves.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides only names and types with no descriptions, and the description never mentions the parameters. 'seed_url' and 'engagement_id' are somewhat self-explanatory, but the 0% schema coverage is not compensated by any parameter-level guidance in the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('propose'), identifies the resource ('waves'), and names the exact wave sequence ('probe_alive then template_scan'). The phrase 'without executing them' clearly distinguishes this tool from the sibling execute_wave.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies use when you want to preview or plan waves before execution, and 'without executing them' effectively rules out execute_wave. It does not explicitly name alternatives or state when to switch to execution, but the intended context is reasonably clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

engagement_healthC

Report registered engagements, cooldowns, and pending waves.

ParametersJSON Schema
NameRequiredDescriptionDefault
engagement_idNo

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. 'Report' suggests a read-only style operation, but the description does not explicitly state that no state changes occur, does not mention auth requirements or side effects, and provides no detail about what 'registered' or 'pending' statuses mean. With no output schema, return behavior is also undisclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single compact sentence with no filler or repetition. It front-loads the verb and the key reported categories, which makes it easy to scan, although it is terse enough that it contributes to under-specification in other dimensions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, no output schema, and no parameter descriptions, this one-line description is not enough for a fully informed call. A no-argument health check is guessable, but the behavior of engagement_id, the meaning of 'registered,' and the relationship to sibling status tools like query_cards and state_impact are left unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description never mentions engagement_id. The agent cannot tell whether the optional parameter filters the report to one engagement, scopes the results, or is required for a valid call. The description adds no meaning beyond the raw property name in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the specific verb 'Report' and names concrete resources: registered engagements, cooldowns, and pending waves. This makes the tool's core purpose clear and distinguishes it from the execution-focused siblings like execute_wave and begin_engagement, though it does not explicitly differentiate it from query-oriented siblings like query_cards.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives. The description states what it reports but does not mention prerequisites, exclusions, or a preferred context such as 'check status before executing a wave.' An agent would have to infer usage from the tool name and sibling names.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

execute_proofC

Run an allowlisted proof. Requires allow_safe_proof and operator_confirmed.

ParametersJSON Schema
NameRequiredDescriptionDefault
card_idYes
session_aYes
session_bNo
playbook_idYes
engagement_idYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description bears the full burden of behavioral disclosure, and it does reveal a meaningful precondition — an allowlisted proof and operator confirmation — implying an approval gate beyond the schema's surface, which is useful. However, it says nothing about side effects, return values, reversibility, or whether execution is long-running, a notable gap for an 'execute' tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no filler and the preconditions are stated directly. It is concise to the point of thinness — the efficiency is real, but the brevity reflects under-specification rather than disciplined economy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given zero annotations, no output schema, and 0% schema description coverage, the description leaves critical information uncovered: the meaning of a 'proof', expected parameter values, and the outcome of execution. For a 5-parameter (4 required) tool, this is incomplete and would leave an agent uncertain how to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description offers no compensatory explanation of card_id, session_a, session_b, playbook_id, or engagement_id, or how they interrelate. With five unannotated string parameters, the agent is left guessing at values and formats, which the description failed to address.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('run') and a resource ('allowlisted proof'), with the 'allowlisted' qualifier adding an authorization constraint that helps set context. However, it never defines what a 'proof' is or differentiates this from close siblings like execute_wave and draft_proof, leaving the agent to infer the distinction on its own.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no guidance on when to choose this tool over its siblings, despite obvious ambiguity with execute_wave, draft_proof, and draft_waves. The 'Requires allow_safe_proof and operator_confirmed' line reads as a precondition rather than a usage context, and no alternatives or exclusions are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

execute_waveC

Run one admitted wave behind the per-host token bucket.

ParametersJSON Schema
NameRequiredDescriptionDefault
wave_idYes
engagement_idYes

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of disclosing side effects. 'Run one admitted wave' hints at a mutating action but does not state whether it is idempotent, what happens to the wave, what errors occur, or what the rate limit entails. The token bucket reference suggests throttling but lacks concrete behavioral details expected for an execute-style operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, dense sentence that front-loads the core action with no filler words. Every word contributes to the intended meaning, and it is brief. However, its extreme brevity sacrifices clarity—conciseness is not a substitute for explaining 'admitted' or the token bucket without further context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Even though the tool is simple (2 string params, no output schema), the description fails to cover key aspects like return values, side effects, or the meaning of 'admitted' and 'per-host token bucket.' For a mutating tool with no annotations, more behavioral context is necessary. The lack of any output or error information makes it incomplete for an agent to call this safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not mention either 'wave_id' or 'engagement_id'. There is no explanation of how the parameters influence execution or what 'admitted' means for them. The description provides zero help in understanding parameter semantics, leaving the agent completely reliant on parameter names alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description provides a verb ('run') and a resource ('wave') with additional context about a token bucket, but the meaning of 'admitted wave' is jargon-heavy and unclear without domain knowledge. It does not clearly differentiate from the sibling 'execute_proof'—both suggest executing something. It is not a tautology, but it fails to concretely state what the tool does or what a 'wave' is.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus the siblings. It does not mention alternatives like 'execute_proof' or conditions under which a wave is 'admitted.' The token bucket hint implies rate limiting but does not explain when a user should call this versus other execution tools. No exclusions or prerequisites are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

query_cardsB

Return hunter-relevant cards. Informational and hardening are hidden by default.

ParametersJSON Schema
NameRequiredDescriptionDefault
engagement_idYes
include_noiseNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description must carry the full burden of behavioral disclosure. It does state a key behavior: informational and hardening cards are hidden by default, which tells the agent about default filtering. However, it does not mention whether the tool is read-only, any permission requirements, rate limits, or failure modes. The non-mutating nature of a 'query' is implied but not explicitly stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely compact at two short sentences, leading with the primary purpose. It avoids redundancy and wastes no words, though it could have used the available space to clarify parameters or usage since it is so brief.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema and no annotations, the description is incomplete. It does not describe the return format, possible results, pagination, error cases, or what constitutes 'hunter-relevant'. The single behavioral note about default hiding is helpful but does not make the tool safely callable by an agent that needs to know what to expect or how to interpret the output.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain the parameters. It does not explain what engagement_id is or what include_noise does beyond default false. The text 'Informational and hardening are hidden by default' indirectly suggests include_noise might control showing those, but it never explicitly links the parameter to that behavior. The agent is left guessing about the meaning and usage of both parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool returns 'hunter-relevant cards', which is a clear verb (return/query) and resource (cards). It does not formally distinguish itself from sibling tools, but the action-oriented siblings (execute_proof, draft_waves, etc.) are clearly different, so the purpose is recognizable without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies this is a query tool for retrieving cards, but it gives no explicit guidance on when to use it versus alternatives. The note 'Informational and hardening are hidden by default' hints at the include_noise parameter, but it does not explicitly say 'use include_noise when you need these types of cards' or provide any exclusions relative to other tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

second_lookC

Re-run a bounded template scan on a single card URL.

ParametersJSON Schema
NameRequiredDescriptionDefault
card_idYes
engagement_idYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description must carry all behavioral information. It only says 're-run' and 'bounded template scan,' which hints at a read-only operation but does not disclose side effects, auth requirements, rate limits, or return behavior. This is sparse coverage that leaves significant behavioral uncertainty.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, tight sentence with no filler. It is front-loaded with the core action ('re-run') and scope ('bounded template scan'). While it is not verbose, its brevity comes at the cost of missing crucial details, so it earns a 4 for clarity of structure but not a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the absence of annotations, output schema, and parameter descriptions, the description is severely under-informed. It does not explain what an 'engagement' or 'card' is, what a 'template scan' yields, or how to interpret results. For a tool with two required parameters and no output schema, this is insufficient for an agent to call it correctly without additional context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% (parameters have no descriptions), and the description does not compensate. It mentions 'single card URL,' implying card_id is a URL, but leaves engagement_id unexplained. The agent must infer parameter purpose and types from names alone, which is insufficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('re-run'), resource ('bounded template scan'), and object ('single card URL'), making the tool's core function clear. It distinguishes implicitly from siblings like execute_proof or execute_wave by emphasizing a 'second look' on a single card, but it does not explicitly name alternatives, so it falls short of full differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 're-run' implies a use case where a previous scan already occurred and a refresh is needed, offering some contextual guidance. However, there is no explicit mention of when to choose this tool over siblings, and no exclusions are stated. The guidance is present but implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

state_impactC

Record hunter impact_class and preconditions on a card.

ParametersJSON Schema
NameRequiredDescriptionDefault
impactYes
card_idYes
hunter_whyYes
engagement_idYes
preconditionsYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses that this is a write operation ('Record'), but with no annotations and no output schema, that's all it reveals. It doesn't specify whether this creates a new record, updates an existing state, requires any authentication, or what happens on repeated calls. For a mutation tool, this is a substantial gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, about eight words, with no filler. It leads with the action and object, making it easy to parse and free of redundant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With five required parameters, no annotations, and no output schema, this description is too minimal to support correct invocation. It doesn't explain what a valid 'preconditions' string looks like, what 'hunter_why' is for, or what the tool returns. The agent would need to inspect external docs or guess.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, but it only mentions two of the five required parameters (impact, preconditions) and doesn't explain formats, constraints, or how they relate. engagement_id, card_id, and hunter_why are absent from the description, and the schema only labels them as strings. This leaves the agent to guess at their meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a concrete verb ('Record') with a specific resource ('a card') and identifies the payload ('hunter impact_class and preconditions'). It distinguishes itself from the sibling tools, none of which address recording impact state. However, it introduces the term 'impact_class' that doesn't appear in the schema ('impact'), and omits the other required fields from the description.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus the siblings. It doesn't state prerequisites, whether it should be called before/after other tools like execute_proof or draft_waves, or any alternative to use instead. The only inference is from the verb 'record', but that's not enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 9 tool updatesv0.1.1
    • First observedbegin_engagement
    • First observeddraft_proof
    • First observeddraft_waves
    • First observedengagement_health
    • First observedexecute_proof
    • First observedexecute_wave
    • First observedquery_cards
    • First observedsecond_look
    • First observedstate_impact

TDQS

C2.9/5.0

Scored across 9 tools

Disambiguation5/5

Each tool targets a distinct resource and action: engagements, waves, proofs, and cards are cleanly separated. The two execute tools are disambiguated by proof vs wave, and the two draft tools by waves vs proof, so an agent should not confuse them.

Naming Consistency4/5

Most tools follow a clear verb_noun snake_case pattern like execute_proof, begin_engagement, draft_waves, and query_cards. engagement_health and second_look break the pattern by using noun phrases, but the overall convention remains readable and predictable.

Tool Count5/5

Nine tools is a well-scoped size for this domain, covering engagement setup, wave and proof execution, and card interaction without unnecessary redundancy or bloat.

Completeness3/5

The core loop is represented, but there is no explicit admission step for drafted waves before execute_wave, and engagements lack update/close lifecycle operations. These are notable workflow gaps that could stall end-to-end operations.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    A
    maintenance
    Autonomous pentests from one command: real security tools, working PoCs, and audit-ready reports, all driven via MCP.
    271 PyPI
    1,716
    MIT
  • A
    license
    B
    quality
    D
    maintenance
    An MCP server for authorized bug bounty work that enforces an evidence-driven workflow with session management, preflight checks, surface discovery, and verified scanning.
    12
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables automated bug bounty hunting and security research with tools for reconnaissance, web vulnerability scanning, API testing, binary analysis, and mobile app analysis through an MCP interface.
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables authorized penetration testing through MCP, providing parallel reconnaissance, vulnerability scanning, attack path analysis, and self-contained HTML reporting with compliance tagging.
    MIT